# Environments

> Task sets with graders for reinforcement learning and evaluation, from the Nodus catalog or your own image.

Source: https://nodus-platform-site.pages.dev/docs/guides/environments/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

An Environment is a versioned set of tasks with a grader: a prompt for each task, and a program that decides whether a completion is correct. TrainingJobs use Environments for reinforcement learning and evaluation, and agent evaluations use the same ones. Environments are generally available (`nodus.dev/v1`).

## Browse the catalog

Terminal window

```console
$ nodus get environments -n nodus
NAME               VERSION   CATEGORY    MODES             PHASE
graph-coloring     1.0.0     Reasoning   Train, Evaluate   Ready
arithmetic         2.0.0     Math        Train, Evaluate   Ready
gsm8k              1.0.0     Math        Train, Evaluate   Ready
reasoning-gym      1.0.0     Reasoning   Train, Evaluate   Ready
python-functions   1.0.0     Code        Train, Evaluate   Ready
$ nodus describe environment/graph-coloring -n nodus
```

`describe` shows the summary, licenses, the size of each split, a sample task, the graders and the examples: ready TrainingJob templates with measured results. The console’s Training page lists the same catalog, and “Run example” starts a TrainingJob from one; in Python, `nodus.TrainingJob.from_example("nodus/graph-coloring@1.0.0")` does the same.

## Use one in a TrainingJob

Name the environment and version, and how many tasks to train and evaluate on:

```yaml
spec:
  runtime: nodus/grpo-lora
  environment:
    name: nodus/graph-coloring@1.0.0
    trainTasks: 50       # from the train split
    heldOutTasks: 64     # from the test split, never trained on
    seed: 42
```

The same name, version and seed always give the same tasks in the same order, so two runs are comparable. The train and test splits never share a task.

## How grading works

* Tasks carry a prompt and non-secret metadata only. Expected answers stay with the grader; the trainer, the model and your code never see them.
* Completions are graded by Nodus, not by the trainer. Program graders run in grading Sandboxes that Nodus creates for your TrainingJob or agent evaluation: no network access, deleted after five minutes idle, at most `grading.maxParallel` at once. Code from a completion runs as a separate user that cannot read the answers. `ExactMatch` graders compare the completion with the answer inside Nodus, without a Sandbox.
* Every verdict is `Correct`, `Incorrect`, `InvalidOutput` or `InfrastructureFailure`, with a reward and evidence recorded by the host that ran the grader: the process, its exit code, its duration and a hash of its output.
* `InvalidOutput` is a completion the grader cannot parse, such as an answer without the expected tag. It is a failed task with a reward of 0.
* `InfrastructureFailure` means the grader could not run. It carries no reward and does not count toward a pass rate, so an outage never looks like a wrong answer.
* Grading Sandboxes are billed to the TrainingJob and count toward its `maxCostUSD`. They are deleted when the run finishes or is suspended.

## Train on a reward function in one call

When your tasks and reward are in Python, `rl.train` is all you need. To see it work first, run the example that ships with the SDK:

Terminal window

```sh
python -m nodus.examples.rl
```

It is one file, `nodus/examples/rl.py`, and it is the template for your own task:

```python
import re
from nodus.recipes import rl

ANSWER = re.compile(r"<answer>\s*([A-Za-z]+)\s*</answer>")
tasks = [(f"Spell the word backwards inside <answer></answer>.\n\nWord: {w}", w[::-1]) for w in WORDS]

def reward(completion, answer):
    found = ANSWER.findall(completion)
    return None if not found else float(found[-1].lower() == answer)

run = rl.train(tasks, reward, max_cost=2)
run.watch()                          # each stage, then the reward, loss and KL of every step
run.outputs.download("./outputs")    # the LoRA adapter and the before-and-after comparison
```

`run.tasks(phase="Evaluation", outcome="Failed")` lists the graded tasks the run has scored so far, each with its `taskId`, `reward` and outcome, so you can read which tasks the trained model still gets wrong.

* **Tasks** are `(prompt, answer)` pairs or `{"prompt", "answer", "metadata"}` dicts. Only `reward` sees the answer.
* **Model.** The base model is `Qwen/Qwen3-0.6B` unless you pass `model=`, for example `model="Qwen/Qwen3-4B"`.
* **Held-out tasks.** A fifth of the tasks, and at least 16, is held out. The model is graded on them before and after training. Pass `test_tasks=` to choose them yourself.
* **What gets packaged.** The reward goes with the lines of its own file that it uses: imports, constants and helper functions. A name it cannot take along fails before anything is built.
* **Checked before it runs.** `rl.train` grades one task’s answer and an empty reply on your machine first. A reward that raises, or returns something other than a number, a bool or `None`, fails there instead of on a GPU. Return `None` or `0` when a reply holds no answer.
* **Reuse.** The same code and tasks reuse the same Environment, so a second run starts at once.
* **No network.** The reward runs in Nodus’s grader, which has no network access. Pass `pip=["package==1.2.3"]` for packages it imports.
* **How it learns.** Each step samples `group_size=8` replies to each of `groups_per_step=4` tasks and scores every reply against the others in its group. Replies are capped at `max_tokens=256`, and the learning rate is `learning_rate=4e-5`. You can change any of them.
* **Format rewards.** To reward the shape of an answer as well as its value, fold both into the reward, for example `return float(correct) - 0.1 * (not well_formed)`.
* **Other options.** `steps`, `gpu`, `max_cost`, `lora` and the other `grpo_lora` parameters are optional. The run appears under **Runs › Post-training** in the console, with its curves, results and cost.

To train on a catalog Environment, pass its name instead of tasks and a reward:

```python
run = rl.train("nodus/gsm8k@1.0.0", max_cost=5)
```

## Bring your own environment

Your own tasks and reward are a Python module with two functions. `tasks(split)` lists the prompts of the `train` or `test` split with their answers, and `reward(completion, answer)` scores one completion:

```python
def tasks(split):
    return [{"prompt": "Spell the word backwards...\n\nWord: valley", "answer": "yellav"}, ...]

def reward(completion, answer):
    found = ANSWER.findall(completion)
    if not found:
        return None                      # no answer in the completion: InvalidOutput
    return 1.0 if found[-1].lower() == answer else 0.0
```

A reward of 1 (or `True`) is `Correct`, anything lower `Incorrect`, and `None` `InvalidOutput`. If `reward` raises, the verdict is `InfrastructureFailure`, never a reward of 0. Answers reach only `reward`; the trainer and the model see the prompt and the optional `metadata`. A test prompt never appears in train, even if your lists repeat it.

The quickest way in needs no Docker: publish the module as a pip package whose `load_environment()` returns an object (or the module itself) with `tasks` and `reward`, and name the exact version. Nodus builds it into an image when you apply the Environment; if the package does not load, the build fails and the Image `env-<name>-<version>` shows the step and its log:

```yaml
apiVersion: nodus.dev/v1
kind: Environment
metadata:
  name: reverse-words
spec:
  version: 1.0.0
  package:
    pip: {name: reverse-words, version: 1.0.0}   # index: for a private index
    loader: load_environment                     # the default; or module:function
  category: Custom
  modes: [Train, Evaluate]
  rewardType: Binary
```

## Run an environment from the Environments Hub

An environment published on [Prime Intellect’s Environments Hub](https://app.primeintellect.ai/dashboard/environments) is a pip package, so it runs on Nodus as it is. Name it with the Hub’s index for its owner and the version the Hub lists:

```yaml
apiVersion: nodus.dev/v1
kind: Environment
metadata:
  name: reverse-text
spec:
  version: 0.1.4
  package:
    pip:
      name: reverse-text
      version: 0.1.4
      index: https://hub.primeintellect.ai/primeintellect/simple/
  category: Custom
  modes: [Train, Evaluate]
```

Its `load_environment()` returns a `verifiers` environment, and Nodus trains on it directly: the `dataset` is the train split, the `eval_dataset` the held-out one (with only one of them, every fifth task is held out), and each completion is scored by the environment’s own rubric, with `verifiers` 0.1 through 0.3. A multi-turn environment, such as a game like `wordle`, runs too: see [Train on a multi-turn environment](https://nodus-platform-site.pages.dev/docs/guides/environments/#train-on-a-multi-turn-environment), and so does one where the model calls tools: see [Train on a tool-calling environment](https://nodus-platform-site.pages.dev/docs/guides/environments/#train-on-a-tool-calling-environment). An environment whose tools run in a remote sandbox, a browser or an MCP server fails the build with that reason, as does one scored by another model (a `JudgeRubric`), since grading has no network and holds no credential for that model. The console marks each kind in its Hub search from the Hub’s own tags. The package comes from the Hub index alone and its dependencies may also come from PyPI.

Grading has no network. Nodus fetches the datasets an environment loads while it builds the image, and every task list and grade after that reads that copy, so a run always sees the data its build saw.

## Train on a multi-turn environment

In a multi-turn environment the model and the environment take turns: the model replies, the environment answers with feedback, and the episode goes on until the environment ends it. Set `maxTurns` on the TrainingJob (**Turns per episode** under **Advanced** in the console) to let an episode run that many model turns:

```yaml
spec:
  runtime: nodus/grpo-lora
  environment: {name: wordle@0.1.7, trainTasks: 512, heldOutTasks: 64, seed: 42}
  parameters:
    maxTurns: 6              # model turns per episode
    maxCompletionLength: 256 # tokens per reply
    maxEpisodeTokens: 2048   # tokens of the whole episode after the prompt, the model's and the environment's
```

The reward is the one the environment gives the whole episode, and training learns only from the model’s own tokens. An episode that reaches either cap ends there and is graded as it stands. With the default `maxTurns: 1`, only the first reply is graded. Baseline and final evaluation play the same episodes greedily.

An [OpenEnv](https://github.com/meta-pytorch/OpenEnv) environment runs in process from a package whose `load_environment()` returns it. Each task is a seed: Nodus resets the environment with it, and every action is the model’s reply as the one text field of the environment’s action class. The environment must play the same episode for the same seed. Seeds 0 to 799 train and 800 to 999 are held out, unless the environment sets `train_seeds` and `test_seeds`:

```python
import nltk
from textarena_env.server.environment import TextArenaEnvironment

def load_environment():
    try:  # the build downloads NLTK's word lists; grading reads them offline
        nltk.data.find("corpora/words")
        nltk.data.find("taggers/averaged_perceptron_tagger_eng")
        cached = True
    except LookupError:
        cached = False
    return TextArenaEnvironment("Wordle-v0", download_nltk=not cached)
```

A module of your own can also be multi-turn: give it `step(turns, answer)` in place of `reward`. It gets every model turn so far and returns the environment’s next message, or `None` once the episode ended, and the reward so far. Grading keeps no state between turns, so `step` replays the turns from the start.

## Train on a tool-calling environment

In a `verifiers` `ToolEnv` the model calls the environment’s tools and reads their results. Each task’s prompt offers the tools through the model’s chat template, the model calls one by writing a `<tool_call>{"name": "...", "arguments": {...}}</tool_call>` block (the form Qwen and most open chat templates teach), and the environment runs the call in its own code. The results are the next turn, and the episode ends when the model answers without calling a tool. Set `maxTurns` to the rounds of calls an episode may make:

```yaml
spec:
  runtime: nodus/grpo-lora
  environment: {name: tool-test@0.1.1, trainTasks: 21, heldOutTasks: 43, seed: 42}
  parameters:
    maxTurns: 3
    maxCompletionLength: 256
```

The tools run inside the grader, which has no network, so a tool that searches the web or runs code in a remote sandbox cannot train. Pick a model whose chat template supports tools.

## Run any Reasoning Gym family

The catalog’s `nodus/reasoning-gym` serves five reviewed families. To train on any other [Reasoning Gym](https://github.com/open-thought/reasoning-gym) family, or a mix of them, publish a package whose `load_environment()` returns the dataset; its own `score_answer` scores each completion:

```python
import reasoning_gym

def load_environment():
    return reasoning_gym.create_dataset("knights_knaves", size=2000, seed=7)
```

Every fifth entry is held out. The model’s last `<answer>…</answer>` is its answer, or the whole completion without one; partial credit is the reward and only a full score is `Correct`. Each family’s own licence applies.

## Run a verl or SkyRL dataset

A dataset prepared for [verl](https://github.com/volcengine/verl) or [SkyRL](https://github.com/NovaSky-AI/SkyRL) runs with its reward unchanged. Publish a module that names the parquet files with verl’s own keys, and your verl reward function under its usual name:

```python
train_files = "hf://datasets/BytedTsinghua-SIA/DAPO-Math-17k/data/dapo-math-17k.parquet"
val_files = "hf://datasets/BytedTsinghua-SIA/AIME-2024/data/aime-2024.parquet"   # optional

def compute_score(data_source, solution_str, ground_truth, extra_info=None):
    ...                                  # verl's custom reward signature; a dict's "score" also works
```

Each row is verl’s: `prompt` (a system message and one user message at most), `data_source`, `reward_model.ground_truth` and `extra_info`. Without `compute_score`, each row’s `env_class` names the SkyRL-gym environment that scores it, built from the row as SkyRL builds it; add `skyrl-gym` to the package’s dependencies. Without `val_files`, every fifth row is held out. The package needs `datasets` as a dependency, and a SkyRL environment of more than one turn fails the build with that reason.

To ship your own system packages or files, build the module into an image on the `env-base` image instead, push it and name its digest:

```dockerfile
FROM ghcr.io/nodus-compute/catalog/env-base:1.0.0
COPY reverse_words.py /opt/environment/
ENV NODUS_ENVIRONMENT_LOADER=reverse_words NODUS_ENVIRONMENT=reverse-words NODUS_ENVIRONMENT_VERSION=1.0.0
RUN nodus-env info
ENV HF_HUB_OFFLINE=1 HF_DATASETS_OFFLINE=1
```

```yaml
apiVersion: nodus.dev/v1
kind: Environment
metadata:
  name: reverse-words
spec:
  version: 1.0.0
  package: {image: registry.example.com/acme/reverse-words@sha256:...}
  category: Custom
  modes: [Train, Evaluate]
  rewardType: Binary
```

Terminal window

```console
$ nodus apply -f environment.yaml
$ nodus get environment/reverse-words -w      # Ready once the image digest is verified
```

A TrainingJob names it without the `nodus/` prefix (`environment: {name: reverse-words@1.0.0}`), and in Python `rl.grpo_lora(environment="reverse-words@1.0.0", ...)`. The whole example, with a GRPO TrainingJob, is in [`examples/training/custom-reward`](https://github.com/nodus-compute/nodus-platform/tree/main/examples/training/custom-reward). Nodus pulls the image from a public registry or from your organization’s space in the Nodus registry; other private registries are not supported for Environments yet.

Any image works if it provides the two commands Nodus runs, as uid 10001 with no network:

* `nodus-env tasks --split train|test --seed N` writes one JSON line per task: `{"taskId", "prompt", "metadata"}`.
* `nodus-env grade` reads `{"taskId", "completion"}` lines and writes one `{"taskId", "verdict", "reward", "evidence"}` line for each, in order.

The Environment becomes `Ready` once its image is verified, with the declared split sizes in `status.splits`. The first TrainingJob that uses a split and seed runs `nodus-env tasks` in one of its own grading Sandboxes and Nodus keeps that manifest for every later run of your organization, so the tasks never change between runs. A version’s image and graders cannot change: publish a new version instead, so earlier results stay reproducible.
