Skip to content

Environments

View Markdown

An Environment is a versioned set of tasks with a grader: a prompt for each task, and a program that decides whether a completion is correct. TrainingJobs use Environments for reinforcement learning and evaluation, and agent evaluations use the same ones. Environments are generally available (nodus.dev/v1).

Terminal window
$ nodus get environments -n nodus
NAME VERSION CATEGORY MODES PHASE
graph-coloring 1.0.0 Reasoning Train, Evaluate Ready
arithmetic 2.0.0 Math Train, Evaluate Ready
gsm8k 1.0.0 Math Train, Evaluate Ready
reasoning-gym 1.0.0 Reasoning Train, Evaluate Ready
python-functions 1.0.0 Code Train, Evaluate Ready
$ nodus describe environment/graph-coloring -n nodus

describe shows the summary, licenses, the size of each split, a sample task, the graders and the examples: ready TrainingJob templates with measured results. The console’s Training page lists the same catalog, and “Run example” starts a TrainingJob from one; in Python, nodus.TrainingJob.from_example("nodus/graph-coloring@1.0.0") does the same.

Name the environment and version, and how many tasks to train and evaluate on:

spec:
runtime: nodus/grpo-lora
environment:
name: nodus/graph-coloring@1.0.0
trainTasks: 50 # from the train split
heldOutTasks: 64 # from the test split, never trained on
seed: 42

The same name, version and seed always give the same tasks in the same order, so two runs are comparable. The train and test splits never share a task.

  • Tasks carry a prompt and non-secret metadata only. Expected answers stay with the grader; the trainer, the model and your code never see them.
  • Completions are graded by Nodus, not by the trainer. Program graders run in grading Sandboxes that Nodus creates for your TrainingJob or agent evaluation: no network access, deleted after five minutes idle, at most grading.maxParallel at once. Code from a completion runs as a separate user that cannot read the answers. ExactMatch graders compare the completion with the answer inside Nodus, without a Sandbox.
  • Every verdict is Correct, Incorrect, InvalidOutput or InfrastructureFailure, with a reward and evidence recorded by the host that ran the grader: the process, its exit code, its duration and a hash of its output.
  • InvalidOutput is a completion the grader cannot parse, such as an answer without the expected tag. It is a failed task with a reward of 0.
  • InfrastructureFailure means the grader could not run. It carries no reward and does not count toward a pass rate, so an outage never looks like a wrong answer.
  • Grading Sandboxes are billed to the TrainingJob and count toward its maxCostUSD. They are deleted when the run finishes or is suspended.

When your tasks and reward are in Python, rl.train is all you need. To see it work first, run the example that ships with the SDK:

Terminal window
python -m nodus.examples.rl

It is one file, nodus/examples/rl.py, and it is the template for your own task:

import re
from nodus.recipes import rl
ANSWER = re.compile(r"<answer>\s*([A-Za-z]+)\s*</answer>")
tasks = [(f"Spell the word backwards inside <answer></answer>.\n\nWord: {w}", w[::-1]) for w in WORDS]
def reward(completion, answer):
found = ANSWER.findall(completion)
return None if not found else float(found[-1].lower() == answer)
run = rl.train(tasks, reward, max_cost=2)
run.watch() # each stage, then the reward, loss and KL of every step
run.outputs.download("./outputs") # the LoRA adapter and the before-and-after comparison

run.tasks(phase="Evaluation", outcome="Failed") lists the graded tasks the run has scored so far, each with its taskId, reward and outcome, so you can read which tasks the trained model still gets wrong.

  • Tasks are (prompt, answer) pairs or {"prompt", "answer", "metadata"} dicts. Only reward sees the answer.
  • Model. The base model is Qwen/Qwen3-0.6B unless you pass model=, for example model="Qwen/Qwen3-4B".
  • Held-out tasks. A fifth of the tasks, and at least 16, is held out. The model is graded on them before and after training. Pass test_tasks= to choose them yourself.
  • What gets packaged. The reward goes with the lines of its own file that it uses: imports, constants and helper functions. A name it cannot take along fails before anything is built.
  • Checked before it runs. rl.train grades one task’s answer and an empty reply on your machine first. A reward that raises, or returns something other than a number, a bool or None, fails there instead of on a GPU. Return None or 0 when a reply holds no answer.
  • Reuse. The same code and tasks reuse the same Environment, so a second run starts at once.
  • No network. The reward runs in Nodus’s grader, which has no network access. Pass pip=["package==1.2.3"] for packages it imports.
  • How it learns. Each step samples group_size=8 replies to each of groups_per_step=4 tasks and scores every reply against the others in its group. Replies are capped at max_tokens=256, and the learning rate is learning_rate=4e-5. You can change any of them.
  • Format rewards. To reward the shape of an answer as well as its value, fold both into the reward, for example return float(correct) - 0.1 * (not well_formed).
  • Other options. steps, gpu, max_cost, lora and the other grpo_lora parameters are optional. The run appears under Runs › Post-training in the console, with its curves, results and cost.

To train on a catalog Environment, pass its name instead of tasks and a reward:

run = rl.train("nodus/gsm8k@1.0.0", max_cost=5)

Your own tasks and reward are a Python module with two functions. tasks(split) lists the prompts of the train or test split with their answers, and reward(completion, answer) scores one completion:

def tasks(split):
return [{"prompt": "Spell the word backwards...\n\nWord: valley", "answer": "yellav"}, ...]
def reward(completion, answer):
found = ANSWER.findall(completion)
if not found:
return None # no answer in the completion: InvalidOutput
return 1.0 if found[-1].lower() == answer else 0.0

A reward of 1 (or True) is Correct, anything lower Incorrect, and None InvalidOutput. If reward raises, the verdict is InfrastructureFailure, never a reward of 0. Answers reach only reward; the trainer and the model see the prompt and the optional metadata. A test prompt never appears in train, even if your lists repeat it.

The quickest way in needs no Docker: publish the module as a pip package whose load_environment() returns an object (or the module itself) with tasks and reward, and name the exact version. Nodus builds it into an image when you apply the Environment; if the package does not load, the build fails and the Image env-<name>-<version> shows the step and its log:

apiVersion: nodus.dev/v1
kind: Environment
metadata:
name: reverse-words
spec:
version: 1.0.0
package:
pip: {name: reverse-words, version: 1.0.0} # index: for a private index
loader: load_environment # the default; or module:function
category: Custom
modes: [Train, Evaluate]
rewardType: Binary

Run an environment from the Environments Hub

Section titled “Run an environment from the Environments Hub”

An environment published on Prime Intellect’s Environments Hub is a pip package, so it runs on Nodus as it is. Name it with the Hub’s index for its owner and the version the Hub lists:

apiVersion: nodus.dev/v1
kind: Environment
metadata:
name: reverse-text
spec:
version: 0.1.4
package:
pip:
name: reverse-text
version: 0.1.4
index: https://hub.primeintellect.ai/primeintellect/simple/
category: Custom
modes: [Train, Evaluate]

Its load_environment() returns a verifiers environment, and Nodus trains on it directly: the dataset is the train split, the eval_dataset the held-out one (with only one of them, every fifth task is held out), and each completion is scored by the environment’s own rubric, with verifiers 0.1 through 0.3. A multi-turn environment, such as a game like wordle, runs too: see Train on a multi-turn environment, and so does one where the model calls tools: see Train on a tool-calling environment. An environment whose tools run in a remote sandbox, a browser or an MCP server fails the build with that reason, as does one scored by another model (a JudgeRubric), since grading has no network and holds no credential for that model. The console marks each kind in its Hub search from the Hub’s own tags. The package comes from the Hub index alone and its dependencies may also come from PyPI.

Grading has no network. Nodus fetches the datasets an environment loads while it builds the image, and every task list and grade after that reads that copy, so a run always sees the data its build saw.

In a multi-turn environment the model and the environment take turns: the model replies, the environment answers with feedback, and the episode goes on until the environment ends it. Set maxTurns on the TrainingJob (Turns per episode under Advanced in the console) to let an episode run that many model turns:

spec:
runtime: nodus/grpo-lora
environment: {name: wordle@0.1.7, trainTasks: 512, heldOutTasks: 64, seed: 42}
parameters:
maxTurns: 6 # model turns per episode
maxCompletionLength: 256 # tokens per reply
maxEpisodeTokens: 2048 # tokens of the whole episode after the prompt, the model's and the environment's

The reward is the one the environment gives the whole episode, and training learns only from the model’s own tokens. An episode that reaches either cap ends there and is graded as it stands. With the default maxTurns: 1, only the first reply is graded. Baseline and final evaluation play the same episodes greedily.

An OpenEnv environment runs in process from a package whose load_environment() returns it. Each task is a seed: Nodus resets the environment with it, and every action is the model’s reply as the one text field of the environment’s action class. The environment must play the same episode for the same seed. Seeds 0 to 799 train and 800 to 999 are held out, unless the environment sets train_seeds and test_seeds:

import nltk
from textarena_env.server.environment import TextArenaEnvironment
def load_environment():
try: # the build downloads NLTK's word lists; grading reads them offline
nltk.data.find("corpora/words")
nltk.data.find("taggers/averaged_perceptron_tagger_eng")
cached = True
except LookupError:
cached = False
return TextArenaEnvironment("Wordle-v0", download_nltk=not cached)

A module of your own can also be multi-turn: give it step(turns, answer) in place of reward. It gets every model turn so far and returns the environment’s next message, or None once the episode ended, and the reward so far. Grading keeps no state between turns, so step replays the turns from the start.

In a verifiers ToolEnv the model calls the environment’s tools and reads their results. Each task’s prompt offers the tools through the model’s chat template, the model calls one by writing a <tool_call>{"name": "...", "arguments": {...}}</tool_call> block (the form Qwen and most open chat templates teach), and the environment runs the call in its own code. The results are the next turn, and the episode ends when the model answers without calling a tool. Set maxTurns to the rounds of calls an episode may make:

spec:
runtime: nodus/grpo-lora
environment: {name: tool-test@0.1.1, trainTasks: 21, heldOutTasks: 43, seed: 42}
parameters:
maxTurns: 3
maxCompletionLength: 256

The tools run inside the grader, which has no network, so a tool that searches the web or runs code in a remote sandbox cannot train. Pick a model whose chat template supports tools.

The catalog’s nodus/reasoning-gym serves five reviewed families. To train on any other Reasoning Gym family, or a mix of them, publish a package whose load_environment() returns the dataset; its own score_answer scores each completion:

import reasoning_gym
def load_environment():
return reasoning_gym.create_dataset("knights_knaves", size=2000, seed=7)

Every fifth entry is held out. The model’s last <answer>…</answer> is its answer, or the whole completion without one; partial credit is the reward and only a full score is Correct. Each family’s own licence applies.

A dataset prepared for verl or SkyRL runs with its reward unchanged. Publish a module that names the parquet files with verl’s own keys, and your verl reward function under its usual name:

train_files = "hf://datasets/BytedTsinghua-SIA/DAPO-Math-17k/data/dapo-math-17k.parquet"
val_files = "hf://datasets/BytedTsinghua-SIA/AIME-2024/data/aime-2024.parquet" # optional
def compute_score(data_source, solution_str, ground_truth, extra_info=None):
... # verl's custom reward signature; a dict's "score" also works

Each row is verl’s: prompt (a system message and one user message at most), data_source, reward_model.ground_truth and extra_info. Without compute_score, each row’s env_class names the SkyRL-gym environment that scores it, built from the row as SkyRL builds it; add skyrl-gym to the package’s dependencies. Without val_files, every fifth row is held out. The package needs datasets as a dependency, and a SkyRL environment of more than one turn fails the build with that reason.

To ship your own system packages or files, build the module into an image on the env-base image instead, push it and name its digest:

FROM ghcr.io/nodus-compute/catalog/env-base:1.0.0
COPY reverse_words.py /opt/environment/
ENV NODUS_ENVIRONMENT_LOADER=reverse_words NODUS_ENVIRONMENT=reverse-words NODUS_ENVIRONMENT_VERSION=1.0.0
RUN nodus-env info
ENV HF_HUB_OFFLINE=1 HF_DATASETS_OFFLINE=1
apiVersion: nodus.dev/v1
kind: Environment
metadata:
name: reverse-words
spec:
version: 1.0.0
package: {image: registry.example.com/acme/reverse-words@sha256:...}
category: Custom
modes: [Train, Evaluate]
rewardType: Binary
Terminal window
$ nodus apply -f environment.yaml
$ nodus get environment/reverse-words -w # Ready once the image digest is verified

A TrainingJob names it without the nodus/ prefix (environment: {name: reverse-words@1.0.0}), and in Python rl.grpo_lora(environment="reverse-words@1.0.0", ...). The whole example, with a GRPO TrainingJob, is in examples/training/custom-reward. Nodus pulls the image from a public registry or from your organization’s space in the Nodus registry; other private registries are not supported for Environments yet.

Any image works if it provides the two commands Nodus runs, as uid 10001 with no network:

  • nodus-env tasks --split train|test --seed N writes one JSON line per task: {"taskId", "prompt", "metadata"}.
  • nodus-env grade reads {"taskId", "completion"} lines and writes one {"taskId", "verdict", "reward", "evidence"} line for each, in order.

The Environment becomes Ready once its image is verified, with the declared split sizes in status.splits. The first TrainingJob that uses a split and seed runs nodus-env tasks in one of its own grading Sandboxes and Nodus keeps that manifest for every later run of your organization, so the tasks never change between runs. A version’s image and graders cannot change: publish a new version instead, so earlier results stay reproducible.