Environments
View MarkdownAn Environment is a versioned set of tasks with a grader: a prompt for each task, and a program that decides
whether a completion is correct. TrainingJobs use Environments for reinforcement learning and evaluation, and agent
evaluations use the same ones. Environments are generally available (nodus.dev/v1).
Browse the catalog
Section titled “Browse the catalog”$ nodus get environments -n nodusNAME VERSION CATEGORY MODES PHASEgraph-coloring 1.0.0 Reasoning Train, Evaluate Readyarithmetic 2.0.0 Math Train, Evaluate Readygsm8k 1.0.0 Math Train, Evaluate Readyreasoning-gym 1.0.0 Reasoning Train, Evaluate Readypython-functions 1.0.0 Code Train, Evaluate Ready$ nodus describe environment/graph-coloring -n nodusdescribe shows the summary, licenses, the size of each split, a sample task, the graders and the examples: ready
TrainingJob templates with measured results. The console’s Training page lists the same catalog, and “Run example”
starts a TrainingJob from one; in Python, nodus.TrainingJob.from_example("nodus/graph-coloring@1.0.0") does the
same.
Use one in a TrainingJob
Section titled “Use one in a TrainingJob”Name the environment and version, and how many tasks to train and evaluate on:
spec: runtime: nodus/grpo-lora environment: name: nodus/graph-coloring@1.0.0 trainTasks: 50 # from the train split heldOutTasks: 64 # from the test split, never trained on seed: 42The same name, version and seed always give the same tasks in the same order, so two runs are comparable. The train and test splits never share a task.
How grading works
Section titled “How grading works”- Tasks carry a prompt and non-secret metadata only. Expected answers stay with the grader; the trainer, the model and your code never see them.
- Completions are graded by Nodus, not by the trainer. Program graders run in grading Sandboxes that Nodus creates
for your TrainingJob or agent evaluation: no network access, deleted after five minutes idle, at most
grading.maxParallelat once. Code from a completion runs as a separate user that cannot read the answers.ExactMatchgraders compare the completion with the answer inside Nodus, without a Sandbox. - Every verdict is
Correct,Incorrect,InvalidOutputorInfrastructureFailure, with a reward and evidence recorded by the host that ran the grader: the process, its exit code, its duration and a hash of its output. InvalidOutputis a completion the grader cannot parse, such as an answer without the expected tag. It is a failed task with a reward of 0.InfrastructureFailuremeans the grader could not run. It carries no reward and does not count toward a pass rate, so an outage never looks like a wrong answer.- Grading Sandboxes are billed to the TrainingJob and count toward its
maxCostUSD. They are deleted when the run finishes or is suspended.
Train on a reward function in one call
Section titled “Train on a reward function in one call”When your tasks and reward are in Python, rl.train is all you need. To see it work first, run the example that
ships with the SDK:
python -m nodus.examples.rlIt is one file, nodus/examples/rl.py, and it is the template for your own task:
import refrom nodus.recipes import rl
ANSWER = re.compile(r"<answer>\s*([A-Za-z]+)\s*</answer>")tasks = [(f"Spell the word backwards inside <answer></answer>.\n\nWord: {w}", w[::-1]) for w in WORDS]
def reward(completion, answer): found = ANSWER.findall(completion) return None if not found else float(found[-1].lower() == answer)
run = rl.train(tasks, reward, max_cost=2)run.watch() # each stage, then the reward, loss and KL of every steprun.outputs.download("./outputs") # the LoRA adapter and the before-and-after comparisonrun.tasks(phase="Evaluation", outcome="Failed") lists the graded tasks the run has scored so far, each with its
taskId, reward and outcome, so you can read which tasks the trained model still gets wrong.
- Tasks are
(prompt, answer)pairs or{"prompt", "answer", "metadata"}dicts. Onlyrewardsees the answer. - Model. The base model is
Qwen/Qwen3-0.6Bunless you passmodel=, for examplemodel="Qwen/Qwen3-4B". - Held-out tasks. A fifth of the tasks, and at least 16, is held out. The model is graded on them before and
after training. Pass
test_tasks=to choose them yourself. - What gets packaged. The reward goes with the lines of its own file that it uses: imports, constants and helper functions. A name it cannot take along fails before anything is built.
- Checked before it runs.
rl.traingrades one task’s answer and an empty reply on your machine first. A reward that raises, or returns something other than a number, a bool orNone, fails there instead of on a GPU. ReturnNoneor0when a reply holds no answer. - Reuse. The same code and tasks reuse the same Environment, so a second run starts at once.
- No network. The reward runs in Nodus’s grader, which has no network access. Pass
pip=["package==1.2.3"]for packages it imports. - How it learns. Each step samples
group_size=8replies to each ofgroups_per_step=4tasks and scores every reply against the others in its group. Replies are capped atmax_tokens=256, and the learning rate islearning_rate=4e-5. You can change any of them. - Format rewards. To reward the shape of an answer as well as its value, fold both into the reward,
for example
return float(correct) - 0.1 * (not well_formed). - Other options.
steps,gpu,max_cost,loraand the othergrpo_loraparameters are optional. The run appears under Runs › Post-training in the console, with its curves, results and cost.
To train on a catalog Environment, pass its name instead of tasks and a reward:
run = rl.train("nodus/gsm8k@1.0.0", max_cost=5)Bring your own environment
Section titled “Bring your own environment”Your own tasks and reward are a Python module with two functions. tasks(split) lists the prompts of the train or
test split with their answers, and reward(completion, answer) scores one completion:
def tasks(split): return [{"prompt": "Spell the word backwards...\n\nWord: valley", "answer": "yellav"}, ...]
def reward(completion, answer): found = ANSWER.findall(completion) if not found: return None # no answer in the completion: InvalidOutput return 1.0 if found[-1].lower() == answer else 0.0A reward of 1 (or True) is Correct, anything lower Incorrect, and None InvalidOutput. If reward raises,
the verdict is InfrastructureFailure, never a reward of 0. Answers reach only reward; the trainer and the model
see the prompt and the optional metadata. A test prompt never appears in train, even if your lists repeat it.
The quickest way in needs no Docker: publish the module as a pip package whose load_environment() returns an
object (or the module itself) with tasks and reward, and name the exact version. Nodus builds it into an image
when you apply the Environment; if the package does not load, the build fails and the Image
env-<name>-<version> shows the step and its log:
apiVersion: nodus.dev/v1kind: Environmentmetadata: name: reverse-wordsspec: version: 1.0.0 package: pip: {name: reverse-words, version: 1.0.0} # index: for a private index loader: load_environment # the default; or module:function category: Custom modes: [Train, Evaluate] rewardType: BinaryRun an environment from the Environments Hub
Section titled “Run an environment from the Environments Hub”An environment published on Prime Intellect’s Environments Hub is a pip package, so it runs on Nodus as it is. Name it with the Hub’s index for its owner and the version the Hub lists:
apiVersion: nodus.dev/v1kind: Environmentmetadata: name: reverse-textspec: version: 0.1.4 package: pip: name: reverse-text version: 0.1.4 index: https://hub.primeintellect.ai/primeintellect/simple/ category: Custom modes: [Train, Evaluate]Its load_environment() returns a verifiers environment, and Nodus trains on it directly: the dataset is the
train split, the eval_dataset the held-out one (with only one of them, every fifth task is held out), and each
completion is scored by the environment’s own rubric, with verifiers 0.1 through 0.3. A multi-turn environment,
such as a game like wordle, runs too: see Train on a multi-turn environment,
and so does one where the model calls tools: see Train on a tool-calling environment.
An environment whose tools run in a remote sandbox, a browser or an MCP server fails the build with that reason, as
does one scored by another model (a JudgeRubric), since grading has no network and holds no credential for that
model. The console marks each kind in its Hub
search from the Hub’s own tags. The package comes from the Hub index alone and its dependencies may also come from
PyPI.
Grading has no network. Nodus fetches the datasets an environment loads while it builds the image, and every task list and grade after that reads that copy, so a run always sees the data its build saw.
Train on a multi-turn environment
Section titled “Train on a multi-turn environment”In a multi-turn environment the model and the environment take turns: the model replies, the environment answers
with feedback, and the episode goes on until the environment ends it. Set maxTurns on the TrainingJob (Turns per
episode under Advanced in the console) to let an episode run that many model turns:
spec: runtime: nodus/grpo-lora environment: {name: wordle@0.1.7, trainTasks: 512, heldOutTasks: 64, seed: 42} parameters: maxTurns: 6 # model turns per episode maxCompletionLength: 256 # tokens per reply maxEpisodeTokens: 2048 # tokens of the whole episode after the prompt, the model's and the environment'sThe reward is the one the environment gives the whole episode, and training learns only from the model’s own
tokens. An episode that reaches either cap ends there and is graded as it stands. With the default maxTurns: 1,
only the first reply is graded. Baseline and final evaluation play the same episodes greedily.
An OpenEnv environment runs in process from a package whose
load_environment() returns it. Each task is a seed: Nodus resets the environment with it, and every action is
the model’s reply as the one text field of the environment’s action class. The environment must play the same
episode for the same seed. Seeds 0 to 799 train and 800 to 999 are held out, unless the environment sets
train_seeds and test_seeds:
import nltkfrom textarena_env.server.environment import TextArenaEnvironment
def load_environment(): try: # the build downloads NLTK's word lists; grading reads them offline nltk.data.find("corpora/words") nltk.data.find("taggers/averaged_perceptron_tagger_eng") cached = True except LookupError: cached = False return TextArenaEnvironment("Wordle-v0", download_nltk=not cached)A module of your own can also be multi-turn: give it step(turns, answer) in place of reward. It gets every
model turn so far and returns the environment’s next message, or None once the episode ended, and the reward so
far. Grading keeps no state between turns, so step replays the turns from the start.
Train on a tool-calling environment
Section titled “Train on a tool-calling environment”In a verifiers ToolEnv the model calls the environment’s tools and reads their results. Each task’s prompt
offers the tools through the model’s chat template, the model calls one by writing a
<tool_call>{"name": "...", "arguments": {...}}</tool_call> block (the form Qwen and most open chat templates
teach), and the environment runs the call in its own code. The results are the next turn, and the episode ends when
the model answers without calling a tool. Set maxTurns to the rounds of calls an episode may make:
spec: runtime: nodus/grpo-lora environment: {name: tool-test@0.1.1, trainTasks: 21, heldOutTasks: 43, seed: 42} parameters: maxTurns: 3 maxCompletionLength: 256The tools run inside the grader, which has no network, so a tool that searches the web or runs code in a remote sandbox cannot train. Pick a model whose chat template supports tools.
Run any Reasoning Gym family
Section titled “Run any Reasoning Gym family”The catalog’s nodus/reasoning-gym serves five reviewed families. To train on any other
Reasoning Gym family, or a mix of them, publish a package whose
load_environment() returns the dataset; its own score_answer scores each completion:
import reasoning_gym
def load_environment(): return reasoning_gym.create_dataset("knights_knaves", size=2000, seed=7)Every fifth entry is held out. The model’s last <answer>…</answer> is its answer, or the whole completion without
one; partial credit is the reward and only a full score is Correct. Each family’s own licence applies.
Run a verl or SkyRL dataset
Section titled “Run a verl or SkyRL dataset”A dataset prepared for verl or SkyRL runs with its reward unchanged. Publish a module that names the parquet files with verl’s own keys, and your verl reward function under its usual name:
train_files = "hf://datasets/BytedTsinghua-SIA/DAPO-Math-17k/data/dapo-math-17k.parquet"val_files = "hf://datasets/BytedTsinghua-SIA/AIME-2024/data/aime-2024.parquet" # optional
def compute_score(data_source, solution_str, ground_truth, extra_info=None): ... # verl's custom reward signature; a dict's "score" also worksEach row is verl’s: prompt (a system message and one user message at most), data_source,
reward_model.ground_truth and extra_info. Without compute_score, each row’s env_class names the SkyRL-gym
environment that scores it, built from the row as SkyRL builds it; add skyrl-gym to the package’s dependencies.
Without val_files, every fifth row is held out. The package needs datasets as a dependency, and a SkyRL
environment of more than one turn fails the build with that reason.
To ship your own system packages or files, build the module into an image on the env-base image instead, push it
and name its digest:
FROM ghcr.io/nodus-compute/catalog/env-base:1.0.0COPY reverse_words.py /opt/environment/ENV NODUS_ENVIRONMENT_LOADER=reverse_words NODUS_ENVIRONMENT=reverse-words NODUS_ENVIRONMENT_VERSION=1.0.0RUN nodus-env infoENV HF_HUB_OFFLINE=1 HF_DATASETS_OFFLINE=1apiVersion: nodus.dev/v1kind: Environmentmetadata: name: reverse-wordsspec: version: 1.0.0 package: {image: registry.example.com/acme/reverse-words@sha256:...} category: Custom modes: [Train, Evaluate] rewardType: Binary$ nodus apply -f environment.yaml$ nodus get environment/reverse-words -w # Ready once the image digest is verifiedA TrainingJob names it without the nodus/ prefix (environment: {name: reverse-words@1.0.0}), and in Python
rl.grpo_lora(environment="reverse-words@1.0.0", ...). The whole example, with a GRPO TrainingJob, is in
examples/training/custom-reward.
Nodus pulls the image from a public registry or from your organization’s space in the Nodus registry; other private
registries are not supported for Environments yet.
Any image works if it provides the two commands Nodus runs, as uid 10001 with no network:
nodus-env tasks --split train|test --seed Nwrites one JSON line per task:{"taskId", "prompt", "metadata"}.nodus-env gradereads{"taskId", "completion"}lines and writes one{"taskId", "verdict", "reward", "evidence"}line for each, in order.
The Environment becomes Ready once its image is verified, with the declared split
sizes in status.splits. The first TrainingJob that uses a split and seed runs nodus-env tasks in one of its own
grading Sandboxes and Nodus keeps that manifest for every later run of your organization, so the tasks never
change between runs. A version’s image and graders cannot change: publish a new version instead, so earlier
results stay reproducible.