Training from the console, start to finish
A model is now named by its Hugging Face repository (Qwen/Qwen3-1.7B, or hf://Qwen/Qwen3-1.7B@main), and
Nodus pins the commit it resolves to at launch, so every run records exactly what it trained. New training offers
the template’s models as one-click choices, a method choice (full, LoRA, QLoRA) and, for fine-tuning, a dataset
picker that lists your Volumes or imports a Hugging Face dataset in place.
An environment’s page now has Start training: a first run with the environment’s own example, or GRPO with a LoRA adapter on the smallest model, priced before you start and launched in one click.
A run’s page plots training loss, mean reward, KL and learning rate by step while it trains. Each graded task now opens to its prompt, the model’s answer and the grader’s verdict, filterable by outcome. Post-training lists every run’s held-out pass rate before and after training, so a sweep’s best run is visible at a glance.
Environments can be created from the console from a pip package, a container image, or Prime Intellect’s
Environments Hub: search the Hub, pick a version and Nodus builds it and trains on it. A single-turn verifiers
environment runs unchanged, graded by its own rubric; the Hub’s own tags mark the ones Nodus refuses before you build.
A build that fails now says what refused the package, rather than only that a step failed.
A verl or SkyRL dataset runs too: a module naming verl’s train_files and val_files and its compute_score, or
rows whose env_class names a single-turn SkyRL-gym environment. Datasets are fetched while the image builds, and
grading reads that copy offline.
A package whose load_environment() returns a Reasoning Gym dataset trains on any family, scored by its own
score_answer.
Multi-turn environments train too: set maxTurns (Turns per episode in the console) and the model and the
environment take turns, capped by turns and by maxEpisodeTokens, with the reward the environment gives the whole
episode. Multi-turn Hub environments such as wordle and in-process OpenEnv environments run this way. Tool-calling
verifiers environments run the same way: the prompt offers the environment’s tools, the model calls them, and
their results are the next turn. Only environments whose tools need a sandbox or the web stay refused.
nodus cp trainingjob/NAME:outputs/ ./outputs downloads every output, the adapter included, into a folder, each
file verified against its SHA-256; outputs/adapter/ downloads only the adapter. In Python,
job.outputs.download("./outputs") does the same, and prefix="adapter/" narrows it.