{"version":"211ad9f836655b1c3a2668c4693e442471f28614","builtAt":"2026-10-06T18:59:23-07:00","site":"https://nodus-platform-site.pages.dev","pages":[{"id":"docs","url":"https://nodus-platform-site.pages.dev/docs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/index.md","title":"Build with Nodus","description":"Run your first job, choose a compute feature, and get from code to results.","stage":"GA","headings":[{"depth":2,"slug":"what-are-you-building","text":"What are you building?"},{"depth":2,"slug":"bring-your-coding-agent","text":"Bring your coding agent"}],"text":"Build with Nodus\n\n Run your first job, choose a compute feature, and get from code to results.\n\nSource: https://nodus-platform-site.pages.dev/docs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nRun your code on cloud CPUs and GPUs. Start with a command, a Python function, or an interactive environment.\n\n From code to your first result \n \n Install the CLI, sign in, and run a job in about five minutes.\n \n Run your first job →\n\nWhat are you building?\n\nChoose a starting point. Each guide covers the essentials and a working example.\n\n Jobs Run a script or batch command to completion.\n Workspaces Develop with SSH, VS Code, or Jupyter.\n Functions Call Python remotely. Run calls in parallel.\n Sandboxes Give agent-written code an isolated place to run.\n Inference Call a hosted model from your application.\n Agents Run agents with tools and saved progress.\n\nLooking for training (Beta), storage, or billing? Browse all guides →\n\nBring your coding agent\n\nConnect over MCP to work from your editor, or give your agent the Markdown task map. Read one relevant guide at a time; use the reference for exact commands and API fields."},{"id":"docs/concepts","url":"https://nodus-platform-site.pages.dev/docs/concepts/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts.md","title":"How Nodus works","description":"Resources, orgs and projects, how your work is placed and kept alive, and how prepaid billing measures it.","stage":"GA","headings":[{"depth":2,"slug":"everything-is-a-resource","text":"Everything is a resource"},{"depth":2,"slug":"orgs-and-projects","text":"Orgs and projects"},{"depth":2,"slug":"how-your-work-runs","text":"How your work runs"},{"depth":2,"slug":"how-billing-works","text":"How billing works"}],"text":"How Nodus works\n\n Resources, orgs and projects, how your work is placed and kept alive, and how prepaid billing measures it.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nThis page is the map. Each section links to the concept page that goes deeper.\n\nEverything is a resource\n\nYou describe work as a resource : a small declarative object with a kind , a metadata.name and a spec , the same shape Kubernetes uses. You create it with the CLI, the Python SDK, the console or the HTTP API, and Nodus reports progress in its status . The same object reads the same everywhere, so nodus get , kubectl get and the console show one truth.\n\n You want to Kind Everyday command \n - - - \n Run a command to completion Job (and Pipeline , Sweep to chain or fan out Jobs) nodus run , nodus apply -f job.yaml \n Keep an isolated container for untrusted code Sandbox nodus create sandbox , nodus exec -it \n Call Python remotely and fan out App , Function , FunctionCall nodus deploy app.py \n Run a durable agent Agent , AgentRun , AgentGroup nodus create agentrun \n Serve or call a model Model , InferenceEndpoint nodus get models -n nodus \n Train or evaluate with a recipe TrainingJob , TrainingRuntime , Environment nodus create trainingjob \n Develop on a remote machine Workspace nodus ssh workspace/ \n Store data, images and secrets Volume , Image , Secret , Connection nodus volume put \n Bring your own machines Pool , Node , EnrollmentToken , CloudAccount nodus create pool \n\n TrainingJob and TrainingRuntime are Beta. Every other kind above is generally available. Run nodus api-resources for the full list and nodus explain .spec for any field.\n\nOrgs and projects\n\nAn org is your billing and access boundary: members, API keys, credit and budgets belong to it. Inside an org, projects group resources (every org starts with default ). Pass -p to the CLI, or set NODUS PROJECT for a whole shell. Names are unique within a project, and labels such as team=nlp let you select resources and break down cost across projects.\n\nHow your work runs\n\nYou state requirements (an accelerator and count, memory, a region class, a deadline or a cost ceiling) and Nodus chooses an offering that satisfies them, such as h100-sxm-80g-x8-us . You see offerings and Nodus ids, never the machines behind them. Each placement of your container on a machine is an Attempt . When capacity is reclaimed, Nodus starts a new Attempt and restores the files your program saved to its checkpoint directory ( NODUS CHECKPOINT DIR ). Your program reloads its own model, optimizer and progress from those files; Nodus restores files, not process memory.\n\nHow billing works\n\nNodus sells prepaid credit , and four rules decide every charge:\n\n1. A hold comes first. Before any paid machine starts, a hold reserves enough credit on your org, and on every budget that applies, to cover the expected run. A launch that cannot be funded is refused with the amounts and a fix, and a running Job stops gracefully, inside its reserved amount, when money runs out.\n\n2. The rate is frozen at launch. nodus run and the console show the estimate before launch; the rate chosen when the machine is acquired is the rate for the whole run, and it is never above the published list price.\n\n3. You pay what the provider bills for your machine, and nothing Nodus caused. Rented capacity bills from the moment the machine is created until its deletion is confirmed, at provider cost ÷ 0.875 (Nodus keeps 12.5 % of what you pay). Capacity Nodus chose and discarded, orphaned machines and Nodus failures are never charged.\n\n4. Usage is itemized by segment. Every compute usage line names the part of the machine’s billed time it covers:\n\n Segment Covers \n - - \n Boot From the machine’s creation to your command starting: start-up, readiness and image pull \n Running Your command, until it stops \n Restore A replacement machine’s start-up when your work resumes from a checkpoint \n Teardown From stop to confirmed deletion, plus the provider’s rounding increment \n\n nodus get usage --group-by segment and the console’s Cost tab show the breakdown for any object.\n\nNodus-operated capacity (Sandbox nodes, warm pools, CPU nodes) bills from the pricebook’s published per-vCPU, per-GiB and per-disk rates instead. Storage above 10 GB per org and egress above 10 GiB per org per day are metered; logs are free.\n\nThe starter grant\n\nThe first org a verified user creates receives $30 of credit that expires 30 days after it is granted . Grant credit is spent before purchased credit, soonest-expiring first. Until your first purchase, the org has starter limits (one Nodus node, three live Sandboxes, Sandbox lifetimes up to two hours), which lift when you buy credit.\n\nSee the pricing reference for every published rate."},{"id":"docs/concepts/attempts-and-recovery","url":"https://nodus-platform-site.pages.dev/docs/concepts/attempts-and-recovery/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/attempts-and-recovery.md","title":"Attempts and recovery","description":"How a run survives a reclaimed or lost machine, what each continuity mode keeps, and who pays for the time a recovery takes.","stage":"GA","headings":[{"depth":2,"slug":"attempts-and-epochs","text":"Attempts and epochs"},{"depth":2,"slug":"continuity-modes","text":"Continuity modes"},{"depth":2,"slug":"what-survives-a-preemption","text":"What survives a preemption"},{"depth":2,"slug":"recovery-limits","text":"Recovery limits"},{"depth":2,"slug":"suspending","text":"Suspending"},{"depth":2,"slug":"who-pays-for-recovery","text":"Who pays for recovery"}],"text":"Attempts and recovery\n\n How a run survives a reclaimed or lost machine, what each continuity mode keeps, and who pays for the time a recovery takes.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/attempts-and-recovery/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nEvery run on Nodus executes as one or more attempts . An attempt is one incarnation of your command on one machine. When that machine is reclaimed, loses its network or fails, Nodus starts a new attempt somewhere else and your run continues from what it saved. This page explains what carries over, how Nodus decides a machine is gone, and what you pay for along the way.\n\nAttempts and epochs\n\nA Job index, a Sandbox or a worker slot holds one running attempt at a time . Each new attempt gets the next epoch , a number that only grows. Nodus accepts reports, checkpoints and outputs only from the current epoch, so a machine that comes back after it was replaced can never overwrite the work of its successor.\n\nYou see a run’s attempts with nodus get attempts -l nodus.dev/job= . A run’s attempts share its name with an epoch suffix, and gang members add a rank suffix ( -r1 , -r2 ).\n\n Attempt phase Meaning \n - - \n Pending , Placing , Acquiring Choosing and preparing capacity \n Starting The machine is ready; the image is pulled and inputs or a checkpoint are restored \n Running Your command is running \n Succeeded Your command exited 0 \n Failed Your command or its machine failed; reason says which ( NodeLost , Preempted , OOMKilled , …) \n Cancelled The attempt was stopped on purpose: a suspend, a cancel, a budget or lifetime limit \n\nContinuity modes\n\n recovery.continuity says what a new attempt starts from:\n\n Mode A new attempt starts with Use it for \n - - - \n Checkpointed The latest committed checkpoint of your declared state paths ( NODUS CHECKPOINT DIR by default) Training and long jobs that save their own model, optimizer and progress files \n Restartable A cold start plus the progress cursor you reported ( NODUS CURSOR COMPLETED , NODUS CURSOR TOTAL ) Batch work that can skip what it already finished \n Ephemeral A cold start Short or idempotent work \n Snapshotted The latest filesystem snapshot Sandboxes \n\nRestoring files restores files only , never process memory. Your program loads its own checkpoint when it starts; Nodus never adds resume flags to your command.\n\nAn empty checkpoint never counts as saved progress and never replaces an earlier useful one. If every checkpoint a run commits is empty, the run shows Checkpointed=False, reason=NotCheckpointable , and Nodus plans and prices it as Ephemeral from then on.\n\nWhat survives a preemption\n\nWhen a provider reclaims interruptible capacity or a machine stops answering, Nodus recovers make-before-break :\n\n1. On a reclaim notice, Nodus asks your attempt for an urgent checkpoint and, at the same time, starts preparing a replacement machine.\n2. The replacement is prepared up to the point where it could start, but it does not start while the old machine might still be writing.\n3. It starts only once the old machine is provably gone: the old attempt acknowledged its stop, the provider confirmed the machine terminated, or the old attempt’s lease ran out.\n4. The new attempt restores according to the continuity mode above.\n\nSo a Checkpointed run loses at most the work since its last committed checkpoint, and a checkpoint committed during the reclaim notice still counts. If the old machine comes back before it is replaced, the replacement is released and your run simply continues; you are not charged for the replacement.\n\nA machine that stops sending heartbeats gets a replacement prepared after 20 seconds. After 60 seconds of silence an attempt on shared capacity is declared lost and replaced. An attempt on a machine dedicated to your organization may keep running up to the edge of its funding while its replacement waits, because such a machine can only lose its own work.\n\nRecovery limits\n\nRecovery stops, and the run fails with a typed reason, when:\n\n Reason Rule \n - - \n NoProgress Two attempts in a row made no progress. Progress means the command started and either ran for recovery.minProgressDuration (default 2 minutes) or committed a non-empty checkpoint. A preemption after progress never counts \n RestoreFailed The same checkpoint failed to restore twice \n ImagePullFailed The image failed to pull on a second machine (a pull failure is retried once elsewhere) \n MaxAttemptsExceeded recovery.maxAttempts recoveries were used (default 8; 3 for distributed Jobs) \n Preempted , NodeLost recovery.onInterruption: Fail was set, so the first interruption ends the run \n\nFailures caused by your command (a non-zero exit, out of memory, an invalid checkpoint) fail fast and are not retried. At most three recoveries per organization prepare capacity at the same time; the others wait their turn and show an Event.\n\nSuspending\n\n nodus suspend job/ stops the run after a final checkpoint and releases its compute; nodus resume continues at a new epoch from that checkpoint. If the final checkpoint fails or takes longer than max(10 minutes, 2 × the shutdown reserve) , the run keeps running with Suspended=False, reason=SuspendFailed and a SuspendFailed Event, so a suspend never throws work away. Stops caused by money (credits, budgets, the maximum cost) or by a lifetime limit do not wait: they stop at the funded edge with whatever checkpoint exists.\n\nA stop always wins over recovery. If the machine is lost while a suspend, cancel or money stop is in progress, the attempt ends Cancelled with the stop’s reason and nothing is restarted or billed again; nodus resume starts the next epoch as usual.\n\nWho pays for recovery\n\nYou pay for the machines your run uses, at the rate frozen when each was acquired, from the moment the provider starts billing through the confirmed deletion of the machine. The bill splits each machine’s time into segments:\n\n Segment Covers \n - - \n Boot Start-up: boot, image pull, readiness, and for gangs the wait at the start barrier \n Restore Start-up of an attempt that resumes: a recovery, a resume after a suspend, and for gangs the surviving machines’ wait from the restart to the next epoch’s start \n Running Your command running \n Teardown Stop to confirmed deletion, plus the provider’s billing increment \n\nA recovery therefore costs you the replacement’s Boot or Restore time and the lost machine’s Teardown . A machine the provider refuses, or that never appears, before it is ready is replaced with another one, up to six tries per attempt; after that the attempt fails with ReadinessFailed and the run’s retry rule applies. Nodus pays, and never charges you, for:\n\n a replacement released because the old machine came back;\n extra machines prepared to start faster (hedges) that did not win;\n machines that failed Nodus’s own checks after creation, and Nodus internal failures;\n for distributed Jobs, probe failures on a qualified network path and outages of the Nodus mesh.\n\n nodus billing usage --group-by segment and the Cost tab show every segment of every attempt."},{"id":"docs/concepts/billing","url":"https://nodus-platform-site.pages.dev/docs/concepts/billing/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/billing.md","title":"How billing works","description":"Prepaid credits, holds before every paid action, exact per-second charges, and what happens when money runs out.","stage":"GA","headings":[{"depth":2,"slug":"your-balance","text":"Your balance"},{"depth":2,"slug":"holds-and-captures","text":"Holds and captures"},{"depth":2,"slug":"what-is-charged","text":"What is charged"},{"depth":2,"slug":"limits","text":"Limits"},{"depth":2,"slug":"low-balance","text":"Low balance"},{"depth":2,"slug":"arrears","text":"Arrears"}],"text":"How billing works\n\n Prepaid credits, holds before every paid action, exact per-second charges, and what happens when money runs out.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/billing/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNodus is prepaid. You add credits, every paid action reserves funds before it starts, and charges come from the rate frozen when the capacity was acquired. Work never runs on credit you do not have, and it stops gracefully, with its progress saved, when money runs out.\n\nYour balance\n\nYour org has one balance, made of buckets:\n\n Purchased credit from top-ups you pay for by card.\n Credit grants : the starter credit, promo codes and credits from Nodus support. Each grant is its own bucket and may expire.\n\nCharges draw from grants first, soonest expiry first, then from grants that never expire, then from purchased credit. That way a grant is used before it expires, and purchased credit, the only kind that can be refunded, lasts longest.\n\n nodus billing and Usage & billing → Overview show:\n\n Field Meaning \n - - \n Available What you can spend now: your buckets minus open holds \n Reserved Open holds, with the objects that hold them \n Purchased, credits What is left in each kind of bucket, and the next grant to expire \n Arrears Unpaid charges; new work waits until a top-up settles them \n\nHolds and captures\n\nBefore a Job, Sandbox, Workspace, Function worker, agent run or build starts, Nodus places a hold : enough funds for the first stretch of work plus the cost of stopping it cleanly. While the work runs, Nodus captures what it used every 5 minutes and renews the hold for the next stretch. When the work ends, the final capture charges the exact amount and the rest of the hold is released.\n\nThe estimate before launch shows the hold, the expected cost range and the minimum charge. A create that cannot be funded fails at once with the amounts, instead of queuing and failing later:\n\nError from server (InsufficientCredits): job \"train-a\" needs a $3.20 hold to start (released when it ends); available $1.10.\n fix: nodus billing top-up 20, or lower spec.maxCostUSD\n\nWhat is charged\n\nMachines Nodus acquires for you are charged per second from the moment the provider starts billing until the machine is confirmed deleted, at the rate frozen when it was acquired. Usage itemizes each machine’s time by segment ( Boot , Restore , Running , Teardown ). What you pay for lists every kind of time and who pays for it, and Pricing explains how rates are set.\n\nLimits\n\nThree limits apply to every hold, and the tightest wins:\n\n Your balance. A hold never exceeds what is available.\n Budgets. A Block Budget over the org, a project or a label selector stops new holds and renewals in its scope when it is exhausted.\n Object caps. spec.maxCostUSD on an object bounds that object and everything it owns.\n\nWhen a limit is reached, running work stops gracefully: Jobs suspend after a checkpoint, Sandboxes and Workspaces stop, and agent runs wait. Each stopped object shows Funded=False with the reason, and resumes when funds return. If a charge would still go past the limit, Nodus absorbs the difference.\n\nLow balance\n\nYou get a low-balance warning by email, webhook and a console banner when your available balance falls below $5.00 (configurable) or below what your open holds need for their next renewal. Turn on auto-recharge to top up automatically below a threshold.\n\nArrears\n\nA charge that arrives after its hold is gone, such as daily storage, can leave unpaid charges. While arrears are open, new holds and uploads are refused with 402 ArrearsOutstanding ; your next top-up or grant settles them first."},{"id":"docs/concepts/core-resource-model","url":"https://nodus-platform-site.pages.dev/docs/concepts/core-resource-model/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/core-resource-model.md","title":"The resource model","description":"How every Nodus object is named, stored, changed, watched and deleted, and how that maps to Kubernetes tools.","stage":"GA","headings":[{"depth":2,"slug":"the-shape-of-an-object","text":"The shape of an object"},{"depth":2,"slug":"desired-state","text":"Desired state"},{"depth":2,"slug":"create-review-apply","text":"Create, review, apply"},{"depth":2,"slug":"watch-instead-of-polling","text":"Watch instead of polling"},{"depth":2,"slug":"deleting","text":"Deleting"},{"depth":2,"slug":"errors","text":"Errors"},{"depth":2,"slug":"kubernetes-tools","text":"Kubernetes tools"}],"text":"The resource model\n\n How every Nodus object is named, stored, changed, watched and deleted, and how that maps to Kubernetes tools.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/core-resource-model/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nEverything you run on Nodus is an object : a Job, a Sandbox, a Function, a Volume, a Budget. Every object has the same shape and the same verbs, so once you know one kind you know them all, and the CLI, the Python SDK, the console, MCP clients and kubectl all work the same way.\n\nThe shape of an object\n\napiVersion: nodus.dev/v1\nkind: Job\nmetadata:\n name: train-a # unique per project and kind\n namespace: default # the project\n labels:\n app: trainer\nspec: # what you want\n image: nodus/pytorch:2.8-cuda12.8\n command: [python, train.py]\n maxCostUSD: \"40.00\"\nstatus: # what Nodus observed; written by Nodus only\n phase: Running\n\n metadata.name is a DNS label of at most 63 characters, unique within its project and kind. Set metadata.generateName instead to get a unique suffix; such a create needs an Idempotency-Key header.\n metadata.namespace is the project. Every org starts with the project default . The read-only project nodus holds what Nodus publishes, such as base images, which you reference as nodus/ : .\n metadata.uid is a stable id such as job 01j9… , and metadata.resourceVersion changes on every write. Send the resourceVersion you read with an update to make it conditional: if someone changed the object in between, the update fails with error code Conflict and you read it again.\n Labels select objects ( -l app=trainer ). Labels and annotations under the nodus.dev/ prefix are set by Nodus; the label nodus.dev/created-by records who created an object and nodus.dev/launched-by from where.\n\nDesired state\n\nYou change what an object does by changing its spec . Stopping, suspending and cancelling are values of spec.state , not separate actions: nodus suspend job/train-a sets spec.state: Suspended , and a later nodus apply of the same file keeps it suspended. Actions that change nothing in spec , such as restarting a Sandbox’s service, are requests: nodus request restart sandbox/dev sets the annotation nodus.dev/restart-requested-at to the current time, and Nodus acts once per value it has not handled yet.\n\nMost spec fields cannot change after create. An update that changes one fails with error code FieldImmutable , and the response lists each field. The fields that can change, such as spec.state and a raise of spec.maxCostUSD , are listed in each kind’s reference.\n\n status.phase uses the same words on every kind, so --field-selector status.phase=Running means the same thing everywhere. Run-to-completion kinds move through Queued , Provisioning , Running and end in Succeeded , Failed or Cancelled ; long-running kinds move through Pending , Starting , Running and Stopped . status.conditions explain the details, such as the condition Ready .\n\nCreate, review, apply\n\n Create by name is safe to retry. Creating an object that already exists with the same spec returns the existing object. A different spec fails with error code AlreadyExists and lists the differences.\n Review before you pay. Add ?dryRun=All ( --dry-run=server in the CLI) to run every check without creating anything. The response shows the object with its defaults and, for kinds that use compute, status.estimate : the expected cost, the first hold and the start time. Send the ETag of that response as If-Match on the real create to launch exactly what you reviewed: if the spec or the price book changed, the create fails with error code PreconditionFailed and you review again. If only the hourly rate moved, the create goes ahead and never runs above the reviewed rate plus 10 % (unless you set placement.maxRateUSDPerHour ); when nothing fits, the object waits in Queued and its status says the price is above the estimate.\n Retries never double-create. The CLI and SDK send an Idempotency-Key with every create. Retrying with the same key returns the first response with the header Idempotent-Replayed: true ; reusing a key for a different request fails with error code IdempotencyKeyReused .\n\nWatch instead of polling\n\nEvery list can be watched: nodus get jobs -w , or ?watch=true&resourceVersion= on the API. A watch delivers every change after that point, in order, and never skips one. Watches send newline-separated events ( ADDED , MODIFIED , DELETED ) as JSON, or as application/x-ndjson when you ask for it; with allowWatchBookmarks=true you also get a bookmark every 30 seconds to resume from. Nodus keeps 24 hours of changes: resuming from an older point fails with error code Expired , and you list again.\n\nLists return at most 500 objects by default ( limit , up to 1,000) and a continue token for the next page.\n\nDeleting\n\n nodus delete job/train-a stops the work first and removes the object once cleanup is done: while it waits the object shows a deletionTimestamp , and the phase Cancelling or Terminating . Nodus tracks that cleanup with finalizers such as the finalizer nodus.dev/billing , which clears when the final charge has posted, so you never see an object gone while it is still costing money. Deleting an object also deletes what it created (a Pipeline’s Jobs, for example); with --cascade=orphan those stay.\n\nErrors\n\nEvery error has the same body: a Kubernetes Status whose reason is a stable error code, plus fix (what to do next), docs (a page for that code) and requestId (quote it to support). The error reference lists every code.\n\nKubernetes tools\n\nThe API speaks the Kubernetes wire protocol for these objects, so kubectl , k9s and client-go work against it: projects are namespaces, kubectl get jobs.nodus.dev lists Jobs, and kubectl explain , apply , diff , wait and -w behave as they do on a cluster. The nodus CLI adds what kubectl does not have, such as logs, exec and file transfer for Nodus kinds."},{"id":"docs/concepts/durable-execution","url":"https://nodus-platform-site.pages.dev/docs/concepts/durable-execution/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/durable-execution.md","title":"Durable execution","description":"How an agent run survives restarts, and what it needs from your code to do so.","stage":"Beta","headings":[{"depth":2,"slug":"what-this-guarantees","text":"What this guarantees"},{"depth":2,"slug":"what-survives-a-restart","text":"What survives a restart"},{"depth":2,"slug":"the-determinism-contract","text":"The determinism contract"},{"depth":2,"slug":"where-the-record-lives","text":"Where the record lives"},{"depth":2,"slug":"when-an-external-outcome-is-uncertain","text":"When an external outcome is uncertain"}],"text":"Durable execution\n\n How an agent run survives restarts, and what it needs from your code to do so.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/durable-execution/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n Beta: this feature may change.\n\nAn agent run is recorded as it goes. Each effect it has is a step : creating its sandbox, choosing a model for a prompt, each model call and each command in the sandbox. Nodus stores the result of every step the first time it runs. If Nodus restarts, or a machine fails, the run is picked up again and replays from the top: every step that was recorded returns its stored result, and the run continues with the first step that was not.\n\nWhat this guarantees\n\n A recorded step never runs again for the same run: no second model call, no second command, no second charge.\n Pure and idempotent steps can retry after an interruption. Nodus model calls reuse their inference key.\n An interrupted sandbox command or model call using your own key can have an uncertain outcome. Its run parks in Waiting with reason NeedsResolution instead of repeating the effect.\n A run is picked up again by one dispatcher at a time. If two start, one stops without writing.\n A run that reaches a different step than the one recorded at the same position fails with NonDeterministicReplay , and its record stays readable.\n\nWhat survives a restart\n\nThe run’s conversation, its answer so far and its status survive, because they come from the recorded steps. The sandbox’s files are not part of the record. If a sandbox is lost, the run creates a new one from the agent’s image and the files are not restored, so keep what must outlive a sandbox in the run’s answer.\n\nThe determinism contract\n\nCode that runs between steps must give the same result for the same inputs and step results. Clock reads, random numbers and network calls belong inside steps. Agents on Nodus’s Claude follow this for you: the loop that runs the model and its commands is Nodus’s.\n\nWhere the record lives\n\nStep results are encrypted with your organisation’s key. A run’s recorded payloads expire 30 days after it ends; its status, its cost and its step list stay.\n\nWhen an external outcome is uncertain\n\nInspect the run’s steps and the affected files or external system before starting replacement work. A Started step means the intent was saved but its outcome was not recorded; it does not prove the command failed or succeeded. Sending another message does not resume this run. Cancel it with nodus cancel agentrun/ to release its sandbox. The sandbox can continue billing while the run waits. A configured deadline still ends the run.\n\nThe managed step-resolution API is not available yet. Do not treat a new run as a safe retry until you have checked the earlier effect."},{"id":"docs/concepts/gang-networking","url":"https://nodus-platform-site.pages.dev/docs/concepts/gang-networking/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/gang-networking.md","title":"Gang networking","description":"How the members of a distributed Job reach each other across machines and providers, how Nodus checks the path before training starts, and what each path can carry.","stage":"Beta","headings":[{"depth":2,"slug":"path-classes","text":"Path classes"},{"depth":3,"slug":"private","text":"Private"},{"depth":3,"slug":"direct","text":"Direct"},{"depth":3,"slug":"relayed","text":"Relayed"},{"depth":2,"slug":"what-every-rank-sees","text":"What every rank sees"},{"depth":2,"slug":"one-network-per-epoch","text":"One network per epoch"},{"depth":2,"slug":"the-probe","text":"The probe"},{"depth":2,"slug":"measured-throughput","text":"Measured throughput"},{"depth":2,"slug":"billing","text":"Billing"},{"depth":2,"slug":"caveats","text":"Caveats"}],"text":"Gang networking\n\n How the members of a distributed Job reach each other across machines and providers, how Nodus checks the path before training starts, and what each path can carry.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/gang-networking/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n Beta: this feature may change.\n\nBeta\n\nDistributed Jobs ( Job.spec.distributed ) are in beta behind per-org access. Gangs have 2 to 8 members, and Relayed gangs have 2 until multi-member Relayed runs are qualified. Ask for access from the console.\n\nA distributed Job runs as a gang : one member per machine, every member started together, sometimes on machines from different providers. Before your command starts, Nodus joins the members into a private network that belongs to that gang alone. Each rank gets the addresses of the others in its environment, so torchrun , Ray and plain torch.distributed work without any networking code of yours.\n\nYou choose two things in spec.distributed :\n\n network : how close the members must be. Colocated keeps them in one provider region, Regional allows any provider inside one region class, and Global allows anywhere.\n transport : which paths you accept. Direct (the default) accepts only Private and Direct paths. Auto also accepts Relayed paths, which lets members on machines without a kernel network device join.\n\nNodus picks the path. Every member of a gang uses the same path class, so every rank sees the same addressing and the same NCCL settings.\n\nPath classes\n\n Path When Nodus uses it How traffic flows Encryption Bandwidth \n - - - - - \n Private Every member is in one provider region with a private network The provider’s private network, opened only between the gang’s members None added (the provider’s network) Provider native \n Direct Every member can create a WireGuard device and the probe finds a direct path between every pair WireGuard between the members, on a network device named nodus0 WireGuard Measured per pair (see measured throughput) \n Relayed A member has no network device of its own (container-only offerings) and transport: Auto WireGuard in user space, carried through Nodus relays WireGuard end to end; relays see only ciphertext Low: small models only \n\nPrivate\n\nMembers in the same provider region talk over the provider’s private network. Nodus opens the rendezvous and NCCL ports between the gang’s members only and closes them when the gang ends. The advertised addresses are the members’ private IPs.\n\nDirect\n\nEach member runs a WireGuard endpoint next to your container, never inside it: your container gets no extra privileges. The members find each other through the Nodus mesh control and connect peer to peer, punching through NAT where needed. The advertised addresses are the members’ mesh addresses ( 100.64.0.0/10 ), and NCCL and Gloo use the nodus0 interface.\n\n Direct means direct. If the probe finds that one pair can only connect through a relay, a transport: Direct gang is placed again without that pair, and a transport: Auto gang continues as Relayed . If a running pair loses its direct path, its traffic falls back to a relay without interrupting training, and the Job’s NetworkDegraded condition becomes true.\n\nRelayed\n\nContainer-only offerings give your container no network device, no CAP NET ADMIN and no UDP. Their members can still join a gang: each runs WireGuard in user space and sends peer traffic through two relays that Nodus operates. To reach that WireGuard endpoint, your processes start with a small preloaded library, libnodus-netshim.so , which redirects only connections to the other members of your gang. Everything else, such as downloads, datasets, object storage and model APIs, goes out directly as usual.\n\nThe advertised addresses are the members’ own eth0 addresses, so MASTER ADDR , the torchrun rendezvous and the listeners NCCL opens on ephemeral ports all work unchanged, in both directions. The relays mesh with each other: if one goes down, members move to the other within seconds and the TCP connections inside WireGuard survive.\n\nWhat every rank sees\n\nThe path decides a few variables; the rest of the distributed environment is the same on every path.\n\n Variable Private Direct Relayed \n - - - - \n NODUS GANG TRANSPORT private direct relayed \n MASTER ADDR , PET RDZV ENDPOINT , NODUS NODE IPS Private IPs Mesh addresses eth0 addresses \n NCCL SOCKET IFNAME , GLOO SOCKET IFNAME The private interface nodus0 eth0 \n LD PRELOAD , NODUS NETSHIM PEERS , NODUS NETSHIM SOCKS Not set Not set The shim first, then your image’s own LD PRELOAD \n\nYour own NCCL values, such as NCCL DEBUG=INFO , win, except NCCL NET , NCCL SOCKET IFNAME , NCCL SOCKET FAMILY and NCCL IB DISABLE , which the path fixes. A spec LD PRELOAD is rejected when transport: Auto , because Nodus needs to place the shim first on a Relayed path. An LD PRELOAD set in the image, such as jemalloc or tcmalloc, is kept after the shim.\n\nOne network per epoch\n\nA gang’s network lives exactly as long as one epoch of the gang: one attempt at running it with a fixed set of members. Each epoch gets a fresh mesh identity with single-use join keys that expire after 15 minutes, and its members can reach only each other. When Nodus restarts the gang, for example to replace a lost member, it fences the old epoch first: the old network is torn down, and a member left behind on it can no longer reach the new gang, nor can the new gang reach it. Members of other gangs, in your org or anyone else’s, are never reachable.\n\nThe probe\n\nBefore your command starts on any member, every member checks its path to every other member. Nodus bills this time as boot time.\n\n1. Readiness. Each member opens three fresh TCP connections to every peer, on ephemeral ports, within 60 seconds, through the same path your processes will use.\n2. Measurement. For each pair, the member records the round-trip time, a 10-second TCP throughput sample and the path class it actually got: direct , or derp: when the pair goes through a relay.\n3. Shim self-test ( Relayed only). The member runs a check program with your image’s own loader to confirm that the shim loads.\n\nThe slowest pair is published in status.gang.probe as tcpGbps and rttMs , with its path, and sets the NetworkQualified condition. A pair that cannot connect, or a path worse than your transport allows, fails the epoch and Nodus places the gang again; that time is not billed to you. The one exception is an image whose programs cannot load the shim: placing the gang elsewhere cannot fix it, so the Job fails at once with ShimNotLoaded and the time is billed (see caveats).\n\n tcpGbps is a TCP throughput sample between two members, not NCCL bus bandwidth. Read it as an upper bound on what one connection between that pair can carry.\n\nTerminal window\n\n$ nodus get job llama-ft -o jsonpath='{.status.gang.probe}'\n{\"tcpGbps\":0.41,\"rttMs\":38.2,\"path\":\"derp:nodus-us\",\"measuredTime\":\"…\"}\n\nMeasured throughput\n\nNodus publishes only measured numbers: the probe’s TCP throughput and RTT, and NCCL all-reduce bus bandwidth from a separate benchmark on 2 × 1 H100. Nodus places a gang only on combinations of offerings it has qualified for the path, and a row appears here once its qualification run passes.\n\n Path Members TCP throughput (probe) RTT NCCL bus bandwidth \n - - - - - \n Private One provider region Not yet published Not yet published Not yet published \n Direct VMs in different providers Not yet published Not yet published Not yet published \n Relayed A container-only offering with a VM Not yet published Not yet published Not yet published \n\nUntil the benchmark publishes, treat Relayed as low bandwidth: the scheduler assumes it is four times slower than Direct when it compares placements, and the estimate warns about it.\n\nBilling\n\n Private and Direct traffic costs nothing beyond the members’ own time. Relayed traffic crosses Nodus relays and is billed per GiB on the mesh-relay line of your usage, measured on the relays as the bytes each member sends. Mesh traffic never counts as container egress. Each Relayed gang may send up to 1 Gbit/s through the relays, split evenly across its members; the relays are shared and best effort within that limit.\n\nCaveats\n\n Synchronous training across providers is bound by the WAN. Expect cross-provider data-parallel training to be limited by network bandwidth and latency, not by the GPUs. Relayed suits small models, algorithms that communicate little, reinforcement learning with separate rollouts, and getting N GPUs now wherever they are.\n Relayed needs dynamically linked glibc programs. The shim is a preloaded library, so it loads only into programs that use the system’s glibc loader. Launchers built on musl (Alpine images) or linked statically never load it and cannot reach their peers. The probe detects this and the Job fails with ShimNotLoaded . Fix it by using a glibc-based image (any Debian, Ubuntu or CUDA image works), or by setting transport: Direct .\n Relayed is low bandwidth. All gang traffic crosses the relays, under the per-gang limit above. Nodus makes no throughput claim for Relayed until the benchmark is published.\n Advertised addresses must be distinct and exclusive. On Relayed , each member is reached at its own eth0 address, so no two members of a gang may share one, and no two running Relayed gangs may route the same address. When placement produces such a collision, Nodus places the members again; it shows as IPCollision in the GangReady condition while it does.\n Only connections to peers go through the mesh. On Relayed , connections to other members are redirected; UDP and connections to any other address are not. Tools that need UDP between members, or connect to peers through a hostname that does not resolve to their advertised address, do not work on Relayed .\n The probe measures TCP. A good tcpGbps does not guarantee NCCL performance; NCCL’s numbers come only from the benchmark in the table above."},{"id":"docs/concepts/lifecycles","url":"https://nodus-platform-site.pages.dev/docs/concepts/lifecycles/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/lifecycles.md","title":"Lifecycles","description":"The phases a Job, Pipeline or Sweep moves through, what moves it, and how spec.state and conditions relate.","stage":"GA","headings":[{"depth":2,"slug":"phases","text":"Phases"},{"depth":2,"slug":"transitions","text":"Transitions"},{"depth":2,"slug":"specstate","text":"spec.state"},{"depth":2,"slug":"conditions","text":"Conditions"},{"depth":2,"slug":"pipelines-and-sweeps","text":"Pipelines and Sweeps"}],"text":"Lifecycles\n\n The phases a Job, Pipeline or Sweep moves through, what moves it, and how spec.state and conditions relate.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/lifecycles/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nJobs, Pipelines and Sweeps share one lifecycle: they run until they finish. Each shows where it is in status.phase , why in status.reason and status.message , and the details in status.conditions . You steer it with spec.state .\n\nPhases\n\n Phase Meaning Billed \n - - - \n Queued Admitted; waiting for capacity that fits and for a funded hold No \n Provisioning Capacity acquired; the image, source, inputs and any saved state are being prepared Yes \n Running The command is running Yes \n Recovering The capacity was lost; the work is moving to new capacity For new capacity, once acquired \n Suspending State is being saved and compute released Yes \n Suspended Paused with its state saved; no compute is held No compute \n Cancelling Stopping and releasing compute Until released \n Succeeded Finished; outputs collected No \n Failed Finished without success; status.reason says why No \n Cancelled Stopped by spec.state: Cancelled or a delete No \n\n Succeeded , Failed and Cancelled are final: once there, the phase never changes and nothing more is billed.\n\nTransitions\n\n From To When \n - - - \n (new) Queued The object is admitted \n Queued Provisioning Capacity is placed and acquired \n Queued Failed Nothing fits within placement.queueTimeout ( CapacityUnavailable ) \n Provisioning Running The command starts \n Provisioning Failed The container cannot start ( LaunchFailed , ImagePullFailed ) \n Running Recovering The capacity is lost \n Recovering Running The work restarts on new capacity \n Recovering Failed Recovery limits are used up ( RecoveryLimitExceeded , NoProgress , RestoreFailed ) \n Running Succeeded The work is done and its outputs are collected \n Running Failed The command failed more than backoffLimit allows, or timeout elapsed \n Queued , Provisioning , Running , Recovering Suspending spec.state: Suspended , credits or a budget ran out, or maxCostUSD was reached \n Suspending Suspended State is saved and compute released \n Suspending Running A suspend you asked for could not save state ( SuspendFailed ) \n Suspending Failed Saving state failed permanently or timed out \n Suspended Queued Resumed, and funds cover a new hold \n any unfinished phase Cancelling spec.state: Cancelled , or the object is deleted \n Cancelling Cancelled Compute is released and the final charge posted \n\nspec.state\n\n spec.state is what you want; status.phase is what is happening. nodus suspend , nodus resume and nodus cancel set it, and so can a manifest:\n\n spec.state Effect \n - - \n Running (default) Run, or resume from Suspended \n Suspended Save state, release compute and stop billing \n Cancelled Stop for good \n\nA suspend caused by money (credits, a budget or maxCostUSD ) leaves spec.state as it is and resumes on its own once the work is funded again or the cap is raised. Time spent suspended does not count against timeout .\n\nConditions\n\n Condition Meaning \n - - \n Admitted The object passed admission \n Scheduled Capacity is placed; while false, its message says what it is waiting for \n Funded Credits and budgets cover the work; false while it is stopped for money \n Ready The command is running \n Checkpointed The latest state save succeeded; false with CheckpointFailed when it did not \n Suspended The work is suspended; SuspendFailed when a suspend could not save state \n OutputsCommitted Every declared output is collected \n SinksLoaded Every output sink has loaded into its table \n\nWait on a phase or a condition from the CLI:\n\nTerminal window\n\n$ nodus wait job/train --for=jsonpath='{.status.phase}'=Succeeded --timeout 2h\n$ nodus wait job/train --for=condition=OutputsCommitted\n\nPipelines and Sweeps\n\nA Pipeline or a Sweep takes its phase from its child Jobs. It is Queued until its first Job exists, Running while any runs, Suspended when every unfinished Job is suspended, and final once every Job is final: Succeeded if all succeeded, otherwise Failed with reason StageFailed (Pipeline) or CellsFailed (Sweep). Its spec.state passes to every unfinished child, and its maxCostUSD caps the spend of all its children together."},{"id":"docs/concepts/nodusd-container-contract","url":"https://nodus-platform-site.pages.dev/docs/concepts/nodusd-container-contract/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/nodusd-container-contract.md","title":"Container runtime contract","description":"The directories, environment variables and sockets your code sees inside every Nodus container.","stage":"GA","headings":[{"depth":2,"slug":"directories","text":"Directories"},{"depth":2,"slug":"environment-variables","text":"Environment variables"},{"depth":2,"slug":"network","text":"Network"},{"depth":2,"slug":"gpus","text":"GPUs"},{"depth":2,"slug":"sidecars-and-the-init-command","text":"Sidecars and the init command"},{"depth":2,"slug":"commands-you-run-in-a-container","text":"Commands you run in a container"},{"depth":2,"slug":"the-events-socket","text":"The events socket"},{"depth":3,"slug":"checkpoint-handshake","text":"Checkpoint handshake"},{"depth":2,"slug":"the-api-socket","text":"The API socket"},{"depth":2,"slug":"signals-and-exit-codes","text":"Signals and exit codes"}],"text":"Container runtime contract\n\n The directories, environment variables and sockets your code sees inside every Nodus container.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/nodusd-container-contract/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nEvery Job, Sandbox and Function worker runs your image in an isolated container. The same contract holds on every offering: your code can rely on the paths and variables below wherever Nodus places it.\n\nDirectories\n\n Path Variable What it is for \n - - - \n /nodus/state NODUS STATE DIR Recovery state. Write your model, optimizer and progress files here; Nodus checkpoints this directory and restores it on the next attempt. \n /nodus/inputs/ NODUS INPUT Each declared input, read-only, downloaded before your command starts. The variable names the file (a URL or one object) or the directory (an object prefix, such as an earlier stage’s outputs). An input may set its own path instead. \n /nodus/outputs NODUS OUTPUT DIR Results. Every regular file here is uploaded when your command exits with code 0, with its SHA-256, and appears as the outputs output. Symbolic links are not collected. \n /dev/shm Shared memory for NCCL between GPUs and for data-loader workers: half the container’s memory limit, or half the machine’s memory without one. It counts against the memory limit. \n /etc/nodus/hostfile Multi-node Jobs (Beta): one slots= line per node in rank order, for DeepSpeed and MPI tools. \n /run/secrets/ / Secret values as read-only files (mode 0400 ) on a memory-backed filesystem, so they never reach a disk. The environment also carries each value, named by its key. \n /run/nodus/events.sock NODUS EVENTS SOCKET Progress, metrics and the checkpoint handshake (below). \n /run/nodus/api.sock NODUS RUNTIME SOCKET The Nodus API and model calls, authenticated as your Job, Sandbox or Function (below). \n\n NODUS CHECKPOINT DIR is kept as another name for NODUS STATE DIR for existing programs.\n\nEnvironment variables\n\n Variable Set for Meaning \n - - - \n NODUS ATTEMPT Every container The attempt id. A retried or recovered run gets a new attempt. \n NODUS JOB , NODUS INDEX , JOB COMPLETION INDEX Jobs The Job and, for indexed Jobs, this index. \n NODUS RESTORED Recovered attempts 1 when /nodus/state was restored from a checkpoint before your command started. \n NODUS CURSOR COMPLETED , NODUS CURSOR TOTAL Restartable Jobs The progress cursor your program last reported. \n NODUS PARAM Sweep cells The cell’s parameters. \n PET NPROC PER NODE Single-node GPU Jobs The GPU count, which torchrun reads as its --nproc-per-node default, so torchrun train.py starts one process per GPU. Your own value or flag wins. \n\nRestoring a checkpoint restores files, never process memory: your program starts from the beginning and reads its own progress files from /nodus/state . Nodus never adds resume flags to your command.\n\nNetwork\n\nJobs, Functions and Workspaces reach the public Internet by default; Sandboxes and Agents have no outbound network unless their spec opens it. Either way a container never reaches private addresses ( 10.0.0.0/8 , 172.16.0.0/12 , 192.168.0.0/16 ), cloud metadata services, other containers or the machine it runs on, and has no IPv6. localhost always works inside the container. /etc/resolv.conf points at public resolvers, and /etc/hosts resolves localhost .\n\nWith an allow list ( egress: AllowList ) the container has no direct route out. HTTPS PROXY and HTTP PROXY point at a proxy on the container’s loopback ( http://127.0.0.1:3128 ) that reaches only the listed hosts, on ports 80 and 443. A .example.com entry allows every subdomain of example.com . Most HTTP clients ( pip , npm , curl , git over HTTPS, the Python and Node SDKs) use the proxy on their own. A listed name that resolves to a private address is still refused.\n\nCPU Sandboxes run under gVisor, whose localhost belongs to the sandbox alone, so there the proxy and the inference proxy listen on the sandbox’s gateway address instead of 127.0.0.1 . Read the address from HTTPS PROXY and OPENAI BASE URL rather than writing 127.0.0.1 into your code; NO PROXY already covers it.\n\nOn a multi-node Job whose nodes share a private network, each rank’s eth0 lists the node’s private address first, so NCCL, Gloo and torchrun advertise an address the other ranks reach. Ranks reach each other only on the rendezvous and NCCL ports, whatever the Job’s outbound setting.\n\nGPUs\n\nA container on a GPU offering sees exactly the GPUs assigned to it, numbered from 0 as CUDA and PyTorch see them, with the NVIDIA driver libraries mounted read-only. It never sees another container’s GPUs. Use nvidia-smi or torch.cuda.device count() to check what you have; do not set CUDA VISIBLE DEVICES yourself.\n\nSidecars and the init command\n\nSidecars start before your command, in the order you list them, in the same container: they share its files, network, environment and logs. Each must answer its readiness probe (an HTTP GET of its path, any status below 400, or a TCP connection to its port) within 15 minutes before the next starts, and they stop when your command exits. A sidecar runs from your container’s image.\n\n initCommand runs after the sidecars and before your command, for preflight checks such as imports, free disk or the GPU count. Your command starts only if it exits with code 0; any other code fails the attempt with that code, and running past its timeout (30 minutes unless you set one) fails it with code 124. Both run inside your billed time.\n\nCommands you run in a container\n\nCommands you start in a running container ( nodus exec , Processes) run with the same environment, secrets and directories as your main command, plus any variables you add. Their output is streamed to you and, for Processes, kept with your logs with secret values masked. A command started from a terminal session stops when you disconnect; a Process keeps running until it exits, you cancel it, its timeout passes or the container stops.\n\nPort forwarding and preview URLs reach a server listening on the container’s localhost , whatever its outbound network setting. In a CPU Sandbox they reach the sandbox’s own address instead, so listen on all interfaces ( 0.0.0.0 ); the same holds for a sidecar’s readiness port. File operations ( nodus cp , the console Files tab) read, write and watch paths as your container sees them, including /nodus/state and /nodus/outputs , with the permissions of your container’s user: files written this way belong to that user and appear only once complete. /proc , /sys and /dev are not available to them.\n\nThe events socket\n\nThe socket speaks newline-delimited JSON, one object of at most 16 KiB per line. Send telemetry as objects with a type and, to make retries safe, a unique id :\n\n{\"id\": \"evt-10\", \"type\": \"progress\", \"completed\": 120, \"total\": 1000}\n{\"id\": \"evt-11\", \"type\": \"log.metrics\", \"step\": 1200, \"loss\": 0.41}\n\nWhen a program cannot reach the socket, it can print the same object on standard output after the prefix nodus.event . The line also stays in your logs.\n\nA program that sends no log.metrics events still gets its training metrics charted: Nodus reads loss, learning rate, epoch, step and eval values from Hugging Face Trainer and PyTorch Lightning log lines until the program sends a log.metrics event of its own. With recovery.checkpoint.integration: HFTrainer , the Trainer helper package is mounted at /.nodus/python and put first on PYTHONPATH , so its sitecustomize registers the Nodus callback in images without the Nodus SDK.\n\nCheckpoint handshake\n\nNodus decides when to checkpoint. To save a consistent state first, send {\"type\": \"checkpoint.subscribe\"} once. Before each checkpoint Nodus writes a request on the connection:\n\n{\"type\": \"checkpoint.request\", \"requestId\": \"ck-1790000000000\", \"seq\": 1790000000000, \"urgent\": true}\n\nFinish writing your files to /nodus/state , then answer with {\"type\": \"checkpoint.ready\", \"requestId\": \"ck-1790000000000\"} . urgent means the capacity is about to go away: save quickly. A program that never subscribes is checkpointed without being asked, so write your state files atomically (write to a temporary name, then rename).\n\nA checkpoint of an empty state directory never replaces an earlier checkpoint that had files in it.\n\nThe API socket\n\n /run/nodus/api.sock is HTTP over a Unix socket. Requests under /apis/nodus.dev/ reach the Nodus API, and requests under /v1/ (OpenAI- and Anthropic-compatible routes) reach Nodus inference. Both carry the service account token of the Job, Sandbox or Function the container belongs to: Nodus adds it to each request, so the token is never in your environment or files, and any Authorization header you send is replaced.\n\nTerminal window\n\ncurl --unix-socket \"$NODUS RUNTIME SOCKET\" http://nodus/apis/nodus.dev/v1/...\n\nWith the inference proxy enabled, the same model routes are also served on http://127.0.0.1:7777 inside the container (on the gateway address in a CPU Sandbox), even with outbound network off, and OPENAI BASE URL , ANTHROPIC BASE URL , OPENAI API KEY and ANTHROPIC API KEY are set for it. The keys are placeholders, so SDKs and tools that take a base URL work unchanged. That port serves nothing but model calls.\n\nSignals and exit codes\n\nYour command runs under a small init process ( tini , at /.nodus/bin ) that forwards signals to your command’s process group and reaps finished child processes, so your command does not need to be written as PID 1.\n\nA stop sends SIGTERM to your command, then SIGKILL after the stop grace period (30 s unless your spec sets another). Exit code 0 completes the attempt; any other code, or a signal, fails it with your exit code and the last lines of your logs, with secret values masked."},{"id":"docs/concepts/pricing","url":"https://nodus-platform-site.pages.dev/docs/concepts/pricing/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/pricing.md","title":"Pricing","description":"How Nodus sets your rate, what each second and token costs, and how to see the price before you launch.","stage":"GA","headings":[{"depth":2,"slug":"see-the-price-before-you-launch","text":"See the price before you launch"},{"depth":2,"slug":"machines-nodus-rents-for-you","text":"Machines Nodus rents for you"},{"depth":2,"slug":"sandboxes-functions-and-builds","text":"Sandboxes, Functions and builds"},{"depth":2,"slug":"models","text":"Models"},{"depth":2,"slug":"storage-egress-and-credits","text":"Storage, egress and credits"},{"depth":2,"slug":"rounding","text":"Rounding"}],"text":"Pricing\n\n How Nodus sets your rate, what each second and token costs, and how to see the price before you launch.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/pricing/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNodus is prepaid: you buy credits, every run shows its rate before it starts, and the rate stays frozen for that run. The numbers on this page come from the same price list the API uses, so the pricing reference and your estimates always agree. This page explains how those numbers are set.\n\nSee the price before you launch\n\n nodus run --dry-run and the console review screen show the estimate for a run: the hourly rate, the expected cost range, the expected boot and teardown time as their own lines, and the minimum charge. The estimate is valid for up to 31 minutes. Launching with the estimate’s ETag holds you to it: if prices change in between, the launch is refused and you review the new estimate.\n\nMachines Nodus rents for you\n\nA Job, a GPU Sandbox, a GPU Workspace or a GPU Function worker runs on a machine Nodus rents for you. For that machine you pay exactly what the provider bills, divided by 0.875, so Nodus keeps 12.5 % of what you pay:\n\n From creation to deletion. The clock runs from the moment the provider starts billing until the machine is confirmed deleted. Your usage is itemized by segment: Boot (start-up, image pull, readiness), Restore (resuming from a checkpoint), Running and Teardown (stop to deletion, plus the provider’s billing increment, charged once per machine).\n Never above list. Every accelerator and count has a list price. Nodus never places your work on a machine whose rate would exceed it.\n Frozen for the run. The rate is fixed when the machine is acquired. A later price change applies only to runs launched after it takes effect.\n\nSpare machines Nodus starts to finish your work sooner, and failures Nodus causes, are never charged to you.\n\nThe list price is the ceiling; the “from” rate is the lowest rate a machine is available at right now, refreshed every minute. Your estimate shows the rate for your run.\n\n Accelerator ×1 list ×1 from ×8 list ×8 from \n - - - - - \n A10 $0.89 List only $7.12 List only \n A100 40G $1.49 List only $11.92 List only \n A100 40G PCIE $1.39 List only $11.12 List only \n A100 80G $1.99 List only $15.92 List only \n A100 80G PCIE $1.89 List only $15.12 List only \n B200 $6.49 List only $51.92 List only \n H100 PCIE $2.99 List only $23.92 List only \n H100 SXM $3.29 List only $26.32 List only \n H200 $4.29 List only $34.32 List only \n L4 $0.89 List only $7.12 List only \n L40S $1.29 List only $10.32 List only \n RTX 3090 $0.59 List only $4.72 List only \n RTX 4090 $0.69 List only $5.52 List only \n RTX 6000 ADA $1.09 List only $8.72 List only \n RTX A6000 $0.79 List only $6.32 List only \n\nUSD per hour for the whole machine. Pricebook 2026.10.8, effective 2026-10-01.\n\nSandboxes, Functions and builds\n\nCPU Sandboxes, Functions, agent workers, image builds and CPU Jobs placed on Nodus nodes are billed per second from placement to release, at published rates per vCPU, per GiB of memory and per GiB of disk above 10 GiB per vCPU. The smallest billable shape is 0.25 vCPU with 512 MiB of memory; smaller requests are billed at that shape.\n\nModels\n\nEach model’s price is the model’s cost divided by 0.95. When a model is served from more than one source, it is listed once, at the price of the cheapest source available now, and a request is charged the price it was accepted at even if another source finishes it. A request is priced once: every input, cached, output, audio or speech unit is added up exactly, then rounded up to the next micro-dollar. A model is served only for the operations it has a price for. If Nodus cannot confirm the outcome of a request, you are not charged for it.\n\nIndra ( nodus/indra ) picks one of 10 models for each request. On the Indra plan you pay the chosen model’s rates plus the routing call (0.044211 USD per 1M input tokens), in one price per request. Without the plan, free Indra routes among the three cheapest models at no charge, up to 100 requests and 200,000 tokens per organization per UTC day.\n\nThe Indra plan costs 20.00 USD a month and includes 20.00 USD of nodus/indra usage each month at the listed prices. The allowance expires at the end of each month, and usage beyond it draws from your credits.\n\nAgents that run on Nodus-managed models are billed from your credits at the model’s cost divided by 0.875, on the agent run. The plan allowance does not apply to them.\n\nStorage, egress and credits\n\n Storage: 10 GB per org is included; beyond it, retained bytes are billed per GB-month.\n Egress: 10 GiB per org per day is included; beyond it, egress is billed per GiB up to your daily egress quota.\n New organizations receive 30.00 USD of credit, valid for 30 days.\n\nRounding\n\nMoney is counted in micro-dollars. A running machine is charged as it goes, rounded down, and settled when it stops, rounded up once, so charging in many small windows costs the same as charging once. Current rates for every line are in the pricing reference."},{"id":"docs/concepts/sandboxes-isolation","url":"https://nodus-platform-site.pages.dev/docs/concepts/sandboxes-isolation/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/sandboxes-isolation.md","title":"Sandbox isolation","description":"What keeps a Sandbox apart from the machine, from other Sandboxes and from the network, and what is kept when it stops.","stage":"GA","headings":[{"depth":2,"slug":"where-sandboxes-run","text":"Where Sandboxes run"},{"depth":2,"slug":"layers","text":"Layers"},{"depth":3,"slug":"1-a-user-space-kernel-gvisor","text":"1. A user-space kernel (gVisor)"},{"depth":3,"slug":"2-a-network-namespace-and-a-firewall-per-sandbox","text":"2. A network namespace and a firewall per Sandbox"},{"depth":3,"slug":"3-resource-limits","text":"3. Resource limits"},{"depth":3,"slug":"4-an-unprivileged-process","text":"4. An unprivileged process"},{"depth":2,"slug":"continuity-what-survives-a-stop","text":"Continuity: what survives a stop"},{"depth":2,"slug":"where-gvisor-is-not-available","text":"Where gVisor is not available"},{"depth":2,"slug":"what-a-sandbox-does-not-protect-against","text":"What a Sandbox does not protect against"}],"text":"Sandbox isolation\n\n What keeps a Sandbox apart from the machine, from other Sandboxes and from the network, and what is kept when it stops.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/sandboxes-isolation/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Sandbox runs code you did not write, such as the output of a model, so the boundary around it matters more than it does for your own jobs. This page lists the layers of that boundary, in the order a piece of code meets them.\n\nWhere Sandboxes run\n\nSandboxes run on CPU machines Nodus operates , not on GPU machines rented for your jobs. Nodus packs machines per organization : a machine serves one organization at a time, your Sandboxes share it with your other Sandboxes, agent workers and CPU jobs, and with nobody else’s. An empty machine is destroyed after 10 minutes and never handed to another organization, so one organization’s files and memory never sit on a machine another organization uses next.\n\nLayers\n\n1. A user-space kernel (gVisor)\n\nEach Sandbox runs under gVisor ( runsc , in its systrap mode). gVisor answers the Sandbox’s system calls in its own kernel written in a memory-safe language and passes the host kernel only a small, filtered set of calls. A bug in the host kernel’s handling of an unusual system call is reachable from a Sandbox only through that narrow filter, not directly.\n\n2. A network namespace and a firewall per Sandbox\n\nEvery Sandbox gets its own network namespace with a default-deny nftables firewall:\n\n Private ranges (RFC 1918), carrier-grade NAT, link-local addresses and cloud metadata addresses are always blocked, whatever the policy, so a Sandbox cannot reach the machine, its neighbours or the cloud’s metadata service.\n Policy Deny (the default) has no route out at all.\n Policy Open translates traffic to public addresses only.\n\nNothing is allowed in from the network. Bytes in and out are counted, and a blocked connection raises an EgressDenied Event on the Sandbox.\n\n3. Resource limits\n\nEach Sandbox runs in its own cgroup (v2) that limits CPU and memory to what you asked for and the number of processes to 256, so a fork bomb or a memory leak stops at the Sandbox’s own limits instead of slowing its neighbours. The root filesystem is backed by a per-Sandbox file under a disk quota ( resources.disk ), not by a directory shared with other Sandboxes.\n\n4. An unprivileged process\n\nCommands run as a non-root user with no Linux capabilities and no new privs set, so the process cannot gain privileges through a set-uid program. The user comes from the image. The Nodus helper that serves commands and files runs outside the user’s process tree and is read-only to it.\n\nContinuity: what survives a stop\n\n spec.continuity.mode decides what a Sandbox keeps when it stops or when its machine is lost:\n\n Mode A stop keeps If the machine is lost \n - - - \n Snapshotted (default) The filesystem changes the Sandbox made, /workspace first among them The Sandbox restarts on another machine from its last snapshot; work since that snapshot is gone \n Ephemeral Nothing: the next start is empty The Sandbox moves to Failed with the reason NodeLost \n\nProcesses never survive a stop: a start runs a fresh container from the image with the saved files restored, so start long-lived servers again from your code. Keep your work in /workspace , the directory the Sandbox starts in. Snapshotted is the right choice for an agent that works in a repository; Ephemeral suits a fresh sandbox per task. Snapshots are taken at every stop and about every 10 minutes while the Sandbox runs.\n\nWhere gVisor is not available\n\nProduction Nodus CPU nodes ship with runsc , and Sandboxes there run under gVisor as described above. On a node without runsc (the local development provider on a laptop, or a Linux host where gVisor is not installed), nodusd runs the Sandbox under runc instead, with the same network namespace, firewall, cgroup limits and unprivileged process. Layers 2 to 4 are identical; layer 1 is not: a runc Sandbox shares the host’s Linux kernel, so it is a boundary for development and trusted code, not for hostile code.\n\nCaution\n\nDo not run untrusted code on a development node that lacks runsc .\n\nWhat a Sandbox does not protect against\n\n Secrets you put in it. Anything in a Sandbox’s environment or files is readable by the code running there. With egress.policy: Open , that code can send it anywhere on the internet. Keep Deny unless the work needs the network, and give a Sandbox only the credentials its task needs.\n Resource use up to your limits. A Sandbox can use all the CPU and memory it asked for, and is billed for them while it runs. Set maxCostUSD to cap the spend of code you do not control."},{"id":"docs/concepts/supply-placement","url":"https://nodus-platform-site.pages.dev/docs/concepts/supply-placement/","markdown":"https://nodus-platform-site.pages.dev/docs/source/concepts/supply-placement.md","title":"Placement and scheduling profiles","description":"How Nodus chooses where a run executes, what the estimate includes, and how profiles, deadlines and budgets change the choice.","stage":"GA","headings":[{"depth":2,"slug":"cost-to-completion","text":"Cost to completion"},{"depth":2,"slug":"profiles","text":"Profiles"},{"depth":2,"slug":"deadlines-and-budgets","text":"Deadlines and budgets"},{"depth":2,"slug":"estimates-and-the-if-match-ceiling","text":"Estimates and the If-Match ceiling"},{"depth":2,"slug":"reuse-before-new-capacity","text":"Reuse before new capacity"},{"depth":2,"slug":"why-a-placement-was-made","text":"Why a placement was made"},{"depth":2,"slug":"multi-node-runs-beta","text":"Multi-node runs (Beta)"}],"text":"Placement and scheduling profiles\n\n How Nodus chooses where a run executes, what the estimate includes, and how profiles, deadlines and budgets change the choice.\n\nSource: https://nodus-platform-site.pages.dev/docs/concepts/supply-placement/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nBy default Nodus places each run on the cheapest offering that can start it now : the lowest hourly rate among the offerings that fit your request, have a machine free and stay within your limits. If that offering cannot be used, the next cheapest takes the run, and so on. You pay the rate of the offering the run starts on, shown in the estimate and frozen for the run. This page explains what the estimate includes and the settings that change the choice.\n\nCost to completion\n\nFor every offering that fits your request, Nodus estimates the whole bill of the run. The estimate shows it for the chosen offering, and the Cost profile ranks offerings on it:\n\n Startup : the machine’s boot, the image pull and any restore from a checkpoint. You pay for these because the capacity bills from the moment it is created.\n Running time : from your declared expectedDuration , a training runtime’s measured estimate, or the history of earlier runs with the same labels on the same accelerator family.\n Teardown : the time until the machine is confirmed deleted.\n Rounding : each machine bills in increments, and the estimate rounds up the same way.\n Interruptions : for interruptible capacity, the expected number of reclaims times the work redone after each one. Work that checkpoints loses only the stretch since its last save; work that saves nothing loses half its run on average.\n\nUnder Cost , a cheap interruptible offering therefore wins only when its expected cost, lost work included, still beats on-demand capacity. A run whose checkpoints are all empty is costed as if it saves nothing.\n\nWhen the running time is unknown, every profile ranks offerings by hourly rate and the estimate shows the rate and the startup and teardown charge but no total.\n\nProfiles\n\n placement.profile picks how the scheduler trades cost against time:\n\n Profile Chooses \n - - \n Balanced (default) The lowest hourly rate that meets your deadline and budget. Among offerings within 2 % of that rate, the healthier one, or the one that starts sooner, fits better or is already warm \n Cost The lowest expected cost to completion, startup, teardown and lost work included, even if it is slower to start. Among offerings within 2 % of that cost, warm capacity and a better fit win \n Speed The fastest expected finish among offerings within 1.5 × the cheapest expected cost \n\n Balanced never chooses an offering whose rate is more than 2 % above the cheapest that meets your limits, and Cost never chooses one whose expected cost is more than 2 % above the cheapest at equal health; interruptible: Prefer gives interruptible capacity a 10 % allowance. Under Cost and Speed , an offering that often fails to start counts as dearer by that risk; under every profile, one that recently failed to start ranks lower for a few minutes within the 2 % band. When two offerings are within 2 % of each other, identical requests are spread across both instead of all taking the same one.\n\nDeadlines and budgets\n\n placement.completeByTime : offerings whose p90 finish is later are not used.\n maxCostUSD , Budgets and your balance: offerings whose expected cost exceeds the money left are not used.\n placement.maxRateUSDPerHour : offerings above this rate are not used.\n\nIf nothing remains, the run waits in Queued and the estimate says why, for example MissesDeadline or ExceedsRemainingBudget .\n\nEstimates and the If-Match ceiling\n\n --dry-run=server -o estimate returns the expected cost p50 and p90, the startup time (cold, and warm when idle capacity of yours fits), the first hold, the minimum charge and validUntil , which is at most 31 minutes away. Creating with the estimate’s If-Match binds the launch to it: Nodus then uses no offering above the estimated rate plus 10 %, unless you set placement.maxRateUSDPerHour yourself. If prices moved beyond that, the run waits with PriceAboveEstimate instead of costing more than you saw.\n\nReuse before new capacity\n\nBefore buying new capacity, the scheduler considers capacity you already pay for: an idle machine of yours that fits reuses the increment already paid, and CPU work packs onto your Nodus nodes. A BYOC pool named in placement.pool is used first at no hourly charge.\n\nWhy a placement was made\n\n nodus describe shows each attempt’s placement: the profile, the scores of the chosen offering (fit, time to result, cost to complete, recovery value, health), the fallbacks in order and every rejected offering with its reason. Offerings are shown by name, such as h100-sxm-80g-x8-us .\n\nMulti-node runs (Beta)\n\nFor spec.distributed , Nodus first resolves the topology (nodes × GPUs per node) and then places every node together:\n\n network: Colocated (the default) keeps every node in one location on one network; Regional keeps them in one region class; Global allows anywhere.\n transport: Direct uses only private or direct paths between nodes; Auto also allows the relayed mesh, with lower bandwidth.\n The whole gang is priced together, from the slowest node’s start. The estimate also shows the assembly bound: the most a failed assembly can cost.\n\nIf no set of offerings satisfies these rules, the run waits with GangInfeasible . See multi-node training."},{"id":"docs/for-agents","url":"https://nodus-platform-site.pages.dev/docs/for-agents/","markdown":"https://nodus-platform-site.pages.dev/docs/source/for-agents.md","title":"For coding agents","description":"Connect a coding agent over MCP, and the machine-readable versions of these docs, the API and the setup.","stage":"GA","skill":"nodus-for-agents","headings":[{"depth":2,"slug":"read-in-this-order","text":"Read in this order"},{"depth":2,"slug":"connect-over-mcp","text":"Connect over MCP"},{"depth":2,"slug":"machine-readable-outputs","text":"Machine-readable outputs"},{"depth":2,"slug":"use-the-current-contract","text":"Use the current contract"}],"text":"For coding agents\n\n Connect a coding agent over MCP, and the machine-readable versions of these docs, the API and the setup.\n\nSource: https://nodus-platform-site.pages.dev/docs/for-agents/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nCoding agents use Nodus the way you do: the same account, projects, budgets and confirmations. Connect one over MCP, or point it at the machine-readable outputs below.\n\nRead in this order\n\n1. Read the compute decision guide to choose the feature that fits the request.\n2. If this is a first run, read the quickstart.\n3. Use the guide directory to find the relevant guide, then fetch its Markdown.\n4. Look up exact commands in the CLI reference, signatures in the Python reference, or request fields in OpenAPI.\n\nFetch individual pages to keep context small. For example, /docs/guides/jobs/ has its Markdown at /docs/source/guides/jobs.md . Each docs page has a View Markdown link. Use the full corpus only for tasks that need many parts of the product.\n\nConnect over MCP\n\nNodus runs one MCP server with generic tools over every resource ( get , describe , logs , estimate , apply , exec and more). Writes return a dry run first and run only once confirmed.\n\n Client type Configuration \n - - \n Hosted (Claude Code, Codex, Cursor and other HTTP clients) Add the server URL from /mcp-hosted.json and sign in with your browser \n Local stdio clients Run nodus mcp ; the configuration is /mcp.json \n\nStep-by-step setup for each client is at /connect, and as Markdown at /connect.md .\n\nMachine-readable outputs\n\n Output What it holds \n - - \n /docs/pages.json Lightweight page manifest: titles, summaries, status and source URLs \n /docs/llms.txt An index of these docs for language models \n Start , compute , models , data , billing Focused bundles with complete pages, examples and source URLs \n /llms-full.txt Guides and concepts in one Markdown file; use the index for exact references \n /docs/source/ .md Each page’s Markdown, linked from the page with rel=\"alternate\" \n /docs/index.json Every page with its headings and text, versioned by build \n /docs/openapi.json The HTTP API contract \n /skills/ /SKILL.md Task instructions for agents that support skills \n /install , /install.ps1 The CLI installer for macOS, Linux and Windows \n\nA good first prompt for an agent that can read URLs:\n\nRead https://nodus-compute.ai/connect.md and help me connect Nodus to this agent. Reuse any existing Nodus\nconnection. Verify setup by listing my jobs. Do not start paid compute.\n\nUse the current contract\n\nFetch the relevant guide and its linked reference before writing commands. Keep the resource’s API version and Beta status explicit. Use the published OpenAPI schemas for field names and the CLI or SDK reference for the installed interface; do not invent flags from examples for another tool.\n\nEach Markdown page includes its canonical source URL and build revision. Cite the source page when explaining behavior. Read billing and recovery limits before creating work, and verify the resource status, logs and outputs before reporting success. A submitted request is not evidence that the run completed."},{"id":"docs/getting-started","url":"https://nodus-platform-site.pages.dev/docs/getting-started/","markdown":"https://nodus-platform-site.pages.dev/docs/source/getting-started.md","title":"Get started","description":"Install the Nodus CLI, sign in, run a command on a GPU, follow it, get its results and see exactly what it cost.","stage":"GA","skill":"nodus-quickstart","headings":[{"depth":2,"slug":"next-steps","text":"Next steps"}],"text":"Get started\n\n Install the Nodus CLI, sign in, run a command on a GPU, follow it, get its results and see exactly what it cost.\n\nSource: https://nodus-platform-site.pages.dev/docs/getting-started/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nThis page takes you from nothing to a finished GPU job and its bill. You need a terminal and a browser. New accounts start with a $30 starter grant , so the first runs need no card.\n\n1. Install the CLI. \n\n pip \n\n Terminal window\n\n pip install nodus-compute\n \n\n The Python package (Python 3.10 or newer) includes the nodus CLI.\n\n macOS and Linux \n\n Terminal window\n\n curl -fsSL https://nodus-compute.ai/install sh\n \n\n Windows \n\n Terminal window\n\n irm https://nodus-compute.ai/install.ps1 iex\n \n\n2. Sign in. \n\n Terminal window\n\n nodus login\n \n\n Your browser opens to confirm the sign-in. A new account gets an org with a default project, an API key for this machine, and the $30 starter grant, which expires 30 days after it is granted. On a machine without a browser, run nodus login --device and approve the code from any other device.\n\n3. Run a command on a GPU. \n\n Terminal window\n\n $ nodus run --gpu L4 --image nodus/pytorch -- python -c \"import torch; print(torch.cuda.get device name())\"\n job/run-4kq7z created · est. $0.01–0.03 · starts in ~2–4 min (cold) · Ctrl+C to cancel, -d to detach\n ✓ Scheduled l4-24g-x1-us · $0.52/h (rate frozen)\n ✓ Provisioning 1m48s\n ✓ Pulling image 22s\n ▶ Running\n NVIDIA L4\n ✓ Succeeded in 2m31s · $0.02 · kept warm 60 s · nodus describe job/run-4kq7z\n \n\n nodus run uploads the current directory, prints the estimate before anything is charged, streams the logs and exits with your command’s exit code. The rate is fixed when the machine is chosen and holds for the whole run. Add -d to return right away and let the job run on its own.\n\n4. Follow it. Every run is a Job you can come back to.\n\n Terminal window\n\n nodus get jobs -w # live list of your jobs\n nodus logs -f job/run-4kq7z # stream the logs again\n nodus describe job/run-4kq7z # status, events and the cost so far\n \n\n5. Get the results. Declare an output path when you run, then copy it back. Downloads are checked against their SHA-256 digest.\n\n Terminal window\n\n nodus run --gpu L4 --image nodus/pytorch --output model=/outputs/model -- python train.py\n nodus cp job/ :outputs/model ./model\n \n\n6. See what it cost. \n\n Terminal window\n\n $ nodus billing\n $ nodus get usage --field-selector object.name=run-4kq7z --group-by segment\n SEGMENT AMOUNT\n Boot $0.018778\n Running $0.003033\n Teardown $0.002889\n \n\n You pay for the machine from the moment it is created until it is deleted: starting up ( Boot ), your command ( Running ) and shutting down ( Teardown ), all at the rate frozen at launch. Credit is prepaid: a hold is reserved before a machine starts, and only what was used is charged.\n\nUsing a coding agent?\n\nConnect Claude Code, Codex or Cursor to Nodus at /connect, then ask it to run your command. The agent uses the same account, projects and spending limits as the CLI.\n\nNext steps\n\n Concepts: resources, projects, and how billing is measured.\n Pricing reference: list prices for every GPU, CPU shape, storage and egress.\n Guides: one guide per feature."},{"id":"docs/getting-started/install","url":"https://nodus-platform-site.pages.dev/docs/getting-started/install/","markdown":"https://nodus-platform-site.pages.dev/docs/source/getting-started/install.md","title":"Install the CLI","description":"Install the nodus CLI on macOS, Linux or Windows, verify the download, turn on shell completion and sign in.","stage":"GA","headings":[{"depth":2,"slug":"sign-in","text":"Sign in"},{"depth":2,"slug":"shell-completion","text":"Shell completion"},{"depth":2,"slug":"verify-a-download","text":"Verify a download"},{"depth":2,"slug":"where-the-cli-keeps-things","text":"Where the CLI keeps things"},{"depth":2,"slug":"upgrade-and-uninstall","text":"Upgrade and uninstall"},{"depth":2,"slug":"next-steps","text":"Next steps"}],"text":"Install the CLI\n\n Install the nodus CLI on macOS, Linux or Windows, verify the download, turn on shell completion and sign in.\n\nSource: https://nodus-platform-site.pages.dev/docs/getting-started/install/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nThe nodus CLI is one self-contained binary for macOS, Linux and Windows on x86-64 and ARM64. Pick one way to install it.\n\n pip \n\nTerminal window\n\npip install nodus-compute\n\nThe Python package (Python 3.10 or newer) includes the CLI, so nodus is on your PATH wherever the package is installed. Use this if you also want the Python SDK.\n\n Homebrew \n\nTerminal window\n\nbrew install --cask nodus-compute/tap/nodus\n\n macOS and Linux \n\nTerminal window\n\ncurl -fsSL https://nodus-compute.ai/install sh\n\nThe script installs to ~/.local/bin and checks the archive against the release’s checksums.txt first. NODUS VERSION=1.2.3 pins a release and NODUS INSTALL DIR picks another directory.\n\n Windows \n\nTerminal window\n\nirm https://nodus-compute.ai/install.ps1 iex\n\nCheck that it works:\n\nTerminal window\n\n$ nodus version\nnodus v1.0.0\n\nSign in\n\nTerminal window\n\nnodus login\n\nYour browser opens to confirm the sign-in. If you belong to several orgs, pick the ones this machine should use: the CLI stores one API key per org in the OS keychain and creates one context per org. Switch between them with nodus config use-context , or pass --org to a single command.\n\n Where you are Command \n - - \n A laptop with a browser nodus login \n An SSH session or a machine without a browser nodus login --device , then approve the code from any device \n CI, with a key in a secret echo \"$NODUS API KEY\" \\ nodus login --with-token , or just set NODUS API KEY \n\n nodus whoami shows who you are signed in as, your role, the current project and your available credit. nodus logout revokes this machine’s key and removes the context.\n\nEnvironment variables\n\nThe CLI reads six variables, all optional: NODUS API KEY , NODUS API URL , NODUS ORG , NODUS PROJECT , NODUS CONTEXT and NODUS CONFIG . A variable wins over the config file for that invocation.\n\nShell completion\n\nTerminal window\n\nnodus completion zsh \"${fpath[1]}/ nodus\" # zsh\nnodus completion bash /etc/bash completion.d/nodus # bash (or ~/.local/share/bash-completion/completions/nodus)\nnodus completion fish ~/.config/fish/completions/nodus.fish\n\nCompletion covers every command and flag.\n\nVerify a download\n\nEvery release publishes checksums.txt , a keyless cosign signature over it, and an SBOM per archive. To verify an archive you downloaded yourself:\n\nTerminal window\n\ncosign verify-blob checksums.txt \\\n --signature checksums.txt.sig --certificate checksums.txt.pem \\\n --certificate-identity-regexp '^https://github.com/nodus-compute/nodus-platform/' \\\n --certificate-oidc-issuer https://token.actions.githubusercontent.com\nshasum -a 256 --ignore-missing -c checksums.txt\n\nWhere the CLI keeps things\n\n Path What it holds \n - - \n ~/.nodus/config Contexts: API server, org and project. It is a kubeconfig, so KUBECONFIG=~/.nodus/config kubectl get jobs.nodus.dev works too \n OS keychain ( ~/.nodus/credentials , mode 0600, where there is none) One API key per context \n ~/.nodus/cache/ The cached list of resource kinds; nodus api-resources refreshes it \n\nIf you used Nodus before 1.0, its ~/.nodus/config.toml is renamed to config.0x.bak the first time the new CLI runs; sign in again with nodus login . A 0.x nodus.toml converts into a manifest you can review and apply:\n\nTerminal window\n\nnodus convert nodus.toml job.yaml # fields that do not carry over are listed on stderr\nnodus apply -f job.yaml --dry-run=server -o estimate\n\nUpgrade and uninstall\n\nUpgrade the same way you installed ( pip install -U nodus-compute , brew upgrade --cask nodus , or run the install script again). To uninstall, run nodus logout , remove the binary, and delete ~/.nodus .\n\nNext steps\n\n Quickstart: run your own code on a GPU and download its results.\n Nodus for kubectl users: the verbs you already know.\n CLI reference: every command and flag."},{"id":"docs/getting-started/kubectl-users","url":"https://nodus-platform-site.pages.dev/docs/getting-started/kubectl-users/","markdown":"https://nodus-platform-site.pages.dev/docs/source/getting-started/kubectl-users.md","title":"Nodus for kubectl users","description":"How Nodus maps onto the Kubernetes API model, which kubectl verbs nodus supports, and how to point kubectl itself at Nodus.","stage":"GA","headings":[{"depth":2,"slug":"the-mapping","text":"The mapping"},{"depth":2,"slug":"verbs-you-already-know","text":"Verbs you already know"},{"depth":2,"slug":"what-is-different","text":"What is different"},{"depth":2,"slug":"use-kubectl-itself","text":"Use kubectl itself"},{"depth":2,"slug":"plugins","text":"Plugins"}],"text":"Nodus for kubectl users\n\n How Nodus maps onto the Kubernetes API model, which kubectl verbs nodus supports, and how to point kubectl itself at Nodus.\n\nSource: https://nodus-platform-site.pages.dev/docs/getting-started/kubectl-users/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nIf you know kubectl, you already know most of nodus . Every Nodus resource (Jobs, Sandboxes, Volumes, Secrets, InferenceEndpoints, Budgets and the rest) is a Kubernetes-style object in the nodus.dev API group, with metadata , spec and status , and the CLI speaks the same verbs over all of them.\n\nThe mapping\n\n Kubernetes Nodus \n - - \n Cluster The Nodus API ( https://api.nodus-compute.ai ) \n Namespace ( -n ) Project ( -p ; -n is accepted) \n User and credentials An API key per org, kept in the OS keychain \n kubeconfig context One context per org, in ~/.nodus/config \n kubectl get pods nodus get jobs , nodus get sandboxes , nodus get all \n kubectl exec nodus exec into a Job, Sandbox or Workspace; each command is recorded as a Process \n\nYour org comes from the API key, so there is no org in a manifest. Projects are created in the console or with nodus create -f , and default always exists.\n\nVerbs you already know\n\nTerminal window\n\nnodus get jobs -o wide # tables are rendered by the server\nnodus get jobs -l team=nlp --field-selector status.phase=Running\nnodus get job/train -o jsonpath='{.status.phase}'\nnodus get jobs -o custom-columns=NAME:.metadata.name,GPU:.spec.resources.gpu\nnodus get jobs -w # watch\nnodus describe job/train # includes a Placement section and events\nnodus apply -f job.yaml # client-side three-way merge\nnodus diff -f job.yaml # what apply would change (exit 1 on differences)\nnodus apply -f jobs/ --prune -l app=nightly # delete what was applied before and is gone now\nnodus edit job/train # $EDITOR, saved as a merge patch\nnodus patch job/train --type merge --patch '{\"spec\":{\"maxCostUSD\":\"50\"}}'\nnodus label job/train tier=gold\nnodus wait job/train --for=jsonpath='{.status.phase}'=Succeeded --timeout 30m\nnodus delete job/train --wait\nnodus logs -f job/train\nnodus exec -it sb/dev -- bash\nnodus shell sandbox/dev # create the Sandbox if needed, then exec -it bash\nnodus events -w --for job/train # events as they are recorded\nnodus port-forward job/train 6006\nnodus explain job.spec.resources\nnodus api-resources\nnodus auth can-i create jobs\n\nShort names work as in kubectl: sb for sandboxes, vol for volumes, sec for secrets, sa for service accounts. nodus api-resources lists them all.\n\nWhat is different\n\n Server dry-run returns a cost estimate. nodus apply -f job.yaml --dry-run=server -o estimate prints the estimated cost, start time and the hold that would be placed, without creating anything.\n State verbs instead of scaling. suspend , resume and cancel apply to Jobs, Pipelines and Sweeps; start and stop to Sandboxes, Workspaces, Functions, Agents and InferenceEndpoints. They set spec.state .\n Requests are annotations. nodus request restart sb/dev (or nodus rollout restart sb/dev ) asks a controller to act once, recorded on the object.\n Typed generators. nodus create sandbox dev --cpu 2 , nodus create secret hf --from-literal HF TOKEN=… , nodus create volume weights --size 200Gi , nodus create budget research --limit 2000 . Add --dry-run=client -o yaml to print the manifest instead of creating it. Kinds that show a secret once print it alone on standard output: nodus create apikey ci --scopes jobs:write , nodus create enrollmenttoken --pool lab (the host installer) and nodus create token sa/ci . nodus create sshkey --from-file ~/.ssh/id ed25519.pub adds a public key, or --generate makes the pair.\n Outputs are files you download. nodus cp job/train:outputs/model ./model copies a declared output and checks its SHA-256 digest.\n Merge patches only. Strategic merge patch and server-side apply are not supported; apply does the three-way merge in the client and records the applied manifest in a last-applied annotation on the object.\n Exit codes. Every command exits 0 , 1 on an error or 2 on a usage error, like kubectl. nodus run passes your command’s exit code through instead.\n\nUse kubectl itself\n\n ~/.nodus/config is a real kubeconfig whose user entry runs nodus auth token as an exec credential plugin, so kubectl and any client-go tool (k9s included) can read and write Nodus resources:\n\nTerminal window\n\nexport KUBECONFIG=~/.nodus/config\nkubectl get jobs.nodus.dev\nkubectl apply -f job.yaml\nkubectl get jobs.nodus.dev -w\nkubectl wait jobs.nodus.dev/train --for=jsonpath='{.status.phase}'=Succeeded\n\nUse the full jobs.nodus.dev resource name with kubectl, because it also knows the built-in batch/v1 Jobs.\n\nWhat kubectl cannot do against Nodus\n\nkubectl’s logs , exec and port-forward only work on Pods, so use nodus for those. There are no Pods, Namespaces or other core Kubernetes resources, and kubectl auth can-i is not served: use nodus auth can-i instead.\n\nPlugins\n\nLike kubectl, nodus runs any executable called nodus- on your PATH as nodus , so you can add your own commands."},{"id":"docs/getting-started/modal-users","url":"https://nodus-platform-site.pages.dev/docs/getting-started/modal-users/","markdown":"https://nodus-platform-site.pages.dev/docs/source/getting-started/modal-users.md","title":"Nodus for Modal users","description":"Move a Modal app to Nodus. The Python SDK keeps Modal's names wherever the concept matches, so most code changes by one import.","stage":"GA","headings":[{"depth":2,"slug":"sign-in-and-run","text":"Sign in and run"},{"depth":2,"slug":"side-by-side","text":"Side by side"},{"depth":2,"slug":"what-nodus-adds","text":"What Nodus adds"},{"depth":2,"slug":"differences-to-know","text":"Differences to know"},{"depth":2,"slug":"next-steps","text":"Next steps"}],"text":"Nodus for Modal users\n\n Move a Modal app to Nodus. The Python SDK keeps Modal's names wherever the concept matches, so most code changes by one import.\n\nSource: https://nodus-platform-site.pages.dev/docs/getting-started/modal-users/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nThe Nodus Python SDK keeps Modal’s names wherever the concept is the same: App , @app.function , .remote() , .map() , .spawn() , @app.cls with enter and exit hooks, Image , Volume , Secret and Sandbox . Most apps move by changing import modal to import nodus . This page lists what is the same, what is spelled differently and what Nodus adds.\n\nSign in and run\n\nTerminal window\n\npip install nodus-compute\nnodus login # opens the console, stores a key for each org you pick\nnodus run app.py # like modal run : an ephemeral App, deleted when the entrypoint returns\nnodus deploy app.py # like modal deploy : a persistent App\nnodus serve app.py # like modal serve : redeploys when a file changes\n\nIn CI, set NODUS API KEY instead of running nodus login .\n\nexamples/python/quickstart/app.py\n\n\"\"\"Quickstart: one Function called three ways.\n\nRun it with nodus run examples/python/quickstart/app.py --n 10 .\n\"\"\"\n\nimport nodus\n\napp = nodus.App(\"quickstart\")\n\n@app.function(cpu=1, memory=\"1Gi\", max cost=1)\ndef square(x: int) - int:\n return x x\n\n@app.local entrypoint()\ndef main(n: int = 10) - None:\n print(\"remote:\", square.remote(7)) # one call; blocks for the result\n call = square.spawn(8) # start without waiting\n print(\"spawned:\", call.get(timeout=600))\n print(\"map:\", list(square.map(range(n)))) # one call per input, results in input order\n\nSide by side\n\n Modal Nodus \n - - \n import modal import nodus \n app = modal.App(\"x\") app = nodus.App(\"x\") \n @app.function(gpu=\"H100\", timeout=3600) @app.function(gpu=\"H100\", timeout=\"1h\") (numbers are also seconds) \n f.remote(x) , f.map(xs) , f.spawn(x) , call.get() The same \n f.local(x) The same \n @app.local entrypoint() The same \n modal.Function.from name(\"app\", \"f\") nodus.Function.from name(\"app\", \"f\") ( Function.lookup also works) \n @app.cls() with @modal.enter() , @modal.method() , @modal.exit() @app.cls() with @nodus.enter() , @nodus.method() , @nodus.exit() \n modal.Image.debian slim().pip install(\"torch\") nodus.Image.debian slim().pip install(\"torch\") \n modal.Image.from registry(...) , .apt install , .run commands , .env , .add local dir The same \n modal.Volume.from name(\"v\", create if missing=True) The same; vol.commit() and vol.reload() behave as in Modal \n modal.Secret.from name(\"hf\") , Secret.from dict({...}) The same, plus Secret.from dotenv(\".env\") \n modal.Sandbox.create(app=app, image=...) nodus.Sandbox.create(image=...) (no App needed) \n sb.exec(\"python\", \"-c\", \"...\") , p.stdout.read() , p.wait() The same \n sb.open(path, \"w\") The same \n sb.tunnels() Returns a list of Tunnel(port, url, public) ; sb.tunnels.open(8080) is not available yet \n sb.snapshot filesystem() The same (Beta) \n @modal.experimental.clustered(size=2) @nodus.clustered(size=2) (Beta) \n min containers , max containers , scaledown window The same, or min workers and max workers \n await f.remote.aio(x) The same: every blocking call has an .aio form \n\nGPU strings use Modal’s spellings: \"H100\" , \"H100:2\" , \"A100-80GB\" , \"A10G\" , \"L40S\" . A family such as H100 matches any of its variants; \"H100!\" pins the exact variant. A list such as [\"H100\", \"H200\"] accepts either.\n\nWhat Nodus adds\n\nNodus places every call on the cheapest capacity that finishes it on time, and it stops work before money runs out. These arguments have no Modal equivalent:\n\n@app.function(\n gpu=\"H100\",\n max cost=40, # a hard cap in USD across this Function's workers\n checkpoint=\"/nodus/state\", # files here are saved and restored if capacity is reclaimed\n interruptible=True, # allow cheaper interruptible capacity; progress is kept through the checkpoint\n region=[\"us\", \"eu\"], # region classes, not provider regions\n)\ndef train(lr: float) - dict: ...\n\nprint(train.estimate(3e-4)) # dry-run: expected cost, cold and warm start, the hold it needs\n\n f.estimate(...) returns the expected cost and start time of one call before you run it.\n A cold start on a GPU with no warm worker prints its expected wait, so a long first call is not a surprise.\n Errors are typed: nodus.errors.InsufficientCredits states the amount needed and how to add credit.\n nodus.Job , nodus.Workspace and nodus.llm cover batch jobs, development machines and inference with the same credentials.\n\nDifferences to know\n\n Parametrized classes ( modal.parameter() ) are not supported. Configure the class in its @nodus.enter() hook instead. Calling MyClass(arg=...) raises nodus.errors.Unsupported .\n The Python minor version of the image must equal yours, as in Modal. Image.debian slim() defaults to your version; a mismatch is refused before anything runs.\n Web endpoints ( @modal.web endpoint , @modal.asgi app ) are not available. Use nodus.InferenceEndpoint for model serving.\n modal.Dict and modal.Queue have no equivalent. Pass data through return values, a Volume or your own database.\n Clustered Functions (Beta) run each .remote() or .spawn() as one gang; .map() over a clustered Function raises nodus.errors.Unsupported .\n Timeouts and durations accept Go-style strings ( \"90s\" , \"6h\" ) as well as seconds. Days are not a unit.\n Money is always a decimal amount in USD ( max cost=40 or \"40.00\" ).\n\nNext steps\n\n Python SDK guide\n Functions and classes\n Sandboxes"},{"id":"docs/getting-started/quickstart","url":"https://nodus-platform-site.pages.dev/docs/getting-started/quickstart/","markdown":"https://nodus-platform-site.pages.dev/docs/source/getting-started/quickstart.md","title":"Run your own code","description":"Run a training script from your own directory on a GPU with nodus run, follow it, download its output and read what it cost.","stage":"GA","headings":[{"depth":2,"slug":"leave-it-running","text":"Leave it running"},{"depth":2,"slug":"control-the-cost","text":"Control the cost"},{"depth":2,"slug":"clean-up","text":"Clean up"},{"depth":2,"slug":"next-steps","text":"Next steps"}],"text":"Run your own code\n\n Run a training script from your own directory on a GPU with nodus run, follow it, download its output and read what it cost.\n\nSource: https://nodus-platform-site.pages.dev/docs/getting-started/quickstart/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nThis page runs a script from your own directory on a GPU, then downloads the file it wrote. It assumes you have installed the CLI and signed in. Every command here runs in CI as the cli/quickstart example.\n\n1. Put your code in a directory. Any directory works; this one holds a small PyTorch training loop that saves its final metrics to /nodus/outputs/metrics.json .\n\n train.py\n\n \"\"\"Fit a small linear model, print progress and save metrics.json as the declared output \"metrics\".\"\"\"\n\n import json\n import pathlib\n\n import torch\n\n device = \"cuda\" if torch.cuda.is available() else \"cpu\"\n print(\"device:\", torch.cuda.get device name() if device == \"cuda\" else \"cpu\")\n\n torch.manual seed(0)\n x = torch.randn(1024, 16, device=device)\n y = x @ torch.randn(16, 1, device=device)\n w = torch.zeros(16, 1, device=device, requires grad=True)\n opt = torch.optim.SGD([w], lr=0.1)\n\n for step in range(200):\n loss = ((x @ w - y) 2).mean()\n opt.zero grad()\n loss.backward()\n opt.step()\n if step % 50 == 0:\n print(f\"step {step} loss {loss.item():.4f}\")\n\n print(f\"final loss {loss.item():.6f}\")\n out = pathlib.Path(\"/nodus/outputs\")\n out.mkdir(parents=True, exist ok=True)\n (out / \"metrics.json\").write text(json.dumps({\"final loss\": loss.item(), \"steps\": 200}))\n \n\n2. Run it. From that directory:\n\n Terminal window\n\n nodus run --name cli-quickstart --gpu L4 --image nodus/pytorch \\\n --output metrics=/nodus/outputs/metrics.json -- python train.py\n \n\n Terminal window\n\n job/cli-quickstart created · est. $0.01–0.03 · starts in ~2–4 min (cold) · hold $0.50 · Ctrl+C to cancel, -d to detach\n ✓ Scheduled l4-24g-x1-us · $0.52/h (rate frozen)\n ✓ Provisioning 1m48s\n ▶ Running\n device: NVIDIA L4\n step 0 loss 15.8732\n ...\n final loss 0.000001\n ✓ Succeeded in 2m31s · $0.02 (boot $0.01 · running $0.01) · kept warm 60 s · nodus describe job/cli-quickstart\n \n\n Before anything is charged you see the estimate and a hold: the most this run can cost before you are asked. nodus run then uploads the directory (files listed in .gitignore or .nodusignore stay local), prints each phase, streams your program’s output on stdout and ends with the cost line. The CLI exits with your command’s exit code, so it works in scripts and CI.\n\n3. Get the output. The --output flag declared metrics ; copy it back:\n\n Terminal window\n\n nodus cp job/cli-quickstart:outputs/metrics metrics.json\n \n\n The file is checked against its SHA-256 digest and only written when it matches.\n\nLeave it running\n\nPress Ctrl+C during a run and the CLI asks whether to cancel the Job; answer n to detach and leave it running. Or start with -d to return as soon as the Job exists. A detached Job keeps running:\n\nTerminal window\n\nnodus get jobs # every Job in the project, with phase and cost so far\nnodus get jobs --mine -w # only yours, updating live\nnodus logs -f job/cli-quickstart # stream the output again\nnodus describe job/cli-quickstart # attempts, placement, cost, conditions and events\nnodus cancel job/cli-quickstart # stop it; you pay only for what ran\n\nControl the cost\n\n Flag Effect \n - - \n --max-cost 5 The Job stops gracefully before it spends more than $5 \n --timeout 2h The Job is stopped after two hours of wall-clock time \n --dry-run Print the estimate and exit without creating anything \n\n nodus billing shows your balance and this month’s spend; nodus billing usage --group-by day breaks it down.\n\nExit codes\n\n nodus run exits with your command’s own exit code. It exits 125 when Nodus could not run the command (an API error or an invalid request), 124 on --timeout and 130 when you interrupt it.\n\nClean up\n\nFinished Jobs are kept for 30 days so you can read their logs and outputs, then deleted. To delete one now:\n\nTerminal window\n\nnodus delete job/cli-quickstart\n\nNext steps\n\n Nodus for kubectl users: manifests, apply , get -o , wait and diff .\n CLI reference: every command and flag, with tested examples."},{"id":"docs/guides","url":"https://nodus-platform-site.pages.dev/docs/guides/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides.md","title":"Find a guide","description":"Choose a task, then read the guide you need. Start with running code and add storage, integrations, or billing as needed.","stage":"GA","headings":[{"depth":2,"slug":"run-code","text":"Run code"},{"depth":2,"slug":"models-and-agents","text":"Models and agents"},{"depth":2,"slug":"files-and-state","text":"Files and state"},{"depth":2,"slug":"tools-and-integrations","text":"Tools and integrations"},{"depth":2,"slug":"account-costs-and-your-own-compute","text":"Account, costs, and your own compute"},{"depth":2,"slug":"coming-from-another-tool","text":"Coming from another tool?"}],"text":"Find a guide\n\n Choose a task, then read the guide you need. Start with running code and add storage, integrations, or billing as needed.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNew to Nodus? Start with your first job. Otherwise, choose the task you have now.\n\nNot sure which feature fits? Compare Jobs, Workspaces, Sandboxes and Functions.\n\nRun code\n\n Task Guide \n - - \n Run a command to completion Jobs \n Work in SSH, VS Code, or Jupyter Workspaces \n Run code in an isolated container Sandboxes \n Call Python remotely or in parallel Functions \n Chain jobs together Pipelines \n Try multiple parameter combinations Sweeps \n Choose resources and check availability Compute selection \n\nModels and agents\n\n Task Guide \n - - \n Call a hosted model Inference \n Run agents with tools and saved progress Agents \n Fine-tune or train a model Training (Beta) \n Train across machines Multi-node training (Beta) \n Define training and evaluation tasks Environments \n\nFiles and state\n\n Task Guide \n - - \n Package code and dependencies Images \n Keep files between runs Volumes \n Pass credentials to your code Secrets \n Download results Outputs \n Save and load progress after an interruption Checkpoints \n Connect external data and services Connections \n\nTools and integrations\n\n Python SDK: write and manage runs from Python.\n MCP: connect your coding agent to Nodus.\n Console and Ask Nodus: manage work in the browser.\n Logs and metrics: follow progress and diagnose a run.\n Webhooks and notifications: react to changes.\n\nAccount, costs, and your own compute\n\n Account setup, projects, and members and roles: organize your team.\n API keys: give applications access.\n Billing, usage and costs, and budgets: fund work and control spending.\n Pools and cloud accounts: use your own compute.\n\nComing from another tool?\n\nStart with Nodus for Modal users, for kubectl users, or migrating from Nodus 0.x."},{"id":"docs/guides/access/api-keys-and-scopes","url":"https://nodus-platform-site.pages.dev/docs/guides/access/api-keys-and-scopes/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/api-keys-and-scopes.md","title":"API keys and scopes","description":"Create API keys for scripts and CI, limit what they can do with scopes and projects, and revoke them.","stage":"GA","headings":[{"depth":2,"slug":"scopes","text":"Scopes"},{"depth":2,"slug":"list-and-revoke","text":"List and revoke"},{"depth":2,"slug":"keys-the-cli-creates","text":"Keys the CLI creates"},{"depth":2,"slug":"if-a-key-leaks","text":"If a key leaks"}],"text":"API keys and scopes\n\n Create API keys for scripts and CI, limit what they can do with scopes and projects, and revoke them.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/api-keys-and-scopes/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nAn API key lets a script, a notebook or CI call Nodus as you. Keys look like nodus sk live … and are sent as Authorization: Bearer , or through NODUS API KEY .\n\nTerminal window\n\nnodus create apikey notebook\n\nThe key is printed once , alone on standard output, so KEY=$(nodus create apikey notebook) captures it. Nodus stores only a keyed hash of it, so a lost key cannot be shown again: delete it and create a new one. -o json or -o yaml prints the whole object, key included.\n\nScopes\n\nA key’s scopes limit what it can do. The default, , means “whatever my role allows”, and it follows your role: if your role changes, the key changes with it.\n\nTerminal window\n\nnodus create apikey ci --scopes jobs:write,volumes:read --projects research --expires 720h\nnodus auth can-i create sandboxes # run with the key to check it\n\nScopes are :read or :write (write includes read), plus :read for read-only access to everything. A key can never hold more than its creator’s role: asking for more is refused with 403 naming the scopes you lack. A key created with another key (or by a connected agent) can hold only what that credential holds, so it cannot ask for or :read unless the creating credential has them itself.\n\n --scopes and --projects take comma-separated lists, and --description says what the key is for. To mint a key for CI that survives its creator, bind it to a service account of the project you work in with --service-account NAME .\n\nA key restricted with --projects reaches only those projects, and it cannot change org-level settings. If every project it was restricted to is deleted, it reaches no project at all. A key bound to a service account can only be restricted to that service account’s project.\n\nList and revoke\n\nTerminal window\n\nnodus get apikeys\nnodus delete apikey ci\n\nA deleted key stops working within 30 seconds everywhere. Listing shows the key’s prefix ( nodus sk live 01j9… ), scopes, owner, last use and expiry, never the key itself. Members can delete their own keys. Deleting another member’s key takes an Admin or Owner, and deleting a service account’s key takes permission to manage service accounts in its project.\n\nKeys the CLI creates\n\n nodus login creates one key per org, named cli- - , with every scope your role grants, valid for 90 days and labelled as launched by the CLI. nodus logout revokes it.\n\nIf a key leaks\n\nDelete it. Keys are registered with GitHub secret scanning, so a key pushed to a public repository is reported to us and revoked."},{"id":"docs/guides/access/connected-agents","url":"https://nodus-platform-site.pages.dev/docs/guides/access/connected-agents/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/connected-agents.md","title":"Connected agents","description":"See and disconnect the MCP clients you let act on Nodus for you.","stage":"GA","headings":[{"depth":2,"slug":"seeing-what-is-connected","text":"Seeing what is connected"},{"depth":2,"slug":"disconnecting","text":"Disconnecting"}],"text":"Connected agents\n\n See and disconnect the MCP clients you let act on Nodus for you.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/connected-agents/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nWhen an MCP client such as Claude Code, Cursor or Codex connects to Nodus, the consent page asks you to pick an org and the scopes the client may use. Nodus records that choice as an OAuth grant . The client can do only what the grant allows, and never more than your own role in that org. It cannot create API keys or other grants.\n\nA client is connected to one org at a time. Approving it again, for any org, replaces the earlier grant.\n\nSeeing what is connected\n\nConsole › Settings › Connected agents lists each client with its scopes and when it was last used. Or:\n\nTerminal window\n\nnodus get oauthgrants\n\nYou see your own grants. Admins and Owners see every member’s grants in the org.\n\nDisconnecting\n\nTerminal window\n\nnodus delete oauthgrant claude-code-3fa2c1\n\nThe client stops working within 30 seconds, even with an access token it already holds. A grant also lapses 30 days after its last use, and removing a member from the org disconnects all of their clients there."},{"id":"docs/guides/access/device-login","url":"https://nodus-platform-site.pages.dev/docs/guides/access/device-login/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/device-login.md","title":"Device login","description":"Sign the CLI in on a machine without a browser, or from CI with an existing key.","stage":"GA","headings":[{"depth":2,"slug":"in-ci","text":"In CI"}],"text":"Device login\n\n Sign the CLI in on a machine without a browser, or from CI with an existing key.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/device-login/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nOn a remote server or container with no browser:\n\nTerminal window\n\nnodus login --device\n\nThe CLI prints a code like BDWP-HMTR and a link. Open the link on any device where you are signed in, check that the code matches, pick the orgs and approve. The CLI finishes on its own within a few seconds. The code works for 10 minutes; denying it stops the login.\n\nThe CLI polls every 5 seconds. Every key it receives is the same kind nodus login creates: one per chosen org, 90 days, revocable with nodus logout or nodus delete apikey .\n\nIn CI\n\nPass an existing key instead of signing in:\n\nTerminal window\n\necho \"$NODUS API KEY\" nodus login --with-token\n\nPrefer a service account token for CI, so it keeps working when people leave."},{"id":"docs/guides/access/members-and-roles","url":"https://nodus-platform-site.pages.dev/docs/guides/access/members-and-roles/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/members-and-roles.md","title":"Members and roles","description":"Invite people to your org, choose their role, and remove members safely.","stage":"GA","headings":[{"depth":2,"slug":"roles","text":"Roles"},{"depth":2,"slug":"invite-someone","text":"Invite someone"},{"depth":2,"slug":"let-your-company-join-without-invites","text":"Let your company join without invites"},{"depth":2,"slug":"change-a-role","text":"Change a role"},{"depth":2,"slug":"remove-a-member","text":"Remove a member"}],"text":"Members and roles\n\n Invite people to your org, choose their role, and remove members safely.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/members-and-roles/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nRoles\n\n Role Can \n - - \n Owner Everything, including granting or removing the Owner role \n Admin Everything except the Owner role: members, invites, projects, API keys, service accounts, billing \n Member Create and manage work: jobs, sandboxes, volumes, secrets, their own API keys \n Viewer Read everything, change nothing \n\n nodus auth can-i create jobs asks the server whether you hold a permission.\n\nInvite someone\n\nTerminal window\n\nnodus create invite --email ada@example.com --role Member\n\nThe invite link works for 72 hours and only once. The invitee signs in with that email address (it must be verified) and accepts it in the console, where pending invites also appear as a banner. You can invite up to your seat limit, 10 members by default; pending invites count as seats. An invite nobody accepts stays under Team › Invites after it expires, marked Expired and holding no seat, so you can resend it. Once your org has bought credits, an Owner or Admin can change the seat limit under Team › Members › Change , up to 1,000. When an invite is accepted, or someone is made an Admin or Owner, every Owner and Admin gets an email.\n\nTerminal window\n\nnodus get invites\nnodus request resend invite/inv-01j9abc # a fresh link, at most 3 times, 10 minutes apart\nnodus delete invite inv-01j9abc\n\nYou can invite someone with a role up to your own: Admins invite Admins, Members and Viewers; only Owners invite Owners.\n\nLet your company join without invites\n\nAn Owner or Admin adds your company’s email domain under Team › Company domains , then adds the TXT record shown there at your DNS provider and selects Verify . After that, anyone who signs up with a verified email at that domain is offered your org during sign-up and joins it as a Member in one click. Seats still apply, public mailbox domains such as gmail.com are refused, and a domain belongs to one org at a time. Remove the domain to stop new joins; people who already joined stay members. Someone you remove from the org can only come back through an invite.\n\nChange a role\n\nTerminal window\n\nnodus get members\nnodus edit member usr-01j9abc --role Admin\n\nAn org always keeps at least one Owner: demoting or removing the last Owner is refused with 409 Conflict .\n\nRemove a member\n\nTerminal window\n\nnodus delete member usr-01j9abc\nnodus delete member usr- # leave an org yourself\n\nRemoving a member revokes their API keys and third-party app access in that org at once. Keys bound to a service account keep working, so CI does not break when the person who set it up leaves."},{"id":"docs/guides/access/mfa","url":"https://nodus-platform-site.pages.dev/docs/guides/access/mfa/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/mfa.md","title":"Account security","description":"Sign in and control access with organization roles and scoped credentials.","stage":"GA","headings":[],"text":"Account security\n\n Sign in and control access with organization roles and scoped credentials.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/mfa/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nSign in with your email and password or an enabled Google or GitHub account. You can then use the resources and settings your organization role allows. Nodus does not require QR scanning, an authenticator app or a separate verification code after sign-in. Email verification for new accounts and password recovery still apply.\n\nCLI and device login use your existing signed-in session. API keys and service account tokens keep their own scopes. Review them under API keys."},{"id":"docs/guides/access/projects","url":"https://nodus-platform-site.pages.dev/docs/guides/access/projects/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/projects.md","title":"Projects","description":"Group work inside an org with projects, and restrict members and keys to some of them.","stage":"GA","headings":[{"depth":2,"slug":"restrict-access-to-projects","text":"Restrict access to projects"},{"depth":2,"slug":"cap-a-teams-spend","text":"Cap a team’s spend"}],"text":"Projects\n\n Group work inside an org with projects, and restrict members and keys to some of them.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/projects/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA project is a namespace inside your org. Jobs, Sandboxes, Volumes, Secrets and the other resources you create live in one project. Every org has default , which cannot be deleted.\n\nTerminal window\n\nnodus create project research --display-name \"Research\"\nnodus get projects\nnodus run -p research -- python train.py\n\nIn the console, Team › Projects lists each project with how many members can use it, its spend over the last 30 days and its spend limit, and has New project . Give each team its own project: its runs, storage and keys stay apart, and its spend shows on its own line.\n\nA project name is lowercase letters, digits and dashes. nodus , team , settings , billing , integrations , infrastructure and support are reserved.\n\nCreating, changing and deleting projects needs the projects:write scope, which Admins and Owners hold. Deleting a project deletes everything in it: running work is cancelled first. Usage and billing records are kept.\n\nRestrict access to projects\n\nMembers and API keys can be limited to some projects:\n\nTerminal window\n\nnodus edit member usr-01j9abc --projects research\nnodus create apikey ci --projects research --scopes jobs:write\n\nA member or key with a project list sees and changes only those projects. An empty list means every project.\n\nCap a team’s spend\n\nA budget scoped to a project caps that team. From Team › Projects , open a project’s menu and choose Set a spend limit , or:\n\nTerminal window\n\nnodus create budget vision-monthly --limit 500 --scope-project vision-team\n\nBudgets can also cap one person, or one person inside one project. When several budgets cover the same work, the tightest one applies."},{"id":"docs/guides/access/quotas","url":"https://nodus-platform-site.pages.dev/docs/guides/access/quotas/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/quotas.md","title":"Quotas","description":"See the limits on your org, what counts against them, and how to raise them.","stage":"GA","headings":[],"text":"Quotas\n\n See the limits on your org, what counts against them, and how to raise them.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/quotas/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nQuotas keep one mistake from running away with your credits. See yours with:\n\nTerminal window\n\nnodus get quota default -o yaml\n\n Quota Default before your first purchase After a purchase \n - - - \n members (seats, including pending invites) 10 10 \n apiKeys (live keys, including CLI logins and ServiceAccount tokens) 50 200 \n liveSandboxes (stopped ones count until their delete completes) 10 100 \n storageGiB 50 1024 \n egressGiBPerDay 20 100 \n nodusNodes 1 10 \n gangNodes (multi-node jobs) 0 8 \n inferenceRPM / inferenceTPM 60 / 100,000 600 / 1,000,000 \n assistantUSDPerDay $1 $10 \n\nGoing over a count quota refuses the create with 429 QuotaExceeded , naming the quota. Nothing already running is touched. Daily quotas ( egressGiBPerDay , assistantUSDPerDay ) reset at 00:00 UTC. Expired and revoked API keys do not count against apiKeys .\n\nMulti-node jobs are in Beta: your org needs the distributed training beta, and gangNodes counts the nodes of running multi-node jobs. A job that would go over it waits in the queue instead of failing.\n\nOwners and Admins of an org that has bought credits set their own seats, up to 1,000, under Team › Members › Change , or with PATCH /apis/nodus.dev/v1/quotas/default and {\"hard\": {\"members\": 25}} . Seats never go below the members and pending invites you have. To raise any other org limit, or seats past 1,000, contact support. To cap what one project or one person spends, use a Budget."},{"id":"docs/guides/access/service-accounts","url":"https://nodus-platform-site.pages.dev/docs/guides/access/service-accounts/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/service-accounts.md","title":"Service accounts","description":"Give CI and code running on Nodus an identity that belongs to the project, not to a person.","stage":"GA","headings":[{"depth":2,"slug":"tokens-for-ci","text":"Tokens for CI"},{"depth":2,"slug":"inside-nodus","text":"Inside Nodus"}],"text":"Service accounts\n\n Give CI and code running on Nodus an identity that belongs to the project, not to a person.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/service-accounts/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA service account is an identity for machines. Every project has one named default , used by code running inside Nodus for its own calls. Create more for CI and automation:\n\nTerminal window\n\nnodus create serviceaccount ci -p research --scopes jobs:write,volumes:read\n\nManaging service accounts needs serviceaccounts:write (Admins and Owners).\n\nTokens for CI\n\nTerminal window\n\nnodus create token sa/ci -p research --duration 720h\n\nThis prints a key bound to the service account, once. It is valid for 90 days by default and at most one year. Its permissions are the service account’s scopes, narrowed further by any --scope you pass.\n\nService account keys do not belong to a person: they keep working when the member who created them leaves the org, and they stop working when the service account is deleted. Use them for the GitHub Action and any shared automation.\n\nInside Nodus\n\nCode running in a Job, Sandbox or Agent reaches the API through /run/nodus/api.sock with a short-lived token of its service account. The token is never in the environment or on disk, and it stops working the moment its run is stopped or moved."},{"id":"docs/guides/access/sign-up-and-orgs","url":"https://nodus-platform-site.pages.dev/docs/guides/access/sign-up-and-orgs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/sign-up-and-orgs.md","title":"Sign up and create your org","description":"Create an account, create or join an organization, and switch between the orgs you belong to.","stage":"GA","headings":[{"depth":2,"slug":"sign-up","text":"Sign up"},{"depth":2,"slug":"start-from-the-terminal","text":"Start from the terminal"},{"depth":2,"slug":"create-another-org","text":"Create another org"},{"depth":2,"slug":"switch-orgs","text":"Switch orgs"},{"depth":2,"slug":"org-settings-and-activity","text":"Org settings and activity"}],"text":"Sign up and create your org\n\n Create an account, create or join an organization, and switch between the orgs you belong to.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/sign-up-and-orgs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nEverything you run on Nodus belongs to an organization (org). Credits, projects, members and API keys are all per org.\n\nSign up\n\nSign up at the console with email, Google or GitHub. Nothing is created until you choose:\n\n If someone invited you, the console shows the invite first. Accept it to join their org.\n Otherwise the console asks you to Create your org . The name is filled in for you, and you can change it.\n\nThe first org you create with a verified email address receives the $30 starter credit , valid for 30 days. Joining an org by invite never adds credit, so accepting an invite does not use up your own starter credit.\n\nStart from the terminal\n\nTerminal window\n\nnodus login\n\nThis opens the console in your browser. Sign in, pick the orgs this computer should use, and the CLI stores one key per org in your OS keychain. If you have no org and no pending invite, nodus login creates your org for you, with the starter credit. On a machine without a browser, use device login.\n\nCreate another org\n\nTerminal window\n\nnodus create org acme --display-name \"Acme Research\"\n\nYou become its Owner. Every org starts with a project named default .\n\nSwitch orgs\n\nThe console has an org switcher. On the CLI, each org is a context:\n\nTerminal window\n\nnodus config get-contexts\nnodus config use-context acme\nnodus get jobs --org acme # one command in another org\n\nAPI calls choose an org with the Nodus-Org header (the org name or id). Without it, a signed-in user acts in the org they used last.\n\nOrg settings and activity\n\nAdmins can rename the org; its name (the slug in console URLs) stays the same:\n\nTerminal window\n\nnodus patch org acme --patch '{\"displayName\": \"Acme Labs\"}'\nnodus get org acme -o yaml # role, seats used and your org's enabled features\n\nAdmins and Owners see the org’s activity under Settings → Activity : sign-ins, key, member and invite changes, each with who did it and when, for the last day up to the last year, and Export CSV downloads what you see. The API serves it at GET /apis/nodus.dev/v1/auditevents , oldest first: ?action=auth. selects sign-ins only, since sets the start, and a full page carries metadata.continue to pass as continue for the next one."},{"id":"docs/guides/access/ssh-keys","url":"https://nodus-platform-site.pages.dev/docs/guides/access/ssh-keys/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/access/ssh-keys.md","title":"SSH keys","description":"Add the public keys you use to open a shell in a Workspace.","stage":"GA","headings":[{"depth":2,"slug":"who-can-connect","text":"Who can connect"},{"depth":2,"slug":"removing-a-key","text":"Removing a key"}],"text":"SSH keys\n\n Add the public keys you use to open a shell in a Workspace.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/access/ssh-keys/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nAn SSH key belongs to you, not to an org. Add it once and it works in the Workspaces of every org you belong to. Console › Account › SSH keys lists yours whatever org is selected.\n\nTerminal window\n\nnodus create sshkey laptop --from-file ~/.ssh/id ed25519.pub\nnodus get sshkeys\n\nWithout a key yet, let the CLI make one. --generate writes a new Ed25519 key pair to ~/.ssh/id ed25519 (the public key next to it, with .pub ) and adds the public key. The private key file is readable by you alone, and nodus never sends it anywhere. It refuses to overwrite a file that is there: use --key-file PATH to write elsewhere. Add a passphrase afterwards with ssh-keygen -p -f ~/.ssh/id ed25519 .\n\nTerminal window\n\nnodus create sshkey --generate\n\nLeave the name out and the key is called - . --from-file - reads the key from standard input, and a file that holds a private key is refused.\n\nOr over the API:\n\nTerminal window\n\ncurl -X POST \"$NODUS API URL/apis/nodus.dev/v1/sshkeys\" \\\n -H \"Authorization: Bearer $NODUS API KEY\" -H \"Content-Type: application/json\" \\\n -d \"{\\\"name\\\": \\\"laptop\\\", \\\"publicKey\\\": \\\"$(cat ~/.ssh/id ed25519.pub)\\\"}\"\n\nNodus accepts ssh-ed25519 , ecdsa-sha2- and ssh-rsa keys of at least 3072 bits. It stores the key without its comment and shows its SHA256: fingerprint, which is what ssh-keygen -lf ~/.ssh/id ed25519.pub prints. Each name and each key can be added once, and you can hold up to 20 keys.\n\nWho can connect\n\n nodus ssh workspace/ authenticates with your keys when you are an Owner, Admin or Member of the Workspace’s org and your membership covers its project. Viewers cannot open a shell. When you leave an org, your keys stop working in its Workspaces, and they keep working everywhere else.\n\nRemoving a key\n\nTerminal window\n\nnodus delete sshkey laptop\n\nDeleting a key ends the live sessions opened with it, and the next login with it is refused."},{"id":"docs/guides/agents","url":"https://nodus-platform-site.pages.dev/docs/guides/agents/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/agents.md","title":"Agents","description":"Run an agent on Claude in its own sandbox, send it follow-up prompts, and pay for the model from your credits.","stage":"Beta","headings":[{"depth":2,"slug":"sign-in-and-create-an-agent","text":"Sign in and create an agent"},{"depth":2,"slug":"start-a-run","text":"Start a run"},{"depth":2,"slug":"from-python","text":"From Python"},{"depth":2,"slug":"keep-a-run-open-for-follow-ups","text":"Keep a run open for follow-ups"},{"depth":2,"slug":"run-agents-in-parallel","text":"Run agents in parallel"},{"depth":3,"slug":"watch-progress-and-read-the-answers","text":"Watch progress and read the answers"},{"depth":3,"slug":"how-many-run-at-once","text":"How many run at once"},{"depth":3,"slug":"dependencies","text":"Dependencies"},{"depth":3,"slug":"the-group-cost-cap","text":"The group cost cap"},{"depth":3,"slug":"seal-cancel-and-delete","text":"Seal, cancel and delete"},{"depth":2,"slug":"evaluate-an-agent","text":"Evaluate an agent"},{"depth":2,"slug":"use-your-own-anthropic-key","text":"Use your own Anthropic key"},{"depth":2,"slug":"what-a-run-costs","text":"What a run costs"},{"depth":2,"slug":"limits","text":"Limits"}],"text":"Agents\n\n Run an agent on Claude in its own sandbox, send it follow-up prompts, and pay for the model from your credits.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/agents/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n Beta: this feature may change.\n\nAn Agent is a definition: a system prompt, the Claude access it runs on and a cost cap. An AgentRun is one conversation with it. Each run works in its own sandbox, where the model runs shell commands to check its work, and it survives restarts of Nodus by replaying recorded results. An interrupted command or model call using your own key parks with NeedsResolution when its outcome is uncertain; inspect its effects before starting replacement work.\n\nAgents, parallel AgentGroups and Environment evaluation batches are in Beta. Each run is one agent conversation.\n\nSign in and create an agent\n\nagent.yaml\n\napiVersion: nodus.dev/v1\nkind: Agent\nmetadata:\n name: hello-agent\nspec:\n system: \n You are a careful assistant. Run shell commands in your sandbox to check your work, and answer in one short\n paragraph.\n # Nodus's Claude: each prompt is routed to Haiku, Sonnet, Opus or Fable, and billed from your credits.\n model:\n access: Nodus\n perRunMaxCostUSD: \"0.50\"\n\nTerminal window\n\nnodus login\nnodus apply -f agent.yaml\n\nThe agent uses Nodus’s Claude. Every prompt goes to a small router that picks the right Claude model for it (Haiku for simple requests up to Fable for demanding ones). That model answers the whole prompt, including its commands, so the model never changes in the middle of a turn. If the router is unavailable, Sonnet answers. Restrict the choice with spec.model.families , for example [haiku, sonnet] .\n\nStart a run\n\nrun.yaml\n\napiVersion: nodus.dev/v1\nkind: AgentRun\nmetadata:\n name: hello-agent-run\nspec:\n agent: hello-agent\n input: \"Use the shell to print the Python version, then tell me which one it is.\"\n\nTerminal window\n\nnodus apply -f run.yaml\nnodus wait agentrun/hello-agent-run --for=jsonpath='{.status.phase}'=Succeeded --timeout=8m\nnodus get agentrun/hello-agent-run -o yaml\n\nOr without a file: nodus create agentrun triage --agent hello-agent --prompt \"Summarize the logs\" . To start from Nodus’s ready-made assistant, name the template claude-assistant .\n\n status.turns lists each prompt with the model that served it, how it was chosen ( jev or fallback ) and its cost. status.answer.preview holds the first 4 KiB of the final answer; nodus agentrun answer prints all of it.\n\nTerminal window\n\nnodus agentrun answer hello-agent-run\nnodus agentrun steps hello-agent-run\n\n steps lists what the run did in order: creating its sandbox, each routing and model call, each command. A step that shows Completed is never repeated, even after a restart.\n\nThe console’s Agents page shows the same: each run with the model of its latest turn and its cost, and on a run every turn with the model that served it, how that model was chosen and what the turn cost, then the answer and the recorded steps.\n\nFrom Python\n\nimport nodus\n\nagent = nodus.ClaudeAgent(\"helper\", system=\"You are careful.\", families=[\"haiku\", \"sonnet\"])\nprint(agent.remote(\"Use the shell to print the Python version\")) # runs to the end, returns the answer\n\nrun = agent.submit(\"Summarize the logs\", keep alive=True) # a handle: run.answer(), run.steps(), run.cancel()\nrun.send(\"message\", \"Now the staging logs\")\n\n ClaudeAgent creates the agent on first use. Pass api key secret=\"anthropic-key\" to use your own Anthropic key, or ClaudeAgent.from name(\"claude-assistant\", project=\"nodus\") to run the ready-made template. To run many at once from Python, see Run agents in parallel.\n\nKeep a run open for follow-ups\n\nSet spec.keepAlive: true and the run waits for messages after each turn instead of finishing. While it waits it uses no model and keeps its sandbox, which can still accrue compute charges. Set spec.deadline to bound the run’s lifetime. The deadline also applies while waiting for a message: the run fails with DeadlineExceeded and its sandbox is deleted without another model call.\n\nTerminal window\n\nnodus apply -f run.yaml # with keepAlive: true\nnodus agentrun send hello-agent-run \"Now do the same for the staging logs\" --key staging-1\n\nSending the same --key again delivers the message once. A run that is still working reads the message after its current turn; a finished run refuses it with AgentRunFinished . Cancel a run with nodus cancel agentrun/NAME ; its sandbox is deleted.\n\nThe same calls are REST: POST …/agentruns/NAME/messages with {\"payload\": \"…\", \"messageKey\": \"…\"} , and GET on …/steps and …/answer .\n\nRun agents in parallel\n\nAn AgentGroup runs many runs of one agent at once. You choose how many run at the same time and how much the whole group may spend. The other runs wait their turn, and one command cancels or deletes all of them. Create the agent and the group, then one run for each task:\n\nagent.yaml\n\napiVersion: nodus.dev/v1\nkind: Agent\nmetadata:\n name: parallel-agents-worker\nspec:\n system: \n Follow the task's requested answer format; otherwise answer in one short sentence.\n # Nodus's Claude, limited to Haiku so each run stays far below its cap.\n model:\n access: Nodus\n families: [haiku]\n maxTokens: 1024\n perRunMaxCostUSD: \"0.10\"\n\ngroup.yaml\n\napiVersion: nodus.dev/v1\nkind: AgentGroup\nmetadata:\n name: parallel-agents\nspec:\n agent: parallel-agents-worker\n limits:\n maxActive: 2 # at most two runs go at once\n maxCostUSD: \"0.60\" # a run starts only while its own cap still fits under what is left\n\nruns.yaml\n\napiVersion: nodus.dev/v1\nkind: AgentRun\nmetadata:\n name: parallel-agents-a\nspec:\n agent: parallel-agents-worker\n group: parallel-agents\n taskKey: a\n input: \"What is the capital of France?\"\n---\napiVersion: nodus.dev/v1\nkind: AgentRun\nmetadata:\n name: parallel-agents-b\nspec:\n agent: parallel-agents-worker\n group: parallel-agents\n taskKey: b\n input: \"What is the capital of Japan?\"\n---\napiVersion: nodus.dev/v1\nkind: AgentRun\nmetadata:\n name: parallel-agents-c\nspec:\n agent: parallel-agents-worker\n group: parallel-agents\n taskKey: c\n dependsOn: [a, b] # c waits until a and b have succeeded\n input: \"Say in one sentence that both questions have been answered.\"\n\nTerminal window\n\nnodus apply -f agent.yaml\nnodus apply -f group.yaml\nnodus apply -f runs.yaml\nnodus patch ag/parallel-agents --patch '{\"spec\":{\"sealed\":true}}'\n\nOr without a file: nodus create agentgroup parallel-agents --agent parallel-agents-worker --max-active 2 --max-cost 0.60 .\n\nEach run is a task of the group: spec.group names the group and spec.taskKey names the task (a DNS label that is unique in the group). The run is named - ; leave metadata.name out or set exactly that. The run uses the group’s agent, so spec.agent must match it. The last command seals the group, which says that no more runs are coming (see below).\n\nWatch progress and read the answers\n\nTerminal window\n\nnodus wait ag/parallel-agents --for=jsonpath='{.status.phase}'=Succeeded --timeout=12m\nnodus get ag\nnodus get ar --field-selector spec.group=parallel-agents\nnodus agentrun answer parallel-agents-c\n\n nodus get ag lists each group with its phase, RUNS (finished successfully out of all) and ACTIVE runs, and its cost. status.counts splits the runs into queued , active , waiting (for a message or for funds), succeeded , failed and cancelled . status.blockedReasons says why queued runs have not started and how many wait for each reason: DependencyWait , MaxActiveReached or MaxCostReached . A queued run’s own status.reason is DependencyWait , or GroupWait while it waits for a free slot or for room under the cost cap. The runs you list are ordinary runs, so status.answer , nodus agentrun answer and nodus agentrun steps work on each of them.\n\nA group is Running until it is sealed and every run has finished. It then becomes Succeeded , or Failed with reason RunsFailed when a run failed or was cancelled because a dependency failed. A sealed group with no runs succeeds.\n\nHow many run at once\n\n spec.limits.maxActive is how many runs the group lets run at the same time (default 10, 1 to 1000). Change it on a live group:\n\nTerminal window\n\nnodus patch ag/parallel-agents --patch '{\"spec\":{\"limits\":{\"maxActive\":4}}}'\n\n spec.limits.maxPending is how many unfinished runs the group takes in (default 10000); a run beyond it is refused.\n\nDependencies\n\n spec.dependsOn lists the task keys a run waits for. It starts only after every one of them has succeeded, and until then it is queued with reason DependencyWait . A dependency must already be in the group, so create the runs it depends on first (earlier in the same file works). When a dependency fails, the run is cancelled with reason DependencyFailed ; when a dependency was cancelled, the run is cancelled with reason DependencyCancelled .\n\nThe group cost cap\n\n spec.maxCostUSD is one cap for the whole group. A run starts only when its own cap (the agent’s perRunMaxCostUSD ) still fits under what is left: the cap minus what the group has spent and minus the caps of the runs that have started and not finished. The members together can therefore never spend more than the cap, and runs that do not fit yet wait with MaxCostReached . In the example, the $0.60 cap has room for the caps of six runs at $0.10 each before any run has spent anything. Raise the cap on a live group with nodus patch ; it cannot be lowered. If the cap is smaller than one run’s cap, no run starts: raise it or lower the agent’s perRunMaxCostUSD . status.cost shows totalUSD and limitUSD . The cap covers the runs’ Claude usage (model and routing calls), the spend each run’s perRunMaxCostUSD limits; the sandbox each run works in is billed on its own and is not counted against either cap.\n\nSeal, cancel and delete\n\nA group takes new runs until you seal it with spec.sealed: true . Sealing cannot be undone, and a sealed group refuses new runs with a conflict. A group finishes only once it is sealed, so seal it after you create the last run.\n\nTerminal window\n\nnodus cancel ag/parallel-agents\nnodus delete ag/parallel-agents\n\n nodus cancel cancels every run that has not finished and leaves the finished runs, with their answers, as they are. status.phase goes through Cancelling to Cancelled . nodus delete removes the group together with all of its runs.\n\nEvaluate an agent\n\nCreate an evaluation group with an Environment at an explicit version. Nodus generates the fixed task batch, runs it through ordinary AgentRuns, and grades completed answers using the Environment’s isolated grader. There is no manual submit or seal step. Set a project Budget to limit total spend: the group cap covers model usage, while agent and grading Sandbox compute is billed separately.\n\nAfter deploying the example’s parallel-agents-worker Agent above, run this Python example. It creates two held-out tasks, waits for a terminal result, prints per-case outcomes and deletes the evaluation group afterward.\n\nevaluation.py\n\n\"\"\"Evaluate the deployed example agent on a fixed, versioned task batch.\"\"\"\n\nimport nodus\n\ndef main() - None:\n # Model caps exclude Sandbox compute; set a project Budget before running this example.\n evaluation = nodus.AgentGroup.create(\n \"parallel-agents-eval\",\n agent=\"parallel-agents-worker\",\n max active=2,\n max cost=\"0.20\",\n evaluation={\n \"environment\": \"nodus/arithmetic-v2@2.0.0\",\n \"split\": \"test\",\n \"tasks\": 2,\n \"repetitions\": 1,\n \"seed\": 42,\n \"timeout\": \"10m\",\n },\n )\n try:\n status = evaluation.wait(timeout=660)\n results = evaluation.results()\n print(\n f\"{status.phase}: cases={len(results)}, pending={status.evaluation.pending}\"\n )\n print(\"pass rate:\", status.evaluation.get(\"passRate\"))\n print(\"mean reward:\", status.evaluation.get(\"meanReward\"))\n for result in results:\n print(result.task id, result.state, result.get(\"verdict\"))\n finally:\n evaluation.delete()\n\nif name == \" main \":\n main()\n\nThe batch is at most 100 cases including repetitions. Every member pins the same Agent revision; results record the Environment version, resolved image digest, split and seed. Only task prompts reach the agent, never hidden answers. The timeout defaults to 30 minutes, may be 1 minute to 24 hours, and includes manifest preparation, queueing and funding waits from group creation. Each member receives that absolute deadline. Grading also stops at the deadline; completed answers that could not be scored are shown as evaluation timeouts, excluded from pass rate and mean reward.\n\nOpen the group in the console to see scored cases, pass rate, mean reward, pending cases and execution/grading failures separately. The Results view links each case to its AgentRun and shows grader evidence. A failed run, missing answer or grader failure is unscored and does not become a zero reward. A wrong or invalid answer is a scored result. Per-case records share the member journal’s 30-day payload retention; deleting the group removes its runs and records. Canceling or finishing the group cleans up its grading pool.\n\nUse your own Anthropic key\n\nStore the key in a Secret under the key ANTHROPIC API KEY and point the agent at it. Nodus then charges nothing for the model, and your Anthropic account pays for it.\n\nspec:\n model:\n access: BYOK\n name: claude-sonnet-5-5\n apiKeySecret: anthropic-key\n\nWhat a run costs\n\nOn Nodus’s Claude each model call, and the routing call before each turn, is billed from your credits at cost divided by 0.875. spec.perRunMaxCostUSD (default $1.00) caps what a run spends on the model. Before each call the run checks that the call, at its largest (the whole conversation as input and spec.maxTokens of output), still fits under the cap. When it does not, the run waits ( reason: MaxCostReached ) and makes no more calls. The cap is copied from the agent when the run is created, so changing the agent does not change a run that has started: to continue the work, cancel the run and start a new one. If your credits run out, the run waits instead and continues when you add credits. A run on your own Anthropic key is not charged by Nodus and has no cap to reach. Sandbox compute is billed as for any sandbox. status.cost shows the model spend so far. Runs in a group also share the group’s cap (see Run agents in parallel).\n\nLimits\n\n One run works one prompt at a time, with at most spec.maxToolCalls commands per turn (default 50).\n The input and each message are at most 256 KiB. A run’s recorded payloads are kept 30 days after it ends.\n A run created for a stopped agent ( spec.state: Stopped ) is refused with AgentStopped .\n A group takes at most 10000 unfinished runs. Create runs in batches of up to 100 by posting an AgentRunList to …/agentruns with an Idempotency-Key header; a batch is all-or-nothing, and its runs are created in list order, so list each run after the runs it depends on.\n\nSee Durable execution for what survives a restart."},{"id":"docs/guides/assistant","url":"https://nodus-platform-site.pages.dev/docs/guides/assistant/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/assistant.md","title":"Ask Nodus","description":"Ask the console assistant why a Job failed, what it cost, or to draft a manifest. It cites the docs, and changes only after you confirm.","stage":"Beta","headings":[{"depth":2,"slug":"ask-a-question","text":"Ask a question"},{"depth":2,"slug":"changes-need-your-confirmation","text":"Changes need your confirmation"},{"depth":2,"slug":"drafts","text":"Drafts"},{"depth":2,"slug":"checkpoint-suggestions","text":"Checkpoint suggestions"},{"depth":2,"slug":"your-preferences-and-followed-objects","text":"Your preferences and followed objects"},{"depth":2,"slug":"chats-and-privacy","text":"Chats and privacy"},{"depth":2,"slug":"limits","text":"Limits"}],"text":"Ask Nodus\n\n Ask the console assistant why a Job failed, what it cost, or to draft a manifest. It cites the docs, and changes only after you confirm.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/assistant/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n Beta: this feature may change.\n\n Ask Nodus is the assistant in the console. It answers questions about your org: why a Job restarted, what a Sandbox cost this week, which flags a manifest needs. It uses the same tools as the MCP server, with your own role and scopes, so it can see and do only what you can.\n\nAsk a question\n\nOpen Ask Nodus in the console and type. Some things to try:\n\n “Why did job/train fail?”\n “What did my Sandboxes cost this month?”\n “How do I make train survive being preempted?”\n “Write a manifest that runs python train.py on an L4 with a $5 cap.”\n\nWhen an answer relies on the docs, it links the pages, and the links are the pages the assistant actually read.\n\nChanges need your confirmation\n\nThe assistant can propose a change, but it cannot make one. When an answer would create, delete, suspend or otherwise change something, the turn stops and shows exactly what will run. For a new object that includes the server’s dry-run: the object as it would be stored, its estimated cost and anything that blocks it.\n\nChoose Confirm to run it, or Cancel . Confirming runs the held change once, with your credential, so the API checks your permissions again. Asking a new question cancels a change you did not confirm. The assistant proposes at most one change per answer.\n\nDrafts\n\nWhen you ask for a manifest, the assistant writes one and checks it with the same dry-run that nodus apply --dry-run=server uses. A draft is shown to you only after that check passes. If the check fails, the assistant sees the error and fixes the manifest first. Each draft comes with the equivalent CLI command and Python, and its estimate. A draft is not created until you apply it, or ask the assistant to and confirm.\n\nCheckpoint suggestions\n\nFor a Job that would start over after a preemption, the assistant can suggest a manifest that checkpoints, with that manifest’s dry-run. Recovery settings cannot change on a running Job, so the suggestion is a new manifest. Your program still has to write its state to the declared checkpoint paths.\n\nYour preferences and followed objects\n\nYou can give the assistant a short note about how you like answers, such as “short answers, show the CLI command”. It reads the note as your preference, and the note cannot change the assistant’s rules. You can also follow up to 8 objects so the assistant knows what you are working on.\n\nBoth are yours alone in each org. In the API:\n\nTerminal window\n\n$ curl -s \"$NODUS API URL/assistant/v1/profile\" -H \"Authorization: Bearer $NODUS API KEY\"\n$ curl -s -X PUT \"$NODUS API URL/assistant/v1/profile\" -H \"Authorization: Bearer $NODUS API KEY\" \\\n -H \"Content-Type: application/json\" -d '{\"note\": \"Short answers.\", \"revision\": 0}'\n\nThe note is at most 8 KiB. Send back the revision you read, and the update fails with a conflict if someone else changed the note in between. /assistant/v1/watches takes {\"watches\": [\" \", …]} .\n\nChats and privacy\n\nYour chats are stored per org, for you only. A chat holds up to 200 questions. Each answer is stored with a checksum, and deleting a chat deletes its answers. A tool the assistant ran is recorded by name with checksums of its arguments and its result, never their content, so file contents and command output are not kept. Chats are kept at most 400 days.\n\nNodus chooses a Claude model for each question and uses it for the whole answer. Nodus pays for the model calls; they do not use your credits.\n\nLimits\n\n Limit Value \n - - \n Questions 1 per second per user, with bursts of up to 20 \n Steps for one question 12 model rounds and 32 tool calls \n Time for one question 180 seconds \n Drafts 1 draft request every 5 seconds per user \n Daily allowance Each org has a daily allowance for the assistant, which resets at 00:00 UTC \n\nWhen a limit is reached the assistant says so and what happens next, instead of failing. Your Jobs, the docs and the CLI are not affected.\n\nNote\n\nThe assistant reads logs and files as data, never as instructions. Text inside a log cannot make it run anything, and anything it proposes still needs your confirmation."},{"id":"docs/guides/billing","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing.md","title":"Usage and billing","description":"Add credits, redeem a code, cap spending with Budgets, and see exactly what every run cost.","stage":"GA","headings":[{"depth":2,"slug":"check-your-balance","text":"Check your balance"},{"depth":2,"slug":"add-credits","text":"Add credits"},{"depth":2,"slug":"auto-recharge","text":"Auto-recharge"},{"depth":2,"slug":"redeem-a-promo-code","text":"Redeem a promo code"},{"depth":2,"slug":"see-what-you-spent","text":"See what you spent"},{"depth":2,"slug":"cap-spending-with-a-budget","text":"Cap spending with a Budget"},{"depth":2,"slug":"when-work-is-refused-for-money","text":"When work is refused for money"},{"depth":2,"slug":"payment-methods-and-billing-details","text":"Payment methods and billing details"},{"depth":2,"slug":"learn-more","text":"Learn more"}],"text":"Usage and billing\n\n Add credits, redeem a code, cap spending with Budgets, and see exactly what every run cost.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNodus is prepaid: you add credits, and every run reserves funds before it starts. New orgs get starter credit to try things out. This guide covers the everyday tasks; the billing concept page explains holds, captures and limits.\n\nCheck your balance\n\nTerminal window\n\n$ nodus billing\nAvailable $17.42 purchased $15.80 · credits $3.52 (starter, expires Oct 31)\nReserved $1.90 job/train-a, sandbox/sb-3, function/embed\nThis month $232.58 budget research-monthly 42 % of $500.00\nAuto-recharge on add $50.00 below $10.00 · visa •••• 4242\n\nIn the console, open Usage & billing . The header shows your available balance everywhere, in amber when it is low.\n\nAdd credits\n\nTerminal window\n\n$ nodus billing top-up 20\nOpening https://checkout.stripe.com/c/pay/cs live ... (expires in 60 min)\nWaiting for payment... added $20.00. Available $37.42.\nReceipt: https://invoice.stripe.com/i/...\n\nTop-ups are between $5.00 and $1,000.00 in whole cents, up to $5,000 per org per day. Payment happens on a Stripe-hosted page; Nodus never sees your card number. In the console, Add credits offers $10, $20, $50, $100 or a custom amount and brings you back to Billing when the payment completes. A top-up settles any unpaid charges first. Every top-up has a receipt and an invoice PDF under Receipts & invoices and in nodus billing receipts .\n\nTopping up needs the billing:write permission, which org Owners and Admins have.\n\nAuto-recharge\n\nAuto-recharge adds credit when your available balance drops below a threshold, so long runs never stop for money:\n\nTerminal window\n\n$ nodus billing auto-recharge --threshold 10 --amount 50\n\nThe first time, this opens a card setup page. Each recharge is a Stripe invoice with its own receipt. At most 5 recharges or $5,000 run per day, and three failed charges in a row turn auto-recharge off and email you. Turn it off with nodus billing auto-recharge --off .\n\n Save updates the amounts and warning level without turning auto-recharge on or off. Use Turn on or Turn off to change that setting. If another admin or failed payments change it while you edit, the console refreshes the settings and asks you to review them before saving again.\n\nRedeem a promo code\n\nTerminal window\n\n$ nodus billing redeem LAUNCH25\nRedeemed CH25: $25.00 of credit, expires 2026-11-30.\n\nSee Promo codes for limits and errors.\n\nSee what you spent\n\nTerminal window\n\n$ nodus get usage --group-by project,label:owner --since 30d\nPROJECT OWNER AMOUNT\nresearch ml $212.41\ndefault - $20.17\n\n$ nodus get usage --group-by project,meter --since 30d -o csv usage.csv\n$ nodus get transactions --since 7d\n\nGroup usage by project , kind , label: , meter , day , segment or rank . segment splits machine time into Boot , Restore , Running and Teardown , and rank itemizes the members of a multi-node run. The console Usage tab shows the same data as a daily chart and a table, and every table exports CSV.\n\nTransactions list every change to your balance, with the balance after it: top-ups, grants, captures, storage, egress, refunds and adjustments. What you pay for lists every kind of time and who pays for it.\n\nCap spending with a Budget\n\nA Budget is an enforced limit over the org, a project or a label selector, per month or in total. When a Block Budget is exhausted, work in its scope stops gracefully and new work is refused until the next period:\n\nbudget.yaml\n\napiVersion: nodus.dev/v1\nkind: Budget\nmetadata:\n name: examples-billing-monthly\nspec:\n limitUSD: \"25.00\"\n period: Monthly\n scope:\n project: examples\n action: Block\n thresholds: [50, 80, 100]\n\nTerminal window\n\n$ nodus apply -f budget.yaml\n$ nodus get budgets\n\nYou get an email at each threshold. The console Budgets tab creates and edits Budgets and previews which objects a scope matches now.\n\nTo cap one person, scope a Budget to them by email. It counts everything they start, with any of their API keys, including inference requests:\n\nTerminal window\n\n$ nodus create budget ada-monthly --limit 300 --period Monthly --scope-member ada@example.com\n\nWhen work is refused for money\n\nA create that cannot be funded fails with 402 and the exact amounts, and the console shows an Add credits action:\n\n Error What to do \n - - \n InsufficientCredits Add credits, or lower spec.maxCostUSD \n BudgetExceeded Raise the Budget’s spec.limitUSD , or wait for its next period \n ArrearsOutstanding Add credits; the top-up settles the unpaid charges first \n PaymentDisputed Contact support; new work waits until the dispute closes \n\nPayment methods and billing details\n\n nodus billing portal opens the Stripe Customer Portal, where you update cards, your billing email, address and tax ID, and download past invoices.\n\nLearn more\n\n Credits: the starter credit, grants and expiry\n Promo codes\n Refunds\n What you pay for"},{"id":"docs/guides/billing/add-a-card","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/add-a-card/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/add-a-card.md","title":"Add a card before running work","description":"Save a payment method without buying credits.","stage":"GA","headings":[],"text":"Add a card before running work\n\n Save a payment method without buying credits.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/add-a-card/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nOpen Billing → Payment method → Add a card , or run:\n\nTerminal window\n\nnodus billing portal --setup\n\nStripe saves your card without charging it or enabling auto-recharge. Your $30 signup credit keeps its original 30-day expiration. New work needs a verified card even when promotional credit covers its entire cost. An organization billing administrator must add the card. Other members can view its status.\n\nAfter returning from Stripe, wait for Billing to show that the card is ready. Returning to the page alone does not complete verification. If you canceled setup, choose Add a card again. CLI and SDK requests receive PaymentMethodRequired until verification completes. PaymentVerificationUnavailable means to retry shortly.\n\nRemoving or letting the card expire blocks new work. Existing running work continues within its available credits and budgets. A restart or new allocation checks the payment method again. Adding a card does not let work exceed your prepaid balance."},{"id":"docs/guides/billing/auto-recharge","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/auto-recharge/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/auto-recharge.md","title":"Auto-recharge","description":"Top up automatically from a saved card when your balance falls below a threshold, so running work never stops for money.","stage":"GA","headings":[{"depth":2,"slug":"turn-it-on","text":"Turn it on"},{"depth":2,"slug":"when-it-recharges","text":"When it recharges"},{"depth":2,"slug":"see-what-it-did","text":"See what it did"},{"depth":2,"slug":"when-a-recharge-fails","text":"When a recharge fails"},{"depth":2,"slug":"turn-it-off","text":"Turn it off"}],"text":"Auto-recharge\n\n Top up automatically from a saved card when your balance falls below a threshold, so running work never stops for money.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/auto-recharge/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nAuto-recharge buys credits for you when your available balance falls below a threshold you choose. Each recharge is charged to a saved card and emailed to you as a receipt, exactly like a top-up you buy yourself. It is off until you turn it on.\n\nTurn it on\n\nTerminal window\n\nnodus billing auto-recharge --threshold 10 --amount 50\n\nThis recharges $50 whenever your available balance drops below $10. If no card is on file yet, the command first opens a Stripe page to save one, then turns auto-recharge on. You can also save a card while buying credits with nodus billing top-up 20 --save-card , and the console billing page has an auto-recharge card that does both.\n\nThe same settings live on your org’s BillingAccount :\n\napiVersion: nodus.dev/v1\nkind: BillingAccount\nmetadata:\n name: default\nspec:\n autoRecharge:\n enabled: true\n thresholdUSD: \"10.00\" # at least 5.00\n amountUSD: \"50.00\" # 5.00 to 1000.00\n\n Setting Default Allowed \n - - - \n thresholdUSD \"10.00\" $5.00 or more \n amountUSD \"20.00\" $5.00 to $1,000.00 \n\nTurning it on without a saved card is refused; check PaymentMethodPresent in nodus billing , or save a card through the billing portal ( nodus billing portal ).\n\nWhen it recharges\n\nNodus checks your balance each time it renews the hold of running work, settles inference or storage, or refuses a new hold. It recharges when all of these are true:\n\n your available balance is below thresholdUSD ;\n a card is on file;\n no other auto-recharge is still being paid;\n the org has had fewer than 5 auto-recharges, and less than $5,000 of them, today (UTC);\n the last 3 auto-recharges did not all fail.\n\nA recharge usually lands in seconds. Work whose funding is running short keeps running while it is paid, so a successful recharge within that window means nothing stops.\n\nSee what it did\n\nEach recharge is a TopUp with spec.origin: AutoRecharge :\n\nTerminal window\n\nnodus get topups -o wide\nnodus billing receipts\n\n nodus billing shows the auto-recharge state from status.autoRecharge : consecutiveFailures , lastAttemptTime , lastResult and, when it has been switched off, disabledReason .\n\nWhen a recharge fails\n\nIf the card is declined, the recharge TopUp ends Failed ( PaymentFailed ), nothing is credited, and the billing email and the org’s owners and admins get an email. Auto-recharge tries again the next time the balance is checked.\n\nAfter 3 failures in a row, Nodus turns auto-recharge off, sets status.autoRecharge.disabledReason to ConsecutiveFailures , shows a banner in the console and emails the same people. To turn it back on:\n\n1. Update the card in the billing portal: nodus billing portal .\n2. Run nodus billing auto-recharge --threshold 10 --amount 50 again, or set enabled: true .\n\nIf your balance runs out while auto-recharge is off, work stops gracefully as described in When money runs out.\n\nTurn it off\n\nTerminal window\n\nnodus billing auto-recharge --off\n\nThe saved card stays on file for later; remove it in the billing portal."},{"id":"docs/guides/billing/budgets","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/budgets/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/budgets.md","title":"Budgets and spending caps","description":"Put an enforced limit on what an org, a project or a labelled group of work can spend, and cap any single run.","stage":"GA","headings":[{"depth":2,"slug":"create-a-budget","text":"Create a Budget"},{"depth":2,"slug":"what-a-budget-covers","text":"What a Budget covers"},{"depth":2,"slug":"what-happens-at-the-limit","text":"What happens at the limit"},{"depth":2,"slug":"notices","text":"Notices"},{"depth":2,"slug":"cap-a-single-run","text":"Cap a single run"}],"text":"Budgets and spending caps\n\n Put an enforced limit on what an org, a project or a labelled group of work can spend, and cap any single run.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/budgets/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNodus enforces two kinds of spending limits. Both are checked before any paid work starts and again every few minutes while it runs, so a limit always stops spend instead of reporting it afterwards.\n\n A Budget limits what matching work across your org can spend in a month or in total.\n maxCostUSD caps one object: a Job, Pipeline, Sweep, TrainingJob, Sandbox, Workspace, Function, Agent or InferenceEndpoint. A cap on a Pipeline or Sweep is shared by every run it creates.\n\nCreate a Budget\n\napiVersion: nodus.dev/v1\nkind: Budget\nmetadata:\n name: research-monthly\nspec:\n limitUSD: \"500.00\"\n period: Monthly # a UTC calendar month; Total counts from creation\n scope:\n project: research # leave out for the whole org\n selector:\n matchLabels: {team: nlp}\n action: Block # Notify only sends notices\n thresholds: [50, 80, 100]\n notify:\n emails: [research-leads@example.com]\n\nTerminal window\n\nnodus apply -f budget.yaml\nnodus get budget research-monthly\n\n nodus get budget shows what is spent, what is held for running work, what remains and the thresholds already notified this period. An org can have up to 50 Budgets.\n\nWhat a Budget covers\n\nA Budget counts every hold and charge of work in its project whose labels matched the selector when the work was created . Changing a running Job’s labels does not move it out of a Budget. A Budget you create while work is already running starts covering that work within 5 minutes.\n\nWhat happens at the limit\n\nWith action: Block :\n\n New work in scope is refused with 402 BudgetExceeded , naming the Budget, what is left and what the work needs:\n\n budget \"research-monthly\" has $0.40 left of $500.00 this month; job \"train-a\" needs $3.20.\n fix: raise spec.limitUSD on budget/research-monthly, or wait until 2026-11-01\n \n\n Running work in scope stops gracefully, as described in When money runs out, with Funded=False and reason BudgetExceeded .\n\n Raising spec.limitUSD , switching to Notify , deleting the Budget or the start of the next month resumes it.\n\nWith action: Notify , nothing is refused; you only get the notices.\n\nNotices\n\nEach threshold is sent once per period as a BudgetThreshold Event, a budget.threshold webhook and an email to the billing email, org owners and admins, and the addresses in spec.notify.emails . Reaching a Block limit also sends BudgetExceeded (webhook budget.exceeded ).\n\nCap a single run\n\nkind: Job\nspec:\n maxCostUSD: \"25.00\"\n\nA Job never holds more than its cap. When its spend reaches the cap, it checkpoints and becomes Suspended with reason MaxCostReached ; raising maxCostUSD resumes it from the checkpoint. You can raise a cap at any time, but not lower it.\n\nA Pipeline, Sweep or TrainingJob shares its cap with the Jobs it creates. Each child must fit both its own cap and its parent’s remaining cap, along with every matching Budget. A dry-run reports an insufficient spending limit in status.estimate.blockingReasons without reserving funds."},{"id":"docs/guides/billing/buy-credits","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/buy-credits/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/buy-credits.md","title":"Buy credits","description":"Add prepaid credits with a card through Stripe Checkout from the console, the CLI or the API, and add the $20 Indra plan.","stage":"GA","headings":[{"depth":2,"slug":"buy-from-the-console","text":"Buy from the console"},{"depth":2,"slug":"buy-from-the-cli","text":"Buy from the CLI"},{"depth":2,"slug":"buy-through-the-api","text":"Buy through the API"},{"depth":2,"slug":"limits","text":"Limits"},{"depth":2,"slug":"follow-a-top-up","text":"Follow a top-up"},{"depth":2,"slug":"when-the-credit-arrives","text":"When the credit arrives"},{"depth":2,"slug":"the-indra-plan","text":"The Indra plan"}],"text":"Buy credits\n\n Add prepaid credits with a card through Stripe Checkout from the console, the CLI or the API, and add the $20 Indra plan.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/buy-credits/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNodus is prepaid. You buy credits, and paid work draws on them through a funded hold placed before it starts. You pay on a Stripe-hosted page, so Nodus never sees your card number. Every purchase is a TopUp object that you can list and follow until the credit is on your balance.\n\nBuy from the console\n\nOpen Usage & billing , choose Add credits , and pick $10, $20, $50, $100 or a custom amount. The console sends you to Stripe Checkout and brings you back to the billing page, which shows the top-up until it is paid.\n\nBuy from the CLI\n\nTerminal window\n\nnodus billing top-up 20\n\nThe CLI opens the payment page in your browser, waits until the payment is credited and prints your new balance.\n\n Flag Does \n - - \n --save-card Saves the card for auto-recharge \n --no-open Prints the payment link instead of opening a browser, for a remote shell \n --no-wait Returns as soon as the payment link exists \n\nBuy through the API\n\nA top-up is an ordinary resource. Create it, then open its status.checkoutURL , which appears within about a second:\n\napiVersion: nodus.dev/v1\nkind: TopUp\nmetadata:\n name: october-credits\nspec:\n amountUSD: \"50.00\"\n savePaymentMethod: false # true also saves the card for auto-recharge\n\nTerminal window\n\nnodus apply -f topup.yaml\nnodus get topup october-credits -o yaml\n\nThe REST path is /apis/nodus.dev/v1/topups . By default Checkout returns to the console billing page; set spec.successURL and spec.cancelURL to return somewhere else on the console, or to an http://127.0.0.1 address for a local tool. A top-up cannot be edited or deleted; an unpaid one expires by itself.\n\nLimits\n\n Limit Value \n - - \n Amount of one top-up $5.00 to $1,000.00, in whole cents \n Top-ups per org per day $5,000 in total, per UTC day \n Payment page Payable for 1 hour after it is created \n\nAn amount outside the range is refused with 422 Invalid before anything is charged.\n\nFollow a top-up\n\nTerminal window\n\nnodus get topups -o wide\n\n Phase Reason Meaning \n - - - \n Queued Nodus is creating the payment page \n Running The payment page is open and waiting for you \n Running PaymentPending You paid with a method that confirms later, such as a bank debit \n Succeeded Paid and credited; status.creditedUSD is on your balance and status.receiptURL links the receipt \n Failed Expired Nobody paid within the hour; nothing was charged \n Failed PaymentFailed The payment did not complete; nothing was credited \n\nA declined card does not fail the top-up: the payment page stays open so you can try another card. When a payment fails after you submit it, such as a bank debit that does not clear, Nodus emails your billing email.\n\nWhen the credit arrives\n\nNodus credits each payment once, as soon as Stripe confirms it, which usually takes a few seconds. If a confirmation from Stripe is delayed or lost, Nodus asks Stripe about the payment itself within the hour and again in a nightly check, so a paid top-up is always credited, and never twice.\n\nIf your org owes arrears, the top-up pays them first and the rest goes to your balance. Work that stopped because money ran out resumes by itself; see When money runs out.\n\nStripe emails a receipt for every paid top-up to your billing email. See Receipts and invoices.\n\nThe Indra plan\n\nThe Indra plan costs $20 a month. Each paid month adds a $20 credit that only Indra ( nodus/indra ) requests can spend, at list price. Indra requests use that credit before your other credits, and everything else draws on your regular balance. Unused plan credit expires when the month it belongs to ends; it does not carry over.\n\nStart the plan by opening an Indra plan checkout session (type ComposerSubscription ) and paying on the page it returns:\n\nTerminal window\n\ncurl -X POST \"https://api.nodus-compute.ai/apis/nodus.dev/v1/billingaccounts/default/sessions\" \\\n -H \"Authorization: Bearer $NODUS TOKEN\" -H \"Idempotency-Key: $(uuidgen)\" \\\n -d '{\"type\": \"ComposerSubscription\"}'\n\nThe response is {url, expirationTime} . An org can have one Indra plan; a second request while the plan is active is refused with 409 Conflict .\n\n Renewal. Stripe charges the card each month and emails the invoice. A new month’s credit is added only once its invoice is paid; while a renewal payment is failing, Stripe retries it and Indra requests draw on your regular credits.\n Cancel. Open the billing portal with nodus billing portal and cancel the plan there. It ends at the end of the paid month: that month’s credit stays usable until then, and nothing further is charged or added."},{"id":"docs/guides/billing/credits","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/credits/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/credits.md","title":"Credits","description":"The starter credit, credit grants, the order credits are spent in, and when they expire.","stage":"GA","headings":[{"depth":2,"slug":"starter-credit","text":"Starter credit"},{"depth":2,"slug":"where-grants-come-from","text":"Where grants come from"},{"depth":2,"slug":"spending-order","text":"Spending order"},{"depth":2,"slug":"expiry","text":"Expiry"}],"text":"Credits\n\n The starter credit, credit grants, the order credits are spent in, and when they expire.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/credits/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nCredit grants are credit Nodus gives you, as opposed to credit you buy. Each grant is a CreditGrant you can read but not change:\n\nTerminal window\n\n$ nodus get creditgrants\nNAME SOURCE AMOUNT REMAINING EXPIRES\ngrt 01j9x... Starter $30.00 $12.40 2026-10-31\ngrt 01j9y... Promo $25.00 $25.00 2026-11-30\n\nThe console lists them under Usage & billing → Credits , with what each has spent and when it expires.\n\nStarter credit\n\nThe first org you create gets $30.00 of starter credit once your email address is verified. It expires 30 days after the org is created. Each person gets starter credit once, however many orgs they create.\n\nOrgs that only have starter credit keep tighter limits until their first top-up: 20 GiB of egress a day, $1 a day of console assistant use and no multi-node runs.\n\nWhere grants come from\n\n Source How you get it \n - - \n Starter Your first org, once your email is verified \n Promo Redeeming a promo code \n Admin Credit added by Nodus support, for example a pilot or a migrated balance \n Goodwill Credit from Nodus support for a problem on our side \n\nSome plan grants pay only for Indra ( nodus/indra ) model calls; they show nodus/auto , Indra’s other name, as their scope and are never spent on anything else.\n\nSpending order\n\nCharges draw from grants first, soonest expiry first, then from grants that never expire, and only then from purchased credit. You never have to pick which credit pays.\n\nExpiry\n\nYou get an email 7 days and 1 day before a grant expires. At its expiry the unspent remainder leaves your balance at once, even if a run is still using it; the run keeps going on its other credit. Grants are never refunded or turned into cash.\n\nA grant also settles unpaid charges first, the same way a top-up does."},{"id":"docs/guides/billing/partner-offers","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/partner-offers/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/partner-offers.md","title":"Partner offers","description":"Apply an eligible partner offer and understand promotional matching.","stage":"GA","headings":[],"text":"Partner offers\n\n Apply an eligible partner offer and understand promotional matching.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/partner-offers/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nOpen Billing → Partner offers , enter your code and optionally add your company or batch. An organization admin submits the request; a platform administrator verifies eligibility before credit is granted. Organization members can view progress.\n\nFor the YC offer, promotional credit reaches $100 total , including earlier grants even if spent or expired. A previous $30 starter grant therefore leaves $70 extra . You keep the existing 10 GB storage inclusion .\n\nAfter approval, earn 10% back on the first $10,000 of eligible usage , for a maximum $1,000 bonus . Only captured usage paid from purchased credit qualifies. Usage paid by starter credit, promotional grants or earned bonuses does not. Earned credit does not expire. Usage corrections adjust the matching benefit, including after an offer is revoked.\n\nYou can also use the CLI:\n\nTerminal window\n\nnodus billing deal\nnodus billing deal YC --note 'Company and batch'\n\nA request snapshots the terms you applied for. Later offer edits do not change them. Approval keeps the usual card prerequisite and usage prices. If an offer is revoked, future matching stops and existing earned credit remains, subject to corrections."},{"id":"docs/guides/billing/promo-codes","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/promo-codes/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/promo-codes.md","title":"Promo codes","description":"Redeem a promo code for credit, and what each redemption error means.","stage":"GA","headings":[{"depth":2,"slug":"errors","text":"Errors"},{"depth":2,"slug":"expiry","text":"Expiry"}],"text":"Promo codes\n\n Redeem a promo code for credit, and what each redemption error means.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/promo-codes/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA promo code adds a credit grant to your org. Redeem it from the CLI or from Usage & billing → Credits → Redeem code in the console:\n\nTerminal window\n\n$ nodus billing redeem LAUNCH25\nRedeemed CH25: $25.00 of credit, expires 2026-11-30.\n\nCodes are not case-sensitive, and spaces around them are ignored. Redeeming needs the billing:write permission. The response shows the code masked to its last four characters; Nodus stores only a keyed digest of each code.\n\nErrors\n\n Error Meaning \n - - \n 404 NotFound The code does not exist, or it has been withdrawn \n 410 Expired The code’s last redemption date has passed \n 409 Conflict Your org has already redeemed this code, or the code has been fully redeemed \n 429 TooManyRequests More than 5 attempts in a minute; wait a minute and try again \n\nRetrying a redemption with the same Idempotency-Key returns the first result and never adds credit twice. The CLI and console set the key for you.\n\nExpiry\n\nA promo grant expires on the date the code sets, counted from when you redeem it. Codes without one never expire."},{"id":"docs/guides/billing/receipts","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/receipts/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/receipts.md","title":"Receipts and invoices","description":"Find the receipt and invoice for every top-up and plan payment, change where receipts are sent, and manage your card and billing details.","stage":"GA","headings":[{"depth":2,"slug":"where-receipts-go","text":"Where receipts go"},{"depth":2,"slug":"find-a-receipt","text":"Find a receipt"},{"depth":2,"slug":"the-billing-portal","text":"The billing portal"},{"depth":2,"slug":"refunds","text":"Refunds"},{"depth":2,"slug":"questions-about-a-charge","text":"Questions about a charge"}],"text":"Receipts and invoices\n\n Find the receipt and invoice for every top-up and plan payment, change where receipts are sent, and manage your card and billing details.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/receipts/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nEvery payment you make to Nodus, whether a top-up you buy, an auto-recharge or an Indra plan month, has a Stripe-hosted invoice with a PDF, and Stripe emails a receipt for it. Nodus does not send postpaid bills: credits are prepaid, and what they paid for is in your usage records.\n\nWhere receipts go\n\nReceipts go to the billing email of your org’s BillingAccount , which defaults to the email of the owner who created the org. To send them somewhere else, such as a finance mailbox, change spec.billingEmail :\n\napiVersion: nodus.dev/v1\nkind: BillingAccount\nmetadata:\n name: default\nspec:\n billingEmail: finance@example.com\n\nTerminal window\n\nnodus apply -f billing.yaml\n\nThe new address applies to the next payment and is also where Nodus sends money notices, such as a failed auto-recharge, together with the org’s owners and admins.\n\nFind a receipt\n\nTerminal window\n\nnodus billing receipts\n\nThis lists your top-ups, Checkout and auto-recharge alike, with a link to each invoice page and its PDF. The same links are on each TopUp :\n\nTerminal window\n\nnodus get topup october-credits -o yaml\n\n Field Links to \n - - \n status.receiptURL The Stripe-hosted invoice page, which shows the payment and the card used \n status.invoicePDFURL The invoice as a PDF \n\nThe links appear a few seconds after the payment succeeds, once Stripe has finalized the invoice.\n\nThe billing portal\n\nTerminal window\n\nnodus billing portal\n\nThis opens the Stripe billing portal for your org, where you can:\n\n see and download every invoice and receipt, including Indra plan invoices;\n add, replace or remove the card used for auto-recharge and the Indra plan;\n change the name, address and other details printed on your invoices;\n cancel the Indra plan.\n\nThe portal link is valid for 5 minutes; run the command again for a new one. From the API, create a session with POST /apis/nodus.dev/v1/billingaccounts/default/sessions and {\"type\": \"Portal\"} .\n\nNodus does not charge sales tax or VAT on top-ups.\n\nRefunds\n\nRefunds cover unspent purchased credit and are made by Nodus support. A refunded top-up shows the amount in status.refundedUSD , the credit leaves your balance when the refund is approved, and the refund appears on the top-up’s invoice page. Credit from grants and promo codes is not refundable.\n\nQuestions about a charge\n\nContact Nodus support before disputing a charge with your bank. While a dispute is open, new work that needs credit is refused with 402 PaymentDisputed ; running work is not stopped. When the dispute closes in your favor the credit returns to your balance."},{"id":"docs/guides/billing/refunds","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/refunds/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/refunds.md","title":"Refunds","description":"How to get unspent purchased credit back to your card, and what can be refunded.","stage":"GA","headings":[{"depth":2,"slug":"ask-for-a-refund","text":"Ask for a refund"},{"depth":2,"slug":"what-can-be-refunded","text":"What can be refunded"},{"depth":2,"slug":"disputes","text":"Disputes"}],"text":"Refunds\n\n How to get unspent purchased credit back to your card, and what can be refunded.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/refunds/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nYou can get unspent purchased credit back to the card that paid for it. Credit grants, including starter credit and promo codes, are never refunded.\n\nAsk for a refund\n\nContact support with your org and the top-up you want refunded. Support checks the request and starts the refund; refunds above $500 are approved by a second person at Nodus. You do not need to stop your work first.\n\nThe amount leaves your available balance as soon as the refund is approved, so it cannot be spent while the card refund is processing. The top-up shows the refund under Usage & billing → Receipts & invoices , and nodus get transactions shows a RefundRequest and then a Refund . Your card is usually credited within 5 to 10 business days. If the card refund fails, a RefundFailed transaction returns the amount to your balance.\n\nWhat can be refunded\n\n At most what you paid on that top-up, less anything already refunded on it.\n At most your unspent purchased credit at approval: credit already used for runs is not refunded.\n One refund per top-up at a time, in whole cents.\n\nDisputes\n\nIf you dispute a payment with your bank instead, the disputed amount is removed from your balance and new work is paused with 402 PaymentDisputed until the dispute closes. Running work continues. Contacting support first is usually faster."},{"id":"docs/guides/billing/usage-and-costs","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/usage-and-costs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/usage-and-costs.md","title":"Usage and costs","description":"See what each run cost, split by project, label, meter and segment, and export usage as CSV.","stage":"GA","headings":[{"depth":2,"slug":"see-what-a-run-cost","text":"See what a run cost"},{"depth":2,"slug":"what-each-segment-covers","text":"What each segment covers"},{"depth":2,"slug":"list-and-group-usage","text":"List and group usage"},{"depth":2,"slug":"export-as-csv","text":"Export as CSV"},{"depth":2,"slug":"inference-and-indra","text":"Inference and Indra"},{"depth":2,"slug":"agents","text":"Agents"},{"depth":2,"slug":"storage-and-egress","text":"Storage and egress"},{"depth":2,"slug":"meters-reference","text":"Meters reference"}],"text":"Usage and costs\n\n See what each run cost, split by project, label, meter and segment, and export usage as CSV.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/usage-and-costs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nEvery charge Nodus makes is backed by usage records: one line per meter, per object, per window. This page shows how to read them, group them and export them.\n\nSee what a run cost\n\nEach compute object shows its cost so far in status.cost :\n\nTerminal window\n\nnodus get job train-llama -o jsonpath='{.status.cost}'\n\n{\"totalUSD\": \"4.182500\", \"heldUSD\": \"0.750000\", \"bySegment\": {\"bootUSD\": \"0.121000\", \"runningUSD\": \"3.980000\", \"teardownUSD\": \"0.081500\"}}\n\n heldUSD is reserved but not yet charged. The total becomes final once the capacity behind the run is confirmed deleted, which can be a minute or two after the run itself ends.\n\nWhen a run fails because of Nodus, the teardown after the failure and any of its start not yet billed are not charged: bySegment.coveredByNodus names those segments, and nodus describe shows them on the cost line, for example $0.00 (covered by Nodus: boot, teardown) .\n\nWhat each segment covers\n\nCompute usage from the moment the capacity starts billing until it is confirmed deleted is split into segments:\n\n Segment Covers \n - - \n Boot From the start of billing to the moment your command starts: boot, registration, readiness checks and pulling your image \n Running Your command running, including reading inputs, lazy image fetches, checkpoint writes and the snapshot when the run stops \n Restore For a run that resumes from a checkpoint or snapshot: from the start of billing to the moment your command starts again, including the restore. It replaces Boot \n Teardown From the stop to confirmed deletion, plus the capacity’s billing increment (lines with detail: IncrementRounding ) \n\nTime between runs that reuse the same capacity is billed on the warm idle seconds meter and has no segment. When Nodus itself delays a deletion, that time is not charged: it shows as a zero-amount line with detail: NodusBorne .\n\nList and group usage\n\nTerminal window\n\nnodus get usage --since 2d\nnodus get usage --group-by project,meter\nnodus get usage --group-by member --since 30d\nnodus get usage --group-by label:team,day\nnodus get usage --group-by rank,segment --field-selector object.uid= \n\n --group-by accepts project , kind , member , meter , day , segment , rank and label: . member is the person who started the work, whichever of their API keys they used. Labels are the ones the object carried when it was admitted, so label your work before you submit it:\n\nmetadata:\n name: train-llama\n labels:\n team: research\n\nGrouped totals always add up to the ungrouped total: every line lands in exactly one group, and lines without the grouped value share an empty group.\n\nExport as CSV\n\nTerminal window\n\nnodus get usage --since 30d -o csv usage.csv\ncurl -H \"Authorization: Bearer $NODUS TOKEN\" -H \"Accept: text/csv\" \\\n \"https://api.nodus-compute.ai/v1/usagerecords?since=30d\"\n\nAmounts are in US dollars with six decimals. quantity is in the meter’s unit (seconds, tokens, GiB or characters), and rate micros is the rate in micro-dollars per rate basis units.\n\nInference and Indra\n\nInference requests are rolled up into one line per model, meter and hour. A request to nodus/indra (Indra) shows the model that served it and, beside it, the routing call’s input and output tokens on lines whose SKU ends in :routing:input and :routing:output . The request is charged once, rounded up to the micro-dollar over all of its lines together.\n\nAgents\n\nAn Agent’s idle workers bill as warm idle on the Agent. When a run claims a worker, the worker’s time from that moment bills on the AgentRun until the run releases it, so each second of worker time appears once. Model calls a run makes with Nodus’s Claude are billed per request on the AgentRun, with the routing call beside the model that answered; calls made with your own Anthropic key carry no Nodus model charge.\n\nStorage and egress\n\nRetained storage (checkpoints, Volumes, images you build, and Job and Agent outputs) is sampled every hour per object. The first 10 GB across your organization are included and shared across your objects in proportion to their size; each object’s line shows only its billable part. The hours of a UTC day are charged together shortly after midnight UTC, as one Storage transaction.\n\nEgress from your containers and interactive sessions is counted per object per UTC day. The first 10 GiB a day across your organization are included, and the rest is charged the next morning as one Egress transaction, split across objects by bytes.\n\nIf your balance cannot cover a daily storage or egress charge, the remainder becomes arrears. While arrears are outstanding, new runs and uploads are refused with ArrearsOutstanding ; your next top-up pays them first.\n\nStorage left in arrears gets an email at 7, 21 and 28 days. At 30 days Nodus proposes deleting the stored data beyond the included 10 GB: finished Jobs’ checkpoints and outputs first, then older Volume revisions, then the latest revisions, oldest first. Nothing is deleted until two Nodus administrators approve the list, and paying your arrears before then cancels it. Deleting data does not clear the arrears.\n\nMeters reference\n\n quantity is in the meter’s unit. rate applies to rateBasis units of quantity, so a line’s amount is quantity × rate ÷ rateBasis , rounded down, except that the last line of a run rounds up to the capacity’s billing increment.\n\n Meter Billed for Unit Rate is per rateBasis SKU \n - - - - - - \n compute seconds Dedicated capacity for Jobs, GPU Sandboxes, GPU Workspaces, Functions and training members, with a segment seconds hour 3 600 rented: \n warm idle seconds Warm capacity kept between runs, idle Function and agent workers seconds hour 3 600 rented: \n node vcpu seconds vCPU on shared nodes (Sandboxes, agent runs), with a segment milli-vCPU seconds vCPU-hour 3 600 000 node:vcpu \n node gib seconds Memory on shared nodes, with a segment MiB seconds GiB-hour 3 686 400 node:memory \n node disk gib seconds Requested disk above 10 GiB per vCPU on shared nodes GiB seconds GiB-hour 3 600 node:disk \n build vcpu seconds Image builds, with a segment milli-vCPU seconds vCPU-hour 3 600 000 build:vcpu \n storage gb hours Retained storage above 10 GB per organization MB hours GB-month (30 days) 720 000 storage:retained \n egress gib Egress above 10 GiB per organization per UTC day; relayed training traffic from the first byte MiB GiB 1 024 egress:gib , egress:mesh-relay , egress:supplier \n inference input tokens Input tokens tokens million tokens 1 000 000 model: :input \n inference output tokens Output tokens, reasoning included tokens million tokens 1 000 000 model: :output \n inference cache read tokens Cached input read tokens million tokens 1 000 000 model: :cache read \n inference cache write tokens Cache writes, 5-minute or 1-hour tokens million tokens 1 000 000 model: :cache write 5m , model: :cache write 1h \n inference audio seconds Transcription and translation, 10 s minimum per request milliseconds audio-hour 3 600 000 model: :audio \n inference speech characters Speech synthesis input characters million characters 1 000 000 model: :speech \n platform device hours Your own devices assigned to Nodus-scheduled work device seconds device-hour 3 600 platform:device-hour \n platform predict seconds Predict on one of your pools seconds 30-day month 2 592 000 platform:predict \n\nIndra requests add a routing line beside the model’s lines, with SKU :routing:input or :routing:output . An egress:supplier line passes through an egress charge billed for your dedicated capacity at the same margin as the capacity itself."},{"id":"docs/guides/billing/what-you-pay-for","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/what-you-pay-for/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/what-you-pay-for.md","title":"What you pay for","description":"Every kind of time and cost, who pays for it, and the usage segment it appears under.","stage":"GA","headings":[{"depth":2,"slug":"the-table","text":"The table"},{"depth":2,"slug":"limits-stop-work-before-it-overruns","text":"Limits stop work before it overruns"}],"text":"What you pay for\n\n Every kind of time and cost, who pays for it, and the usage segment it appears under.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/what-you-pay-for/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nYou pay for every second a provider bills for a machine Nodus acquired for your work, from the moment billing starts until the machine is confirmed deleted. Nodus pays for everything it chose or caused. Your usage shows each machine’s time by segment, so you can see where every second went:\n\n Boot : from billing start to your command starting.\n Restore : a replacement machine’s time until your command starts again after a recovery, including the checkpoint restore.\n Running : from your command starting to its stop.\n Teardown : from stop to confirmed deletion, plus the provider’s billing increment, rounded once per machine.\n\nRun nodus get usage --group-by segment or open Usage & billing → Usage in the console to see the split. For a multi-node run, --group-by rank,segment itemizes each member.\n\nThe table\n\n What Segment Who pays \n - - - \n Time from billing start to your command starting: boot, registration, readiness checks, image pull and checkpoint restore Boot, Restore You \n A boot that fails on the provider’s side, until the machine is confirmed deleted Boot You \n Spare machines Nodus starts to finish your work sooner, machines Nodus releases or replaces on its own decision, machines that fail Nodus’s checks, and failures Nodus causes None: never on your usage Nodus \n From your command starting to its stop, including lazy image fetches, checkpoint writes, the snapshot on stop and work lost after a preemption Running You \n Downloading inputs after your command starts, and the initCommand preflight Running You \n Imports, exports and sink loads that run on your org’s capacity Running You \n Disk above 10 GiB per vCPU on Nodus nodes Running You \n A network partition your work tolerates, until it reconnects or until the old machine is fenced and confirmed deleted Running, Teardown You \n The provider’s billing increment, rounded once per machine Teardown You \n Warm idle time beyond an increment you already paid for, minimum workers, and idle Function and agent workers Warm idle You \n The idle share of Nodus-operated capacity, built into its rates and never metered on its own Included in rates You \n From stop to confirmed deletion Teardown You \n Time between stop and the delete request beyond 60 seconds when Nodus caused the delay None: never on your usage Nodus \n Any cost beyond your tightest limit: maxCostUSD , a Budget or your balance None: never on your usage Nodus \n A model request whose outcome Nodus cannot confirm, and upstream charges beyond the usage observed on a cut-off stream None: never on your usage Nodus \n Console assistant and command generation calls, within the daily cap None Nodus \n Grading Sandboxes, evaluations, dataset previews, task manifests and model calls made from inside your containers, under your caps Running You \n Mirroring a private image so it can run on hosted containers, and the mirror’s storage None Nodus \n Nodus’s own test runs None Nodus \n Image builds that fail Build You \n Storage above 10 GB per org: Volumes, checkpoints, outputs, images and session state Storage You \n Logs None Nodus \n Egress within 10 GiB per org per day None Nodus \n Egress beyond the included 10 GiB a day, up to your quota, and relayed traffic between multi-node members Egress You \n Your own pool hosts and cloud accounts None Free \n Pool devices running work Nodus schedules, Predict per pool, and capacity Nodus acquires for you Platform You \n Differences between a provider’s invoice and what Nodus metered None Nodus \n\nThis table is checked against the billing reference model on every change, so the rows here are the rows Nodus charges by.\n\nLimits stop work before it overruns\n\nEvery paid action holds funds before it starts, and work stops gracefully when a hold cannot be renewed. If a charge would still go past your tightest limit, Nodus absorbs the difference: your balance never pays more than the limit you set. See How billing works for holds, captures and the low-balance stop."},{"id":"docs/guides/billing/when-money-runs-out","url":"https://nodus-platform-site.pages.dev/docs/guides/billing/when-money-runs-out/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/billing/when-money-runs-out.md","title":"When money runs out","description":"How Nodus warns you, stops work gracefully and resumes it when credits, a raised Budget or a raised cap return.","stage":"GA","headings":[{"depth":2,"slug":"before-you-start-fail-fast","text":"Before you start: fail fast"},{"depth":2,"slug":"warnings","text":"Warnings"},{"depth":2,"slug":"how-each-kind-stops","text":"How each kind stops"},{"depth":2,"slug":"getting-going-again","text":"Getting going again"}],"text":"When money runs out\n\n How Nodus warns you, stops work gracefully and resumes it when credits, a raised Budget or a raised cap return.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/billing/when-money-runs-out/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nCredits are prepaid. Before paid work starts, Nodus places a funded hold on your balance, and it renews that hold every 5 minutes while the work runs. When a renewal cannot be funded, from your balance, a Budget or a maxCostUSD cap, the work stops gracefully inside time that is already paid for. Nothing runs past what you funded, and nothing is lost that the work saved.\n\nBefore you start: fail fast\n\nA create that cannot be funded is refused at once with the amounts and the fix:\n\n402 InsufficientCredits\njob \"train-a\" needs a $3.20 hold to start (released when it ends); available $1.10.\nfix: nodus billing top-up 20, or lower spec.maxCostUSD\n\nOther refusals are BudgetExceeded (a Budget or a cap has too little left), ArrearsOutstanding (unpaid charges) and PaymentDisputed . A distributed Job checks all of its nodes together: it starts with every node funded or not at all.\n\nWarnings\n\nWhen your available balance falls below $5, or below 20 % of what the next renewal of your running work needs, you get a LowBalance Event, a billing.low balance webhook, at most one email a day and a console banner. Running work whose funds will run out soon still shows Funded=True , and the condition’s message gives the time it will stop. A top-up, a credit grant or an auto-recharge before that time keeps it running.\n\nWhen a charge is larger than what is left, the rest becomes unpaid charges. You get an ArrearsPosted Event, a billing.arrears posted webhook and at most one email a day, and new work waits until a top-up or credit grant pays them.\n\nHow each kind stops\n\n Kind What happens State Resumes \n - - - - \n Job (checkpointed) Takes an urgent checkpoint, then stops Suspended , Funded=False Automatically once funded, from the checkpoint \n Job (restartable or ephemeral) Stops at the funded edge Suspended , Funded=False Automatically; ephemeral Jobs restart from the start \n Job at maxCostUSD Checkpoints, then stops Suspended , reason MaxCostReached When you raise maxCostUSD \n Distributed Job Every node checkpoints together, then all stop Suspended , Funded=False Automatically, with all nodes funded \n Sandbox Snapshots its filesystem and state, then stops Stopped , stopReason: InsufficientCredits ; spec.state stays Running On the next exec or connect once funded \n Workspace Saves the home volume, then stops Stopped with the money reason On nodus start , SSH connect or the schedule once funded \n AgentRun Finishes or interrupts the current step Waiting for funds Automatically \n Agent idle workers Scale to zero Funded=False on the Agent Automatically \n Function Stops taking calls, finishes in-flight calls, scales to zero New calls queue with Funded=False Automatically \n TrainingJob Its Jobs suspend; grading Sandboxes are deleted Suspended Automatically \n Image build The build stops Image Pending , Funded=False Automatically rebuilds \n Volume import or export The transfer stops ImportFailed(InsufficientCredits) nodus request reimport volume/ \n Inference New requests return an OpenAI- or Anthropic-shaped 402 ; in-flight requests finish — Immediately \n\nNodus never changes the spec.state you set. It records the stop in the object’s phase and Funded condition, emits FundingLost , and emits FundingRestored when it resumes. You get at most one email an hour listing stopped work.\n\nGetting going again\n\nTerminal window\n\nnodus billing top-up 20\nnodus get jobs --watch\n\nA top-up resumes waiting work oldest first, as each hold fits. Work stopped by a Budget resumes when you raise the Budget’s limit, switch it to Notify , delete it or the next month starts; work stopped at maxCostUSD resumes when you raise the cap."},{"id":"docs/guides/checkpoints","url":"https://nodus-platform-site.pages.dev/docs/guides/checkpoints/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/checkpoints.md","title":"Checkpoints","description":"Keep a Job's progress through suspends, stops and lost capacity by saving it to the state directory, and answer checkpoint requests so every save is consistent.","stage":"GA","headings":[{"depth":2,"slug":"save-your-progress-to-the-state-directory","text":"Save your progress to the state directory"},{"depth":2,"slug":"choose-what-is-saved","text":"Choose what is saved"},{"depth":2,"slug":"when-nodus-checkpoints","text":"When Nodus checkpoints"},{"depth":2,"slug":"answer-checkpoint-requests","text":"Answer checkpoint requests"},{"depth":2,"slug":"hugging-face-trainer","text":"Hugging Face Trainer"},{"depth":2,"slug":"gang-checkpoints-beta","text":"Gang checkpoints (Beta)"},{"depth":2,"slug":"run-the-example","text":"Run the example"}],"text":"Checkpoints\n\n Keep a Job's progress through suspends, stops and lost capacity by saving it to the state directory, and answer checkpoint requests so every save is consistent.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/checkpoints/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA checkpoint is a copy of your Job’s state directory that Nodus takes while the Job runs. When the Job is suspended, or the capacity under it is reclaimed, the next attempt starts with that directory restored and your program carries on from what it saved. Nodus decides when to checkpoint and where to store it; your program decides what to write and how to load it.\n\nCheckpoints hold files, not memory. A restored attempt starts your command again from the beginning, with the state directory as it was at the last checkpoint, so the program must read its own progress back.\n\nSave your progress to the state directory\n\nWrite everything you need to continue, such as model weights, optimizer state and the current step, under /nodus/state . The path is also in NODUS STATE DIR ( NODUS CHECKPOINT DIR is the older name for the same directory). On start, load what is there:\n\nexamples/checkpoints/resume/train.py\n\n\"\"\"A training loop that saves its progress where Nodus checkpoints it and resumes from it.\n\nIt uses only the standard library, so it runs in any image. With the Nodus SDK installed,\n nodus.state dir() , nodus.checkpoint.on request() and nodus.restored() do the same work.\n\"\"\"\n\nimport json\nimport os\nimport socket\nimport threading\nimport time\nfrom pathlib import Path\n\nSTATE DIR = Path(os.environ.get(\"NODUS STATE DIR\", \"/nodus/state\"))\nSTATE = STATE DIR / \"progress.json\"\nSTEPS = int(os.environ.get(\"STEPS\", \"120\"))\nSAVE EVERY = 20 # like Hugging Face Trainer's save steps: a regular save even when Nodus does not ask\n\npending = threading.Event() # set while Nodus waits for a consistent checkpoint\nrequest seq = None\nevents = None\n\ndef save(step):\n \"\"\"Write the state atomically, so a snapshot never sees a half-written file.\"\"\"\n STATE DIR.mkdir(parents=True, exist ok=True)\n tmp = STATE.with suffix(\".tmp\")\n tmp.write text(json.dumps({\"step\": step}))\n os.replace(tmp, STATE)\n\ndef listen():\n \"\"\"Subscribe to checkpoint requests on the events socket and flag each one for the loop.\"\"\"\n global events, request seq\n path = os.environ.get(\"NODUS EVENTS SOCKET\", \"/run/nodus/events.sock\")\n try:\n events = socket.socket(socket.AF UNIX, socket.SOCK STREAM)\n events.connect(path)\n except OSError:\n return # outside Nodus there is nobody to ask for checkpoints\n events.sendall(b'{\"type\":\"checkpoint.subscribe\"}\\n')\n for line in events.makefile(\"r\"):\n message = json.loads(line)\n if message.get(\"type\") == \"checkpoint.request\":\n request seq = message[\"seq\"]\n pending.set()\n\ndef ack():\n \"\"\"Tell Nodus the files in the state directory are complete.\"\"\"\n events.sendall((json.dumps({\"type\": \"checkpoint.ack\", \"seq\": request seq}) + \"\\n\").encode())\n pending.clear()\n\ndef main():\n start = 0\n if STATE.exists():\n start = json.loads(STATE.read text())[\"step\"]\n print(f\"resumed from step {start}\", flush=True)\n threading.Thread(target=listen, daemon=True).start()\n for step in range(start + 1, STEPS + 1):\n time.sleep(1) # one step of work\n print(f\"step {step}/{STEPS}\", flush=True)\n if pending.is set():\n save(step)\n ack()\n elif step % SAVE EVERY == 0:\n save(step)\n print(\"training complete\", flush=True)\n\nif name == \" main \":\n main()\n\nTwo habits keep checkpoints useful:\n\n Write atomically. Write to a temporary file and rename it over the old one, as save() does. A checkpoint can then never capture a half-written file.\n Keep outputs separate. Results you want to download go to /nodus/outputs . The state directory is recovery state: it is restored into the next attempt, not offered as a download.\n\n NODUS RESTORED=1 is set in an attempt that started from a checkpoint, if your program wants to log the difference. An empty state directory never replaces an earlier checkpoint that had files in it, so an attempt that fails before it saves anything cannot erase progress.\n\nChoose what is saved\n\n recovery.checkpoint in the Job spec controls the checkpoint. The defaults suit most programs:\n\nrecovery:\n continuity: Checkpointed # restore the latest checkpoint into each new attempt\n checkpoint:\n paths: [/nodus/state] # up to 64 absolute paths, all saved in one checkpoint\n interval: auto # or a fixed interval from 1m to 6h\n maxSize: 1Ti # larger checkpoints fail with CheckpointTooLarge\n retainAfterFinish: 168h # after the Job finishes, keep only the final checkpoint\n\nOn the command line, nodus run --checkpoint /nodus/state sets paths . Saving a whole folder such as your working directory is possible by listing it, but it makes every checkpoint larger and slower; list only what you need to continue.\n\n continuity: Restartable is the lighter choice for programs that track their position as a counter: Nodus keeps NODUS CURSOR COMPLETED and NODUS CURSOR TOTAL from your progress reports and skips the file copy. continuity: Ephemeral starts every attempt from scratch.\n\nWhen Nodus checkpoints\n\nWith interval: auto , Nodus sets the cadence from how often the capacity your Job runs on is interrupted and how long a checkpoint takes to save. A Job gets at least four checkpoints over its expected run time, and checkpointing takes no more than about a tenth of it. Capacity that is rarely interrupted is checkpointed less often.\n\nNodus also checkpoints, whatever the interval:\n\n when you run nodus suspend job/NAME , before the compute is released, so nodus resume continues from there;\n when the capacity gives notice that it is about to be reclaimed, so the replacement attempt loses as little work as possible.\n\n nodus describe job/NAME shows the latest checkpoint, and GET …/jobs/NAME/checkpoints lists each one with its sequence number, attempt, time, size and file count.\n\nAnswer checkpoint requests\n\nA checkpoint taken while your program is halfway through writing its files would restore a broken state. To avoid that, a program can ask to be told before each checkpoint and say when its files are complete. This is the request/ack handshake , and the example above implements it with the standard library.\n\nIt runs over the events socket at /run/nodus/events.sock ( NODUS EVENTS SOCKET ), one JSON object per line:\n\n1. Your program connects and sends {\"type\": \"checkpoint.subscribe\"} once.\n2. Before each checkpoint, Nodus sends {\"type\": \"checkpoint.request\", \"seq\": 7, \"urgent\": false} .\n3. Your program finishes the current step, writes its state and replies {\"type\": \"checkpoint.ack\", \"seq\": 7} .\n4. Nodus copies the state directory, then your program carries on. It does not need to pause while the copy runs.\n\n urgent: true means the capacity is about to go away: save at the next safe point and skip optional work. Nodus waits for the ack for as long as the shutdown allows and then takes the checkpoint anyway, so a program that hangs cannot block a suspend.\n\nYou choose whether to use the handshake with recovery.checkpoint.integration :\n\n Value Behaviour \n - - \n Auto (default) Use the handshake when the program subscribes; otherwise checkpoint without asking \n None Never ask; checkpoint the paths as they are \n HFTrainer Answer requests from inside Hugging Face Trainer, with no change to your image (below) \n\nWith the Python SDK installed ( pip install nodus-compute ), nodus.checkpoint.on request(save) registers a callback and sends the ack after it returns, nodus.checkpoint.requested() lets a loop poll instead, and nodus.state dir() returns the directory. All of them do nothing outside Nodus, so the same script runs on your laptop.\n\nHugging Face Trainer\n\nTrainer already saves and resumes; point it at the state directory and it works with Nodus checkpoints:\n\n set output dir to the state directory ( os.environ[\"NODUS STATE DIR\"] ), or a folder inside it;\n set save steps to how often Trainer saves on its own, and save total limit (for example 2 ) so old checkpoints do not fill the directory;\n call trainer.train(resume from checkpoint=True) when NODUS RESTORED is 1 , and trainer.train() on a first start, because Trainer refuses to resume from an empty directory.\n\nTo also save when Nodus asks, add nodus.checkpoint.HFTrainerCallback() to the Trainer’s callbacks, or set integration: HFTrainer and Nodus registers the same callback for you, even in an image without the SDK. Trainer then saves at the end of the current step and Nodus checkpoints once the save is written.\n\nGang checkpoints (Beta)\n\nBeta\n\nMulti-node Jobs ( distributed ) are Beta. Their checkpoints use a different format.\n\nA Job with distributed set defaults to recovery.checkpoint.format: Dcp . Every rank writes its shard with torch.distributed.checkpoint to $NODUS CHECKPOINT URI , and rank 0 commits it, which nodus.checkpoint.dcp.save and .load do for you. After a restart, $NODUS RESTORE URI names the latest committed checkpoint. A rank that lost its place in the gang cannot commit, so a restored gang always loads a checkpoint every rank finished.\n\nRun the example\n\nThe example suspends the Job mid-run and resumes it; the second attempt prints resumed from step N and finishes the remaining steps:\n\nTerminal window\n\n$ cd examples/checkpoints/resume\n$ nodus run --name checkpoints-resume --cpu 2 --checkpoint /nodus/state -d -- python train.py\n$ nodus suspend job/checkpoints-resume\n$ nodus resume job/checkpoints-resume\n$ nodus logs -f job/checkpoints-resume\nresumed from step 12\nstep 13/120\n…\n\nCheckpoints are deleted with their Job. After a Job finishes, Nodus keeps only its final checkpoint once retainAfterFinish (seven days by default) has passed."},{"id":"docs/guides/choose-compute","url":"https://nodus-platform-site.pages.dev/docs/guides/choose-compute/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/choose-compute.md","title":"Which Nodus compute feature should I use?","description":"Compare Jobs, Workspaces, Sandboxes and Functions. Choose the right way to run code, call models, run agents or train a model on Nodus.","stage":"GA","headings":[{"depth":2,"slug":"choose-by-the-work-you-need-to-do","text":"Choose by the work you need to do"},{"depth":2,"slug":"what-is-the-difference-between-a-job-and-a-function","text":"What is the difference between a Job and a Function?"},{"depth":2,"slug":"what-is-the-difference-between-a-workspace-and-a-sandbox","text":"What is the difference between a Workspace and a Sandbox?"},{"depth":2,"slug":"do-i-need-agents-to-use-claude-code-codex-or-cursor","text":"Do I need Agents to use Claude Code, Codex or Cursor?"},{"depth":2,"slug":"what-survives-an-interruption","text":"What survives an interruption?"},{"depth":2,"slug":"how-do-i-control-spending","text":"How do I control spending?"},{"depth":2,"slug":"where-should-an-agent-look-up-exact-syntax","text":"Where should an agent look up exact syntax?"}],"text":"Which Nodus compute feature should I use?\n\n Compare Jobs, Workspaces, Sandboxes and Functions. Choose the right way to run code, call models, run agents or train a model on Nodus.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/choose-compute/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nUse a Job when you have a command that should run to completion. Use a Workspace for interactive GPU or CPU development, a Sandbox for isolated CPU code execution, and a Function to call Python remotely or in parallel. Nodus also provides hosted inference, managed agents and training workflows.\n\nChoose by the work you need to do\n\n I need to… Start with Why it fits \n - - - \n Run a script, batch task or existing training command Jobs Runs a command to completion with logs, outputs and resource and cost limits \n Develop with SSH, VS Code or JupyterLab Workspaces An interactive GPU or CPU machine with a saved home directory \n Execute agent-generated or untrusted code Sandboxes An isolated CPU container driven by command and file requests, with idle stopping \n Call Python remotely or map it over many inputs Functions Remote calls with workers that scale with demand \n Call a hosted language model Inference (Beta) An OpenAI-compatible model API billed per token \n Run an agent conversation with shell tools Agents (Beta) An AgentRun with its own sandbox and recorded execution history \n Fine-tune or train through managed runtimes Training (Beta) A TrainingJob defines the model, data, runtime and training parameters \n\nFor a first run, follow your first job. It walks through signing in, running a command, following logs and downloading results.\n\nWhat is the difference between a Job and a Function?\n\nA Job runs an executable command and finishes when that command exits. A Function is a Python callable deployed in an App; your code invokes it with .remote() , .map() or .spawn() . Choose Jobs for an existing script or batch process. Choose Functions when remote calls should be part of your Python application.\n\nBoth can use GPU or CPU workers. Function workers can stay warm between calls; read Function billing before choosing idle and scaling settings.\n\nWhat is the difference between a Workspace and a Sandbox?\n\nA Workspace is for interactive development with SSH, VS Code or JupyterLab on GPU or CPU compute. Its home directory is backed by a Volume. A Sandbox is an isolated CPU container for programmatic commands and file operations, including code produced by an agent. Its network is closed unless you open it, and it can stop after an idle period.\n\nRead Workspace storage and Sandbox isolation before deciding what state and access your task needs.\n\nDo I need Agents to use Claude Code, Codex or Cursor?\n\nNo. Connect your existing coding agent to Nodus through MCP using the client setup guide. It can work with Nodus resources through that connection. The Agents feature is for running an agent conversation inside Nodus itself.\n\nWhat survives an interruption?\n\nRecovery depends on the feature and its configuration. For checkpointed Jobs, your program writes and loads its own files in NODUS CHECKPOINT DIR ; Nodus saves and restores that directory. This does not restore arbitrary process or GPU memory. Restartable Jobs start over, while ephemeral Jobs keep no state.\n\nRead Checkpoints for application recovery, Volumes for persistent files and Outputs for results you need to download.\n\nHow do I control spending?\n\nCheck the estimate before starting work, set the resource’s supported cost and time limits, and configure project budgets. Compute usage, saved storage and model tokens have different billing rules; the billing guide explains them. Use the pricing reference for published amounts instead of copying prices from an old example.\n\nWhere should an agent look up exact syntax?\n\nUse the CLI reference for commands, Python SDK reference for signatures and OpenAPI for v1 request and response fields. Beta resources use the v1beta1 contract.\n\nEvery authored guide has a Markdown version with its examples. The agent guide links a lightweight page manifest and focused topic bundles for retrieval."},{"id":"docs/guides/connections","url":"https://nodus-platform-site.pages.dev/docs/guides/connections/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/connections.md","title":"Connections","description":"Connect databases, S3 buckets, Weights & Biases and GitHub once, verified, and use them from Jobs, Sandboxes, imports and outputs.","stage":"GA","headings":[{"depth":2,"slug":"create-a-connection","text":"Create a Connection"},{"depth":2,"slug":"use-it","text":"Use it"},{"depth":2,"slug":"connect-github","text":"Connect GitHub"}],"text":"Connections\n\n Connect databases, S3 buckets, Weights & Biases and GitHub once, verified, and use them from Jobs, Sandboxes, imports and outputs.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/connections/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Connection is an external system that Nodus talks to for you: a Postgres database for query imports and outputs, an S3 bucket for inputs, a Weights & Biases project for live run tracking, or your GitHub repositories for private sources. Its credentials live in a Secret; the Connection says what they are for, and Nodus checks that they work before anything uses them.\n\nCreate a Connection\n\nPut the credentials in a Secret, then point the Connection at it:\n\nTerminal window\n\nnodus secret create analytics-db --from-literal DATABASE URL=postgres://reader:...@db.example.com:5432/app\n\nconnection.yaml\n\napiVersion: nodus.dev/v1\nkind: Connection\nmetadata:\n name: ex-connections-postgres\nspec:\n type: Postgres\n secret: ex-connections-postgres # holds DATABASE URL\n scope: Read # verified: the role can read, and is not asked to write\n\nTerminal window\n\nnodus apply -f connection.yaml\nnodus wait connection/ex-connections-postgres --for=condition=Verified\nnodus describe connection/ex-connections-postgres\n\n type Secret keys Settings Verified by \n - - - - \n Postgres DATABASE URL scope : Read , Write or ReadWrite Connecting, a query, and for write scopes the right to create tables \n Neon , Supabase DATABASE URL scope ; neon.branch The same checks, on the provider’s own host \n S3 ROLE ARN , or AWS ACCESS KEY ID and AWS SECRET ACCESS KEY s3.bucket , s3.region , s3.prefix , s3.endpoint Reaching the bucket with short-lived credentials \n WandB WANDB API KEY wandb.entity , wandb.project , wandb.live The key’s access to the entity \n\nPrefer an IAM role ( ROLE ARN ) for S3: Nodus assumes it for minutes at a time, so no long-lived key is stored. Every S3 Connection has its own external id in status.externalId , and Nodus sends exactly that id each time it assumes the role, so a role your trust policy grants to one Connection can’t be used through any other Connection or organization. Create the Connection, read the id, and require it in the role’s trust policy:\n\nTerminal window\n\nnodus get connection/datasets -o jsonpath='{.status.externalId}'\n\n{\n \"Effect\": \"Allow\",\n \"Principal\": { \"AWS\": \" \" },\n \"Action\": \"sts:AssumeRole\",\n \"Condition\": { \"StringEquals\": { \"sts:ExternalId\": \" \" } }\n}\n\nThe Connection shows Failed until the trust policy names the id; the next check, within the hour, turns it Ready . You don’t need EXTERNAL ID in the Secret. If you set it, it must equal status.externalId , and any other value is refused.\n\nA Connection is Ready once verified, and status.egressHosts lists the hosts it may reach. Nodus checks it again every day, and every hour after a failure. When a check fails the Connection shows Failed , the Verified condition says why, and anything that needs it fails with ConnectionNotReady until you fix the Secret.\n\nUse it\n\nspec:\n connections: [tracking] # injects WANDB and allows its hosts through the egress policy\n inputs:\n - name: shards\n bucket: {uri: s3://acme-data/shards/, connection: datasets}\n\nA WandB Connection with wandb.live: true sets WANDB API KEY , WANDB ENTITY , WANDB PROJECT , WANDB RUN GROUP , WANDB NAME and WANDB RUN ID for the Job, so wandb.init() needs no arguments. Each index of the Job logs to one run, and a run that Nodus restarts after lost capacity resumes it rather than starting a new one. A Job with one index shows the run’s page in status.links . A variable you set yourself keeps your value; setting WANDB RUN ID , WANDB ENTITY or WANDB PROJECT yourself means Nodus doesn’t link the run. A Job that names a Connection that is missing or not Ready runs without it and gets a ConnectionNotReady Event.\n\nConnect GitHub\n\nA GitHub Connection installs the Nodus GitHub App on your account or organization. It needs no Secret:\n\nTerminal window\n\nnodus create connection github --type GitHub\nnodus get conn github -o jsonpath='{.status.installURL}'\n\nOpen the link, choose the repositories to share and install the App. GitHub returns you to Nodus, which checks that you can see the installation, and the Connection becomes Ready with status.github naming the account and repositories. GET …/connections/github/repositories?limit=100 pages through all of them.\n\nA Sandbox’s init.git can then clone a private repository: Nodus resolves the branch or tag to a commit when you create the Sandbox and clones that commit. The clone uses a short-lived token that can only read that one repository, and the token never appears in the Sandbox’s environment or files. If someone uninstalls the App, the Connection shows Failed with InstallationRemoved and a fresh status.installURL to install it again.\n\nA Volume can import from a bucket ( source.s3 ) or a query ( source.connectionQuery ); see Volumes. A Job or Sandbox reaches only the hosts its Connections were verified against."},{"id":"docs/guides/console","url":"https://nodus-platform-site.pages.dev/docs/guides/console/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/console.md","title":"Use the console","description":"Sign in, find your work, follow it live and copy the matching CLI command from any page.","stage":"GA","headings":[{"depth":2,"slug":"sign-in","text":"Sign in"},{"depth":2,"slug":"sign-in-the-cli-from-the-browser","text":"Sign in the CLI from the browser"},{"depth":2,"slug":"find-your-work","text":"Find your work"},{"depth":2,"slug":"follow-runs-live","text":"Follow runs live"},{"depth":2,"slug":"copy-the-command","text":"Copy the command"},{"depth":2,"slug":"create-something","text":"Create something"},{"depth":2,"slug":"when-credits-run-out","text":"When credits run out"},{"depth":2,"slug":"appearance","text":"Appearance"}],"text":"Use the console\n\n Sign in, find your work, follow it live and copy the matching CLI command from any page.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/console/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nThe console at console.nodus-compute.ai shows the same objects the CLI and the Python SDK work with. Every page has the command that does the same thing, so you can move between them at any time.\n\nSign in\n\n1. Open the console and sign in with your email and password, or with GitHub or Google.\n2. New accounts confirm their email with a 6-digit code.\n3. If someone invited you, the console shows the invite first. Accept it to join their organization.\n4. Otherwise create your organization. Its URL name (for example research ) appears in every link.\n\nYour first organization receives starter credit once your email is verified.\n\nSign in the CLI from the browser\n\nRun nodus login . The CLI opens the console, which asks which organizations the CLI may use and names the key it creates, for example cli-laptop-2026-09-30 . Choose Authorize and return to your terminal. On a machine without a browser, run nodus login --device and enter the code it prints on the console’s device page.\n\nIf you have no organization yet, authorizing creates one for you. If someone invited you, both pages show the invite first: accept it, and the CLI signs in to their organization.\n\nFind your work\n\n The URL always names the organization and project: /research/default/jobs/lora-fine-tune/logs . Links you share open in the same place for anyone with access, and two browser tabs can work in two organizations.\n Switch organization or project from the top bar.\n Press Cmd+K (or Ctrl+K ) to search objects, pages and actions. Type job/ or sb/ and a name to jump straight to an object.\n Press ? for every keyboard shortcut. g then a letter goes to a section, for example g b for Billing.\n\nFollow runs live\n\nLists and detail pages update as things change, without reloading. The note Live · 12s ago shows the time since the last update. Choose it and Pause live updates to freeze what you are reading; Resume applies the changes at once. Running work shows its cost so far, marked with ≈ between updates.\n\nCopy the command\n\nThe terminal icon in each page header copies the command for what you are looking at, with the project and organization spelled out:\n\nTerminal window\n\nnodus describe job/lora-fine-tune -p default --org research\n\n View YAML in the actions menu shows the live object as a manifest you can download and apply elsewhere.\n\nCreate something\n\nCreate pages show four tabs: Form , YAML , CLI and Python . They always show the same manifest. Before you launch, the estimate shows the expected cost range, the hourly rate, the first hold and how long the price holds. Launch uses exactly that estimate; if prices move first, the console estimates again and asks you to confirm. Your form is saved as a draft in this browser until you launch or discard it.\n\nIf the network drops while you launch, the console says it could not confirm the result. Check status looks for the object by name, and Retry is safe: it can never create a second copy.\n\nWhen credits run out\n\nIf your balance cannot fund the next renewal, running work stops at its funded edge. The project Overview then shows Paused for funds , grouped by what adding credits does:\n\n Resumes automatically : jobs and agent runs continue by themselves once a hold fits.\n Needs a start : workspaces stay stopped until you choose Start (or Start all ).\n Wakes on next use : sandboxes start again on the next command, file request or preview.\n Rejecting requests : inference endpoints answer with 402 until funded.\n Failed imports : choose Retry import on the volume.\n Needs a higher limit : work that reached its own spend limit resumes only when you raise that limit.\n\nChoose Add credits on the panel to top up.\n\nAppearance\n\nChoose System , Light or Dark , and Comfortable or Compact density, from the account menu. The theme follows you across the console and the docs."},{"id":"docs/guides/environments","url":"https://nodus-platform-site.pages.dev/docs/guides/environments/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/environments.md","title":"Environments","description":"Task sets with graders for reinforcement learning and evaluation, from the Nodus catalog or your own image.","stage":"GA","headings":[{"depth":2,"slug":"browse-the-catalog","text":"Browse the catalog"},{"depth":2,"slug":"use-one-in-a-trainingjob","text":"Use one in a TrainingJob"},{"depth":2,"slug":"how-grading-works","text":"How grading works"},{"depth":2,"slug":"train-on-a-reward-function-in-one-call","text":"Train on a reward function in one call"},{"depth":2,"slug":"bring-your-own-environment","text":"Bring your own environment"},{"depth":2,"slug":"run-an-environment-from-the-environments-hub","text":"Run an environment from the Environments Hub"},{"depth":2,"slug":"train-on-a-multi-turn-environment","text":"Train on a multi-turn environment"},{"depth":2,"slug":"train-on-a-tool-calling-environment","text":"Train on a tool-calling environment"},{"depth":2,"slug":"run-any-reasoning-gym-family","text":"Run any Reasoning Gym family"},{"depth":2,"slug":"run-a-verl-or-skyrl-dataset","text":"Run a verl or SkyRL dataset"}],"text":"Environments\n\n Task sets with graders for reinforcement learning and evaluation, from the Nodus catalog or your own image.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/environments/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nAn Environment is a versioned set of tasks with a grader: a prompt for each task, and a program that decides whether a completion is correct. TrainingJobs use Environments for reinforcement learning and evaluation, and agent evaluations use the same ones. Environments are generally available ( nodus.dev/v1 ).\n\nBrowse the catalog\n\nTerminal window\n\n$ nodus get environments -n nodus\nNAME VERSION CATEGORY MODES PHASE\ngraph-coloring 1.0.0 Reasoning Train, Evaluate Ready\narithmetic 2.0.0 Math Train, Evaluate Ready\ngsm8k 1.0.0 Math Train, Evaluate Ready\nreasoning-gym 1.0.0 Reasoning Train, Evaluate Ready\npython-functions 1.0.0 Code Train, Evaluate Ready\n$ nodus describe environment/graph-coloring -n nodus\n\n describe shows the summary, licenses, the size of each split, a sample task, the graders and the examples: ready TrainingJob templates with measured results. The console’s Training page lists the same catalog, and “Run example” starts a TrainingJob from one; in Python, nodus.TrainingJob.from example(\"nodus/graph-coloring@1.0.0\") does the same.\n\nUse one in a TrainingJob\n\nName the environment and version, and how many tasks to train and evaluate on:\n\nspec:\n runtime: nodus/grpo-lora\n environment:\n name: nodus/graph-coloring@1.0.0\n trainTasks: 50 # from the train split\n heldOutTasks: 64 # from the test split, never trained on\n seed: 42\n\nThe same name, version and seed always give the same tasks in the same order, so two runs are comparable. The train and test splits never share a task.\n\nHow grading works\n\n Tasks carry a prompt and non-secret metadata only. Expected answers stay with the grader; the trainer, the model and your code never see them.\n Completions are graded by Nodus, not by the trainer. Program graders run in grading Sandboxes that Nodus creates for your TrainingJob or agent evaluation: no network access, deleted after five minutes idle, at most grading.maxParallel at once. Code from a completion runs as a separate user that cannot read the answers. ExactMatch graders compare the completion with the answer inside Nodus, without a Sandbox.\n Every verdict is Correct , Incorrect , InvalidOutput or InfrastructureFailure , with a reward and evidence recorded by the host that ran the grader: the process, its exit code, its duration and a hash of its output.\n InvalidOutput is a completion the grader cannot parse, such as an answer without the expected tag. It is a failed task with a reward of 0.\n InfrastructureFailure means the grader could not run. It carries no reward and does not count toward a pass rate, so an outage never looks like a wrong answer.\n Grading Sandboxes are billed to the TrainingJob and count toward its maxCostUSD . They are deleted when the run finishes or is suspended.\n\nTrain on a reward function in one call\n\nWhen your tasks and reward are in Python, rl.train is all you need. To see it work first, run the example that ships with the SDK:\n\nTerminal window\n\npython -m nodus.examples.rl\n\nIt is one file, nodus/examples/rl.py , and it is the template for your own task:\n\nimport re\nfrom nodus.recipes import rl\n\nANSWER = re.compile(r\" \\s ([A-Za-z]+)\\s \")\ntasks = [(f\"Spell the word backwards inside .\\n\\nWord: {w}\", w[::-1]) for w in WORDS]\n\ndef reward(completion, answer):\n found = ANSWER.findall(completion)\n return None if not found else float(found[-1].lower() == answer)\n\nrun = rl.train(tasks, reward, max cost=2)\nrun.watch() # each stage, then the reward, loss and KL of every step\nrun.outputs.download(\"./outputs\") # the LoRA adapter and the before-and-after comparison\n\n run.tasks(phase=\"Evaluation\", outcome=\"Failed\") lists the graded tasks the run has scored so far, each with its taskId , reward and outcome, so you can read which tasks the trained model still gets wrong.\n\n Tasks are (prompt, answer) pairs or {\"prompt\", \"answer\", \"metadata\"} dicts. Only reward sees the answer.\n Model. The base model is Qwen/Qwen3-0.6B unless you pass model= , for example model=\"Qwen/Qwen3-4B\" .\n Held-out tasks. A fifth of the tasks, and at least 16, is held out. The model is graded on them before and after training. Pass test tasks= to choose them yourself.\n What gets packaged. The reward goes with the lines of its own file that it uses: imports, constants and helper functions. A name it cannot take along fails before anything is built.\n Checked before it runs. rl.train grades one task’s answer and an empty reply on your machine first. A reward that raises, or returns something other than a number, a bool or None , fails there instead of on a GPU. Return None or 0 when a reply holds no answer.\n Reuse. The same code and tasks reuse the same Environment, so a second run starts at once.\n No network. The reward runs in Nodus’s grader, which has no network access. Pass pip=[\"package==1.2.3\"] for packages it imports.\n How it learns. Each step samples group size=8 replies to each of groups per step=4 tasks and scores every reply against the others in its group. Replies are capped at max tokens=256 , and the learning rate is learning rate=4e-5 . You can change any of them.\n Format rewards. To reward the shape of an answer as well as its value, fold both into the reward, for example return float(correct) - 0.1 (not well formed) .\n Other options. steps , gpu , max cost , lora and the other grpo lora parameters are optional. The run appears under Runs › Post-training in the console, with its curves, results and cost.\n\nTo train on a catalog Environment, pass its name instead of tasks and a reward:\n\nrun = rl.train(\"nodus/gsm8k@1.0.0\", max cost=5)\n\nBring your own environment\n\nYour own tasks and reward are a Python module with two functions. tasks(split) lists the prompts of the train or test split with their answers, and reward(completion, answer) scores one completion:\n\ndef tasks(split):\n return [{\"prompt\": \"Spell the word backwards...\\n\\nWord: valley\", \"answer\": \"yellav\"}, ...]\n\ndef reward(completion, answer):\n found = ANSWER.findall(completion)\n if not found:\n return None # no answer in the completion: InvalidOutput\n return 1.0 if found[-1].lower() == answer else 0.0\n\nA reward of 1 (or True ) is Correct , anything lower Incorrect , and None InvalidOutput . If reward raises, the verdict is InfrastructureFailure , never a reward of 0. Answers reach only reward ; the trainer and the model see the prompt and the optional metadata . A test prompt never appears in train, even if your lists repeat it.\n\nThe quickest way in needs no Docker: publish the module as a pip package whose load environment() returns an object (or the module itself) with tasks and reward , and name the exact version. Nodus builds it into an image when you apply the Environment; if the package does not load, the build fails and the Image env- - shows the step and its log:\n\napiVersion: nodus.dev/v1\nkind: Environment\nmetadata:\n name: reverse-words\nspec:\n version: 1.0.0\n package:\n pip: {name: reverse-words, version: 1.0.0} # index: for a private index\n loader: load environment # the default; or module:function\n category: Custom\n modes: [Train, Evaluate]\n rewardType: Binary\n\nRun an environment from the Environments Hub\n\nAn environment published on Prime Intellect’s Environments Hub is a pip package, so it runs on Nodus as it is. Name it with the Hub’s index for its owner and the version the Hub lists:\n\napiVersion: nodus.dev/v1\nkind: Environment\nmetadata:\n name: reverse-text\nspec:\n version: 0.1.4\n package:\n pip:\n name: reverse-text\n version: 0.1.4\n index: https://hub.primeintellect.ai/primeintellect/simple/\n category: Custom\n modes: [Train, Evaluate]\n\nIts load environment() returns a verifiers environment, and Nodus trains on it directly: the dataset is the train split, the eval dataset the held-out one (with only one of them, every fifth task is held out), and each completion is scored by the environment’s own rubric, with verifiers 0.1 through 0.3. A multi-turn environment, such as a game like wordle , runs too: see Train on a multi-turn environment, and so does one where the model calls tools: see Train on a tool-calling environment. An environment whose tools run in a remote sandbox, a browser or an MCP server fails the build with that reason, as does one scored by another model (a JudgeRubric ), since grading has no network and holds no credential for that model. The console marks each kind in its Hub search from the Hub’s own tags. The package comes from the Hub index alone and its dependencies may also come from PyPI.\n\nGrading has no network. Nodus fetches the datasets an environment loads while it builds the image, and every task list and grade after that reads that copy, so a run always sees the data its build saw.\n\nTrain on a multi-turn environment\n\nIn a multi-turn environment the model and the environment take turns: the model replies, the environment answers with feedback, and the episode goes on until the environment ends it. Set maxTurns on the TrainingJob ( Turns per episode under Advanced in the console) to let an episode run that many model turns:\n\nspec:\n runtime: nodus/grpo-lora\n environment: {name: wordle@0.1.7, trainTasks: 512, heldOutTasks: 64, seed: 42}\n parameters:\n maxTurns: 6 # model turns per episode\n maxCompletionLength: 256 # tokens per reply\n maxEpisodeTokens: 2048 # tokens of the whole episode after the prompt, the model's and the environment's\n\nThe reward is the one the environment gives the whole episode, and training learns only from the model’s own tokens. An episode that reaches either cap ends there and is graded as it stands. With the default maxTurns: 1 , only the first reply is graded. Baseline and final evaluation play the same episodes greedily.\n\nAn OpenEnv environment runs in process from a package whose load environment() returns it. Each task is a seed: Nodus resets the environment with it, and every action is the model’s reply as the one text field of the environment’s action class. The environment must play the same episode for the same seed. Seeds 0 to 799 train and 800 to 999 are held out, unless the environment sets train seeds and test seeds :\n\nimport nltk\nfrom textarena env.server.environment import TextArenaEnvironment\n\ndef load environment():\n try: # the build downloads NLTK's word lists; grading reads them offline\n nltk.data.find(\"corpora/words\")\n nltk.data.find(\"taggers/averaged perceptron tagger eng\")\n cached = True\n except LookupError:\n cached = False\n return TextArenaEnvironment(\"Wordle-v0\", download nltk=not cached)\n\nA module of your own can also be multi-turn: give it step(turns, answer) in place of reward . It gets every model turn so far and returns the environment’s next message, or None once the episode ended, and the reward so far. Grading keeps no state between turns, so step replays the turns from the start.\n\nTrain on a tool-calling environment\n\nIn a verifiers ToolEnv the model calls the environment’s tools and reads their results. Each task’s prompt offers the tools through the model’s chat template, the model calls one by writing a {\"name\": \"...\", \"arguments\": {...}} block (the form Qwen and most open chat templates teach), and the environment runs the call in its own code. The results are the next turn, and the episode ends when the model answers without calling a tool. Set maxTurns to the rounds of calls an episode may make:\n\nspec:\n runtime: nodus/grpo-lora\n environment: {name: tool-test@0.1.1, trainTasks: 21, heldOutTasks: 43, seed: 42}\n parameters:\n maxTurns: 3\n maxCompletionLength: 256\n\nThe tools run inside the grader, which has no network, so a tool that searches the web or runs code in a remote sandbox cannot train. Pick a model whose chat template supports tools.\n\nRun any Reasoning Gym family\n\nThe catalog’s nodus/reasoning-gym serves five reviewed families. To train on any other Reasoning Gym family, or a mix of them, publish a package whose load environment() returns the dataset; its own score answer scores each completion:\n\nimport reasoning gym\n\ndef load environment():\n return reasoning gym.create dataset(\"knights knaves\", size=2000, seed=7)\n\nEvery fifth entry is held out. The model’s last … is its answer, or the whole completion without one; partial credit is the reward and only a full score is Correct . Each family’s own licence applies.\n\nRun a verl or SkyRL dataset\n\nA dataset prepared for verl or SkyRL runs with its reward unchanged. Publish a module that names the parquet files with verl’s own keys, and your verl reward function under its usual name:\n\ntrain files = \"hf://datasets/BytedTsinghua-SIA/DAPO-Math-17k/data/dapo-math-17k.parquet\"\nval files = \"hf://datasets/BytedTsinghua-SIA/AIME-2024/data/aime-2024.parquet\" # optional\n\ndef compute score(data source, solution str, ground truth, extra info=None):\n ... # verl's custom reward signature; a dict's \"score\" also works\n\nEach row is verl’s: prompt (a system message and one user message at most), data source , reward model.ground truth and extra info . Without compute score , each row’s env class names the SkyRL-gym environment that scores it, built from the row as SkyRL builds it; add skyrl-gym to the package’s dependencies. Without val files , every fifth row is held out. The package needs datasets as a dependency, and a SkyRL environment of more than one turn fails the build with that reason.\n\nTo ship your own system packages or files, build the module into an image on the env-base image instead, push it and name its digest:\n\nFROM ghcr.io/nodus-compute/catalog/env-base:1.0.0\nCOPY reverse words.py /opt/environment/\nENV NODUS ENVIRONMENT LOADER=reverse words NODUS ENVIRONMENT=reverse-words NODUS ENVIRONMENT VERSION=1.0.0\nRUN nodus-env info\nENV HF HUB OFFLINE=1 HF DATASETS OFFLINE=1\n\napiVersion: nodus.dev/v1\nkind: Environment\nmetadata:\n name: reverse-words\nspec:\n version: 1.0.0\n package: {image: registry.example.com/acme/reverse-words@sha256:...}\n category: Custom\n modes: [Train, Evaluate]\n rewardType: Binary\n\nTerminal window\n\n$ nodus apply -f environment.yaml\n$ nodus get environment/reverse-words -w # Ready once the image digest is verified\n\nA TrainingJob names it without the nodus/ prefix ( environment: {name: reverse-words@1.0.0} ), and in Python rl.grpo lora(environment=\"reverse-words@1.0.0\", ...) . The whole example, with a GRPO TrainingJob, is in examples/training/custom-reward . Nodus pulls the image from a public registry or from your organization’s space in the Nodus registry; other private registries are not supported for Environments yet.\n\nAny image works if it provides the two commands Nodus runs, as uid 10001 with no network:\n\n nodus-env tasks --split train test --seed N writes one JSON line per task: {\"taskId\", \"prompt\", \"metadata\"} .\n nodus-env grade reads {\"taskId\", \"completion\"} lines and writes one {\"taskId\", \"verdict\", \"reward\", \"evidence\"} line for each, in order.\n\nThe Environment becomes Ready once its image is verified, with the declared split sizes in status.splits . The first TrainingJob that uses a split and seed runs nodus-env tasks in one of its own grading Sandboxes and Nodus keeps that manifest for every later run of your organization, so the tasks never change between runs. A version’s image and graders cannot change: publish a new version instead, so earlier results stay reproducible."},{"id":"docs/guides/functions","url":"https://nodus-platform-site.pages.dev/docs/guides/functions/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/functions.md","title":"Functions","description":"Run Python functions on Nodus workers, deploy them as an App that stays up, and look them up from anywhere.","stage":"GA","headings":[{"depth":2,"slug":"run-an-app","text":"Run an App"},{"depth":2,"slug":"deploy-an-app","text":"Deploy an App"},{"depth":2,"slug":"what-nodus-creates","text":"What Nodus creates"},{"depth":2,"slug":"stop-restart-and-delete","text":"Stop, restart and delete"}],"text":"Functions\n\n Run Python functions on Nodus workers, deploy them as an App that stays up, and look them up from anywhere.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/functions/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Function is a Python function that runs on Nodus workers. You decorate it, call it from your own code with .remote() , .map() or .spawn() , and Nodus starts workers when calls arrive, keeps them warm for a while and scales them back down. The pages of this guide cover calling Functions, scaling them and what is billed while a worker waits.\n\nRun an App\n\nAn App is the group of Functions in one file. nodus run creates an ephemeral App, runs your main on your machine and deletes the App when main returns. Every .remote() call runs on a Nodus worker.\n\nexamples/functions/map/app.py\n\n\"\"\"One Function called three ways, then mapped over a thousand inputs.\n\nRun it with nodus run examples/functions/map/app.py .\n\"\"\"\n\nimport nodus\n\napp = nodus.App(\"fn-map\")\n\n@app.function(cpu=1, memory=\"1Gi\", max workers=4, target concurrency=8, max cost=1)\ndef square(x: int) - int:\n return x x\n\n@app.function(cpu=1, memory=\"1Gi\", max cost=1)\ndef divide(a: int, b: int) - float:\n return a / b\n\n@app.local entrypoint()\ndef main(n: int = 1000) - None:\n print(\"remote:\", square.remote(7)) # one call; blocks for the result\n call = square.spawn(8) # starts a call and returns a handle\n print(\"spawned:\", call.get(timeout=600))\n results = list(square.map(range(n))) # one call per input, results in input order\n print(\"map:\", len(results), \"ordered:\", results == [i i for i in range(n)])\n try:\n divide.remote(1, 0)\n except ZeroDivisionError as exc: # the remote exception comes back as its own type\n print(\"raised:\", type(exc). name )\n\nTerminal window\n\n$ nodus run examples/functions/map/app.py\nremote: 49\nspawned: 64\nmap: 1000 ordered: True\nraised: ZeroDivisionError\n\nThe directory of the file (minus what .gitignore and .nodusignore exclude) is uploaded once per content hash, so the workers import the same code you ran. The App renews itself while main runs and is deleted shortly after main stops, even when your machine loses its connection.\n\nDeploy an App\n\n nodus deploy keeps the App. Its Functions stay available after your program exits, and any program can call them.\n\nexamples/functions/warm-pool/app.py\n\n\"\"\"A Function that keeps one worker warm, so a call never waits for a start.\n\nDeploy it with nodus deploy examples/functions/warm-pool/app.py . The idle worker is billed at the Function's\nworker rate for as long as it stays warm; set min workers=0 to pay only while calls run.\n\"\"\"\n\nimport nodus\n\napp = nodus.App(\"fn-warm\")\n\n@app.function(cpu=1, memory=\"1Gi\", min workers=1, max workers=3, scaledown window=\"2m\", max cost=1)\ndef ping() - str:\n return \"pong\"\n\nTerminal window\n\n$ nodus deploy examples/functions/warm-pool/app.py\n\nimport nodus\n\nping = nodus.Function.from name(\"fn-warm\", \"ping\")\nprint(ping.remote()) # pong\n\nA deploy updates the App in place. A Function whose code, image or resources changed rolls its workers once, after their in-flight calls finish. A Function you removed from the file is deleted, and the calls it still had queued end as Failed with the reason FunctionDeleted , so no caller waits for a call nobody will run.\n\nWhat Nodus creates\n\n Object What it is Look at it with \n - - - \n App The Functions of one file; ephemeral for run , persistent for deploy nodus get apps \n Function One decorated function or class, with its image, resources and scaling nodus get functions \n FunctionCall One invocation, kept for seven days after it ends nodus get functioncalls -l nodus.dev/function=fn-warm-ping \n\nA Function reports its state in status.phase , how many workers it has in status.workers , how many calls wait in status.queue , and the expected start time in status.estimate .\n\nTerminal window\n\n$ nodus get function fn-warm-ping\nNAME PHASE WORKERS QUEUED COST AGE\nfn-warm-ping Running 1/3 0 $0.02 3m\n\nStop, restart and delete\n\nTerminal window\n\n$ nodus stop function/fn-warm-ping # drain the workers; new calls wait in the queue\n$ nodus start function/fn-warm-ping # start workers again for the calls that waited\n$ nodus restart function/fn-warm-ping # replace every worker once, after in-flight calls finish\n$ nodus delete app/fn-warm\n\nA stopped Function keeps accepting calls and holds them in the queue, so work you submit while it is stopped runs once it starts. Workers also stop by themselves when your balance cannot cover another renewal; calls queue until you add credit and then run.\n\nNote\n\nA call that runs for longer than its timeout (five minutes by default, 24 hours at most) ends as Failed with the reason DeadlineExceeded , and its worker restarts, because the thread that ran it cannot be interrupted."},{"id":"docs/guides/functions/billing","url":"https://nodus-platform-site.pages.dev/docs/guides/functions/billing/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/functions/billing.md","title":"What is billed while warm","description":"How Function workers are billed when they run calls, when they wait, and when they scale to zero.","stage":"GA","headings":[{"depth":2,"slug":"when-a-worker-is-billed","text":"When a worker is billed"},{"depth":2,"slug":"reading-the-bill","text":"Reading the bill"},{"depth":2,"slug":"spending-limits","text":"Spending limits"}],"text":"What is billed while warm\n\n How Function workers are billed when they run calls, when they wait, and when they scale to zero.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/functions/billing/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nYou pay for worker time, by the second, at the worker’s rate. A call does not add a charge of its own. A Function shows the worker’s rate before you run anything, in status.estimate.rateUSDPerHour and in f.estimate(...) .\n\nWhen a worker is billed\n\n State Billed? \n - - \n A worker starting: placing, pulling its image, running enter hooks Yes, from the moment it is placed. \n A worker running at least one call Yes. \n An idle worker inside min workers Yes, always. This is warm capacity. \n An idle worker above min workers , inside scaledown window Yes, until the window ends. \n A worker that was released No. \n A Function with no workers No. \n\nSo a Function with min workers=0 costs nothing while nothing is calling it, and each burst costs its workers’ time plus their scaledown window . A Function with min workers=1 costs one worker’s rate every hour, and its calls never wait for a start.\n\nReading the bill\n\nEvery worker is billed under one Function and shows up as two lines of usage: the time the worker spent running calls and the time it spent waiting ( warm idle seconds ). Both appear under the Function in your usage records, so the cost of keeping workers warm is never mixed into the calls themselves.\n\nEach call’s status.costUSD is the worker time attributed to that call: the worker’s rate divided by target concurrency , for the whole seconds the call spent on a worker, at least one. A worker that runs four calls at once shows each of them a quarter of the rate. Call costs show where the busy time went. The bill is the workers’.\n\nSpending limits\n\nA Function does not take a spend cap yet: setting max cost is refused when you deploy, so no cap can be set and silently ignored. Your credit balance is the limit on what its workers spend, and an idle min workers worker bills until you stop the Function.\n\nWhen your balance cannot cover another renewal, the workers drain inside the reserve and the Function scales to zero. New calls queue and show Funded=False on the Function. Add credit and the Function starts workers again for the calls that waited. Nothing the calls were doing is lost: a call whose worker drained finishes first, and one that could not finish goes back to the queue."},{"id":"docs/guides/functions/calls","url":"https://nodus-platform-site.pages.dev/docs/guides/functions/calls/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/functions/calls.md","title":"Call a Function","description":"Run one call, start calls without waiting, or map a Function over thousands of inputs, and read what comes back.","stage":"GA","headings":[{"depth":2,"slug":"map-over-many-inputs","text":"Map over many inputs"},{"depth":2,"slug":"arguments-and-results","text":"Arguments and results"},{"depth":2,"slug":"errors-retries-and-timeouts","text":"Errors, retries and timeouts"},{"depth":2,"slug":"look-a-call-up-later","text":"Look a call up later"}],"text":"Call a Function\n\n Run one call, start calls without waiting, or map a Function over thousands of inputs, and read what comes back.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/functions/calls/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Function has three ways to start a call. Each call is a FunctionCall object, so you can look it up later, wait for it from another process or cancel it.\n\n Call What it does \n - - \n f.remote(x) Runs one call and blocks until it returns its result. \n f.spawn(x) Starts one call and returns a handle: handle.get(timeout=...) waits, handle.cancel() stops it. \n f.map(xs) Starts one call per input and yields the results in input order. \n\nwith app.run():\n print(square.remote(7)) # 49\n handle = square.spawn(8)\n print(handle.get(timeout=600)) # 64\n print(list(square.map(range(1000)))) # 1,000 results, in order\n\nMap over many inputs\n\n .map() creates its calls in batches of up to 1,000 per request, and a batch is all or nothing: a request that fails leaves no calls behind, and sending it again does not create them twice. Results come back in input order whichever worker finishes first. order outputs=False yields them as they finish, and return exceptions=True yields a failed input’s exception instead of stopping the loop.\n\nresults = list(process.map(files, order outputs=False, return exceptions=True))\n\nSeveral callers can map into one Function at the same time. Workers take calls from each caller in turn, so a caller that submits 10 inputs is not stuck behind another caller’s 10,000.\n\nArguments and results\n\nArguments and results are serialized with cloudpickle. Values up to 64 KiB travel inside the call. Larger ones are uploaded once, by content hash, and the call carries a reference, so a large argument costs nothing extra to send again. The caller and the worker image must run the same Python minor version, 3.10 to 3.13.\n\nErrors, retries and timeouts\n\nAn exception raised inside the Function is raised again in your process as its own type when that type can be imported there. Otherwise you get nodus.errors.RemoteError . Either way the remote traceback is attached.\n\ntry:\n divide.remote(1, 0)\nexcept ZeroDivisionError:\n ...\n\nTwo things can send a call back to the queue, and they are counted separately:\n\n The function raised. With retries=nodus.Retries(max retries=3) the call is retried with exponential backoff, up to ten times. Without retries the exception ends the call.\n The worker was lost. A worker that disappears or is replaced does not count as a retry. Its calls go back to the queue and run on another worker, up to recovery.maxAttempts times (eight by default), then end as Failed with the reason RecoveryLimitExceeded .\n\nA call can finish only once. A result that arrives from a worker that no longer holds the call is refused, so a call that was re-dispatched never ends with two results.\n\nLook a call up later\n\ncall = nodus.FunctionCall.from name(\"fn-warm-ping-bcdfghjklm\")\nprint(call.get(timeout=60))\n\nTerminal window\n\n$ nodus get functioncalls -l nodus.dev/map= \n$ nodus get functioncall -o yaml # phase, result reference, retries, recoveries, costUSD\n\nA call keeps its result for seven days after it ends. Each call reports status.costUSD , the worker time attributed to it; the billing page explains how it relates to what you pay."},{"id":"docs/guides/functions/scaling","url":"https://nodus-platform-site.pages.dev/docs/guides/functions/scaling/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/functions/scaling.md","title":"Scaling and warm workers","description":"Set how many workers a Function keeps, how long an idle one stays, and how many calls each runs at once.","stage":"GA","headings":[{"depth":2,"slug":"cold-and-warm-starts","text":"Cold and warm starts"},{"depth":2,"slug":"classes-set-up-once-per-worker","text":"Classes: set up once per worker"},{"depth":2,"slug":"changing-a-function","text":"Changing a Function"}],"text":"Scaling and warm workers\n\n Set how many workers a Function keeps, how long an idle one stays, and how many calls each runs at once.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/functions/scaling/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Function’s scaling bounds its worker pool. Nodus sizes the pool from the calls that are waiting and running.\n\n Setting Default What it does \n - - - \n min workers 0 Workers kept warm even when idle. They are billed. \n max workers 10 The most workers the Function ever has, 1 to 1,000. \n scaledown window 1m How long an idle worker above min workers stays before it is released, up to 20 minutes. \n target concurrency 1 Calls one worker runs at once, 1 to 1,000. Use more than 1 for I/O-bound functions. \n\n@app.function(cpu=2, memory=\"4Gi\", min workers=1, max workers=20, scaledown window=\"5m\", target concurrency=4)\ndef embed(text: str) - list[float]:\n ...\n\nA pool with target concurrency=4 holds one worker for every four calls that are waiting or running, and never more than max workers . A worker leaves only after it has been idle for the whole scaledown window , so a burst that comes back inside the window finds its workers still there.\n\nCold and warm starts\n\nA call that lands on a warm worker with a free slot starts at once. A call that needs a new worker waits for the worker to place, pull its image and start. Nodus shows both before you run anything:\n\nprint(embed.estimate(\"hello\")) # expected cost, cold and warm start times, and the hold\n\nTerminal window\n\n$ nodus get function embed -o jsonpath={.status.estimate}\n\n status.estimate.startup holds the cold and warm bands, and status.estimate.rateUSDPerHour the worker’s rate. While workers are starting, the Function’s Ready condition says WorkersStarting . With no workers and nothing queued it says ScaledToZero , and with an image that is still building it says ImagePending .\n\nClasses: set up once per worker\n\nA class keeps its state for the life of a worker. @nodus.enter() methods run once when the worker starts, before its first call. @nodus.exit() methods run once when the worker drains. Methods marked @nodus.method() get .remote() , .map() and .spawn() .\n\nexamples/python/classes/app.py\n\n\"\"\"A class whose model loads once per worker, then serves many calls.\n\nRun it with nodus run examples/python/classes/app.py .\n\"\"\"\n\nimport nodus\n\napp = nodus.App(\"classes\")\n\n@app.cls(cpu=2, memory=\"4Gi\", scaledown window=\"5m\", max cost=1)\nclass Greeter:\n @nodus.enter()\n def load(self) - None:\n # Runs once when a worker starts, before its first call: load weights or open connections here.\n self.greeting = \"hello\"\n\n @nodus.method()\n def greet(self, name: str) - str:\n return f\"{self.greeting}, {name}\"\n\n @nodus.exit()\n def close(self) - None:\n self.greeting = \"\"\n\n@app.local entrypoint()\ndef main() - None:\n greeter = Greeter()\n print(greeter.greet.remote(\"Ada\"))\n print(list(greeter.greet.map([\"Grace\", \"Linus\"])))\n\nLoad models and open connections in enter , so a call pays for them once per worker instead of once per call. A failing exit hook is logged and does not stop the others.\n\nChanging a Function\n\nEditing scaling takes effect on the next reconcile without restarting any worker. Changing the code, image, Python version or resources rolls the workers once: new workers start, and the old ones finish their calls and leave. nodus restart rolls them without a change.\n\nA Function whose image is still building keeps the workers it already has and keeps serving calls until the new image is ready."},{"id":"docs/guides/gpus","url":"https://nodus-platform-site.pages.dev/docs/guides/gpus/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/gpus.md","title":"Choose GPUs and check availability","description":"See which accelerators are available right now, what they cost, and how to ask for exactly the hardware your run needs.","stage":"GA","headings":[{"depth":2,"slug":"see-what-is-available","text":"See what is available"},{"depth":2,"slug":"read-an-offering","text":"Read an offering"},{"depth":2,"slug":"region-classes","text":"Region classes"},{"depth":2,"slug":"ask-for-the-hardware-you-need","text":"Ask for the hardware you need"},{"depth":2,"slug":"use-every-gpu-of-one-machine","text":"Use every GPU of one machine"},{"depth":2,"slug":"check-before-you-launch","text":"Check before you launch"},{"depth":2,"slug":"when-a-run-waits-in-queued","text":"When a run waits in Queued"}],"text":"Choose GPUs and check availability\n\n See which accelerators are available right now, what they cost, and how to ask for exactly the hardware your run needs.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/gpus/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nYou describe the hardware a run needs; Nodus finds capacity that fits and finishes it for the lowest expected cost. This guide shows how to see what is available, read an offering, and write a GPU request that says exactly what you mean.\n\nSee what is available\n\nTerminal window\n\n$ nodus get gpus\nNAME TYPE COUNT VCPU MEMORY REGION PRICE/H AVAILABILITY AGE\na100-sxm-80g-x1-any A100-80G 1 30 200Gi unknown $1.27 Available 14s\nh100-sxm-80g-x8-us H100-SXM 8 208 1800Gi us $18.29 Available 14s\nl4-24g-x1-eu-int L4 1 8 32Gi eu $0.31 Limited 14s\n\n PRICE/H is the from rate per machine-hour and AGE is how long ago the price and availability were confirmed.\n\nFilter with flags; each one narrows the list:\n\nTerminal window\n\n$ nodus get gpus --gpu H100 --count 8 --region us\n$ nodus get gpus --interruptible\n$ nodus get gpus --gpu L4 -o wide # typical and list rates, startup time and interruption rate\n\nThe console’s GPU picker shows the same list.\n\nRead an offering\n\nAn offering is a class of capacity: accelerator × count × region class, and whether it is interruptible. Its name spells that out:\n\n Name Meaning \n - - \n h100-sxm-80g-x8-us Eight H100 SXM 80 GB GPUs per machine, in the us region class \n l4-24g-x1-eu-int One L4 per machine in eu , interruptible \n a100-sxm-80g-x1-any Capacity with no guaranteed location ( any ) \n cpu-8c-32g-us An 8 vCPU, 32 GiB CPU machine in us \n\nEach offering shows:\n\n From and typical rates per hour. The from rate is that of the cheapest machine free now, the one a run starts on by default, or of the cheapest machine when none is free; the host shape shown is that machine’s. Your rate is shown in the estimate before launch and frozen for the run. No rate is ever above the list price.\n Availability : Available , Limited (a few machines, or a count that is not reported) or Unavailable .\n Startup : how long a machine usually takes to be ready (p50 and p90), from what Nodus has measured.\n Interruption rate for interruptible offerings: how often such capacity is taken back per hour.\n\nPrices and availability refresh continuously. The public pricing page shows the same from rates.\n\nRegion classes\n\n placement.regions takes region classes, not data centers: us , ca , eu , uk , in , apac , me and latam . Capacity whose location is not reported has the class any in its name and never satisfies a non-empty placement.regions , so a run that must stay in a region never lands there.\n\nspec:\n placement:\n regions: [eu, uk]\n\nAsk for the hardware you need\n\n resources.gpu.type lists every accelerator you accept. Order does not matter: the scheduler picks among them by expected cost to finish.\n\n You write Nodus may use \n - - \n H100 Any H100 variant: SXM, PCIe or NVL \n [H100, H200] Any H100 or H200, never another family \n H100-SXM Only the H100 SXM \n H100! or exact: true Only the family’s primary variant (H100 SXM) \n minMemory: 80Gi with no type Any accelerator with at least 80 GiB per GPU \n\nModal spellings work as written: A100-40GB , A100-80GB , A10G , L40S , H100 , H200 , B200 , T4 , L4 .\n\nspec:\n resources:\n gpu: {type: [H100, H200], count: 8}\n placement:\n interruptible: Allow # use interruptible capacity only when it is cheaper to finish\n maxRateUSDPerHour: \"30.00\"\n\n count is per machine (1, 2, 4 or 8). For more GPUs than one machine holds, see multi-node training (Beta).\n\nUse every GPU of one machine\n\n --gpu A6000:2 (or count: 2 ) gives your container both GPUs of one machine, numbered 0 and 1. torchrun starts one process per GPU without flags, because Nodus sets PET NPROC PER NODE to the GPU count; a --nproc-per-node you pass yourself wins.\n\ntrain.py\n\n\"\"\"A DDP smoke run on every GPU of one machine. torchrun starts one process per GPU (PET NPROC PER NODE).\"\"\"\nimport os\n\nimport torch\nimport torch.distributed as dist\nfrom torch.nn.parallel import DistributedDataParallel as DDP\n\ndist.init process group(\"nccl\")\nlocal rank = int(os.environ[\"LOCAL RANK\"])\ndevice = torch.device(\"cuda\", local rank)\ntorch.cuda.set device(device)\n\nEvery process contributes 1, so the sum is the world size.\none = torch.ones(1, device=device)\ndist.all reduce(one)\n\nmodel = DDP(torch.nn.Linear(16, 1).to(device), device ids=[local rank])\nopt = torch.optim.SGD(model.parameters(), lr=0.1)\nfor step in range(20):\n x = torch.randn(32, 16, device=device)\n loss = (model(x) - x.sum(dim=1, keepdim=True)).pow(2).mean()\n opt.zero grad()\n loss.backward()\n opt.step()\n\nif dist.get rank() == 0:\n print(f\"world size={int(one.item())} gpus={torch.cuda.device count()} \"\n f\"device={torch.cuda.get device name(device)} loss={loss.item():.4f}\")\ndist.destroy process group()\n\nrun.sh\n\nnodus run --name torchrun-multi-gpu --gpu A6000:2 --max-cost 1.00 -- torchrun train.py\n\nThe run prints world size=2 gpus=2 from the first process. The processes reach each other over NCCL on the machine itself; /dev/shm is sized for it (half the container’s memory), so you need no --shm-size or --ipc=host .\n\nCheck before you launch\n\nA dry run returns the estimate without starting anything:\n\nTerminal window\n\n$ nodus apply -f job.yaml --dry-run=server -o estimate\n\nIt shows the offering Nodus would use, the expected cost p50 and p90, the startup time, the first hold and the minimum charge, and how long the estimate stays valid. When nothing fits, it says why, for example CapacityUnavailable , NoListPrice , ExceedsRemainingBudget or MissesDeadline , with a fix.\n\nWhen a run waits in Queued\n\nA run with no matching capacity stays Queued and is placed as soon as capacity appears, until placement.queueTimeout (24 hours by default). nodus describe shows the offerings that were considered and why each was rejected. To start sooner, widen resources.gpu.type or placement.regions , allow interruptible capacity, or raise placement.maxRateUSDPerHour .\n\nHow Nodus chooses among offerings is explained in placement and scheduling profiles."},{"id":"docs/guides/images","url":"https://nodus-platform-site.pages.dev/docs/guides/images/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/images.md","title":"Images","description":"Run on the catalog images, build your own from a few steps or a Dockerfile, and use private registries.","stage":"GA","headings":[{"depth":2,"slug":"catalog-images","text":"Catalog images"},{"depth":2,"slug":"build-availability","text":"Build availability"},{"depth":2,"slug":"build-an-image","text":"Build an Image"},{"depth":2,"slug":"private-registries","text":"Private registries"},{"depth":2,"slug":"pull-a-built-image-yourself","text":"Pull a built Image yourself"},{"depth":2,"slug":"status-and-errors","text":"Status and errors"}],"text":"Images\n\n Run on the catalog images, build your own from a few steps or a Dockerfile, and use private registries.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/images/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nEvery Job, Sandbox, Function and Agent runs in a container image. You can use a catalog image, any public or private registry image, or an Image that Nodus builds for you. When you submit work, Nodus resolves the image’s tag to a digest and records it, so a retry or a recovery runs exactly the same bytes.\n\nCatalog images\n\n Image Contents Default for \n - - - \n nodus/python:3.12 (also 3.10 , 3.11 , 3.13 ) Debian slim, Python, uv, git Jobs without a GPU \n nodus/pytorch:2.8-cuda12.8 CUDA 12.8 runtime, Python 3.12, PyTorch 2.8 Jobs with a GPU \n nodus/agent-tools Python 3.12, Node 22, git, ripgrep, the Nodus SDK Sandboxes and Agents \n\nTerminal window\n\nnodus get images -n nodus\nnodus run --gpu L4 --image nodus/pytorch -- python -c \"import torch; print(torch.cuda.get device name())\"\n\nCatalog images run as uid 1000 and include /bin/sh .\n\nBuild availability\n\nCurrent hosted deployments support catalog images and prebuilt public or private registry images. Image builds from steps, Dockerfiles and Sandbox snapshots are unavailable. Those requests return ImageBuildUnavailable with HTTP 503 before accepting a build or charging for one. Publish your image to a registry, then use its reference directly or create an Image with only base and no steps. Environment packages that need a pip build must instead name a prebuilt package.image .\n\nExisting unstarted build requests report Failed with an explanation instead of waiting indefinitely. The build examples below describe the contract for a deployment with a configured build runner and registry.\n\nBuild an Image\n\nAn Image starts from a base and adds steps : aptInstall , pipInstall , uvPipInstall , uvSync , run , copy , env and workdir .\n\nimage.yaml\n\napiVersion: nodus.dev/v1\nkind: Image\nmetadata:\n name: ex-images-build\nspec:\n base: nodus/python:3.12\n steps:\n - aptInstall: [jq]\n - pipInstall: {packages: [\"requests==2.32.5\"]}\n - env: {APP MODE: example}\n\nTerminal window\n\nnodus apply -f image.yaml\nnodus wait image/ex-images-build --for=condition=Ready\n\nThe build log streams from GET …/images/{name}/log?follow=true ; the Python SDK prints it while it builds.\n\nRun on it with imageRef :\n\njob.yaml\n\napiVersion: nodus.dev/v1\nkind: Job\nmetadata:\n name: ex-images-build\nspec:\n imageRef: {name: ex-images-build} # runs the digest the Image built, pinned at admission\n command: [python, -c, \"import os, requests; print('requests', requests. version , os.environ['APP MODE'])\"]\n\nIn Python:\n\nimage = (nodus.Image.from registry(\"nodus/pytorch:2.8-cuda12.8\")\n .apt install(\"git\").pip install(\"transformers==4.57.6\", \"peft\")\n .env({\"HF HUB ENABLE HF TRANSFER\": \"1\"}))\n\nBuilds run on Nodus and are billed per vCPU-second at the CPU rate. An Image whose spec matches one your org has already built, on the same base digest, reuses that digest without building ( status.build.cached: true ). An Image with only a base is ready at once: it is that image’s digest.\n\nYou can also build from a Dockerfile, inline or uploaded with its context:\n\nspec:\n dockerfile:\n inline: \n FROM nodus/python:3.12\n RUN pip install polars==1.33.1\n\nBuild-time credentials go in buildSecrets : they are mounted for the build only and never stored in a layer.\n\nPrivate registries\n\nCreate a Registry Secret (Secrets) and name it under imagePullSecrets , on the Job or Sandbox for image , or on the Image for a private base .\n\nWhen work runs on a provider that pulls containers itself, Nodus copies the image by digest into your org’s space in the Nodus registry first, so your registry credential never reaches the provider. Such providers need /bin/sh in the image; an image without it gets the ProviderContainerNeedsShell warning and runs elsewhere.\n\nPull a built Image yourself\n\nA built Image lives in the Nodus registry under your org; status.reference holds its full reference. Log in with any username and an API key as the password, then pull it:\n\nTerminal window\n\nREF=$(nodus get image ex-images-build -o jsonpath='{.status.reference}')\necho \"$NODUS API KEY\" docker login \"${REF%%/ }\" -u nodus --password-stdin\ndocker pull \"$REF\"\n\nViewers, and keys without the images:write scope, can pull but not push.\n\nStatus and errors\n\n nodus get image ex-images-build shows PHASE , DIGEST and SIZE . A failed build has a reason:\n\n Reason or error Meaning \n - - \n BaseNotFound The base does not exist, or its registry refused the pull Secret \n StepFailed A step exited non-zero; the build log shows which \n BuildTimeout The build ran past its time limit \n ImageNotFound , ImagePullFailed A Job’s image could not be resolved when you submitted it \n ImageNotReady A Job or Sandbox names an Image that has not finished building"},{"id":"docs/guides/inference","url":"https://nodus-platform-site.pages.dev/docs/guides/inference/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/inference.md","title":"Inference","description":"Call hosted models through an OpenAI-compatible API, route with Indra, and pay per token.","stage":"Beta","headings":[{"depth":2,"slug":"sign-in-and-create-a-key","text":"Sign in and create a key"},{"depth":2,"slug":"call-a-model","text":"Call a model"},{"depth":2,"slug":"from-the-cli","text":"From the CLI"},{"depth":2,"slug":"models","text":"Models"},{"depth":2,"slug":"responses-messages-and-embeddings","text":"Responses, Messages and embeddings"},{"depth":2,"slug":"audio","text":"Audio"},{"depth":2,"slug":"indra-nodusindra","text":"Indra: nodus/indra"},{"depth":2,"slug":"named-endpoints-and-limits","text":"Named endpoints and limits"},{"depth":2,"slug":"billing-per-token","text":"Billing per token"},{"depth":2,"slug":"errors","text":"Errors"}],"text":"Inference\n\n Call hosted models through an OpenAI-compatible API, route with Indra, and pay per token.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/inference/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n Beta: this feature may change.\n\nNodus inference is an OpenAI-compatible API over hosted open models. Point any OpenAI SDK at the Nodus base URL, use a Nodus API key, and pay per token from your prepaid credits.\n\nSign in and create a key\n\nSign in to the console, open Settings → API keys , and create a key with the inference:invoke scope.\n\nCall a model\n\nfrom openai import OpenAI\n\nclient = OpenAI(base url=\"https://inference.nodus-compute.ai/v1\", api key=\"YOUR NODUS API KEY\")\nreply = client.chat.completions.create(\n model=\"nodus/gpt-oss-120b\",\n messages=[{\"role\": \"user\", \"content\": \"What is the capital of France?\"}],\n)\nprint(reply.choices[0].message.content)\n\nTerminal window\n\ncurl https://inference.nodus-compute.ai/v1/chat/completions \\\n -H \"Authorization: Bearer $NODUS API KEY\" -H \"Content-Type: application/json\" \\\n -d '{\"model\": \"nodus/indra\", \"stream\": true, \"messages\": [{\"role\": \"user\", \"content\": \"Hello\"}]}'\n\nStreaming works as in the OpenAI API ( stream: true ). Completed chat streams end with data: [DONE] ; an interrupted or errored stream does not receive a generated completion marker. Every response carries a Nodus-Request-Id header.\n\nFrom the CLI\n\nThe CLI calls the same API with the credential of your current context:\n\nTerminal window\n\n$ nodus inference models\nMODEL NAME CONTEXT MAX OUTPUT PREVIEW\nnodus/indra Indra - - false\nopenai/gpt-oss-120b GPT-OSS 120B 128K 64K false\nopenai/gpt-oss-20b GPT-OSS 20B 128K 64K false\n$ nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 \"What is the capital of France?\"\nThe capital of France is Paris.\nrequest ireq 01k6d2x7q9fvjt3y8m0c4r5n2e · openai/gpt-oss-20b · 84 input + 21 output tokens · $0.000015\n$ nodus inference receipt ireq 01k6d2x7q9fvjt3y8m0c4r5n2e\n\n chat prints the answer on standard output and the request id, the model that answered, the tokens and the charge on standard error; -o json prints the whole completion. receipt shows any request’s charge for 30 days.\n\nModels\n\n GET /v1/models lists the models you can call now. Each model has a catalog name such as openai/gpt-oss-120b and a nodus/ alias such as nodus/gpt-oss-120b . Models marked preview have provisional prices.\n\nIn the console, open Inference → Models and select Indra to send a request to nodus/indra , or choose a model from the catalog. The response summary’s Model field shows the model identifier submitted with that request, including nodus/indra when using Indra. The API’s routing metadata and receipt still record which underlying model served the request.\n\nResponses, Messages and embeddings\n\nChat models also accept stateless Responses and Anthropic Messages requests, with the same model ids, credits, endpoint limits and receipt headers. Responses supports text, inline images and client function tools; stored responses, background work and hosted tools are refused. Text Completions accepts one string prompt and one choice, without echo, suffix, logprobs or best-of sampling. Messages token counting is not available.\n\nresponse = client.responses.create(model=\"nodus/gpt-oss-20b\", input=\"Hello\", max output tokens=64, store=False)\nprint(response.output text)\nvectors = client.embeddings.create(model=\"nodus/bge-m3\", input=[\"first document\", \"second document\"])\nprint(vectors.data[0].embedding)\n\nEmbeddings are available when the embedding model appears in GET /v1/models . They charge only input tokens. A model’s operations field identifies the routes it accepts, including Embeddings , and aliases lists its alternative model ids. The console labels embedding models as text → vectors and provides embedding examples in the model panel and endpoint API tab. A request can carry up to 2,048 inputs, each within the model’s context window. Messages accepts the same chat models at /v1/messages , with max tokens , messages and optional stream: true ; it does not expose private models used by managed agents. Streaming client tool arguments remain intact when several calls interleave.\n\nAudio\n\nTranscribe or translate an audio file up to 25 MiB (flac, mp3, mp4, mpeg, m4a, ogg, wav or webm) with openai/whisper-large-v3 , and generate speech with canopylabs/orpheus-v1-english :\n\nwith open(\"meeting.m4a\", \"rb\") as f:\n text = client.audio.transcriptions.create(model=\"nodus/whisper-large-v3\", file=f)\nprint(text.text)\n\nspeech = client.audio.speech.create(model=\"canopylabs/orpheus-v1-english\", voice=\"tara\",\n input=\"Your job finished.\", response format=\"wav\")\nspeech.write to file(\"done.wav\")\n\nUpload the file itself: audio URLs and streamed transcription are not supported. Transcription and translation are billed per audio hour for the file’s length, with a 10-second minimum per request; speech is billed per million input characters, and its audio is limited to 32 MiB.\n\nIndra: nodus/indra \n\nSend model: \"nodus/indra\" and Indra picks the catalog model best suited to each request: small, fast models for simple requests and stronger models for hard ones. The X-Nodus-Routed-Model response header names the model that answered. If routing is unavailable, Indra answers with its default model, so requests never fail because of routing. Indra was first called Composer, and nodus/auto still works as another name for it.\n\nIndra comes in two tiers, and the X-Nodus-Composer-Tier response header ( free or paid ) names the one that answered:\n\n Free Indra is for every organization without the Indra plan. It routes among the three cheapest catalog models that are available (by blended price, three input tokens to one output token) and never draws on your credits. It is limited per organization to 100 requests and 200,000 tokens per UTC day, and each request to 4,096 output tokens (fewer when the day has less left); past either daily limit, requests get 429 with code QuotaExceeded until 00:00 UTC. Free Indra can also pause for the rest of the day when it is in heavy demand.\n Paid Indra comes with the $20 per month Indra plan. It routes across every catalog model, and you pay the routed model’s per-token price plus the small routing charge, from the plan’s monthly Indra allowance first and then from your credits.\n\nCalling a model by name, directly or through a named endpoint, is always billed per token from your credits.\n\nNamed endpoints and limits\n\nAn InferenceEndpoint gives a model its own base URL and access policy:\n\nTerminal window\n\n$ nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-concurrent 8 --max-cost 50\n$ nodus wait ep/support-bot --for=jsonpath={.status.phase}=Running\n$ nodus get ep\nNAME PHASE MODEL READY RPM CONCURRENT MAX-COST AGE\nsupport-bot Running nodus/gpt-oss-120b True 120 8 $50 12s\n$ nodus inference chat --endpoint support-bot \"Where is my order?\"\n\nCall it at https://inference.nodus-compute.ai/endpoints/ /support-bot/v1 , or send model: \"endpoint/support-bot\" on the shared base URL. Endpoints enforce rpm , tpm , maxConcurrent , allowedKeys (API key names, --allowed-key ) and an optional maxCostUSD : once the requests through an endpoint have spent its cap, the next one gets 402 BudgetExceeded , and raising maxCostUSD lets requests through again (it can only be raised). A new endpoint answers 503 for the few seconds until it is Running . nodus stop ep/support-bot makes it answer 503 until nodus start ep/support-bot ; its Ready condition says whether its model can serve now.\n\nBilling per token\n\n Before a request runs, Nodus holds its maximum cost: the input bound plus max tokens at the model’s rates. Lower max tokens to hold less.\n When it finishes, you are charged the tokens used (input, cached input and output) and the rest of the hold is released. A stream you stop early is charged for the tokens it produced.\n Send an Idempotency-Key header to retry safely: a repeat within 24 hours returns the same answer with Idempotent-Replayed: true and no second charge. A request that failed without a charge ( Released ) frees its key, so the retry runs again. Reusing a key for a different request body returns 409 .\n GET /v1/requests/{id} returns the receipt for 30 days: its operation , usage , amountUSD , pricebookVersion and state : Running , Settled , Released (no tokens, no charge) or Unknown . A request whose outcome Nodus never learned is Unknown and is never charged.\n\nErrors\n\nErrors use the OpenAI shape, with the Nodus error code in both type and code :\n\n Status Code Meaning \n - - - \n 400 Unsupported A feature outside per-token billing: built-in or server-side tools, a non-default service tier , store , background , file parts, image URLs or audio URLs, a key repeated in another letter case, or an operation the model does not serve \n 400 Invalid A malformed request, or an upload that is not audio in a supported format \n 401 Unauthorized Missing or invalid API key \n 402 InsufficientCredits , BudgetExceeded Not enough credit for the hold, or a budget or endpoint cap blocks it \n 409 RequestInProgress The same Idempotency-Key is still running; retry after Retry-After \n 409 IdempotencyKeyReused The Idempotency-Key was used for a different request; send a new key \n 413 RequestEntityTooLarge A JSON body over 1 MiB or an audio file over 25 MiB \n 429 TooManyRequests An org or endpoint limit; retry after Retry-After \n 502, 503 Unavailable The model failed, is at capacity or is not available; retry shortly. You are not charged"},{"id":"docs/guides/jobs","url":"https://nodus-platform-site.pages.dev/docs/guides/jobs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/jobs.md","title":"Jobs","description":"Run a command to completion on a GPU or CPU, watch it, and download what it wrote.","stage":"GA","headings":[{"depth":2,"slug":"sign-in","text":"Sign in"},{"depth":2,"slug":"submit-a-job","text":"Submit a Job"},{"depth":2,"slug":"watch-it","text":"Watch it"},{"depth":2,"slug":"get-the-results","text":"Get the results"},{"depth":2,"slug":"run-many-indexes","text":"Run many indexes"},{"depth":2,"slug":"limit-time-and-cost","text":"Limit time and cost"},{"depth":2,"slug":"suspend-resume-and-cancel","text":"Suspend, resume and cancel"},{"depth":2,"slug":"recovery","text":"Recovery"},{"depth":2,"slug":"when-a-job-fails","text":"When a Job fails"},{"depth":2,"slug":"clean-up","text":"Clean up"},{"depth":2,"slug":"multi-node-jobs","text":"Multi-node Jobs"},{"depth":2,"slug":"environment","text":"Environment"},{"depth":2,"slug":"run-code-from-github","text":"Run code from GitHub"}],"text":"Jobs\n\n Run a command to completion on a GPU or CPU, watch it, and download what it wrote.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/jobs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Job runs your command until it finishes. You say what it needs (a GPU type, memory, a time limit, a spending cap) and Nodus runs it on capacity that fits, streams its logs, collects its outputs and charges only for the time it used.\n\nSign in\n\nTerminal window\n\n$ pip install nodus-compute\n$ nodus login\n\n nodus login opens the console in your browser and stores a key for this machine.\n\nSubmit a Job\n\nThe fastest way is nodus run . It uploads the current directory, starts the command and follows it until it exits, passing the exit code through:\n\nTerminal window\n\n$ nodus run --gpu L4 --image nodus/pytorch -- python hello.py\n\nhello.py\n\nimport json\nimport os\n\nimport torch\n\nname = torch.cuda.get device name(0)\nprint(f\"Hello from {name}\")\n\nAnything written under /nodus/outputs is collected when the Job succeeds.\nreport = {\"gpu\": name, \"cuda\": torch.version.cuda, \"index\": os.environ[\"NODUS INDEX\"]}\nwith open(os.path.join(os.environ[\"NODUS OUTPUTS DIR\"], \"report.json\"), \"w\") as f:\n json.dump(report, f)\n\nBefore anything runs, nodus run prints the estimate: the expected cost to completion and when it should start. Add --max-cost 5 to stop the Job if it would spend more than $5, and --timeout 2h to limit its wall time.\n\nTo keep a Job in version control, write it as a manifest and apply it:\n\njob.yaml\n\napiVersion: nodus.dev/v1\nkind: Job\nmetadata:\n name: hello\nspec:\n image: nodus/pytorch\n command:\n - python\n - -c\n - \n import json, os, torch\n name = torch.cuda.get device name(0)\n print(f\"Hello from {name}\")\n with open(os.path.join(os.environ[\"NODUS OUTPUTS DIR\"], \"report.json\"), \"w\") as f:\n json.dump({\"gpu\": name, \"index\": os.environ[\"NODUS INDEX\"]}, f)\n resources:\n gpu: L4\n timeout: 10m\n maxCostUSD: \"0.25\"\n outputs:\n - name: report\n path: /nodus/outputs/report.json\n\nTerminal window\n\n$ nodus apply -f job.yaml\njob.nodus.dev/hello created\n\nAnything under /nodus/outputs is collected when the Job succeeds. Declare the paths you want to download by name under outputs . kubectl apply -f job.yaml works too once your kubeconfig points at Nodus.\n\nYou do not have to name a GPU. Set resources.gpu.minMemory instead of resources.gpu.type and Nodus picks any accelerator with at least that much memory per GPU. If you also leave out minMemory , the annotations nodus.dev/model (for example meta-llama/Llama-3.1-8B ) and nodus.dev/dataset-bytes , which you set under metadata.annotations , set it for you: 24Gi for 7B and 8B models, 32Gi for 13B and 14B, 48Gi for 34B and 40B, 80Gi for 70B and 72B, and 80Gi for a dataset over 1 TiB. nodus apply -f job.yaml --dry-run=server -o yaml shows the floor it picked.\n\nWatch it\n\nIn the console, open Runs , then select a run. Logs shows reported progress, stored outputs and the latest recovery checkpoint alongside the log stream. Distributed and indexed runs can be narrowed to a worker rank or index. The stream shows one worker at a time, starting with rank 0 and index 0. It follows that worker’s latest attempt; it is not a history of every retry. Connection failures show an error and Retry logs , while Download loaded logs saves the lines currently loaded in the browser. Run activity shows scheduling and lifecycle events separately from program output.\n\nThe command field preserves quoted arguments such as python -c \"print(1 + 1)\" . Shell pipelines and redirects need an explicit shell command, for example bash -lc 'python prepare.py && python train.py' .\n\nTerminal window\n\n$ nodus get job/hello -w\n$ nodus logs job/hello -f\n$ nodus describe job/hello\n\n get shows the phase, the GPU, the attempt, the cost so far and the age ( -o wide adds the offering, its rate and progress). describe adds the conditions, the events and the reason for a failure with a suggested fix. A Job moves through these phases:\n\n Phase What is happening \n - - \n Queued Waiting for capacity that fits and for a funded hold \n Provisioning Capacity is acquired; the image, source and inputs are being prepared \n Running Your command is running \n Recovering The capacity was lost; Nodus is moving the Job to new capacity \n Suspending , Suspended Saving state and releasing compute, then paused \n Cancelling Stopping and releasing compute \n Succeeded , Failed , Cancelled Finished; nothing is running or billed \n\nLifecycles lists every transition.\n\nGet the results\n\nOpen Outputs to download stored files and inspect their size and SHA-256. Distributed outputs identify their rank, and indexed outputs identify their index. A failed or cancelled run may have no stored outputs; its logs and activity explain what happened. Details includes attempts and the execution cost breakdown.\n\nTerminal window\n\n$ nodus cp job/hello:outputs/report ./report.json\n./report.json: 214 bytes, sha256 4f1c0e9a2b7d\n\n nodus cp checks every download against the SHA-256 recorded when the Job finished. Outputs stay available until you delete the Job. Outputs covers directories, Indexed Jobs and loading results into a database.\n\nRun many indexes\n\nAn Indexed Job runs the same command completions times, at most parallelism at once. Each run sees its number in NODUS INDEX (and in JOB COMPLETION INDEX , as on Kubernetes):\n\njob.yaml\n\napiVersion: nodus.dev/v1\nkind: Job\nmetadata:\n name: indexed-outputs\nspec:\n image: python:3.12-slim\n # Four indexes, two at a time. Each index sees its number in NODUS INDEX.\n completions: 4\n parallelism: 2\n command:\n - python\n - -c\n - \n import json, os\n i = int(os.environ[\"NODUS INDEX\"])\n rows = [{\"index\": i, \"n\": n, \"square\": n n} for n in range(i 100, (i + 1) 100)]\n os.makedirs(\"/nodus/outputs/shards\", exist ok=True)\n with open(f\"/nodus/outputs/shards/part-{i}.jsonl\", \"w\") as f:\n f.writelines(json.dumps(r) + \"\\n\" for r in rows)\n print(f\"index {i}: wrote {len(rows)} rows\")\n resources:\n cpu: \"2\"\n memory: 4Gi\n backoffLimit: 2\n timeout: 15m\n maxCostUSD: \"0.10\"\n outputs:\n - name: shards\n path: /nodus/outputs/shards\n\nThe Job succeeds when every index has succeeded. backoffLimit is how many failed runs the whole Job tolerates before it fails; a retried index starts fresh on new capacity. status.completedIndexes lists the finished indexes, such as 0-2,5 .\n\nLimit time and cost\n\n Field Flag What it does \n - - - \n maxCostUSD --max-cost The most the Job may spend. At the cap it saves its state and becomes Suspended with reason MaxCostReached ; raise the cap to resume it. The cap can only be raised. \n timeout --timeout Wall-clock limit counted from the first Provisioning ; the Job fails with DeadlineExceeded . Time spent Suspended does not count. You can raise, lower or remove it on a running Job, and status.timeoutTime moves with it. activeDeadlineSeconds is accepted as an alias. \n completeByTime --complete-by When you need the result. Nodus picks capacity that should finish in time. \n expectedDuration --expected-duration Your estimate of the run time, used for the cost estimate before the Job has any history. \n placement.queueTimeout How long to wait for capacity before failing with CapacityUnavailable . The wait starts again when a suspended Job resumes. \n\nWhen credits run out or a budget is reached, a Job saves its state and becomes Suspended with reason InsufficientCredits or BudgetExceeded . It stops by status.stopByTime , the time its funds run out, even if the save is not finished, and resumes on its own once it is funded again.\n\nSuspend, resume and cancel\n\nTerminal window\n\n$ nodus suspend job/train # save state, release compute, stop billing\n$ nodus resume job/train # continue from the saved state\n$ nodus cancel job/train # stop for good\n\nThese set spec.state to Suspended , Running or Cancelled , so the same change works from a manifest. If a suspend cannot save state, the Job goes back to Running and its Suspended condition says SuspendFailed ; nothing is lost. Nodus tries again ten minutes later, up to three tries in all ( status.suspendFailures counts the failures), and then leaves the Job running. A Job with recovery.continuity: Ephemeral keeps no state, so suspending it restarts it from the beginning, and nodus suspend warns first. nodus delete job/x cancels a running Job before removing it.\n\nRecovery\n\nCapacity can be lost while a Job runs, for example when interruptible capacity is reclaimed. What happens next depends on recovery.continuity :\n\n Continuity After a loss \n - - \n Checkpointed (default) Resumes on new capacity from the last saved state in NODUS CHECKPOINT DIR ( /nodus/state ) \n Restartable Starts again from the beginning on new capacity \n Ephemeral Fails; nothing is kept \n\nYour program writes and loads its own state files in NODUS CHECKPOINT DIR ; Nodus saves and restores that directory. Set recovery.onInterruption: Fail to fail instead of recovering. An index that uses up recovery.maxAttempts fails the Job with RecoveryLimitExceeded .\n\nWhen a Job fails\n\n status.reason says why and status.fix says what to change. The common reasons:\n\n Reason Meaning \n - - \n BackoffLimitExceeded Your command exited non-zero more times than backoffLimit allows \n OOMKilled The command ran out of memory; request more resources.memory or a larger GPU \n DeadlineExceeded timeout elapsed \n CapacityUnavailable No capacity fit the request within placement.queueTimeout \n ImagePullFailed The image does not exist or its registry refused the pull \n InvalidOutputPath A declared output was not written \n NoProgress Two recoveries in a row each ran less than recovery.minProgressDuration before losing capacity \n StorageUnavailable Nodus could not issue the storage credentials your command needs; run the Job again later \n\nClean up\n\nFinished Jobs are deleted ttlSecondsAfterFinished seconds after they finish. Jobs from nodus run default to 30 days; pass --keep to keep one. Deleting a Job deletes its outputs.\n\nMulti-node Jobs\n\nBeta\n\n spec.distributed runs one Job across several nodes as a gang, with the rank, world size and rendezvous address in the environment. The Multi-node training guide covers launchers, networking and outputs from each rank.\n\nEnvironment\n\nEvery run sees these variables. Your own env cannot use the NODUS prefix, except NODUS PARAM .\n\n Variable Value \n - - \n NODUS JOB The Job’s name \n NODUS INDEX , JOB COMPLETION INDEX This run’s index, from 0 \n NODUS COMPLETIONS spec.completions \n NODUS OUTPUTS DIR /nodus/outputs \n NODUS CHECKPOINT DIR The first checkpoint path, /nodus/state by default \n NODUS INPUT Where the input is mounted, under /nodus/inputs \n NODUS PARAM A Sweep parameter (see Sweeps) \n\nRun code from GitHub\n\nInstall a GitHub Connection in the Job’s project and wait until it is Ready. Set source.git.repo to owner/repository or its GitHub HTTPS URL, and set ref to a branch, tag or commit. Nodus records the resolved commit and source archive in status.source before placing the Job. Retries and parallel indexes use the same files even when the branch moves.\n\nsource:\n git:\n repo: your-org/your-repository\n ref: main\n\nThe files appear in workingDir before the command starts. Set source.git.path to use another checkout directory; this does not change the command’s working directory. The checkout contains repository files, without .git history, submodule contents or Git LFS downloads. GitHub credentials stay outside the guest. The source archive is limited to 100 MiB compressed, 512 MiB expanded and 10,000 entries. A missing Connection or rejected archive prevents execution."},{"id":"docs/guides/logs","url":"https://nodus-platform-site.pages.dev/docs/guides/logs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/logs.md","title":"Logs and metrics","description":"Follow a Job's output live, read it back after the capacity is gone, and watch GPU, CPU and memory use with nodus top.","stage":"GA","headings":[{"depth":2,"slug":"follow-a-run","text":"Follow a run"},{"depth":2,"slug":"choose-which-lines","text":"Choose which lines"},{"depth":2,"slug":"logs-after-the-run","text":"Logs after the run"},{"depth":2,"slug":"watch-usage-with-nodus-top","text":"Watch usage with nodus top"},{"depth":2,"slug":"charts-and-queries","text":"Charts and queries"}],"text":"Logs and metrics\n\n Follow a Job's output live, read it back after the capacity is gone, and watch GPU, CPU and memory use with nodus top.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/logs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nEverything your program writes to stdout and stderr is kept as its log. You can follow it while the program runs, read it back after it finishes, and filter it by time, attempt, index or rank. Usage samples (GPU, CPU and memory) are kept beside it, for nodus top and the console charts.\n\nFollow a run\n\n nodus run streams the log until the Job finishes. To attach to a Job that is already running, or to come back after you detached:\n\nTerminal window\n\n$ nodus logs -f job/finetune-llama\n\n -f ( --follow ) prints what is already stored, then new lines as they are written, without repeating or skipping any. Press Ctrl+C to stop following; the Job keeps running.\n\nThe same command works for Sandboxes, Workspaces, Functions, Agents, AgentRuns, Images and individual Attempts: nodus logs sb/dev , nodus logs attempt/NAME .\n\nChoose which lines\n\n Flag Shows \n - - \n --tail 100 The last 100 lines, then follows with -f \n --since 10m Lines from the last 10 minutes (any duration: 30s , 2h ) \n --timestamps Each line prefixed with the time it was written \n --attempt 2 One attempt of the Job; the default is the latest one \n --index 3 , --all-indexes One index of an Indexed Job, or all of them \n --rank 1 , --all-ranks One rank of a multi-node Job (Beta), or all of them prefixed with [r ] \n --process 12 The output of one Process in a Sandbox or Workspace, such as a nodus exec session \n\n --tail and --since combine: --since 1h --tail 50 prints at most the last 50 lines of the last hour. Without --attempt , a Job that was recovered shows its newest attempt; earlier attempts stay readable by number, which helps when a recovery followed a crash you want to look at.\n\nFrom Python, job.logs() returns the same text, and the MCP logs tool reads up to 2,000 lines at a time.\n\nLogs after the run\n\nLogs are stored as they are written, not only on the machine that ran the program. When a Job finishes, is suspended, or loses its capacity, nodus logs job/NAME still prints the whole log, and nodus logs --attempt 1 reads an attempt whose machine is long gone.\n\nLogs are kept for 30 days. When capacity disappears without warning, the last few seconds written before the loss may be missing; everything before them is kept. Secret values you pass to the Job ( secrets: ) are replaced with [redacted] before the log is stored or shown.\n\nLogs are your program’s own output, exactly as it wrote it. Nodus’ own messages about scheduling and recovery are Events ( nodus describe job/NAME , nodus events ), never mixed into your log. nodus events -w prints the events already recorded and then each new one as it happens; add --for job/NAME to follow one object, and -o json to get the full Event objects.\n\nWatch usage with nodus top\n\n nodus top shows what running work is using right now:\n\nTerminal window\n\n$ nodus top jobs\nNAME GPU GPU MEM CPU MEMORY $/H SPEND\nfinetune-llama 94% 71.2Gi 3800m 41.0Gi $2.49 $6.12\neval-sweep-3 61% 18.5Gi 1200m 12.3Gi $0.80 $0.35\n$ nodus top sandbox dev\n\nGPU is the average utilization across the Job’s GPUs; GPU memory, CPU and memory are totals. Samples are taken every 15 to 30 seconds, so a value can be up to half a minute old. nodus top sandboxes , workspaces and functions work the same way.\n\nCharts and queries\n\nThe console’s Metrics tab charts the same samples over time. For your own dashboards and scripts, the history of one object is at:\n\nGET /metrics/v1/namespaces/PROJECT/jobs/NAME/history?metric=gpu&since=6h&step=1m\n\n metric is gpu , gpuMemory , cpu or memory , the resource can be jobs , sandboxes , workspaces or functions , and the answer has the shape of a Prometheus query range result. Instead of since , pass start and end as RFC 3339 times.\n\nFor anything else, the project’s samples answer PromQL in the Prometheus HTTP API at /metrics/v1/namespaces/PROJECT/api/v1/query and query range , so Grafana and other Prometheus clients can use it as a data source with your API key. The series are:\n\n Series Labels Value \n - - - \n nodus container gpu utilization percent kind , name , uid , attempt , gpu Utilization of one GPU, 0 to 100 \n nodus container gpu memory bytes kind , name , uid , attempt , gpu Memory in use on one GPU \n nodus container cpu millis kind , name , uid , attempt CPU in use, in thousandths of a core \n nodus container memory bytes kind , name , uid , attempt Memory in use \n\nEvery query is limited to the project in the path: a query that names another org or project is refused. For example, max over time(nodus container gpu utilization percent{kind=\"Job\",name=\"finetune-llama\"}[1h]) gives the peak utilization of each GPU over the last hour.\n\nTraining metrics such as loss and accuracy are separate: report them with nodus.log.metrics(step=..., loss=...) and they appear in the Job’s status under status.progress.metrics ."},{"id":"docs/guides/mcp","url":"https://nodus-platform-site.pages.dev/docs/guides/mcp/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/mcp.md","title":"MCP","description":"Let Claude Code, Cursor, Codex and other agents use Nodus through one MCP server, with confirmation before anything is created.","stage":"GA","skill":"nodus-mcp","headings":[{"depth":2,"slug":"connect-a-client","text":"Connect a client"},{"depth":3,"slug":"claude-code","text":"Claude Code"},{"depth":3,"slug":"cursor","text":"Cursor"},{"depth":3,"slug":"codex","text":"Codex"},{"depth":3,"slug":"other-clients","text":"Other clients"},{"depth":2,"slug":"what-the-agent-can-do","text":"What the agent can do"},{"depth":2,"slug":"nothing-is-created-without-a-confirmation","text":"Nothing is created without a confirmation"},{"depth":2,"slug":"logs-and-output-are-data-not-instructions","text":"Logs and output are data, not instructions"},{"depth":2,"slug":"change-or-remove-access","text":"Change or remove access"},{"depth":2,"slug":"troubleshooting","text":"Troubleshooting"}],"text":"MCP\n\n Let Claude Code, Cursor, Codex and other agents use Nodus through one MCP server, with confirmation before anything is created.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/mcp/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNodus runs one MCP server. An agent that speaks MCP can list your Jobs, read their logs, estimate what a manifest would cost, and, after you confirm, create it. The tools work on every kind of resource with the same names, so a new kind needs no new tools.\n\nThe agent acts as you, with a role and scopes that you choose when you connect it. It can never do more than your own role allows.\n\nConnect a client\n\nThere are two ways to connect. The hosted server runs at https://api.nodus-compute.ai/mcp and signs you in with your browser. The local server is nodus mcp , which the client starts on your machine and which acts with your nodus login .\n\nNote\n\n nodus mcp install CLIENT writes the setup for you. It adds only a nodus entry to the client’s own configuration file, keeps a copy of the original next to it, and refuses to replace a different nodus entry unless you pass --force . Add --dry-run to see what would change.\n\nClaude Code\n\nTerminal window\n\n$ claude mcp add --transport http nodus https://api.nodus-compute.ai/mcp\n\nRun /mcp in Claude Code and choose nodus to sign in. Or for the local server:\n\nTerminal window\n\n$ nodus login\n$ nodus mcp install claude\n\nCursor\n\nUse the Cursor link on the connect page, or:\n\nTerminal window\n\n$ nodus mcp install cursor --hosted\n\nDrop --hosted to run nodus mcp locally instead. Cursor reads ~/.cursor/mcp.json .\n\nCodex\n\nTerminal window\n\n$ nodus mcp install codex --hosted\n$ codex mcp login nodus\n\nCodex reads config.toml in $CODEX HOME ( ~/.codex by default). Without --hosted , the entry starts nodus mcp .\n\nOther clients\n\nAny client that supports streamable HTTP can use the hosted URL. For one that starts a command, use /mcp.json (the local server) or /mcp-hosted.json (the hosted one).\n\nWhat the agent can do\n\nWhen you sign in to the hosted server, a consent page asks which org and which scopes the client may use. The tools it lists follow those scopes, so a client granted only jobs:read sees only the tools that read Jobs.\n\n Tool What it does \n - - \n whoami , api resources , explain Who you are, the kinds and their fields. Agents start here. \n get , describe , events Read objects, a summary with its recent events, and Events. \n logs , wait , outputs Read logs (at most 2,000 lines), wait up to 60 seconds for a state, list outputs or get a download link valid for 10 minutes. \n estimate Dry-run a manifest: the cost estimate, anything blocking it, and the etag that apply needs. \n apply Create an object from a manifest. See below. \n delete , set state , request , record Delete an object, suspend or resume it, ask for an action such as a restart, and create a record such as a webhook replay. \n exec , files Run a command in a running Sandbox, Workspace or Job (up to 300 seconds), and read, write or list its files. \n\nA local server also has three tools that touch your machine’s files: upload source (upload a directory as a code blob; it follows .gitignore and .nodusignore ), download output (save an output, verified against its checksum) and cp (copy a file to or from a Sandbox, Workspace or Job).\n\nNothing is created without a confirmation\n\n apply does not create anything the first time. It returns the dry-run: the object as it would be stored, its estimated cost, and an etag . The agent shows you that. Only when you accept does it call apply again with confirmed: true and the same etag , and the create fails if the estimate changed in between. Objects created this way carry the label nodus.dev/launched-by with the value mcp .\n\nMCP never completes a payment. Applying a TopUp returns a checkout link for you to open yourself.\n\nLogs and output are data, not instructions\n\nLogs, command output and file contents come back marked as untrusted. They can contain anything a program wrote, so an agent must treat them as data. Nodus also stops a tool call that touches more than your grant allows, even if text in a log asks for it.\n\nChange or remove access\n\nConsole › Settings › Connected agents lists each client with its scopes. Narrowing or deleting a grant takes effect within 30 seconds, even for a client that already holds a token. See Connected agents.\n\nTerminal window\n\n$ nodus get oauthgrants\n$ nodus delete oauthgrant claude-code-3fa2c1\n\nTroubleshooting\n\n Symptom What to do \n - - \n The client lists fewer tools than you expect The grant lacks the scope. Reconnect and approve more scopes, or check your role with nodus auth can-i . \n not logged in from nodus mcp Run nodus login , or set NODUS API KEY . \n install says there is a different nodus connection Remove that entry, or pass --force to replace it. \n install says it cannot parse the file Fix the file’s syntax first. Nothing was changed. \n A write was refused with a changed estimate Ask the agent to run estimate again and show you the new numbers."},{"id":"docs/guides/migrate-from-0x","url":"https://nodus-platform-site.pages.dev/docs/guides/migrate-from-0x/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/migrate-from-0x.md","title":"Migrate from Nodus Compute 0.x","description":"What moves to the new Nodus automatically, what you recreate, and how the CLI, SDK, agents and billing change.","stage":"GA","headings":[{"depth":2,"slug":"what-moves-automatically","text":"What moves automatically"},{"depth":2,"slug":"your-30-starter-credit","text":"Your $30 starter credit"},{"depth":2,"slug":"cutover-timeline","text":"Cutover timeline"},{"depth":2,"slug":"what-you-recreate","text":"What you recreate"},{"depth":2,"slug":"update-your-tools","text":"Update your tools"},{"depth":2,"slug":"billing-changes","text":"Billing changes"}],"text":"Migrate from Nodus Compute 0.x\n\n What moves to the new Nodus automatically, what you recreate, and how the CLI, SDK, agents and billing change.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/migrate-from-0x/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNodus 1.0 carries over your account, sign-in and organization memberships. Each migrated user receives $30 in free starter credit once, subject to email verification. Runs, Sandboxes and API keys must be recreated.\n\nWhat moves automatically\n\n From 0.x In Nodus 1.0 \n - - \n Email/password, Google or GitHub sign-in Your existing sign-in method \n Your account and organization memberships An org with a default project and your mapped member role \n An account without an organization Your sign-in is preserved; create your first org during onboarding \n\nThe old billing system used sandbox payments. Its balances, debt, saved cards, payments and invoice history do not carry over. Add a payment method in the new console when required to run work.\n\nYour $30 starter credit\n\nThe grant is once per user across all organizations and expires 30 days after it is granted. If you already received starter credit in the new platform, migration does not grant it again, even if it has expired.\n\nVerified users with an imported organization receive the grant when that organization is released. If your email is not verified, verify it and open an organization you belong to; the console claims your credit automatically. If you have no organization, create one during onboarding. Your entitlement remains available while you complete these steps; its 30-day expiry starts only when the grant is issued.\n\nCutover timeline\n\nFollow the announced cutover dates for pausing work and downloading outputs. The old service is kept available for the announced read-only window; running work and stored outputs are not copied into the new platform.\n\nWhat you recreate\n\n API keys. 0.x keys ( nk … ) stop working. Sign in with nodus login or create a new key in the console.\n Runs, Sandboxes and 0.x workspaces. They are not copied. Download any files you need before the cutover date in your notice; the 0.x API keeps serving reads for 30 days afterwards. Recreate workspaces as Workspace resources.\n BYOC pools and hosts. Recreate pools and enroll hosts in the new platform.\n Cloud accounts. Re-connect AWS and GCP accounts as CloudAccount resources and remove the old role or grant.\n\nUpdate your tools\n\n1. Upgrade the SDK and CLI. pip install -U nodus-compute installs 1.0 under the same package, import and command names. The 1.0 API is new: a 0.x Workload becomes a Job , and nodus run replaces the 0.x submit commands. On first run, 1.0 moves an old ~/.nodus config aside to config.0x.bak .\n2. Sign in again. Run nodus login . NODUS BASE URL still works as an alias of NODUS API URL for this major version, with a warning.\n3. Reconnect coding agents. Hosted MCP clients prompt you to authenticate again; local clients follow the setup at /connect.\n4. Update CI. The 0.x GitHub Action fails with a migration message. Use nodus-compute/run-action@v1 with a new API key.\n\nBilling changes\n\nStart-up and shutdown time is now billed\n\n0.x started the meter when your command started. Nodus 1.0 bills rented machines from the moment they are created until their deletion is confirmed, at provider cost ÷ 0.875, and itemizes each charge as Boot , Running , Restore or Teardown . The same run can cost more than it did on 0.x, and the estimate shows each part before you launch.\n\nLinks to 0.x docs pages redirect to their 1.0 equivalents, and old console links open a page that helps you find the object in the new console."},{"id":"docs/guides/multi-node","url":"https://nodus-platform-site.pages.dev/docs/guides/multi-node/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/multi-node.md","title":"Multi-node training (Beta)","description":"Run one training job across several machines with torchrun, Ray or your own launcher, and understand what it costs to assemble them.","stage":"Beta","headings":[{"depth":2,"slug":"run-a-two-node-job","text":"Run a two-node job"},{"depth":2,"slug":"topology","text":"Topology"},{"depth":2,"slug":"launchers","text":"Launchers"},{"depth":2,"slug":"failures-and-restarts","text":"Failures and restarts"},{"depth":2,"slug":"the-assembly-bound","text":"The assembly bound"},{"depth":2,"slug":"beta-caveats","text":"Beta caveats"}],"text":"Multi-node training (Beta)\n\n Run one training job across several machines with torchrun, Ray or your own launcher, and understand what it costs to assemble them.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/multi-node/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n Beta: this feature may change.\n\nBeta\n\nMulti-node training is in Beta and on for every organization, with no access request and no purchase needed. Beta gangs have at most 8 nodes. When members on different providers connect over their public addresses (the public path), traffic between them, including gradients and the rendezvous, is not encrypted ; each member accepts it only from the other members’ addresses. Keep network: Colocated if your data must not cross the internet unencrypted.\n\nA distributed Job runs your command on several machines at once, called a gang . Nodus acquires every member, connects them, checks that they can reach each other, and only then starts your command on all of them together. You write an ordinary training script; the launcher finds its peers from environment variables Nodus sets.\n\nRun a two-node job\n\ntrain.py\n\n\"\"\"A two-node DDP smoke run. torchrun reads its rendezvous from the PET variables Nodus sets on every rank.\"\"\"\nimport os\n\nimport torch\nimport torch.distributed as dist\nfrom torch.nn.parallel import DistributedDataParallel as DDP\n\nbackend = \"nccl\" if torch.cuda.is available() else \"gloo\"\ndist.init process group(backend)\nlocal rank = int(os.environ[\"LOCAL RANK\"])\ndevice = torch.device(\"cuda\", local rank) if backend == \"nccl\" else torch.device(\"cpu\")\nif backend == \"nccl\":\n torch.cuda.set device(device)\n\nEvery process contributes 1, so the sum is the world size.\none = torch.ones(1, device=device)\ndist.all reduce(one)\n\nmodel = DDP(torch.nn.Linear(16, 1).to(device), device ids=[local rank] if backend == \"nccl\" else None)\nopt = torch.optim.SGD(model.parameters(), lr=0.1)\nfor step in range(20):\n x = torch.randn(32, 16, device=device)\n loss = (model(x) - x.sum(dim=1, keepdim=True)).pow(2).mean()\n opt.zero grad()\n loss.backward()\n opt.step()\n\nif dist.get rank() == 0:\n print(f\"world size={int(one.item())} node={os.environ['NODUS NODE RANK']} \"\n f\"transport={os.environ['NODUS GANG TRANSPORT']} loss={loss.item():.4f}\")\ndist.destroy process group()\n\nrun.sh\n\nnodus run --name torchrun-2-node --gpu H100 --nodes 2 --launcher torchrun --max-cost 2.00 -- torchrun train.py\n\n torchrun needs no flags: its node count, rendezvous address and local address come from the PET variables on every rank. The run prints world size=2 from rank 0 when both nodes joined.\n\nTopology\n\nSay how big the gang is with exactly one of:\n\n Field ( nodus run flag) Meaning \n - - \n distributed.nodes ( --nodes ) The number of machines, 2 to 8 \n distributed.totalGPUs ( --total-gpus ) The total GPU count; Nodus picks the machine shape. A total one machine can hold becomes a single-node Job \n distributed.gpusPerNode ( --gpus-per-node ) GPUs per machine: 1, 2, 4 or 8. With nodes it defaults to the GPU count of resources.gpu ; with totalGPUs Nodus picks it unless you set it \n\nEvery member gets the same GPU count. When resources.gpu.type lists several accelerator types, members may get different ones, each billed at its own rate. The resolved shape is written once to status.topology and never changes on a restart, so the world size stays the same for the whole run. A dry run shows it, with the hold for every member ( gangHoldUSD ) and the assembly bound:\n\nTerminal window\n\n$ nodus run --dry-run --gpu H100 --total-gpus 16 --launcher torchrun -- torchrun train.py\n\n distributed.network bounds where members may be:\n\n network Members are placed \n - - \n Colocated (default) With one provider in one region, on its private network \n Regional With any providers inside one region class \n Global Anywhere. The run shows the WANBound condition, because training between distant machines is limited by the network \n\n distributed.transport says which network paths you accept. Direct (default) admits every path that is not relayed: a provider’s private network, the members’ public addresses between providers (unencrypted in the Beta), and direct encrypted paths. Auto also admits relayed paths, which carry traffic through Nodus relays: they work between providers that cannot reach each other directly, but they are slow and fit small models only. The estimate warns RelayedLowBandwidth when a relayed path is possible.\n\nrun.sh (across providers)\n\nnodus run --name torchrun-2-node-xp --gpu H100 --nodes 2 --launcher torchrun \\\n --network global --transport auto --startup-timeout 15m --max-cost 2.00 -- torchrun train.py\n\nLaunchers\n\n launcher What each rank runs \n - - \n Plain (default) Your command once per node. RANK , LOCAL RANK=0 , LOCAL WORLD SIZE=1 and WORLD SIZE (the node count) make env:// initialization work for one process per node \n Torchrun Your torchrun command on every node; rank 0 hosts the rendezvous. Nodus owns restarts, so PET MAX RESTARTS=0 \n Ray A Ray cluster: rank 0 is the head, the others join it, and your command runs on rank 0 once all nodes are up, with RAY ADDRESS set \n Verl The Ray launcher plus NODUS VERL OVERRIDES ( trainer.nnodes , trainer.n gpus per node ) to append to your verl command \n Accelerate , Deepspeed The environment contract plus ACCELERATE or /etc/nodus/hostfile . Accepted in the Beta, qualified later \n\nEvery rank gets the same variables except the rank-specific ones:\n\n Variable Value \n - - \n NODUS NODE RANK , NODE RANK The Nodus rank of this node; rank 0 hosts the rendezvous \n NODUS NUM NODES , NNODES , NODUS GPUS PER NODE The node count and GPUs per node \n NODUS NODE IPS Every member’s address in rank order \n MASTER ADDR , MASTER PORT Rank 0’s address and 29500 \n NODUS GANG EPOCH , NODUS GANG TRANSPORT The current epoch and its path: private , public , direct or relayed \n NODUS RESTORE URI The checkpoint to resume from after a restart, when one exists \n\n NODUS NODE RANK is Nodus’s rank. torch assigns its own global RANK inside torchrun and may order nodes differently. Setting any NODUS , PET or MASTER variable, NODE RANK or NNODES in your spec is rejected. Your NCCL settings are kept, except the few the network path decides.\n\nFailures and restarts\n\nIf a member’s machine is lost or reclaimed, Nodus stops the whole gang’s current epoch at once and restarts it: surviving machines are kept, only the lost ranks get new machines, and every rank starts again at the next epoch from the latest gang checkpoint ( NODUS RESTORE URI ). A process left over from the old epoch cannot join the new rendezvous. recovery.maxAttempts counts these restarts (default 3). The first failure’s cause is the gang’s reason, so a crash on one rank that brings down the others reports that crash.\n\nIf a machine is refused or never appears while the gang is still assembling, only that rank gets another machine (up to two per assembly) and the registered ranks keep waiting at the start barrier. A suspend or cancel in progress is never turned into a restart: a machine lost during it simply completes the stop.\n\nIf your command exits non-zero on any rank, the run fails; a zero exit on every rank succeeds it.\n\nEach rank’s attempts carry the nodus.dev/rank label, so nodus get attempts -l nodus.dev/job= ,nodus.dev/rank=1 lists rank 1’s attempts across restarts, with each attempt’s placement, boot or restore time and cost.\n\nThe assembly bound\n\nAssembly waits at most distributed.startupTimeout (default 15 minutes, 5 to 60) for every member to become ready and pass the network check. If it does not, every acquired member is released, and Nodus tries again up to maxAssemblyRetries times (default 2) before the run fails with GangAssemblyTimeout .\n\nYou pay for each member from the moment its provider starts billing until it is confirmed deleted, including time spent waiting for the other members. The estimate shows the most a failed assembly can cost, the assembly bound : the sum of the members’ rates plus up to two replacement machines, times the startup timeout plus teardown, times the number of tries. Nodus pays, not you, when assembly fails for a Nodus reason: a network check that fails on a path Nodus lists as qualified, an address collision, or an outage of the Nodus mesh.\n\nBeta caveats\n\n Public paths carry traffic between members’ public addresses in clear text. Only the gang’s own members may connect to a member’s rendezvous and collective ports, but the bytes are not encrypted on the way.\n Relayed paths carry every byte through a Nodus relay, are billed per relayed GiB, and suit small models only.\n Relayed gangs have at most 2 members until larger relayed gangs are qualified.\n On relayed paths your image must use dynamically linked glibc programs; otherwise the run fails with ShimNotLoaded .\n Provider pairs are added as each is qualified; a pair that is not qualified is never offered. The measured TCP throughput and round-trip time per path class are published here as pairs are qualified."},{"id":"docs/guides/notifications","url":"https://nodus-platform-site.pages.dev/docs/guides/notifications/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/notifications.md","title":"Notifications","description":"Which emails Nodus sends, who receives them, how often, and how to switch off the optional ones.","stage":"GA","headings":[{"depth":2,"slug":"who-receives-what","text":"Who receives what"},{"depth":2,"slug":"what-nodus-sends","text":"What Nodus sends"},{"depth":2,"slug":"switch-off-the-optional-ones","text":"Switch off the optional ones"},{"depth":2,"slug":"add-recipients-to-a-budget","text":"Add recipients to a Budget"}],"text":"Notifications\n\n Which emails Nodus sends, who receives them, how often, and how to switch off the optional ones.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/notifications/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNodus emails you when something needs your attention: a balance running low, a Budget threshold, a Job that stopped because of money, a pool host that went offline. Each notice is sent once for the condition that caused it. A retry or a repeat of the same condition never sends a second email.\n\nEvery notice is also an Event, and most are webhook types, so a program can react to them too.\n\nWho receives what\n\n Kind of notice Goes to \n - - \n Money: balance, funding, Budgets, grants, top-ups, arrears, disputes, storage and egress quotas The billing email of the org’s BillingAccount, and every Owner and Admin, once each \n Budget notices The same people, and the addresses in the Budget’s spec.notify.emails (up to 10) \n About an object: for example a Job stopped by its cost cap The person who created it. For an object created by an API key or service account, the key’s owner or the project’s Admins \n Pools: host offline, capacity exhausted, burst approval needed, forecast shortfall, an action proposed or executed, a cloud sync that failed The org’s Owners and Admins, and the person who created the Pool \n Invites The invited address, even before it belongs to a member \n\nWhat Nodus sends\n\n Notice When How often \n - - - \n Low balance Available credit fell below your warning level At most once a day \n Funding lost An object stopped because credit ran out One digest per hour \n Maximum cost reached An object stopped at its maxCostUSD Once per object \n Budget threshold, Budget exceeded A Budget reached a percentage, or refused a new run Once per threshold and period \n Grant expiring A credit grant expires in 7 days and in 1 day Once each \n Top-up failed, auto-recharge failed A charge did not go through Once per attempt \n Arrears A balance went negative At most once a day \n Dispute opened or closed A payment dispute changed Once per change \n Storage or egress quota reached A quota was hit At most once a day \n Storage in arrears Storage is unpaid after 7, 21 and 28 days, with the purge proposed at 30 Once each \n Pool notices A pool event listed above Once per event and window \n Invite Someone invited you to an org Once per invite, and again when it is resent \n\nStripe sends the receipt for a top-up to your billing email. Sign-in and password emails come from your account system, not from these notices.\n\nSwitch off the optional ones\n\nBudget thresholds, low-balance warnings and run-completion emails are optional: each member can turn them off for themselves. Stops, disputes, arrears, quota limits and security notices are not optional.\n\nChange your preferences in your account settings. They are also at GET and PUT /auth/me/settings ; read the current settings first to see the fields.\n\nYour choice applies to you only. Other recipients of the same notice still get it.\n\nAdd recipients to a Budget\n\napiVersion: nodus.dev/v1\nkind: Budget\nmetadata:\n name: research\nspec:\n limitUSD: \"500\"\n period: Monthly\n notify:\n emails: [finance@example.com]\n\nNote\n\nNotices never contain secrets. An invite link is sealed until the email is sent and is erased afterward."},{"id":"docs/guides/outputs","url":"https://nodus-platform-site.pages.dev/docs/guides/outputs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/outputs.md","title":"Outputs","description":"Collect files from a Job, download them with a verified checksum, and load results into Postgres.","stage":"GA","headings":[{"depth":2,"slug":"write-outputs","text":"Write outputs"},{"depth":2,"slug":"download-outputs","text":"Download outputs"},{"depth":3,"slug":"with-the-api","text":"With the API"},{"depth":2,"slug":"stage-outputs","text":"Stage outputs"},{"depth":2,"slug":"load-outputs-into-postgres","text":"Load outputs into Postgres"},{"depth":2,"slug":"how-long-outputs-last","text":"How long outputs last"}],"text":"Outputs\n\n Collect files from a Job, download them with a verified checksum, and load results into Postgres.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/outputs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nAn output is a file or directory a Job writes for you to download. Outputs are collected when the Job succeeds, checksummed, and kept until you delete the Job.\n\nWrite outputs\n\nWrite anything you want to keep under /nodus/outputs (also in NODUS OUTPUTS DIR ). Every file there is collected when the Job succeeds. To give a file or directory a name of your own, declare it:\n\njob.yaml\n\napiVersion: nodus.dev/v1\nkind: Job\nmetadata:\n name: hello\nspec:\n image: nodus/pytorch\n command:\n - python\n - -c\n - \n import json, os, torch\n name = torch.cuda.get device name(0)\n print(f\"Hello from {name}\")\n with open(os.path.join(os.environ[\"NODUS OUTPUTS DIR\"], \"report.json\"), \"w\") as f:\n json.dump({\"gpu\": name, \"index\": os.environ[\"NODUS INDEX\"]}, f)\n resources:\n gpu: L4\n timeout: 10m\n maxCostUSD: \"0.25\"\n outputs:\n - name: report\n path: /nodus/outputs/report.json\n\nA declared path is a file or a directory under /nodus/outputs . Every file is collected as its own output, named by its path below /nodus/outputs , and a declared name is another name for the file or directory it points at. A Job can declare up to 32 outputs; names use lowercase letters, digits, . , and - , and the name outputs and the nodus. prefix are reserved.\n\nOutputs are separate from checkpoints: files in NODUS CHECKPOINT DIR are for resuming the Job, not for download.\n\nDownload outputs\n\nTerminal window\n\n$ nodus cp job/hello:outputs/report ./report.json\n./report.json: 214 bytes, sha256 4f1c0e9a2b7d\n$ nodus get job/hello -o jsonpath='{.status.outputs}'\n\n nodus cp writes the file only after its SHA-256 matches the digest recorded when the output was collected. Name a file by its declared name ( report ) or by its path below /nodus/outputs ( report.json ); the files of a directory output are listed one by one in status.outputs .\n\nFor an Indexed Job, every index writes its own copy of each output. Pick one with --index :\n\nTerminal window\n\n$ nodus cp job/indexed-outputs:outputs/shards/part-2.jsonl ./part-2.jsonl --index 2\n\nWith the API\n\n GET /apis/nodus.dev/v1/namespaces/ /jobs/ /outputs lists the collected outputs with their index, size and sha256 . GET …/outputs/ answers 302 with a short-lived download URL and the digest in the X-Nodus-SHA256 header; add ?index= when several indexes produced the output. Verify the digest of what you download. Errors are Status objects with a reason and a fix : OutputNotFound (404) when the Job committed no such output, OutputIndexRequired (400) when you need to pass ?index= , and Invalid (422) for a malformed index .\n\nStage outputs\n\nIn a Pipeline, a stage reads an earlier stage’s output as an input, mounted read-only at /nodus/inputs/ and named by NODUS INPUT . Any Job can do the same with an output input that names another Job:\n\ninputs:\n - name: data\n output: {job: prepare, name: shards}\n\nLoad outputs into Postgres\n\nAn output with a sink is loaded into a table in your database after the Job succeeds. The database is reached through a Postgres, Neon or Supabase Connection:\n\njob.yaml\n\napiVersion: nodus.dev/v1\nkind: Job\nmetadata:\n name: output-sink\nspec:\n image: python:3.12-slim\n command:\n - python\n - -c\n - \n import json\n with open(\"/nodus/outputs/squares.jsonl\", \"w\") as f:\n for n in range(1000):\n f.write(json.dumps({\"n\": n, \"square\": n n}) + \"\\n\")\n timeout: 10m\n maxCostUSD: \"0.05\"\n outputs:\n - name: squares\n path: /nodus/outputs/squares.jsonl\n # After the Job succeeds, the rows are loaded into this table through the Connection named analytics.\n sink:\n connection: analytics\n table: squares\n mode: Replace\n\n The output must be a .csv , .jsonl or .parquet file.\n mode: Append (the default) adds rows; Replace replaces the table’s rows.\n Each file can be up to 5 GB and 50 million rows, and each record up to 8 MiB.\n The load runs as part of the Job and is billed to it as CPU time.\n\n status.outputs[].sink shows each load’s phase ( Pending , Loading , Loaded or Failed ) and the rows loaded, and the SinksLoaded condition turns true when every load has finished. To retry the failed loads of a finished Job:\n\nTerminal window\n\n$ nodus request reload-sinks job/output-sink\n\nHow long outputs last\n\nOutputs stay downloadable until the Job is deleted, and deleting the Job deletes them. Jobs started with nodus run are deleted 30 days after they finish unless you pass --keep ; set ttlSecondsAfterFinished on a manifest to choose your own retention."},{"id":"docs/guides/pipelines","url":"https://nodus-platform-site.pages.dev/docs/guides/pipelines/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/pipelines.md","title":"Pipelines","description":"Chain Jobs into stages that pass outputs along, run in parallel where they can, and share one spending cap.","stage":"GA","headings":[{"depth":2,"slug":"submit-a-pipeline","text":"Submit a Pipeline"},{"depth":2,"slug":"watch-it","text":"Watch it"},{"depth":2,"slug":"stages-and-dependencies","text":"Stages and dependencies"},{"depth":2,"slug":"when-a-stage-fails","text":"When a stage fails"},{"depth":2,"slug":"cost-cap","text":"Cost cap"},{"depth":2,"slug":"suspend-resume-and-cancel","text":"Suspend, resume and cancel"}],"text":"Pipelines\n\n Chain Jobs into stages that pass outputs along, run in parallel where they can, and share one spending cap.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/pipelines/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Pipeline runs Jobs as stages. A stage starts when the stages it depends on have succeeded, reads their outputs as inputs, and runs in parallel with every other stage that is ready. The whole Pipeline shares one spending cap.\n\nSubmit a Pipeline\n\nThis Pipeline prepares a dataset on a CPU, trains on a GPU, and evaluates the trained model:\n\npipeline.yaml\n\napiVersion: nodus.dev/v1\nkind: Pipeline\nmetadata:\n name: train-eval\nspec:\n # One cap for the whole run: every stage's spend counts against it.\n maxCostUSD: \"0.50\"\n failurePolicy: FailFast\n # Stage templates inherit these fields unless they set their own.\n defaults:\n image: nodus/pytorch\n timeout: 15m\n stages:\n - name: prepare\n template:\n resources:\n cpu: \"2\"\n memory: 4Gi\n command:\n - python\n - -c\n - \n import os, torch\n os.makedirs(\"/nodus/outputs/shards\", exist ok=True)\n x = torch.randn(4096, 16)\n y = (x.sum(dim=1, keepdim=True) 0).float()\n torch.save({\"x\": x, \"y\": y}, \"/nodus/outputs/shards/data.pt\")\n outputs:\n - name: shards\n path: /nodus/outputs/shards\n - name: train\n dependsOn: [prepare]\n inputs:\n # Mounted at /nodus/inputs/data and named by NODUS INPUT DATA.\n - name: data\n fromStage: prepare\n output: shards\n template:\n resources:\n gpu: L4\n command:\n - python\n - -c\n - \n import os, torch\n d = torch.load(os.path.join(os.environ[\"NODUS INPUT DATA\"], \"data.pt\"))\n model = torch.nn.Linear(16, 1).cuda()\n opt = torch.optim.SGD(model.parameters(), lr=0.1)\n x, y = d[\"x\"].cuda(), d[\"y\"].cuda()\n for step in range(200):\n opt.zero grad()\n loss = torch.nn.functional.binary cross entropy with logits(model(x), y)\n loss.backward()\n opt.step()\n print(f\"final loss {loss.item():.4f}\")\n os.makedirs(\"/nodus/outputs/model\", exist ok=True)\n torch.save(model.state dict(), \"/nodus/outputs/model/model.pt\")\n outputs:\n - name: model\n path: /nodus/outputs/model\n - name: eval\n dependsOn: [train]\n inputs:\n - name: model\n fromStage: train\n output: model\n template:\n resources:\n cpu: \"2\"\n memory: 4Gi\n command:\n - python\n - -c\n - \n import os, torch\n model = torch.nn.Linear(16, 1)\n model.load state dict(torch.load(os.path.join(os.environ[\"NODUS INPUT MODEL\"], \"model.pt\")))\n x = torch.randn(1024, 16)\n acc = ((model(x) 0).float() == (x.sum(dim=1, keepdim=True) 0).float()).float().mean()\n print(f\"accuracy {acc.item():.3f}\")\n\nTerminal window\n\n$ nodus apply -f pipeline.yaml\npipeline.nodus.dev/train-eval created\n\nThe console and the Python SDK also offer this prepare, train and evaluate shape as the train-eval template.\n\nWatch it\n\nTerminal window\n\n$ nodus get pipeline/train-eval -w\nNAME PHASE STAGES COST AGE\ntrain-eval Running 1/3 $0.02 3m\n$ nodus get jobs -l nodus.dev/pipeline=train-eval\n$ nodus logs job/train-eval-train -f\n\nEach stage runs as a Job named - , so every Job command works on a stage: logs , exec , describe and cp . status.stages shows each stage’s phase, Job, start and finish times and cost. The pipeline name plus the stage name can be at most 62 characters.\n\nStages and dependencies\n\n dependsOn lists the stages that must succeed first. The stages form a graph with no cycles; a cycle is rejected with the path that closes it.\n inputs hands an output of an earlier stage to this one: fromStage names the stage (it must also be in dependsOn ) and output names one of its declared outputs. The output is mounted read-only at /nodus/inputs/ , and NODUS INPUT holds that path.\n maxParallel limits how many stages run at once. It defaults to the number of stages, so every ready stage starts.\n defaults holds Job fields every stage inherits. A field a stage sets wins; objects merge field by field, and a list a stage sets replaces the default list.\n\nWhen a stage fails\n\n failurePolicy What happens \n - - \n FailFast (default) The running stages are cancelled, the stages not yet started are skipped, and the Pipeline fails \n RunIndependent Stages that do not depend on the failed stage run to the end; the stages that depend on it are skipped; the Pipeline then fails \n\nThe Pipeline’s status.reason is StageFailed and its message names the failed stages. Each stage keeps its own status.reason , so nodus describe job/ - shows why it failed.\n\nCost cap\n\n maxCostUSD caps the whole Pipeline: the spend of every stage counts against it. When the Pipeline reaches the cap, every running stage saves its state and suspends, no new stage starts, and the Pipeline becomes Suspended with reason MaxCostReached . Raise the cap to resume:\n\nTerminal window\n\n$ nodus patch pipeline/train-eval --type merge --patch '{\"spec\":{\"maxCostUSD\":\"1.00\"}}'\n\nThe cap can only be raised. A stage can set a lower maxCostUSD of its own.\n\nSuspend, resume and cancel\n\n nodus suspend pipeline/x , resume and cancel apply to every stage that has not finished. They set the Pipeline’s spec.state , which its stage Jobs follow. Deleting a Pipeline deletes its stage Jobs and their outputs; ttlSecondsAfterFinished deletes a finished Pipeline for you."},{"id":"docs/guides/pools","url":"https://nodus-platform-site.pages.dev/docs/guides/pools/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/pools.md","title":"Run jobs on your own machines","description":"Enroll your hosts into a pool, route jobs to them first, forecast demand and automate routine actions.","stage":"GA","headings":[{"depth":2,"slug":"create-a-pool","text":"Create a pool"},{"depth":2,"slug":"enroll-a-host","text":"Enroll a host"},{"depth":2,"slug":"route-jobs-to-the-pool","text":"Route jobs to the pool"},{"depth":2,"slug":"forecast-demand","text":"Forecast demand"},{"depth":2,"slug":"automate-actions","text":"Automate actions"},{"depth":2,"slug":"see-your-cloud-accounts","text":"See your cloud accounts"}],"text":"Run jobs on your own machines\n\n Enroll your hosts into a pool, route jobs to them first, forecast demand and automate routine actions.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/pools/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA pool is a group of your own machines that Nodus schedules work onto before it rents anything. Nodus does not charge a compute rental fee for your hosts; you continue paying your cloud or hardware costs. Routing a job to a pool costs $0.02 per GPU-hour while a Nodus-scheduled attempt uses a GPU. CPU-only pool work has no device-hour fee. Predict capacity forecasts cost $99.00 per pool per month. You agree to each price when you turn it on.\n\nPools have four parts:\n\n Measure : enroll hosts with a one-line installer and see their inventory, utilization and health.\n Route : send jobs to the pool first, wait for it or burst to the market under rules you set.\n Predict : forecasts of demand and free capacity, and recommendations for sizing and wait policies.\n Act : routine actions such as reclaiming idle hosts or draining in a maintenance window, with approvals.\n\nCreate a pool\n\nTerminal window\n\nnodus create pool lab --routing prefer\n\nThe CLI and the console ask you to agree to the routing fee before the pool is created. In the console, open BYOCompute and choose New pool .\n\nEnroll a host\n\nA host needs Linux on x86-64 or arm64 with a running systemd service manager and cgroup v2 with CPU and memory controllers. It also needs running containerd at /run/containerd/containerd.sock , a healthy overlayfs snapshotter, and ctr , runc , containerd-shim-runc-v2 and timeout on its system path. Follow the containerd setup instructions if these are missing. GPU hosts also need working NVIDIA drivers and the NVIDIA Container Toolkit with CDI.\n\nThe installer checks these prerequisites before downloading the agent or consuming the enrollment token. It does not install system packages, reconfigure Docker or restart containerd. Download the installer from the URL in the enrollment command to install.sh , then check the host without enrolling it:\n\nTerminal window\n\nsudo sh ./install.sh --check\n\nA successful check confirms runtime prerequisites, not a completed GPU Job. Create an enrollment token for the pool; the response shows the installer once:\n\nTerminal window\n\nnodus create enrollmenttoken --pool lab\n\n --ttl 2h shortens the token’s life and --label rack=a copies a label onto the node when it enrolls. The token needs no name; the server gives it one.\n\nRun the command it prints as a user who can sudo on the host. It writes the token to a file only you can read, downloads the signed nodusd agent, checks its sha256 checksum, installs it as a systemd service and joins the pool. It also verifies the signature when cosign is installed; --require-signature makes that check mandatory. The token works once and expires after 24 hours; the console shows the countdown. Wait for the node to report Ready before submitting work. The installer is for enrollment; upgrading an existing host requires draining its work, replacing the verified agent binary and restarting nodusd .\n\nTerminal window\n\nnodus get nodes\nnodus get node/host-1 -o yaml\n\nEach job runs in a network namespace of its own, and nodusd keeps its network rules in the nftables table inet nodus . Outside that table it adds two rules to the host’s firewall, -i ndv+ -j ACCEPT and -o ndv+ -j ACCEPT . They match only its job interfaces, because Docker and ufw drop all forwarded traffic and jobs with open egress could not reach the internet otherwise. With Docker installed they go at the head of the DOCKER-USER chain; without Docker they go at the head of FORWARD , and only when its policy is DROP . Your rules for every other interface are untouched, and inet nodus still blocks private, link-local and metadata addresses for every job. nodusd puts the two rules back if they go missing and removes them when it stops.\n\nA node reports its GPUs, CPU, memory and disk, its utilization and a heartbeat. A node with no heartbeat for 10 minutes shows Offline and the pool’s members get one email about it.\n\nTo stop new work on a host without stopping what runs there, drain it; undrain it to take work again. Deleting a node drains it, then revokes its credential:\n\nTerminal window\n\nnodus node drain host-1\nnodus node undrain host-1\nnodus delete node/host-1\n\nRoute jobs to the pool\n\nName the pool in a job’s placement. With mode: Prefer the pool is used first and the market takes what the pool cannot; with mode: Only work never leaves the pool.\n\nIn the console, open your pool and choose Run a job . The job form selects your private pool under Where to run . Review the image, command and resources before launching. Your pool remains private to your organization.\n\nPrivate execution currently supports single-node Jobs, including indexed jobs. Each attempt reserves one whole host, so another attempt waits until that host is released. A GPU attempt sees only the GPUs it requested and pays the routing fee for those GPUs. Distributed jobs and other resource kinds cannot use this route yet.\n\nGPU jobs wait when the requested devices have existing GPU work, resident GPU memory, or no fresh utilization reading. After Nodus assigns a device, keep other host processes off it until the job ends; the host owner still controls processes started outside Nodus.\n\napiVersion: nodus.dev/v1\nkind: Job\nmetadata:\n name: train\nspec:\n image: nodus/pytorch:2.8-cuda12.8\n command: [python, train.py]\n resources: {gpu: \"H100:1\"}\n placement: {pool: lab}\n\nThe pool’s routing decides what happens when it is full:\n\n Field Values Meaning \n - - - \n waitPolicy Never , Cheaper , Timeout Go to the market at once, wait while waiting costs less than the market, or wait up to waitTimeout . \n burstToMarket Allow , Approve , Deny Burst at market rates, wait for an approval, or never burst. \n\nWith Approve , a job that would burst shows BurstApprovalRequired until someone approves it:\n\nTerminal window\n\nnodus request approve-burst job/train\n\nA job on the pool is billed only the routing fee for its device-hours. A burst to the market is billed like any other job.\n\nForecast demand\n\nPool Predict is currently unavailable while the updated host telemetry is being qualified. Enabling it is rejected before charging. Cloud-spend projections in BYOCompute are available separately.\n\nTurn on Predict on the pool’s Forecast tab or with spec.predict.enabled: true . Forecasts start once the pool has 7 days of utilization history, then refresh hourly; recommendations refresh daily.\n\nTerminal window\n\nnodus pool utilization lab\nnodus pool forecast lab --horizon 7d\nnodus pool recommendations lab\n\nThe forecast gives demand and free capacity in GPUs per hour with p50 and p90 bands. Recommendations cover right-sizing, idle hosts, maintenance windows, fragmentation and wait-policy tuning. Dismissing one records your reason, and the same recommendation is not raised again. Turning Predict off stops the charge at once.\n\nHost utilization history requires an updated nodusd reporting GPU readings. Missing samples and long gaps do not establish idle capacity, and upgrading does not reconstruct past readings. The BYOCompute Forecast section also shows cloud-spend projections; those use your cloud billing history and are separate from Pool Predict capacity forecasts.\n\nAutomate actions\n\n spec.act sets the mode and the policies. Every action is a PoolAction you can list, and every executed action is audited.\n\n Mode What happens \n - - \n Off Nothing is proposed. \n Shadow Records what would run and runs nothing. \n Propose Each action waits for approval and expires after 24 hours. \n Auto Runs actions within the policies at once. \n\nTerminal window\n\nnodus get poolactions\nnodus pool approve lab-idle-1\nnodus pool reject lab-idle-1\nnodus pool revert lab-idle-1\n\nActions only drain and undrain hosts in the pool, adjust its routing within the policy’s bounds, or approve a burst. They never stop or delete your jobs. To stop all actions at once, pause them; the pool shows ActPaused until you resume:\n\nTerminal window\n\nnodus pool pause lab\nnodus pool resume lab\n\nSee your cloud accounts\n\nTo list GPU instances and spend in your AWS or GCP accounts and enroll them into pools, connect them read-only: see Connect your cloud accounts."},{"id":"docs/guides/pools/cloud-accounts","url":"https://nodus-platform-site.pages.dev/docs/guides/pools/cloud-accounts/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/pools/cloud-accounts.md","title":"Connect your cloud accounts","description":"Give Nodus read-only access to AWS or GCP to list GPU instances and spend, enroll instances into pools and revoke access at any time.","stage":"GA","headings":[{"depth":2,"slug":"connect-aws","text":"Connect AWS"},{"depth":2,"slug":"connect-gcp","text":"Connect GCP"},{"depth":2,"slug":"see-inventory-and-spend","text":"See inventory and spend"},{"depth":3,"slug":"forecast-your-spend","text":"Forecast your spend"},{"depth":3,"slug":"compare-a-workload","text":"Compare a workload"},{"depth":3,"slug":"run-on-an-observed-instance","text":"Run on an observed instance"},{"depth":2,"slug":"revoke-access","text":"Revoke access"}],"text":"Connect your cloud accounts\n\n Give Nodus read-only access to AWS or GCP to list GPU instances and spend, enroll instances into pools and revoke access at any time.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/pools/cloud-accounts/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA cloud account gives Nodus read-only access to your AWS account or GCP projects. Nodus lists your instances, their GPUs and your monthly spend, and you can enroll an instance into a pool. Nodus never creates, changes or deletes anything in your cloud account. Connecting a cloud account is free.\n\nConnect AWS\n\nAWS access is a role in your account that only Nodus can assume, and only with an external id unique to your cloud account.\n\n1. In the console, open BYOCompute › Spend › Manage cloud accounts and choose Connect AWS . Enter your 12-digit account ID and choose Continue to AWS . Nodus names the connection and prepares the role for you. Or create it from the CLI:\n\n Terminal window\n\n nodus create cloudaccount aws-main --provider aws --account-id 123456789012\n \n\n2. In the AWS tab, review the prefilled CloudFormation template, acknowledge the IAM permissions and choose Create stack . If the tab did not open, choose Open AWS approval in Nodus. You can also read the template with nodus get cloudaccount/aws-main --subresource onboarding . The stack creates a role whose policy allows only these calls:\n\n Call Used for \n - - \n ec2:DescribeRegions , ec2:DescribeInstances , ec2:DescribeInstanceTypes Instances and their GPUs \n ce:GetCostAndUsage Daily spend \n\n3. Return to Nodus after the stack finishes. Nodus saves the role ARN, checks access automatically and opens your account when it has read your instances. There is nothing to copy. A failed check shows the reason while AWS finishes applying the permissions. Checks pause after ten minutes; Check again resumes the saved connection. Retrying setup reuses the existing account and its trust identity.\n\nConnect GCP\n\nGCP access is an OAuth grant limited to read-only scopes: compute.readonly , bigquery.readonly (for a billing export) and cloud-platform.read-only .\n\n1. Choose Connect GCP , then Continue with Google . No account name, project ID or service account key is required. If you already know the IDs, you can enter them under Advanced instead.\n2. Approve read-only access on Google’s page. Nodus lists the projects you can access. Choose the projects to connect, then choose Connect projects . Only those projects are imported. The person who signs in must complete this selection within 15 minutes. If it expires, sign in again to resume the saved connection. Nodus refuses grants with broader scopes and encrypts the stored authorization.\n3. Inventory is imported independently of billing. For spend reports, enable a Cloud Billing export in the first configured project’s nodus billing BigQuery dataset. The consenting Google user needs BigQuery Job User on that project and BigQuery Data Viewer on the export. Nodus detects the export table; connecting does not create an export or backfill billing history that Google has not supplied.\n\nAzure is coming soon.\n\nSee inventory and spend\n\nTerminal window\n\nnodus get cloudaccounts\nnodus cloud inventory aws-main\n\nNodus syncs inventory hourly and spend daily. Open an account to see its observed month-to-date charges, service breakdown and daily amounts. The first spend sync reads the previous 35 days plus today’s partial day, so connecting in the middle of a month includes the earlier charges. Currencies stay separate; Nodus does not convert them. Missing days remain unknown. The date of the last successful observation stays visible when a refresh fails.\n\nAWS amounts use Cost Explorer’s unblended cost. GCP amounts use cost plus credits from the billing export, filtered to the connected projects. Charges outside those projects, including unassigned charges, are excluded. Neither is a final invoice. You continue paying your cloud directly; these amounts are separate from your Nodus balance.\n\nForecast your spend\n\nThe next-30-day estimate extends the recent daily average across at least seven consecutive usable billing days. Missing days and negative billing adjustments prevent a projection. AWS non-estimated observations are preferred; provisional AWS and GCP projections exclude today and yesterday to allow for reporting lag. Each projection shows its source, history length and last included date. It assumes the recent spending pattern continues.\n\nThe compute-service subtotal includes CPU and related service charges. It is not a GPU-only bill or a measured cost for an individual workload. A separate compute projection appears only when the source supports it. These cloud-bill projections are distinct from a pool’s paid Predict capacity forecasts.\n\nCompare a workload\n\nChoose Compare a workload with Nodus , select matching capacity, and enter your current cost for the same workload after discounts. Enter its expected runtime on Nodus and any additional transfer, storage or ongoing commitment costs. Nodus requests a Job estimate that includes startup and shutdown; it does not start compute. The difference can be a saving or an additional cost. Changing inputs or an expired quote requires a fresh estimate.\n\nThis comparison does not assume that your entire cloud bill can move to Nodus. Check that the hardware, runtime and work performed are comparable before making a migration decision.\n\nRun on an observed instance\n\nThe Inventory section lists each instance and whether it is enrolled. Choose Enroll instance , select a private pool in your current project, and create an install command. Run it as root on that instance. The command contains a single-use token valid for 24 hours and associates the resulting node with the observed cloud account and instance. See the host prerequisites before installing. The installer checks the container runtime and GPU requirements before using the token.\n\nConnecting a cloud account grants observation only. Installing the agent grants execution on that host for your organization. Neither action offers capacity to other Nodus customers.\n\nRevoke access\n\nDisconnecting stops observation at once and deletes the credential Nodus stored:\n\nTerminal window\n\nnodus delete cloudaccount/aws-main\n\nThen revoke it on your side too:\n\n AWS : delete the CloudFormation stack, which deletes the role.\n GCP : remove Nodus from your Google account’s third-party access.\n\nIf you revoke access on your side first, the next sync marks the account Synced=False with GrantRevoked , and the account’s members get one email about it."},{"id":"docs/guides/python","url":"https://nodus-platform-site.pages.dev/docs/guides/python/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/python.md","title":"Python SDK","description":"Install the nodus-compute package, sign in, and run your first Function, Sandbox and Job from Python.","stage":"GA","headings":[{"depth":2,"slug":"install-and-sign-in","text":"Install and sign in"},{"depth":2,"slug":"your-first-function","text":"Your first Function"},{"depth":2,"slug":"blocking-and-asyncio","text":"Blocking and asyncio"},{"depth":2,"slug":"errors","text":"Errors"},{"depth":2,"slug":"guides","text":"Guides"}],"text":"Python SDK\n\n Install the nodus-compute package, sign in, and run your first Function, Sandbox and Job from Python.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/python/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nThe nodus-compute package (import nodus ) runs Python functions, Sandboxes, Jobs and Workspaces on Nodus, and calls models through the inference data plane. It supports Python 3.10 to 3.13.\n\nInstall and sign in\n\nTerminal window\n\npip install nodus-compute\nnodus login\n\n nodus login opens the console in your browser and stores a key for each org you pick in ~/.nodus/config , which the SDK reads. On a server or in CI, create an API key in the console and set it instead:\n\nTerminal window\n\nexport NODUS API KEY=nodus sk live ...\nexport NODUS PROJECT=default # optional; the project to use\n\nThe SDK resolves credentials in this order: arguments to nodus.Client(...) , the variables NODUS API KEY , NODUS API URL , NODUS ORG , NODUS PROJECT , NODUS CONTEXT and NODUS CONFIG , then the current context of ~/.nodus/config .\n\nYour first Function\n\nexamples/python/quickstart/app.py\n\n\"\"\"Quickstart: one Function called three ways.\n\nRun it with nodus run examples/python/quickstart/app.py --n 10 .\n\"\"\"\n\nimport nodus\n\napp = nodus.App(\"quickstart\")\n\n@app.function(cpu=1, memory=\"1Gi\", max cost=1)\ndef square(x: int) - int:\n return x x\n\n@app.local entrypoint()\ndef main(n: int = 10) - None:\n print(\"remote:\", square.remote(7)) # one call; blocks for the result\n call = square.spawn(8) # start without waiting\n print(\"spawned:\", call.get(timeout=600))\n print(\"map:\", list(square.map(range(n)))) # one call per input, results in input order\n\nTerminal window\n\nnodus run examples/python/quickstart/app.py --n 10\n\n nodus run creates an ephemeral App, runs main on your machine and deletes the App when main returns. Each .remote() call runs on Nodus and returns its result. nodus deploy keeps the App so other programs can call it with nodus.Function.from name(\"quickstart\", \"square\") .\n\nBlocking and asyncio\n\nEvery call blocks by default. Each one also has an .aio form for asyncio code:\n\nasync def main():\n async with app.run.aio():\n print(await square.remote.aio(7))\n async for y in square.map.aio(range(10)):\n print(y)\n\nErrors\n\nEvery error is a nodus.errors.NodusError . API errors have one class per code ( NotFound , Invalid , InsufficientCredits , QuotaExceeded , …) and carry message , fix , docs and request id . An exception raised inside a Function is raised again on your side with its own type when that type can be imported, chained from a nodus.errors.RemoteError that holds the remote traceback.\n\nfrom nodus import errors\n\ntry:\n train.remote(3e-4)\nexcept errors.InsufficientCredits as e:\n print(e.needed usd, e.available usd, e.fix)\n\nGuides\n\n Functions and classes: @app.function , .remote , .map , @app.cls , deploys\n Sandboxes: exec, files, tunnels, snapshots\n Jobs: batch containers, outputs, multi-node gangs\n Images, Volumes and Secrets\n Workspaces: development machines\n Agents: durable runs, fan-out and groups\n Inference: the OpenAI and Anthropic clients\n The resource API: nodus.api for any kind"},{"id":"docs/guides/python/agents","url":"https://nodus-platform-site.pages.dev/docs/guides/python/agents/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/python/agents.md","title":"Agents","description":"Define an agent on Claude from Python, submit runs, run many in parallel, and read each run's answer and steps.","stage":"Beta","headings":[{"depth":2,"slug":"run-agents-in-parallel","text":"Run agents in parallel"},{"depth":2,"slug":"not-available-yet","text":"Not available yet"}],"text":"Agents\n\n Define an agent on Claude from Python, submit runs, run many in parallel, and read each run's answer and steps.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/python/agents/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n Beta: this feature may change.\n\nAn Agent is a definition: a system prompt, the Claude access it runs on and a cost cap. Each run is one conversation in its own sandbox. The Agents guide covers the model, the cost cap and keeping a run open for follow-ups; this page is the Python side.\n\nimport nodus\n\nagent = nodus.ClaudeAgent(\"helper\", system=\"You are careful.\", families=[\"haiku\", \"sonnet\"])\nprint(agent.remote(\"Use the shell to print the Python version\")) # runs to the end, returns the answer\n\nrun = agent.submit(\"Summarize the logs\", keep alive=True)\nrun.send(\"message\", \"Now the staging logs\")\nprint(run.answer())\nprint(run.steps())\nrun.cancel()\n\n ClaudeAgent creates the agent on first use. Pass api key secret=\"anthropic-key\" to use your own Anthropic key, or ClaudeAgent.from name(\"claude-assistant\", project=\"nodus\") to run the ready-made template.\n\nRun agents in parallel\n\nAn AgentGroup runs many runs of one agent at once, with a limit on how many run together, a cost cap for the whole group and dependencies between runs. Parallel agents are in Beta, like the rest of Agents. The Agents guide explains each limit and the YAML and CLI side.\n\nexamples/agents/parallel-agents/main.py\n\n\"\"\"Run questions in parallel under one AgentGroup, then fan a list out with agent.map.\n\nRun it with python examples/agents/parallel-agents/main.py .\n\"\"\"\n\nimport nodus\n\ndef main() - None:\n agent = nodus.ClaudeAgent(\n \"parallel-agents-py\",\n system=\"You answer every question in one short sentence.\",\n families=[\"haiku\"],\n max tokens=1024,\n per run max cost=\"0.10\",\n )\n agent.deploy()\n\n group = nodus.AgentGroup.create(\"parallel-agents-py\", agent, max active=2, max cost=\"0.60\")\n try:\n group.submit many(\n [\n {\"key\": \"a\", \"input\": \"What is the capital of France?\"},\n {\"key\": \"b\", \"input\": \"What is the capital of Japan?\"},\n # c starts only after a and b have succeeded.\n {\"key\": \"c\", \"input\": \"Say that both questions are answered.\", \"depends on\": [\"a\", \"b\"]},\n ]\n )\n group.seal()\n status = group.wait()\n print(f\"{status.phase}: {status.counts.succeeded} of {status.counts.total} runs succeeded\")\n for run in group.runs():\n print(run.name, \"- \", run.answer())\n finally:\n group.delete()\n\n # For a plain list of inputs, map does the same in one call and returns the answers in order.\n for answer in agent.map([\"What is 2 + 2?\", \"What is 3 + 3?\"], max active=2):\n print(answer)\n\nif name == \" main \":\n main()\n\n nodus.AgentGroup.create(name, agent, max active=, max pending=, max cost=) creates the group over a deployed agent. max cost caps the runs’ Claude usage together; their sandboxes are billed on their own. submit many takes tasks {key, input, depends on} , puts every task after the tasks it depends on, sends them in batches of up to 100 runs and returns the runs in your order. A cycle or a repeated key raises nodus.errors.Invalid before anything is sent.\n group.seal() says that no more runs are coming. A group finishes only once it is sealed, and group.wait() returns its status when it has. group.runs() returns the member runs, and each run’s answer() is its text.\n group.cancel() cancels the runs that have not finished. group.delete() removes the group and its runs.\n agent.map(inputs, max active=, order outputs=True, return exceptions=False) creates a group, runs one task for each input, yields each answer text and deletes the group at the end. A run that does not succeed raises nodus.errors.AgentRunFailed , or is yielded as the exception when you pass return exceptions=True .\n\nNot available yet\n\n nodus.Agent is the definition of a durable Python program ( @agent.entrypoint , @agent.step , ctx.step ). Runs execute on Claude, not as your code, so a definition that names source , an entrypoint, setup , secrets , env , network , models , cpu , memory , max cost or worker settings raises nodus.errors.Unsupported when it is deployed. Submitting a run with session key= or group= (use AgentGroup.submit many for group runs), AgentGroup.create(max held=, evaluation=) , group.results() , and a run’s outputs , resolve() , retry() , suspend() , resume() and children() raise it too. nodus.Agent(\"name\", image=..., per run max cost=...) and agent.submit(input, deadline=...) work as written, and run.result() and agent.remote(input) return the run’s answer text, the same as run.answer() ."},{"id":"docs/guides/python/api","url":"https://nodus-platform-site.pages.dev/docs/guides/python/api/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/python/api.md","title":"The resource API","description":"Read, watch, apply and patch any Nodus kind from Python with nodus.api, the generic client under the SDK.","stage":"GA","headings":[{"depth":2,"slug":"health","text":"Health"},{"depth":2,"slug":"clients-and-retries","text":"Clients and retries"}],"text":"The resource API\n\n Read, watch, apply and patch any Nodus kind from Python with nodus.api, the generic client under the SDK.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/python/api/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n nodus.api is the generic client the rest of the SDK is built on. It works with every kind, including ones the Modal-shaped layer does not wrap, and takes the same objects as nodus apply and kubectl.\n\nimport nodus\n\njob = nodus.api.get(\"Job\", \"finetune-llama\")\nfor obj in nodus.api.list(\"Job\", label selector={\"team\": \"nlp\"}):\n print(obj[\"metadata\"][\"name\"], obj[\"status\"][\"phase\"])\n\nfor event in nodus.api.watch(\"Job\", field selector={\"metadata.name\": \"finetune-llama\"}):\n print(event[\"type\"], event[\"object\"][\"status\"].get(\"phase\"))\n\nnodus.api.apply(\"job.yaml\") # three-way merge, like nodus apply\nnodus.api.patch(\"Job\", \"finetune-llama\", {\"spec\": {\"maxCostUSD\": \"80.00\"}})\nestimate, etag = nodus.api.estimate(obj) # dry-run\nnodus.api.create(obj, if match=etag) # launch exactly what was estimated\nfor line in nodus.api.logs(\"Job\", \"finetune-llama\", follow=True):\n print(line, end=\"\")\nnodus.api.delete(\"Job\", \"finetune-llama\")\n\nObjects are plain dicts shaped like the kinds in the API reference.\n\nHealth\n\nnodus.api.healthz() # True when the API process answers\nready = nodus.api.readyz()\nprint(ready.ready, ready.release, ready.degraded, ready.failing)\n\n readyz() returns what /readyz reports instead of raising when the server is not ready: release is the commit it runs, degraded lists checks that fail without making it unready, and failing names the check that did. Neither call retries, so a probe sees one answer.\n\nClients and retries\n\nclient = nodus.Client(context=\"acme\", project=\"nlp\") # or api key=, api url=, org=\nclient.get(\"Sandbox\", \"agent-1\")\nnodus.api.set default client(client) # what module-level calls and the rest of the SDK use\n\nReads and watches are retried on connection errors, rate limits and 503 responses, with exponential backoff that honours Retry-After , for up to 5 minutes. Every create sends an Idempotency-Key , so a retried create never makes a duplicate. When the outcome of a write is unknown, the raised error carries idempotency key : retry the same call with that key to finish it safely."},{"id":"docs/guides/python/functions","url":"https://nodus-platform-site.pages.dev/docs/guides/python/functions/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/python/functions.md","title":"Functions and classes","description":"Define Functions with @app.function, call them with .remote, .map and .spawn, keep state in @app.cls classes, and deploy Apps.","stage":"GA","headings":[{"depth":2,"slug":"define-a-function","text":"Define a Function"},{"depth":2,"slug":"call-it","text":"Call it"},{"depth":2,"slug":"deploy-and-look-up","text":"Deploy and look up"},{"depth":2,"slug":"classes","text":"Classes"},{"depth":2,"slug":"clustered-functions-beta","text":"Clustered Functions (Beta)"}],"text":"Functions and classes\n\n Define Functions with @app.function, call them with .remote, .map and .spawn, keep state in @app.cls classes, and deploy Apps.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/python/functions/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Function is a Python function that runs on Nodus workers. Workers start when calls arrive, stay warm for scaledown window and scale between min workers and max workers . Every call is a FunctionCall object, so you can look it up, wait for it later or cancel it.\n\nDefine a Function\n\nimport nodus\n\napp = nodus.App(\"finetune\")\nimage = nodus.Image.debian slim().pip install(\"torch\", \"transformers==4.57.6\")\n\n@app.function(gpu=\"H100\", image=image, timeout=\"6h\", checkpoint=\"/nodus/state\")\ndef train(lr: float) - dict:\n ...\n\n Argument What it sets \n - - \n gpu \"H100\" , \"H100:2\" , \"A100-80GB\" , \"H100!\" (exact variant), a list of alternatives, or nodus.GPU(...) \n cpu , memory , ephemeral disk Floors per worker; bare numbers are vCPUs and MiB \n image A nodus.Image ; defaults to nodus/python at your Python version \n secrets , volumes , env , network [nodus.Secret] , {\"/path\": nodus.Volume} , a dict, nodus.Egress.deny() / .allow(...) / .open() \n timeout , retries Per-call limit ( \"6h\" or seconds); retries for exceptions ( int or nodus.Retries(...) ) \n checkpoint A path (or True for /nodus/state ) saved and restored when a worker is lost \n interruptible , region , profile Allow interruptible capacity; region classes such as [\"us\", \"eu\"] ; Balanced , Cost or Speed \n max cost Not available for Functions yet: deploying with it is refused (Jobs and Sandboxes take it) \n min workers , max workers , scaledown window , target concurrency Worker pool sizing ( min containers and max containers also work) \n name The member name; the Function object is - \n\nDecorating does not contact Nodus, so importing the file has no side effects. The App’s code (the directory of the file, minus what .gitignore and .nodusignore exclude) is uploaded once per content hash when the App runs.\n\nCall it\n\nwith app.run(): # or: nodus run app.py\n result = train.remote(3e-4) # one call; blocks and returns the result\n call = train.spawn(1e-4) # starts a call and returns a handle\n print(call.get(timeout=3600))\n for y in train.map([1e-4, 3e-4]): # one call per input, results in input order\n print(y)\n print(train.estimate(3e-4)) # expected cost, start time and hold, without running\n\n .map( iterables, order outputs=True, return exceptions=False) creates calls in batches of 1,000. With return exceptions=True a failed input yields its exception instead of stopping the loop.\n .starmap(pairs) spreads each tuple into arguments; .for each(xs) runs and discards the results.\n nodus.FunctionCall.from name(name) finds a spawned call again, from any process.\n Arguments and results are serialized with cloudpickle. Values above 64 KiB travel as uploaded blobs.\n\nAn exception raised in the Function is raised again in your process as its own type when it can be imported there; otherwise you get nodus.errors.RemoteError . Either way the remote traceback is attached.\n\nDeploy and look up\n\nTerminal window\n\nnodus deploy app.py\n\ntrain = nodus.Function.from name(\"finetune\", \"train\")\nprint(train.remote(3e-4))\n\nA deploy updates the App in place: changed Functions roll their workers after in-flight calls finish, and Functions you removed from the file are deleted.\n\nClasses\n\n @app.cls turns a class into one Function. @nodus.enter() methods run once per worker, before its first call, so expensive setup such as loading a model happens once; @nodus.exit() methods run when the worker drains. Methods marked @nodus.method() get .remote() , .map() and .spawn() .\n\nexamples/python/classes/app.py\n\n\"\"\"A class whose model loads once per worker, then serves many calls.\n\nRun it with nodus run examples/python/classes/app.py .\n\"\"\"\n\nimport nodus\n\napp = nodus.App(\"classes\")\n\n@app.cls(cpu=2, memory=\"4Gi\", scaledown window=\"5m\", max cost=1)\nclass Greeter:\n @nodus.enter()\n def load(self) - None:\n # Runs once when a worker starts, before its first call: load weights or open connections here.\n self.greeting = \"hello\"\n\n @nodus.method()\n def greet(self, name: str) - str:\n return f\"{self.greeting}, {name}\"\n\n @nodus.exit()\n def close(self) - None:\n self.greeting = \"\"\n\n@app.local entrypoint()\ndef main() - None:\n greeter = Greeter()\n print(greeter.greet.remote(\"Ada\"))\n print(list(greeter.greet.map([\"Grace\", \"Linus\"])))\n\nClasses take no constructor arguments; configure them in the enter hook.\n\nClustered Functions (Beta)\n\nBeta\n\nClustered Functions need the distributed training beta for your org. Checkpointed resume of a gang is not yet qualified.\n\n @nodus.clustered(size=N) below @app.function runs each .remote() or .spawn() as a gang of N nodes. Every member runs the function; rank 0’s return value is the result. nodus.cluster.info() tells each member its rank, the gang size, the member addresses and the rendezvous address, so torch.distributed initializes from the environment.\n\nexamples/functions/clustered/app.py\n\n\"\"\"A clustered Function: each call runs as one gang of two nodes, and rank 0's return value is the result (Beta).\n\nRun it with nodus run examples/functions/clustered/app.py .\n\"\"\"\n\nimport nodus\n\napp = nodus.App(\"clustered\")\nimage = nodus.Image.from registry(\"nodus/pytorch:2.8-cuda12.8\")\n\n@app.function(gpu=\"H100\", image=image, timeout=\"30m\", max cost=1)\n@nodus.clustered(size=2)\ndef whoami() - dict:\n info = nodus.cluster.info() # rank, size, member addresses, rendezvous address and epoch\n print(f\"rank {info.rank} of {info.size}; rendezvous {info.master addr}:{info.master port}\")\n return {\"rank\": info.rank, \"size\": info.size}\n\n@app.local entrypoint()\ndef main() - None:\n result = whoami.remote()\n assert result == {\"rank\": 0, \"size\": 2}, result\n print(result)\n\n network ( Colocated , Regional , Global ) and transport ( Direct , Auto ) choose where members may be placed. Clustered Functions keep no warm workers, and .map() on them raises nodus.errors.Unsupported ."},{"id":"docs/guides/python/inference","url":"https://nodus-platform-site.pages.dev/docs/guides/python/inference/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/python/inference.md","title":"Inference","description":"Call catalog models with the official OpenAI and Anthropic clients through nodus.llm, and create named InferenceEndpoints.","stage":"GA","headings":[{"depth":2,"slug":"named-endpoints","text":"Named endpoints"}],"text":"Inference\n\n Call catalog models with the official OpenAI and Anthropic clients through nodus.llm, and create named InferenceEndpoints.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/python/inference/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nThe inference data plane speaks the OpenAI and Anthropic wire formats. nodus.llm returns the official clients, configured with your key and the Nodus base URL, so every feature of those SDKs works unchanged.\n\nTerminal window\n\npip install \"nodus-compute[openai]\" # or [anthropic]\n\nexamples/python/inference/chat.py\n\n\"\"\"Chat with a catalog model through the OpenAI-compatible inference data plane.\n\nNeeds pip install \"nodus-compute[openai]\" . Run it with python examples/python/inference/chat.py .\n\"\"\"\n\nimport nodus\n\ndef main() - None:\n client = nodus.llm.openai() # the official OpenAI client with your Nodus key and base URL\n reply = client.chat.completions.create(\n model=\"nodus/gpt-oss-20b\",\n messages=[{\"role\": \"user\", \"content\": \"Say hello in five words.\"}],\n max tokens=32,\n )\n print(reply.choices[0].message.content)\n\nif name == \" main \":\n main()\n\nclient = nodus.llm.openai(project=\"nlp\") # usage is attributed to the project\nclaude = nodus.llm.anthropic() # messages API\naclient = nodus.llm.async openai() # AsyncOpenAI\n\nInside a Job, Function or Sandbox, nodus.llm uses the container’s built-in proxy: no key is needed, and the calls are billed to, and capped by, the run that makes them.\n\nNamed endpoints\n\nAn InferenceEndpoint gives a model its own base URL with rate limits, allowed keys and a spending cap.\n\nep = nodus.InferenceEndpoint.create(\"support-bot\", \"nodus/gpt-oss-120b\", rpm=120, max concurrent=8, max cost=50)\nclient = ep.openai()\nprint(ep.usage()) # requests, tokens and cost over the last 24 hours"},{"id":"docs/guides/python/jobs","url":"https://nodus-platform-site.pages.dev/docs/guides/python/jobs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/python/jobs.md","title":"Jobs","description":"Run a container to completion from Python with your code uploaded, follow its logs, download its outputs and run multi-node gangs.","stage":"GA","headings":[{"depth":2,"slug":"create","text":"Create"},{"depth":2,"slug":"watch-and-collect","text":"Watch and collect"},{"depth":2,"slug":"multi-node-gangs-beta","text":"Multi-node gangs (Beta)"}],"text":"Jobs\n\n Run a container to completion from Python with your code uploaded, follow its logs, download its outputs and run multi-node gangs.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/python/jobs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Job runs a command to completion. Nodus picks the capacity, keeps its checkpoints and recovers it if the capacity is reclaimed; max cost caps what it may spend.\n\nexamples/python/jobs/main.py\n\n\"\"\"Upload this directory as a Job's source, follow its logs and download its output.\n\nRun it with python examples/python/jobs/main.py .\n\"\"\"\n\nfrom pathlib import Path\n\nimport nodus\n\nHERE = Path( file ).parent\n\ndef main() - None:\n job = nodus.Job.run(\n image=\"nodus/python:3.12\",\n command=[\"python\", \"train.py\"],\n source=HERE, # uploaded once as a content-addressed blob; .gitignore and .nodusignore apply\n cpu=2,\n memory=\"4Gi\",\n timeout=\"30m\",\n max cost=1,\n outputs={\"report\": \"/nodus/outputs/report.txt\"},\n )\n print(\"estimate:\", job.estimate())\n for line in job.logs(follow=True):\n print(line, end=\"\")\n job.wait() # raises nodus.errors.JobFailed with the exit code and log tail\n print(\"saved\", job.outputs[\"report\"].download(HERE / \"report.txt\"))\n\nif name == \" main \":\n main()\n\nCreate\n\njob = nodus.Job.run(\n name=\"finetune-llama\", # optional; a generated name otherwise\n image=\"nodus/pytorch:2.8-cuda12.8\",\n command=[\"python\", \"train.py\", \"--epochs\", \"3\"],\n source=\".\", # uploads the directory; or {\"repo\": \"acme/trainer\", \"ref\": \"main\"}\n gpu=\"H100\", secrets=[\"hf-token\"],\n max cost=40, timeout=\"12h\", expected duration=\"6h\",\n checkpoint=\"/nodus/state\", interruptible=True, region=[\"us\", \"eu\"],\n outputs={\"adapter\": \"/nodus/outputs/adapter\"},\n)\n\nThe source upload skips what .gitignore and .nodusignore exclude and refuses more than 500 MiB compressed unless you pass allow large source=True . Everything written under /nodus/outputs is collected when the Job succeeds.\n\nWatch and collect\n\nprint(job.estimate()) # expected cost p50 and p90, start time and the hold\nfor line in job.logs(follow=True):\n print(line, end=\"\")\njob.wait() # raises nodus.errors.JobFailed(job, exit code, log tail)\njob.outputs[\"adapter\"].download(\"./adapter\") # checks the sha256, then renames into place\n\n job.suspend() , job.resume() and job.cancel() change the Job’s state; job.attempts() lists every attempt with its placement; job.exec(\"nvidia-smi\") runs a command in the running container.\n\nMulti-node gangs (Beta)\n\nBeta\n\nGangs need the distributed training beta for your org. Checkpointed resume of a gang is not yet qualified.\n\nddp = nodus.Job.run(\n name=\"ddp-smoke\", image=\"nodus/pytorch:2.8-cuda12.8\", command=[\"torchrun\", \"train.py\"],\n gpu=\"H100\", max cost=10,\n distributed=nodus.Distributed(nodes=2, launcher=\"Torchrun\", network=\"Global\", transport=\"Auto\"),\n)\nfor line in ddp.logs(follow=True, rank=\"all\"): # every rank, prefixed [r0], [r1]\n print(line, end=\"\")\n\nEvery rank gets the rendezvous environment ( MASTER ADDR , WORLD SIZE , torchrun’s PET variables), and nodus.cluster.info() reads it for you. nodus.checkpoint.dcp.save(state, step) and .load(state) write and read torch.distributed.checkpoint shards to the gang’s checkpoint storage."},{"id":"docs/guides/python/sandboxes","url":"https://nodus-platform-site.pages.dev/docs/guides/python/sandboxes/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/python/sandboxes.md","title":"Sandboxes","description":"Create isolated containers from Python, run commands in them, move files and snapshot their filesystem.","stage":"GA","headings":[{"depth":2,"slug":"create","text":"Create"},{"depth":2,"slug":"run-commands","text":"Run commands"},{"depth":2,"slug":"files","text":"Files"},{"depth":2,"slug":"stop-start-and-snapshot","text":"Stop, start and snapshot"}],"text":"Sandboxes\n\n Create isolated containers from Python, run commands in them, move files and snapshot their filesystem.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/python/sandboxes/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Sandbox is an isolated, long-running container for agents and untrusted code. You drive it with commands and file operations; it stops when idle and starts again on the next command.\n\nexamples/python/sandbox/main.py\n\n\"\"\"Create a Sandbox, run commands in it, move files and delete it.\n\nRun it with python examples/python/sandbox/main.py .\n\"\"\"\n\nimport nodus\n\ndef main() - None:\n sb = nodus.Sandbox.create(\n cpu=1,\n memory=\"2Gi\",\n idle timeout=\"5m\",\n max cost=1,\n )\n try:\n p = sb.exec(\"python\", \"-c\", \"print(6 7)\")\n print(\"stdout:\", p.stdout.read().strip(), \"exit:\", p.wait())\n with sb.open(\"/workspace/notes.txt\", \"w\") as f:\n f.write(\"written from the SDK\\n\")\n print(sb.exec(\"cat\", \"/workspace/notes.txt\").stdout.read(), end=\"\")\n finally:\n sb.terminate()\n\nif name == \" main \":\n main()\n\nCreate\n\nsb = nodus.Sandbox.create(\n name=\"agent-1\", # optional; creating the same name again reconnects\n image=\"nodus/agent-tools\", # the default: Python 3.12, Node 22, git\n cpu=2, memory=\"4Gi\",\n timeout=\"24h\", idle timeout=\"5m\", on idle=\"stop\",\n network=nodus.Egress.open(), # outbound access; the default is nodus.Egress.deny()\n max cost=5,\n)\n\n create returns as soon as the Sandbox is admitted; commands wait for it to start. Egress is denied unless you open it. nodus.Sandbox.from name(\"agent-1\") reconnects and starts it again if you stopped it.\n\nNot available yet\n\nThese arguments raise nodus.errors.Unsupported before anything is sent: volumes= , ports= (and sb.tunnels.open() ), init= , service= , gpu= and an allow-list egress ( nodus.Egress.allow(...) ). image= takes a published image name, not an nodus.Image that Nodus builds. secrets=[...] is sent when you set it, and raises Unsupported while the API does not accept secrets on a Sandbox.\n\nRun commands\n\np = sb.exec(\"python\", \"-c\", \"print(6 7)\")\nprint(p.stdout.read()) # everything the process wrote to stdout\nassert p.wait() == 0 # the exit code\n\np = sb.exec(\"pip install requests && python app.py\") # one string runs under /bin/sh -c\nfor chunk in p.stdout: # stream output as it arrives\n print(chunk, end=\"\")\n\np = sb.exec(\"bash\", pty=True)\np.write(\"ls\\n\"); p.resize(40, 120); p.signal(\"SIGINT\")\n\nEvery command is recorded as a Process, and its output is kept, so a reader that reconnects continues where it stopped. p.stdin.write(...) and p.stdin.write eof() feed input; p.cancel() stops the process.\n\nFiles\n\nwith sb.open(\"/workspace/notes.txt\", \"w\") as f:\n f.write(\"hi\")\nprint(sb.files.read(\"/workspace/notes.txt\"))\nprint(sb.files.list(\"/workspace\"))\n\nStop, start and snapshot\n\nsb.stop() # saves the filesystem; compute billing stops, the saved files count as storage\nsb.start()\nimage = sb.snapshot filesystem() # Beta: an Image of the current filesystem\nsb.terminate() # deletes it\n\nBeta\n\n snapshot filesystem() is Beta. Memory is not captured: processes start fresh in a Sandbox created from the image. A Sandbox cannot start from the snapshot Image yet."},{"id":"docs/guides/python/storage","url":"https://nodus-platform-site.pages.dev/docs/guides/python/storage/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/python/storage.md","title":"Images, Volumes and Secrets","description":"Build container images from Python, share files between workers with Volumes, and pass credentials with Secrets.","stage":"GA","headings":[{"depth":2,"slug":"images","text":"Images"},{"depth":2,"slug":"volumes","text":"Volumes"},{"depth":2,"slug":"secrets","text":"Secrets"}],"text":"Images, Volumes and Secrets\n\n Build container images from Python, share files between workers with Volumes, and pass credentials with Secrets.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/python/storage/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nImages\n\nimage = (\n nodus.Image.from registry(\"nodus/pytorch:2.8-cuda12.8\")\n .apt install(\"git\")\n .pip install(\"transformers==4.57.6\", \"peft\")\n .env({\"HF HUB ENABLE HF TRANSFER\": \"1\"})\n .add local python source(\"mylib\")\n)\nimage = nodus.Image.from dockerfile(\"Dockerfile\", context=\".\")\nimage = nodus.Image.debian slim(\"3.12\").uv pip install(\"numpy\")\n\nNothing is built until an App that uses the image runs, or until you call image.build() , which streams the build log. An Image is named by the hash of its steps, so the same chain reuses the same build. add local file , add local dir and add local python source upload local files into the image; uv sync(\".\") installs a uv project from its lockfile.\n\nVolumes\n\nvol = nodus.Volume.from name(\"data\", create if missing=True)\nvol.put file(\"./local.csv\", \"/train/local.csv\")\nprint(vol.listdir(\"/train\"))\nwith vol.batch upload() as up:\n up.put directory(\"./dataset\", \"/train\")\nweights = nodus.Volume.import from(\"llama\", huggingface=\"meta-llama/Llama-3.1-8B\", revision=\"0e9e39f\")\n\nA Volume made by create if missing=True is ReadWriteMany , with Modal’s semantics: each worker mounts the latest revision when it starts, vol.commit() publishes that worker’s changed files as a new revision (the last writer wins per path), and vol.reload() mounts the newest revision. Close open files before reload() . Uploads and downloads use the nodus CLI that the package installs.\n\nexamples/python/storage/app.py\n\n\"\"\"Workers share a ReadWriteMany Volume: each commits its result, and a reload sees everyone's.\n\nRun it with nodus run examples/python/storage/app.py .\n\"\"\"\n\nfrom pathlib import Path\n\nimport nodus\n\napp = nodus.App(\"storage\")\nresults = nodus.Volume.from name(\"storage-example\", create if missing=True)\ntoken = nodus.Secret.from dict({\"GREETING\": \"hello\"})\n\n@app.function(volumes={\"/results\": results}, secrets=[token], max workers=4, max cost=1)\ndef work(i: int) - str:\n import os\n\n Path(f\"/results/{i}.txt\").write text(f\"{os.environ['GREETING']} {i}\\n\")\n results.commit() # publish this worker's files as a new revision\n return f\"{i}.txt\"\n\n@app.function(volumes={\"/results\": results}, max cost=1)\ndef collect() - list[str]:\n results.reload() # mount the latest revision\n return sorted(p.name for p in Path(\"/results\").iterdir())\n\n@app.local entrypoint()\ndef main() - None:\n print(list(work.map(range(4))))\n print(collect.remote())\n\nSecrets\n\nhf = nodus.Secret.from name(\"hf-token\") # an existing Secret\ncfg = nodus.Secret.from dict({\"API TOKEN\": \"...\"}) # created with the App and deleted with it\nenv = nodus.Secret.from dotenv(\".env\")\nnodus.Secret.create(\"hf-token\", {\"HF TOKEN\": \"hf ...\"}) # writes a new version if it exists\n\nEvery key arrives in the container as an environment variable and as a file under /run/secrets/ / . Values are never returned by the API, and running containers keep the version they started with."},{"id":"docs/guides/python/workspaces","url":"https://nodus-platform-site.pages.dev/docs/guides/python/workspaces/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/python/workspaces.md","title":"Workspaces","description":"Create a development machine with SSH, VS Code, JupyterLab and a persistent home from Python.","stage":"GA","headings":[],"text":"Workspaces\n\n Create a development machine with SSH, VS Code, JupyterLab and a persistent home from Python.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/python/workspaces/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Workspace is a development machine with a persistent home directory. It stops when idle, and its home is saved when it stops.\n\nlab = nodus.Workspace.create(\"lab\", gpu=\"H100:2\", image=\"nodus/workspace-pytorch-cuda\", idle timeout=\"2h\")\nlab.wait ready()\nurl = lab.open(\"vscode\") # or \"jupyter\"; opens the browser and returns the URL\nlab.ssh() # an SSH session through the nodus CLI\nlab.schedule(ready by=\"2026-10-01T08:45:00-07:00\", stop at=\"2026-10-01T19:00:00-07:00\")\nlab.stop()\nfor s in lab.sessions(): # one entry per running period, with what it cost\n print(s.billed seconds, s.charge usd, s.stop reason)\n\n The home Volume is volume= or -home , created if it does not exist. ephemeral=True skips it: files are lost on stop.\n lab.start() wakes it after an idle or schedule stop; an SSH connection also wakes it.\n lab.exec(...) and lab.files work as on Sandboxes."},{"id":"docs/guides/sandboxes","url":"https://nodus-platform-site.pages.dev/docs/guides/sandboxes/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/sandboxes.md","title":"Sandboxes","description":"Create an isolated container, run commands in it with streaming output, read and write files, and stop it to stop paying for compute.","stage":"GA","headings":[{"depth":2,"slug":"quick-start","text":"Quick start"},{"depth":2,"slug":"create-a-sandbox","text":"Create a Sandbox"},{"depth":3,"slug":"create-by-name-reconnects","text":"Create by name reconnects"},{"depth":2,"slug":"run-commands","text":"Run commands"},{"depth":3,"slug":"open-a-shell","text":"Open a shell"},{"depth":3,"slug":"timeouts","text":"Timeouts"},{"depth":2,"slug":"files","text":"Files"},{"depth":3,"slug":"writing-without-overwriting-someone-elses-change","text":"Writing without overwriting someone else’s change"},{"depth":2,"slug":"stop-start-and-wake","text":"Stop, start and wake"},{"depth":3,"slug":"sandboxstarting-and-retry-after","text":"SandboxStarting and Retry-After"},{"depth":2,"slug":"idle-onidle-and-maxlifetime","text":"Idle, onIdle and maxLifetime"},{"depth":2,"slug":"network-access","text":"Network access"},{"depth":2,"slug":"what-a-sandbox-costs","text":"What a Sandbox costs"},{"depth":2,"slug":"next-steps","text":"Next steps"}],"text":"Sandboxes\n\n Create an isolated container, run commands in it with streaming output, read and write files, and stop it to stop paying for compute.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/sandboxes/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Sandbox is an isolated, long-running container for agents and untrusted code. You drive it with commands and file requests. When nothing uses it for a while it stops, keeps its /workspace directory and stops costing compute; the next command or file request starts it again. You pay per second while it holds compute and nothing while it is stopped.\n\nSandboxes run on CPU machines Nodus operates, each in its own isolated runtime with no network access unless you open it. Sandbox isolation explains what keeps one Sandbox from another.\n\nQuick start\n\nThis manifest is the whole Sandbox. It asks for 1 vCPU and 2 GiB of memory, stops after 5 idle minutes and caps its spending at 1 USD:\n\nsandbox.yaml\n\napiVersion: nodus.dev/v1\nkind: Sandbox\nmetadata:\n name: hello\nspec:\n image: nodus/agent-tools # required: Python 3.12, Node 22 and git\n resources:\n cpu: \"1\"\n memory: 2Gi\n network:\n egress:\n policy: Deny # the default: no outbound traffic\n lifecycle:\n idleTimeout: 5m # stop after 5 minutes with no exec, file request or running process\n onIdle: Stop # keep /workspace and stop paying for compute; Delete removes the Sandbox instead\n maxLifetime: 2h # delete it 2 hours after creation whatever it is doing\n maxCostUSD: \"1.00\"\n\nCreate it, run a command, copy a file in and read it back:\n\nTerminal window\n\n$ nodus apply -f sandbox.yaml\nsandbox/hello created\n$ nodus exec sb/hello -- echo hello from a sandbox\nhello from a sandbox\n$ nodus cp ./notes.txt sb/hello:/workspace/notes.txt\n$ nodus exec sb/hello -- cat /workspace/notes.txt\nwritten from my laptop\n$ nodus stop sb/hello\n$ nodus exec sb/hello -- cat /workspace/notes.txt # starts it again; /workspace is kept\nwritten from my laptop\n$ nodus delete sb/hello\n\nFrom Python, the same steps are in the Python Sandboxes guide.\n\nCreate a Sandbox\n\n nodus apply -f sandbox.yaml , nodus create sandbox and the SDK all create the same object. image is required on the API; the CLI and the SDK fill in nodus/agent-tools (Python 3.12, Node 22 and git) when you leave it out.\n\n Field Default Accepted values \n - - - \n spec.image — (required) Any image reference, or a Nodus catalog image such as nodus/agent-tools \n spec.resources.cpu 1 0.25 to 64 vCPU \n spec.resources.memory 1Gi 128 MiB to 256 GiB \n spec.resources.disk 10Gi 1 GiB to 200 GiB \n spec.network.egress.policy Deny Deny (no outbound traffic) or Open (public internet) \n spec.lifecycle.idleTimeout 5m 0s (never idle) or 1 minute to 24 hours \n spec.lifecycle.onIdle Stop Stop or Delete \n spec.lifecycle.maxLifetime 24h 1 minute to 720 hours \n spec.continuity.mode Snapshotted Snapshotted (keep /workspace across stops) or Ephemeral (start empty) \n spec.workingDir /workspace An absolute path \n spec.maxCostUSD none A USD amount such as \"5.00\" ; the Sandbox stops when it has cost this much \n spec.state Running Running or Stopped \n\n nodus get sb/hello shows PHASE ( Pending , Starting , Running , Stopping , Stopped , Recovering , Terminating , Failed ), ACTIVITY ( Busy or Idle ), the CPU , MEMORY and GPU you asked for, and COST so far.\n\nCreate by name reconnects\n\nCreating a Sandbox whose name already exists in the project does not make a second one:\n\n Same spec. You get the existing Sandbox back, and if it is stopped it starts again. This is how an agent reconnects to its Sandbox after a restart: it creates the same manifest every time.\n A maxCostUSD that is added or raised. The new budget is applied; nothing else changes. A budget can only be raised: a lower one, or none when the Sandbox has one, counts as a different spec.\n Any other change. The request fails with 409 AlreadyExists and lists the fields that differ in details.diff and details.causes . Delete the Sandbox and create it again to change its image, shape or lifecycle.\n\nGPU Sandboxes are not available yet: a spec with resources.gpu is refused.\n\nRun commands\n\n nodus exec runs one command in the running Sandbox and streams its standard output and standard error back as the command writes them. The exit code of nodus exec is the command’s. Every command is recorded as a Process of the Sandbox, whatever started it (the CLI, an agent or the MCP tools): its name, such as sb-hello-12 , comes back with the result, and nodus get processes lists it with how it ended.\n\nTerminal window\n\n$ nodus exec sb/hello -- python -c 'print(6 7)'\n42\n$ nodus exec sb/hello -- sh -c 'pip list 2 /dev/null head -3; exit 3'; echo \"exit $?\"\nPackage Version\n---------- -------\npip 24.2\nexit 3\n\nA command runs in spec.workingDir ( /workspace unless you set it) with the Sandbox’s environment, as a non-root user. Pass a single string to sh -c to use pipes and redirection. Every command counts as activity, so a Sandbox never stops as idle while a command is running.\n\nOpen a shell\n\n nodus shell is nodus create sandbox and nodus exec -it in one step: it opens an interactive terminal in a Sandbox and creates the Sandbox first when the name is new. It takes the same flags as nodus create sandbox , and the image is nodus/agent-tools unless you pass --image .\n\nTerminal window\n\n$ nodus shell sandbox/dev --cpu 2 --memory 4Gi\nsandbox/dev created\nsandboxes/dev is starting; waiting for it\n\nYou are now in a terminal in the Sandbox, in /workspace . When you exit, nodus shell sandbox/dev opens it again with /workspace as you left it.\n\nThe Sandbox stays after you exit, and stops by itself when idle. Pass --rm to delete a Sandbox this command created when the shell exits; --rm refuses a Sandbox that already exists, and so does a flag such as --cpu , because neither changes a Sandbox that is there. nodus shell with no name gives the shell a Sandbox of its own and deletes it on exit unless you pass --keep . Put a command after -- to run it instead of bash : nodus shell sandbox/dev -- zsh . With input that is not a terminal, the shell runs without one, so echo 'make test' nodus shell sandbox/dev works in a script. The exit code is the shell’s.\n\nTimeouts\n\nEach command has its own timeout: 10 minutes unless you set one, and at most 24 hours. When the timeout passes, the process is killed with SIGKILL and the command ends with the reason DeadlineExceeded . The Sandbox keeps running. Set the timeout per command, for example sb.exec(\"make\", \"test\", timeout=\"2m\") in Python or \"timeout\": \"2m\" in the API request.\n\nHow a command ended is in its result:\n\n Reason Meaning \n - - \n none Exit code 0 \n NonZeroExit The command exited with another code \n Signaled A signal ended it; the exit code is 128 plus the signal number \n DeadlineExceeded Its timeout passed and it was killed \n StartFailed It could not start, for example because the program does not exist \n NodeLost The machine running the Sandbox was lost while the command ran \n OutputLimitExceeded It wrote more output than a Process may keep \n ParentStopped The Sandbox stopped or was deleted while the command ran \n Cancelled You cancelled it \n\nA command line holds at most 1,024 arguments and 64 KiB in total, plus up to 256 extra environment variables. A Sandbox runs at most 1,024 commands at once; the next one fails with 429 TooManyRequests until one ends.\n\nFiles\n\n nodus cp copies files in both directions, and works on every image because the file service is part of the Sandbox runtime, not of the image:\n\nTerminal window\n\n$ nodus cp ./notes.txt sb/hello:/workspace/notes.txt # laptop to Sandbox\n$ nodus cp sb/hello:/workspace/result.json ./result.json # Sandbox to laptop\n\nThe same operations are one request each on /apis/nodus.dev/v1/namespaces/{project}/sandboxes/{name}/files :\n\n Request Does \n - - \n GET …/files?path=/workspace/a.txt Returns the file’s bytes \n PUT …/files?path=/workspace/a.txt Replaces the file with the request body, creating parent directories \n DELETE …/files?path=/workspace/a.txt Deletes a file or an empty directory \n\nA write replaces the file atomically: a reader sees the old content or the new content, never half of it. A file is at most 64 MiB per request; a larger one fails with 413 FileTooLarge . Paths are absolute, and a path that does not exist answers 404 NotFound .\n\nWriting without overwriting someone else’s change\n\nWhen two writers share a file, send the digest of the content you started from as expectedSHA256 :\n\nTerminal window\n\n$ curl -X PUT --data-binary @plan.md \"$API/sandboxes/hello/files?path=/workspace/plan.md&expectedSHA256=$(sha256sum old-plan.md cut -d' ' -f1)\"\n\nThe write happens only if the file still has that digest. If it changed, nothing is written and the request fails with 412 SHA256Mismatch : read the file again, merge, and retry. The digest is 64 lowercase hex digits. The digest of empty content ( e3b0c442…b855 ) matches a file that does not exist yet, so it means “create only”.\n\nStop, start and wake\n\nA Sandbox is Running while it holds compute and Stopped while it does not. A stop saves /workspace first, so a later start gives it back, as long as spec.continuity.mode is Snapshotted , the default.\n\nTerminal window\n\n$ nodus stop sb/hello # saves /workspace, releases compute, ends the charge\n$ nodus start sb/hello # starts it again with /workspace restored\n$ nodus get sb/hello\nNAME PHASE ACTIVITY CPU MEMORY GPU COST AGE\nhello Stopped Idle 1 1Gi - $0.0042 12m\n\n nodus stop sets spec.state: Stopped and the Sandbox stays stopped until you start it. Every other stop is made by the platform, shows its reason in status.stopReason , and ends the next time something needs the Sandbox:\n\n stopReason Why it stopped \n - - \n User You ran nodus stop \n Idle Nothing used it for idleTimeout \n InsufficientCredits Your balance could not fund the next period \n BudgetExceeded A Budget on the project reached its limit \n MaxCostReached The Sandbox cost as much as spec.maxCostUSD \n\nA stopped Sandbox starts again on any of these wake triggers:\n\n a command ( nodus exec , or an exec request),\n a file request,\n nodus start ,\n creating a Sandbox with the same name and an identical spec.\n\nCredits and budgets still apply to a wake. A Sandbox stopped for InsufficientCredits , BudgetExceeded or MaxCostReached stays stopped until the balance, the Budget or maxCostUSD allows the next period of compute. Until then a command, a file request and nodus start answer 402 instead of SandboxStarting : InsufficientCredits states the hold the start needs and what you have, and BudgetExceeded the Budget or maxCostUSD that is full. Add credits, raise the Budget or raise maxCostUSD , then send a trigger again. A create by name records its wake too, but returns the Sandbox as it is: still stopped, with Funded false in its conditions.\n\n SandboxStarting and Retry-After \n\nA command or file request that finds the Sandbox stopped starts it and answers 503 SandboxStarting with a Retry-After header (2 seconds), because the container is not ready yet. The CLI and the SDK wait and retry for you, so nodus exec on a stopped Sandbox just takes a few seconds longer. If you call the API directly, retry after the time in Retry-After ; each retry is also a wake trigger, so it does no harm to send them early.\n\nOther reasons a request can fail while the Sandbox is not usable:\n\n Status and reason Meaning \n - - \n 409 SandboxNotRunning The Sandbox is being deleted \n 409 SandboxFailed The Sandbox failed to start or lost its machine; status.reason says why \n 402 InsufficientCredits The Sandbox is stopped for lack of credits and cannot start yet \n 402 BudgetExceeded The Sandbox is stopped because a Budget or its maxCostUSD is full and cannot start yet \n 404 NotFound There is no Sandbox with that name \n\nEach code has a page in the error reference.\n\nIdle, onIdle and maxLifetime \n\nA Sandbox is idle when no command is running, no file request arrived for idleTimeout and no terminal had input in that time. A running process, including one that prints nothing for hours, counts as busy; a terminal that is open but silent does not count as activity.\n\nWhen the idle time passes, onIdle decides what happens:\n\n Stop (the default) saves /workspace , releases compute and sets stopReason: Idle . The next wake trigger starts it again.\n Delete removes the Sandbox with its files. Use it for one-off work that leaves nothing worth keeping.\n\n idleTimeout: 0s turns idle detection off, so the Sandbox runs until you stop it, until maxLifetime or until a budget ends it.\n\n maxLifetime is a hard deadline counted from creation, 24 hours by default. When it passes, the Sandbox is deleted whatever it is doing, running or stopped. It is the backstop that keeps a forgotten Sandbox from living forever. Create a new Sandbox, or one with a longer maxLifetime , to carry on.\n\nNetwork access\n\nA Sandbox has no outbound network access unless you open it. spec.network.egress.policy is one of:\n\n Deny (the default): nothing leaves the Sandbox, so code in it cannot download or upload anything.\n Open : the Sandbox can reach the public internet. It still cannot reach other Sandboxes, the machine it runs on or Nodus’s own services.\n\nA Sandbox accepts no inbound connections; you reach it only through commands and file requests.\n\nWhat a Sandbox costs\n\nA Sandbox is billed per second, from the moment its machine slot is reserved until it is released, at the rate of its shape. The rate is the Nodus machine price for the CPU, memory and disk you asked for, and it already includes Nodus’s 12.5 % (how pricing works). Four things to know:\n\n Starting and stopping are billed. The seconds spent starting (or restoring /workspace ) and the seconds spent saving a stop are part of the reservation, shown as their own lines on the bill.\n Stopped time is free. A stopped Sandbox holds no compute, so no compute cost accrues. The saved /workspace counts as storage, which is metered separately from compute.\n The rate is frozen per start. The rate a start shows is the rate that start pays, even if prices change while it runs.\n Shared capacity has limits. New shared capacity starts for you at most four times an hour, and none starts for 30 minutes after capacity started for you goes unused; until your organization’s first purchase it starts only for Sandboxes of up to 3 vCPU and 11 GiB. Otherwise a Sandbox runs on a machine of its own, at that machine’s rate.\n\n nodus get sb/hello shows the running total in the COST column. Set spec.maxCostUSD to cap one Sandbox: it stops when it has cost that much, and you can raise the cap with another create by name. Without a cap, your credit balance and any Budget on the project are the limits.\n\nNext steps\n\n Sandbox isolation for what separates your Sandbox from the machine and from other Sandboxes.\n Sandboxes from Python for the SDK, including terminals and previews."},{"id":"docs/guides/secrets","url":"https://nodus-platform-site.pages.dev/docs/guides/secrets/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/secrets.md","title":"Secrets","description":"Store API tokens, credentials and registry logins once, and pass them to Jobs, Sandboxes, Functions and Agents as env vars and files.","stage":"GA","headings":[{"depth":2,"slug":"create-a-secret","text":"Create a Secret"},{"depth":2,"slug":"use-it-in-a-job","text":"Use it in a Job"},{"depth":2,"slug":"change-a-value","text":"Change a value"},{"depth":2,"slug":"private-registries","text":"Private registries"},{"depth":2,"slug":"limits-and-errors","text":"Limits and errors"}],"text":"Secrets\n\n Store API tokens, credentials and registry logins once, and pass them to Jobs, Sandboxes, Functions and Agents as env vars and files.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/secrets/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Secret holds named values, such as an API token or a database URL, encrypted with a key that belongs to your org. You create it once and name it in the Jobs and Sandboxes that need it. Nodus never returns a value through the API, the console or the CLI, and it redacts secret values from logs and outputs.\n\nCreate a Secret\n\nTerminal window\n\nnodus secret create hf-token --from-literal HF TOKEN=hf xxx\nnodus secret create app-env --from-env-file .env\nnodus secret create registry --type Registry \\\n --from-literal server=ghcr.io --from-literal username=ada --from-literal password=ghp xxx\n\nKey names start with a letter or , use letters, digits and , and may not start with NODUS . A Secret holds up to 64 keys of at most 4 KiB each. nodus get secret hf-token shows the key names and the current version, never the values.\n\nUse it in a Job\n\nList the Secret under secrets to receive every key as an env var and as a read-only file /run/secrets/ / , or pick one key under another name with valueFrom.secretKeyRef :\n\njob.yaml\n\napiVersion: nodus.dev/v1\nkind: Job\nmetadata:\n name: ex-secrets-env\nspec:\n image: nodus/python:3.12\n secrets: [ex-secrets-env] # every key as an env var and a file under /run/secrets/ex-secrets-env/\n env:\n - name: TOKEN # one key under another name\n valueFrom: {secretKeyRef: {name: ex-secrets-env, key: API TOKEN}}\n command: [python, check.py]\n\ncheck.py\n\nimport os\n\ntoken = os.environ[\"API TOKEN\"]\nwith open(\"/run/secrets/ex-secrets-env/API TOKEN\") as f:\n from file = f.read()\n\nPrint facts about the value, never the value: Nodus redacts secret values from logs anyway.\nprint(f\"token has {len(token)} characters\")\nprint(\"file matches\" if from file == token else \"file differs\")\nprint(\"alias matches\" if os.environ[\"TOKEN\"] == token else \"alias differs\")\n\nTerminal window\n\nnodus apply -f job.yaml\nnodus logs job/ex-secrets-env\n\nThe same secrets and env fields work on Sandboxes, Workspaces, Functions and Agents. In Python:\n\nhf = nodus.Secret.from name(\"hf-token\")\nenv = nodus.Secret.from dotenv(\".env\")\n\nA command whose literal env value equals one of the Secret values in its project is refused with SecretValueInEnv : reference the Secret instead of pasting its value. The same check applies to commands you start in a running Sandbox, Workspace or Job with nodus exec or the Processes API, which also refuse an env that sets the same name twice. Values shorter than 8 characters, and the server and username of a Registry Secret, are not checked.\n\nA Job receives the Secret versions pinned when you submitted it, and a Sandbox the versions pinned when it started. If a Secret it uses is deleted (even if you create a new one with the same name), or a new value replaced a Job’s pinned version before a retry started, that attempt fails with LaunchFailed . Secret values delivered as env vars may total at most 1 MiB per container, and as files at most 8 MiB.\n\nChange a value\n\nWriting a Secret again creates a new version; writing the same values again changes nothing. A Sandbox pins the newest version each time it starts, so it gets the new value on its next start, while a running Sandbox keeps the version it started with. A Job keeps the versions pinned when you submitted it, so every attempt runs with the same values; a retry that starts after you replace a value fails with LaunchFailed , so submit the Job again to use it. Earlier versions are deleted 7 days after they are replaced.\n\nhf-token.yaml\n\napiVersion: nodus.dev/v1\nkind: Secret\nmetadata: {name: hf-token}\nspec:\n stringData: {HF TOKEN: hf new}\n\nTerminal window\n\nnodus apply -f hf-token.yaml\n\nPrivate registries\n\nA Secret with type: Registry holds server , username and password . Name it under imagePullSecrets to run or build from a private image:\n\nspec:\n image: ghcr.io/acme/trainer:1.4\n imagePullSecrets: [{name: registry}]\n\nNodus resolves the tag to a digest when you submit, using the Secret, and the machine that runs the work pulls that digest with the same Secret version. The credential is used for the pull alone: it is not an env var or a file in the container. When the work runs on a provider that pulls containers itself, Nodus first copies the image by digest into your org’s space in the Nodus registry, so your registry credential never leaves Nodus.\n\nLimits and errors\n\n Error Meaning Fix \n - - - \n SecretValueInEnv An env value equals a Secret value Use valueFrom.secretKeyRef or secrets \n EncryptionFailed The value could not be encrypted Retry; the Secret keeps its previous version \n ImagePullFailed The registry refused the pull Secret Check server , username and password"},{"id":"docs/guides/sweeps","url":"https://nodus-platform-site.pages.dev/docs/guides/sweeps/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/sweeps.md","title":"Sweeps","description":"Run one Job across GPU types, regions and parameter values, and compare cost, time and throughput per cell.","stage":"GA","headings":[{"depth":2,"slug":"submit-a-sweep","text":"Submit a Sweep"},{"depth":2,"slug":"watch-it-and-read-the-results","text":"Watch it and read the results"},{"depth":2,"slug":"the-matrix","text":"The matrix"},{"depth":2,"slug":"from-python","text":"From Python"},{"depth":2,"slug":"running-and-failing-cells","text":"Running and failing cells"}],"text":"Sweeps\n\n Run one Job across GPU types, regions and parameter values, and compare cost, time and throughput per cell.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/sweeps/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Sweep runs the same Job once for every combination of the values you list: GPU types, region classes and your own parameters. Each combination is a cell. When the cells finish, the Sweep shows what each one cost, how long it took and how fast it went, and points out the cheapest and the fastest.\n\nSubmit a Sweep\n\nThis Sweep measures training throughput for three batch sizes on two GPU types, six cells in all:\n\nsweep.yaml\n\napiVersion: nodus.dev/v1\nkind: Sweep\nmetadata:\n name: batch-size\nspec:\n maxCostUSD: \"0.75\"\n maxParallel: 2\n # 2 GPU types x 3 batch sizes = 6 cells, each a Job named batch-size- .\n matrix:\n gpu: [L4, A10]\n params:\n BATCH SIZE: [\"32\", \"64\", \"128\"]\n template:\n kind: Job\n spec:\n image: nodus/pytorch\n timeout: 10m\n command:\n - python\n - -c\n - \n import os, time, torch\n bs = int(os.environ[\"NODUS PARAM BATCH SIZE\"])\n model = torch.nn.Sequential(torch.nn.Linear(1024, 4096), torch.nn.ReLU(), torch.nn.Linear(4096, 10)).cuda()\n opt = torch.optim.AdamW(model.parameters())\n x, y = torch.randn(bs, 1024).cuda(), torch.randint(0, 10, (bs,)).cuda()\n start, steps = time.time(), 300\n for in range(steps):\n opt.zero grad()\n torch.nn.functional.cross entropy(model(x), y).backward()\n opt.step()\n torch.cuda.synchronize()\n print(f\"batch {bs}: {steps bs / (time.time() - start):.0f} samples/s\")\n\nTerminal window\n\n$ nodus apply -f sweep.yaml\nsweep.nodus.dev/batch-size created\n\nEach parameter reaches the command as an environment variable: BATCH SIZE arrives as NODUS PARAM BATCH SIZE . Parameter names use uppercase letters, digits and .\n\nWatch it and read the results\n\nTerminal window\n\n$ nodus get sweep/batch-size -w\nNAME PHASE CELLS COST BEST-COST AGE\nbatch-size Succeeded 6/6 $0.31 2 14m\n$ nodus get sweep/batch-size -o yaml\n\nEach cell runs as a Job named - , so nodus logs job/batch-size-3 shows one cell’s output. The Sweep’s status.cells lists every cell with its GPU, region, parameters, phase, costUSD , wallSeconds and unitsPerSecond , and status.best names the cell with the lowest cost ( byCost ) and the highest throughput ( byThroughput ).\n\n unitsPerSecond counts the units your program reports with nodus.log.unit(id, ms) from the Python SDK, such as one call per batch or per request, divided by the cell’s wall time. Cells that report no units have no throughput.\n\nThe matrix\n\n Field Values Limit \n - - - \n matrix.gpu GPU requests, such as L4 , H100:8 or H100! 16 \n matrix.regions Region classes, such as us or eu ; within the template’s placement.regions when it sets them \n matrix.params Parameter name → list of string values 16 names, 64 values each \n repetitions Runs of each combination, to measure variance 64 \n\nThe number of cells is the product of the non-empty dimensions times repetitions , at most 256. Cells are numbered in a fixed order: GPU types vary slowest, then regions, then parameters by name, and repetitions fastest, so the repetitions of one combination sit next to each other.\n\nFrom Python\n\nimport nodus\n\nsweep = nodus.Sweep(\n {\"image\": \"nodus/pytorch\", \"command\": [\"python\", \"bench.py\"]}, # the Job spec every cell runs\n grid={\"gpu\": [\"L4\", \"H100\"], \"BATCH SIZE\": [8, 16, 32]}, # gpu and region are dimensions; the rest are parameters\n repetitions=2, max parallel=6, max cost=25,\n).run()\nreport = sweep.wait() # a Sweep with failed cells is Failed (CellsFailed) and still has its report\nprint(report.phase, report.best.by cost)\nfor cell in sweep.cells():\n print(cell.index, cell.phase, cell.cost usd, cell.wall seconds)\n\n nodus.Sweep.from name(\"batch-size\") reads one that exists, and suspend() , resume() and cancel() apply to every unfinished cell. A Function or a recipe as the target is not available yet and raises nodus.errors.Unsupported .\n\nRunning and failing cells\n\n maxParallel limits how many cells run at once (default 4). Cells start in index order.\n maxCostUSD caps the whole Sweep. At the cap the running cells save their state and suspend, no new cell starts, and the Sweep becomes Suspended with reason MaxCostReached ; raise the cap to continue.\n A failed cell does not stop the others. When every cell has finished and some failed, the Sweep fails with reason CellsFailed , and its status.cells still carries the results of the cells that succeeded.\n nodus suspend sweep/x , resume and cancel apply to every unfinished cell. Deleting a Sweep deletes its cell Jobs."},{"id":"docs/guides/training","url":"https://nodus-platform-site.pages.dev/docs/guides/training/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/training.md","title":"Training (Beta)","description":"Fine-tune, post-train, distill and pretrain models with TrainingJobs on managed runtimes, on one GPU or across several machines.","stage":"Beta","headings":[{"depth":2,"slug":"sign-in","text":"Sign in"},{"depth":2,"slug":"choose-a-runtime","text":"Choose a runtime"},{"depth":2,"slug":"fine-tune-with-lora","text":"Fine-tune with LoRA"},{"depth":2,"slug":"watch-it","text":"Watch it"},{"depth":2,"slug":"download-the-results","text":"Download the results"},{"depth":2,"slug":"full-fine-tuning-and-pretraining","text":"Full fine-tuning and pretraining"},{"depth":2,"slug":"preference-tuning","text":"Preference tuning"},{"depth":2,"slug":"distillation","text":"Distillation"},{"depth":2,"slug":"reinforcement-learning-beta-single-node","text":"Reinforcement learning (Beta, single node)"},{"depth":3,"slug":"reproducing-the-legacy-letter-counting-protocol","text":"Reproducing the legacy letter-counting protocol"},{"depth":2,"slug":"multi-node-training","text":"Multi-node training"},{"depth":2,"slug":"spending-caps","text":"Spending caps"},{"depth":2,"slug":"tracking","text":"Tracking"},{"depth":2,"slug":"bring-your-own-runtime","text":"Bring your own runtime"}],"text":"Training (Beta)\n\n Fine-tune, post-train, distill and pretrain models with TrainingJobs on managed runtimes, on one GPU or across several machines.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/training/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n Beta: this feature may change.\n\nBeta\n\nTrainingJob and TrainingRuntime are nodus.dev/v1beta1 . Fields may still change before they reach v1; a change is announced in the changelog with a migration note. Environments, used for RL and evaluation, are GA.\n\nA TrainingJob says what to train (a base model, a dataset or an environment, and the runtime’s parameters) and Nodus runs it as a Job: it caches the model, picks GPUs from the runtime’s presets, estimates the run before it starts, checkpoints it, and collects the trained weights as outputs.\n\nSign in\n\nTerminal window\n\n$ pip install nodus-compute\n$ nodus login\n\nChoose a runtime\n\nA TrainingRuntime is a pinned trainer image with a parameter schema, presets and measured step times. The managed runtimes live in the nodus catalog project:\n\n Runtime What it trains Data \n - - - \n nodus/sft Supervised fine-tuning, full or LoRA/QLoRA; task: Pretrain for continued or from-scratch pretraining prompt/completion or text JSONL \n nodus/dpo Direct preference optimization prompt, chosen, rejected \n nodus/orpo Odds-ratio preference optimization (no reference model) prompt, chosen, rejected \n nodus/kto Kahneman-Tversky optimization from thumbs-up/down labels prompt, completion, label \n nodus/reward-model A reward model for RLHF prompt, chosen, rejected \n nodus/grpo-lora GRPO reinforcement learning against an Environment (single node) an Environment \n nodus/distill Knowledge distillation from a teacher model prompts or chats \n nodus/evaluate Evaluation only ( mode: Evaluate ) an Environment or a benchmark \n\nTerminal window\n\n$ nodus get trainingruntimes -n nodus\n$ nodus get trainingruntime/sft -n nodus -o yaml # the parameter schema, presets and step times\n\nIn the console, Environments lists the same runtimes under Training templates . You start a run from the SDK or the CLI; the console shows it under Runs › Post-training with its stages, curves, results, cost and outputs.\n\nFine-tune with LoRA\n\nsft.yaml\n\napiVersion: nodus.dev/v1beta1\nkind: TrainingJob\nmetadata:\n name: support-sft\nspec:\n runtime: nodus/sft\n model:\n uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca\n secret: hf-token # only for gated repositories\n data:\n volume: support-chats # a Volume holding chats.jsonl\n path: chats.jsonl\n format: JSONL\n columns: {prompt: question, completion: answer}\n parameters:\n steps: 2000\n lora: {r: 16, alpha: 32}\n maxCostUSD: \"20.00\"\n\nReview it before anything runs. A dry run compiles the TrainingJob into the Job it will create and prices it:\n\nTerminal window\n\n$ nodus apply -f sft.yaml --dry-run -o yaml # status.estimate: cost to completion, first hold, start time\n$ nodus apply -f sft.yaml\n\nWhat the dry run decides:\n\n Resources come from the first runtime preset that matches the model, the method (Full, LoRA or QLoRA, read from parameters ) and quantization. Anything you set under spec.resources wins.\n The estimate is the number of steps times the measured seconds per step on the slowest GPU type you allowed, plus evaluation tasks times seconds per task. Set spec.expectedDuration to use your own figure instead.\n Parameters are merged over the runtime’s defaults and checked against its schema. A typo such as learning rate is rejected with the field path, before you are charged anything.\n\nModel, teacher, and dataset Volume inputs are fixed to their latest committed revision when you create the TrainingJob, unless you select an explicit revision. Upload the files and wait for the Volume to be Ready first. Later uploads do not change the admitted run.\n\nPin Hugging Face model revisions to a 40-character commit. The first TrainingJob that names a revision imports it once into a read-only Volume in your project, and every later run on that revision reuses it. The trainer runs with HF HUB OFFLINE=1 , so billed GPU time is never spent downloading weights.\n\nWatch it\n\nTerminal window\n\n$ nodus get tj -w\nNAME RUNTIME PHASE STAGE CHANGE COST AGE\nsupport-sft nodus/sft Running Training $0.12 6m\n$ nodus describe tj/support-sft\n$ nodus logs job/support-sft -f\n\n STAGE is CachingModel while the model imports, then the stage the trainer reports ( Baseline , Training , Evaluating ), and Exporting while the outputs are written to spec.exportTo . describe adds the step, total steps, loss and tokens per second, the provenance of the run (runtime and environment digests, model revision, and whether the runtime is a managed one) and the cost so far.\n\nIn the console, open the run from Compute : the Overview shows the stages, step, loss and the baseline and final pass rates; Gradings lists every graded task as it is scored; Results has the outputs and the before and after table.\n\nThe TrainingJob creates a Job of the same name and owns it. Logs, checkpoints and outputs are the Job’s, so everything in the Jobs guide applies.\n\nDownload the results\n\nSet spec.exportTo.volume to a Ready ReadWriteOnce Volume in the same project to also copy the complete result there. The runtime must declare an output at /nodus/outputs . The export uses an isolated CPU Job under the TrainingJob’s spending cap, waits up to 5 minutes for capacity, and allows 15 minutes for the copy. It follows suspension and cancellation. Success records the exact committed Volume revision in the Exported condition; files remain downloadable as Job outputs if the copy fails. Source or destination symbolic links are rejected.\n\nWhen the run succeeds, every file the trainer wrote under /nodus/outputs is an output of the TrainingJob, named by its path: results.json (the final metrics), provenance.json , manifest.json (a sha256 per file) and the weights under adapter/ (LoRA and QLoRA) or model/ (full fine-tuning), with the tokenizer beside them.\n\nTerminal window\n\n$ nodus get tj/support-sft -o jsonpath='{.status.outputs[ ].name}'\n$ nodus cp tj/support-sft:outputs/results.json ./results.json\n$ nodus cp tj/support-sft:outputs/adapter/adapter model.safetensors ./adapter model.safetensors\n\n nodus cp verifies each file against the sha256 recorded when it was collected.\n\nFull fine-tuning and pretraining\n\nLeave out the lora block for full fine-tuning; presets for the Full method pick larger GPUs.\n\nPretraining uses the same runtime with task: Pretrain . initialization: Continued (the default) continues from the model’s weights; Scratch uses only its architecture and tokenizer:\n\nspec:\n runtime: nodus/sft\n task: Pretrain\n model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca}\n data: {volume: corpus, format: JSONL}\n parameters: {steps: 10000, initialization: Scratch, packing: true}\n resources: {gpu: {type: [h100-sxm-80g], count: 8}}\n\nPreference tuning\n\n dpo , orpo , kto and reward-model take a dataset whose columns you map onto the names the runtime expects:\n\nspec:\n runtime: nodus/dpo\n model: {volume: {name: my-sft-model, revision: 3}} # a model you trained earlier, by Volume\n data:\n volume: prefs\n format: JSONL\n columns: {prompt: prompt, chosen: chosen, rejected: rejected}\n parameters: {steps: 400}\n\nA model can come from a Hugging Face URI or from one of your Volumes at a pinned revision, so a DPO run can start from the output of an SFT run.\n\nDistillation\n\n distill trains the student in model on the teacher’s distributions. The teacher is cached the same way:\n\nspec:\n runtime: nodus/distill\n model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca}\n teacher: {uri: hf://Qwen/Qwen3-4B@0e9e39f249a16976918f6564b8830bc894c89659}\n data: {volume: chats}\n parameters: {temperature: 2.0}\n resources: {gpu: {type: [h100-sxm-80g], count: 2}}\n\nReinforcement learning (Beta, single node)\n\n grpo-lora trains against an Environment: it samples completions for the environment’s tasks and Nodus grades them. The trainer never sees the expected answers and never grades itself: it submits completions to the TrainingJob’s gradings endpoint, and Nodus runs the environment’s grader in isolated containers on the same GPU host, with no network access, model files or GPU devices and the answers readable only by the grader. Task generation also uses this host; Nodus does not provision a separate CPU machine. A grader that fails is reported as an infrastructure failure and never scores the completion 0.\n\nEnvironment-backed training and evaluation require GPU capacity that supports isolated grading. Grading is serialized per TrainingJob. Baseline and final model evaluation run on the GPU. Custom runtime images must implement the task-fetch endpoint and declare spec.environmentTasks: AttemptAPI ; older runtimes are rejected before GPU acquisition.\n\nspec:\n runtime: nodus/grpo-lora\n model: {uri: hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca}\n environment: {name: nodus/graph-coloring@1.0.0, trainTasks: 50, heldOutTasks: 64, seed: 42}\n evaluation: {baseline: true, final: true}\n parameters: {steps: 50, lora: {r: 8}}\n resources: {gpu: {type: [rtx4090-24g], count: 1}}\n maxCostUSD: \"5.00\"\n\nWith evaluation.baseline and evaluation.final , the run scores the held-out tasks before and after training. status.summary holds both pass rates and a comparison: the change in percentage points, and whether it is comparable at all. A comparison needs the same task set on both sides, at least 16 scored tasks, and a result that is not all-pass or all-fail on both sides; otherwise its reason says why ( TooFewTasks , DifferentTasks , NonePassed , AllPassed , IncompleteEvents ). status.diagnostics.rewardSignal warns when every training reward was the same, which teaches a group-relative method nothing.\n\nReproducing the legacy letter-counting protocol\n\nUse nodus/letter-counting-legacy-eval@1.0.0 for the legacy evaluation protocol: the first 64 tasks from Reasoning Gym’s letter counting generator with seed 42, an answer-only system message, greedy decoding with 256 completion tokens, and the generator’s raw-completion scorer. The letter-counting catalog example pins Qwen3-1.7B and generates one prompt at a time.\n\nUse nodus/letter-counting-legacy-rl@1.0.0 for the legacy RL protocol. Its first 64 tasks are training data; the following 64 are held out for both baseline and final evaluation. The example sets 50 optimizer steps, learning rate 5e-6 , eight generations, generation batch size eight, per-GPU batch size one, gradient accumulation four, temperature 1.2 , top-p 0.95 , top-k 50 , and 64 completion tokens. It uses a rank-eight LoRA adapter with alpha 16, no dropout, all linear layers, bf16, non-reentrant gradient checkpointing, and a linear learning-rate schedule without warmup. Thinking is disabled in both phases. Recovery checkpoints remain enabled.\n\nThese profiles pin Reasoning Gym 0.1.25 and preserve the legacy source’s task slicing, prompt roles and scoring. They do not establish historical binary equivalence or a measured pass rate. The evaluation profile’s 64 tasks are different from the RL profile’s held-out 64 tasks; compare baseline and final within the same RL run. The existing nodus/reasoning-gym profile uses shuffled family-specific seeds and remains a different benchmark. Record the exact runtime, environment and model digests alongside results before claiming a numeric comparison.\n\nRL runs on one node. Multi-node RL and custom environment collections are not available yet.\n\nMulti-node training\n\nSet resources.nodes above 1 to train across several machines, each with resources.gpu.count GPUs. The TrainingJob compiles into one multi-node Job (a gang): every node starts together, torchrun is launched on each with its rank and the rendezvous address, and the run checkpoints with torch.distributed.checkpoint .\n\nspec:\n runtime: nodus/dpo\n model: {volume: {name: my-8b, revision: 3}}\n data: {volume: prefs, format: JSONL, columns: {prompt: prompt, chosen: chosen, rejected: rejected}}\n parameters: {steps: 400}\n resources: {gpu: {type: [h100-sxm-80g], count: 8}, nodes: 2}\n distributed: {network: Global, transport: Auto}\n\n Launchers. A runtime declares Torchrun , Accelerate , Deepspeed or Plain , and the run uses that launcher on every node: torchrun reads the gang’s rendezvous settings, accelerate launch and DeepSpeed’s per-node launcher read the rank, address and world size Nodus sets on each node, and Plain starts the runtime’s command once per node. A single node runs torchrun --standalone for a Torchrun runtime, accelerate launch or deepspeed for the other two.\n Placement. distributed.network is Colocated (one provider and region, the fastest interconnect), Regional or Global ; transport: Auto also allows an encrypted SSH connection between two restricted GPU machines without a separate relay server. Native GPU-to-GPU connectivity is preferred. Wider settings find capacity sooner at some cost in throughput.\n Recovery. If a node is lost, the whole gang restarts from the latest checkpoint every rank committed, up to 3 times; progress since that checkpoint is redone. A single-node run restores its latest snapshot instead. Managed trainers bind recovery state to the full run configuration, runtime image digest, and pinned model and dataset references; incompatible state is refused. GRPO reuses the baseline saved in that state so recovery does not repeat baseline grading.\n Cost. The estimate covers every node. Assembly time (nodes waiting for each other) is billed and bounded; see Multi-node training for the limits.\n\nSpending caps\n\n spec.maxCostUSD caps the whole TrainingJob: its Job, its same-host grading, its output export and its retries. When the cap is reached the run is suspended, not failed, and active grader containers are removed. Raise the cap to resume the same run from its checkpoint:\n\nTerminal window\n\n$ nodus patch tj/support-sft --type merge -p '{\"spec\":{\"maxCostUSD\":\"40\"}}'\n\nThe cap can only be raised. spec.state: Suspended pauses a run yourself, and Running resumes it.\n\nTracking\n\n spec.tracking.mlflow ( uri , secret , experiment ) or spec.tracking.wandb ( connection ) send the trainer’s metrics to your own MLflow server or Weights & Biases project as well. Nodus shows its own stage and metrics either way. The catalog runtimes report to MLflow when the run has MLFLOW TRACKING URI and to W\\&B when it has WANDB API KEY , which is what those two fields set; the Secret holds MLFLOW TRACKING TOKEN (or MLFLOW TRACKING USERNAME and MLFLOW TRACKING PASSWORD ), and the W\\&B Connection’s Secret holds WANDB API KEY .\n\nBring your own runtime\n\nA TrainingRuntime in your project can run any image. The trainer reads:\n\n Variable Contents \n - - \n NODUS INPUT PARAMS a directory that holds params.json , the run: task and mode (for example SFT , Train ), the merged, schema-checked parameters , data (path, format and column mapping), environment (task counts, seed and graders), evaluation , trainingJob and the pinned modelRevision and teacherRevision \n Environment tasks Call POST …/trainingjobs/{name}/tasks through the attempt API socket. The response files maps train.jsonl and test.jsonl to JSONL strings containing {taskId, prompt, metadata} and no hidden answers. Managed runtimes fetch them automatically; custom runtimes must implement this step. \n NODUS INPUT MODEL , NODUS INPUT TEACHER , NODUS INPUT DATA where the model, teacher and data are mounted \n NODUS STATE DIR , NODUS OUTPUT DIR resumable state and results \n NODUS NODE RANK , MASTER ADDR , PET the gang of a multi-node run \n NODUS CHECKPOINT URI , NODUS RESTORE URI where ranks write and restore distributed checkpoints \n\nWrite and load your own model, optimizer and progress files under the state directory; Nodus decides when to checkpoint and where checkpoints are stored, and never adds resume flags to your command."},{"id":"docs/guides/volumes","url":"https://nodus-platform-site.pages.dev/docs/guides/volumes/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/volumes.md","title":"Volumes","description":"Keep datasets, model weights and working files in Volumes, upload and download them, import from Hugging Face, git, URLs or S3, and mount them in Jobs and Sandboxes.","stage":"GA","headings":[{"depth":2,"slug":"create-a-volume-and-upload-files","text":"Create a Volume and upload files"},{"depth":2,"slug":"mount-it-in-a-job","text":"Mount it in a Job"},{"depth":2,"slug":"import-from-elsewhere","text":"Import from elsewhere"},{"depth":2,"slug":"in-python","text":"In Python"}],"text":"Volumes\n\n Keep datasets, model weights and working files in Volumes, upload and download them, import from Hugging Face, git, URLs or S3, and mount them in Jobs and Sandboxes.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/volumes/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Volume is named storage that outlives the work that uses it. Every change is saved as a numbered revision, so you can see what changed, mount an earlier revision read-only, and never lose files to a stopped machine. You pay for the bytes stored, after your org’s included 10 GB.\n\nCreate a Volume and upload files\n\nvolume.yaml\n\napiVersion: nodus.dev/v1\nkind: Volume\nmetadata:\n name: ex-volumes-put-get\nspec:\n accessMode: ReadWriteOnce\n size: 1Gi\n source: {upload: {}}\n\nTerminal window\n\nnodus apply -f volume.yaml # or: nodus create volume data --size 100Gi\nnodus volume put ex-volumes-put-get ./data /data # uploads ./data and prints the new revision\nnodus volume ls ex-volumes-put-get /data\nnodus volume get ex-volumes-put-get /data/hello.txt ./out/\nnodus volume put data ./archive.tar.gz /raw --extract\nnodus volume rm data /raw/old.csv\nnodus volume clear data # an empty revision; the Volume and its history stay\n\nUploads only send what changed: a second put of the same files moves almost nothing. An upload that started from an older revision than the latest is refused with 409 Conflict , so two people uploading at once never overwrite each other silently; run it again on top of the new revision.\n\n put adds to what is there: a directory’s contents merge into the remote path, and a file lands at the remote path, or inside it when the remote path is a directory or ends in / . rm and clear also commit new revisions, so earlier revisions keep the removed files. get and ls read the latest revision, or the one --revision names; get … - prints a file to standard output, and existing local files are kept unless --force . --extract expands a .zip , .tar , .tar.gz , .tgz or .tar.zst archive. It refuses an archive whose paths or links lead outside the remote path, or that has more than 10,000 members.\n\nMount it in a Job\n\njob.yaml\n\napiVersion: nodus.dev/v1\nkind: Job\nmetadata:\n name: ex-volumes-put-get\nspec:\n image: nodus/python:3.12\n volumes:\n - {volume: ex-volumes-put-get, mountPath: /mnt/vol, readOnly: true}\n command: [cat, /mnt/vol/data/hello.txt]\n\n accessMode Who can write How changes are saved \n - - - \n ReadWriteOnce (default) One attempt at a time On exit and every commitInterval (default 5 minutes) \n ReadOnlyMany (default for imports) Nobody; any number of readers Revisions come from imports and uploads \n ReadWriteMany Every worker Each worker’s changes are published when it commits or exits; the last writer of a path wins \n\nA ReadWriteOnce Volume is held by one attempt at a time. Starting a second writer, or deleting the Volume while it is held, fails with 409 VolumeBusy , which names the holder; mount it with readOnly: true to read alongside. Functions and Agents with more than one worker need ReadOnlyMany or ReadWriteMany .\n\nMount an earlier revision with revision: and readOnly: true . nodus get volume data shows the latest REVISION , USED and the current HOLDER ; GET …/volumes/{name}/revisions lists the kept revisions ( revisionHistoryLimit , default 10).\n\nImport from elsewhere\n\nSet exactly one source ; the import runs once on Nodus, billed as CPU time, and produces revision 1:\n\nspec:\n accessMode: ReadOnlyMany\n size: 200Gi\n maxCostUSD: \"2.00\"\n source: {huggingface: {repo: meta-llama/Llama-3.1-8B, revision: 0e9e39f249a16976918f6564b8830bc894c89659, secret: hf-token}}\n\n Source Fields \n - - \n huggingface repo , revision (a commit or tag), files globs, secret holding HF TOKEN \n git repo , ref , lfs \n url an https:// url , sha256 , extract: Auto unpacks .zip , .tar , .tar.gz , .tgz and .tar.zst \n s3 uri and an S3 Connection \n\nThe Volume is Pending while it imports and Ready after. A failed import shows Failed with reason ImportFailed and a message; fix the source or add credits, then run nodus request reimport volume/ .\n\nIn Python\n\nvol = nodus.Volume.from name(\"data\", create if missing=True)\nvol.put file(\"./local.csv\", \"/train/local.csv\")\nprint(vol.listdir(\"/train\"))"},{"id":"docs/guides/webhooks","url":"https://nodus-platform-site.pages.dev/docs/guides/webhooks/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/webhooks.md","title":"Webhooks","description":"Get signed HTTPS callbacks when Jobs finish, balances run low and other events happen, and verify them.","stage":"GA","headings":[{"depth":2,"slug":"create-an-endpoint","text":"Create an endpoint"},{"depth":2,"slug":"what-nodus-sends","text":"What Nodus sends"},{"depth":2,"slug":"event-types","text":"Event types"},{"depth":2,"slug":"verify-the-signature","text":"Verify the signature"},{"depth":2,"slug":"retries-and-delivery-order","text":"Retries and delivery order"},{"depth":2,"slug":"the-delivery-log","text":"The delivery log"},{"depth":2,"slug":"rotate-the-signing-secret","text":"Rotate the signing secret"},{"depth":2,"slug":"check-an-endpoints-health","text":"Check an endpoint’s health"}],"text":"Webhooks\n\n Get signed HTTPS callbacks when Jobs finish, balances run low and other events happen, and verify them.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/webhooks/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA WebhookEndpoint sends an HTTPS request to your server each time an event you care about happens: a Job succeeds or fails, a balance runs low, a Budget crosses a threshold. You can have as many endpoints as you need, each with its own URL, event types and projects.\n\nEvery delivery is signed with the Standard Webhooks scheme, so you can check that it came from Nodus and was not changed in transit.\n\nCreate an endpoint\n\nendpoint.yaml\n\napiVersion: nodus.dev/v1\nkind: WebhookEndpoint\nmetadata:\n name: example-receiver\nspec:\n url: https://hooks.example.com/nodus\n description: Job results for the example receiver\n eventTypes:\n - job.succeeded\n - job.failed\n - billing.low balance\n\nTerminal window\n\n$ nodus apply -f endpoint.yaml -o json jq -r .status.signingSecret\nwhsec ...\n\nThe signing secret is in the response to the create, and nowhere else. Copy it into your receiver’s configuration now. If you lose it, rotate it to get a new one.\n\nNodus sends a webhookendpoint.ping delivery right after the create, so you can see your receiver work before any real event.\n\n Field Meaning \n - - \n url Where deliveries go. https only, with a public address. No redirects are followed. \n eventTypes Event types to send: an exact type ( job.succeeded ), a prefix ( job. , billing. ) or . Defaults to . \n projects Send project events only for these projects. Empty means every project. Org-level events always match. \n enabled false pauses deliveries. Pending deliveries to a paused endpoint are dropped. \n description A note for you. \n\nEvery field can change after the create:\n\nTerminal window\n\n$ nodus patch webhookendpoint ci --type merge -p '{\"spec\":{\"eventTypes\":[\"job. \"]}}'\n\nWhat Nodus sends\n\nA delivery is a POST with a JSON body:\n\n{\n \"type\": \"job.succeeded\",\n \"timestamp\": \"2026-10-01T12:00:00Z\",\n \"data\": {\n \"event\": { \"reason\": \"Succeeded\", \"message\": \"…\", \"count\": 1 },\n \"object\": {\n \"apiVersion\": \"nodus.dev/v1\",\n \"kind\": \"Job\",\n \"metadata\": { \"name\": \"train\", \"namespace\": \"default\", \"uid\": \"job …\" },\n \"status\": { \"phase\": \"Succeeded\" }\n }\n }\n}\n\nand three headers:\n\n Header Meaning \n - - \n webhook-id Names this delivery. It stays the same across retries, so use it to ignore duplicates. \n webhook-timestamp When the attempt was signed, in Unix seconds. \n webhook-signature One or more signatures, separated by spaces, each like v1, . \n\nThe event type is . in lowercase with underscores, such as job.succeeded or agentrun.needs resolution . Billing events use the billing. prefix.\n\nEvent types\n\n Type When \n - - \n job.succeeded , job.failed , job.cancelled A Job reached a final phase. The same reasons exist for other run kinds, such as agentrun.failed . \n job.preempted , job.node lost , job.restored A Job lost its machine, and was restored from its last checkpoint. \n job.checkpointed , job.checkpoint failed A checkpoint was saved, or could not be. \n job.max cost reached , job.funding lost , job.funding restored A Job stopped or resumed because of its cost cap or your balance. The same reasons exist for Sandboxes and other kinds with a cost. \n job.output sink failed A Job’s output could not be loaded into its destination. \n job.gang restarted , job.network degraded A multi-node Job restarted or lost network quality. \n sandbox.started , sandbox.stopped , sandbox.stuck , sandbox.setup failed A Sandbox changed state. \n agentrun.needs resolution An AgentRun is waiting for you to resolve a step. \n billing.low balance Your available balance fell below the warning level. \n billing.auto recharge succeeded , billing.auto recharge failed An auto-recharge added credits, or its charge failed. \n billing.arrears posted A charge was larger than your balance, so new work waits until it is paid. \n billing.dispute opened , billing.dispute closed A payment was disputed, which pauses new work, or the dispute closed. \n billing.storage quota exceeded , billing.egress quota exceeded A quota was reached. \n budget.threshold , budget.exceeded A Budget crossed a threshold, or refused a new run. \n topup.succeeded , topup.failed , topup.refunded A top-up completed, its payment failed, or it was refunded. \n pool.node offline , pool.forecast shortfall , pool.action proposed Events of your own capacity pools. \n webhookendpoint.ping Sent once when an endpoint is created. \n webhookendpoint.deliveries failing Your receiver has failed every recent delivery. \n\n nodus get events shows the events of your org; each webhook type is the kind and reason of one of them.\n\nVerify the signature\n\nCheck every delivery before you trust it. The signature covers the exact bytes of the body, so verify the raw body before you parse it. Use a Standard Webhooks library for your language; it also rejects timestamps more than five minutes old, which stops a captured delivery from being replayed later.\n\nmain.go\n\npackage main\n\nimport (\n \"encoding/json\"\n \"fmt\"\n \"io\"\n \"log\"\n \"net/http\"\n \"os\"\n \"sync\"\n \"time\"\n\n standardwebhooks \"github.com/standard-webhooks/standard-webhooks/libraries/go\"\n)\n\n// maxBody bounds a delivery; Nodus sends at most 64 KiB.\nconst maxBody = 64 << 10\n\n// event is the body of a delivery: the event's type and the Event and object it is about.\ntype event struct {\n Type string json:\"type\" \n Timestamp time.Time json:\"timestamp\" \n Data struct {\n Event json.RawMessage json:\"event\" \n Object struct {\n Kind string json:\"kind\" \n Metadata struct {\n Name string json:\"name\" \n Namespace string json:\"namespace\" \n } json:\"metadata\" \n } json:\"object\" \n } json:\"data\" \n}\n\n// receiver verifies and handles deliveries.\ntype receiver struct {\n hook standardwebhooks.Webhook\n out io.Writer\n\n mu sync.Mutex\n seen map[string]bool\n}\n\nfunc newReceiver(secret string, out io.Writer) ( receiver, error) {\n hook, err := standardwebhooks.NewWebhook(secret)\n if err != nil {\n return nil, fmt.Errorf(\"NODUS WEBHOOK SECRET: %w\", err)\n }\n return &receiver{hook: hook, out: out, seen: map[string]bool{}}, nil\n}\n\n// ServeHTTP answers 2xx once the delivery is verified and recorded, which is all Nodus needs to stop retrying. A\n// delivery that fails verification gets 400 and is never parsed: the signature covers the exact bytes received, so\n// verify before you decode.\nfunc (r receiver) ServeHTTP(w http.ResponseWriter, req http.Request) {\n body, err := io.ReadAll(http.MaxBytesReader(w, req.Body, maxBody))\n if err != nil {\n http.Error(w, \"body too large\", http.StatusRequestEntityTooLarge)\n return\n }\n // Verify checks the signature against every secret in the header and rejects timestamps more than five minutes\n // from now, so a captured delivery cannot be replayed later.\n if err := r.hook.Verify(body, req.Header); err != nil {\n http.Error(w, \"invalid signature\", http.StatusBadRequest)\n return\n }\n // A delivery can arrive more than once (a retry after a slow answer): webhook-id names the delivery, so handle\n // each id once and still answer 200.\n if r.firstTime(req.Header.Get(standardwebhooks.HeaderWebhookID)) {\n var e event\n if err := json.Unmarshal(body, &e); err != nil {\n http.Error(w, \"malformed event\", http.StatusBadRequest)\n return\n }\n r.handle(e)\n }\n w.WriteHeader(http.StatusOK)\n}\n\nfunc (r receiver) firstTime(id string) bool {\n r.mu.Lock()\n defer r.mu.Unlock()\n if r.seen[id] {\n return false\n }\n r.seen[id] = true\n return true\n}\n\n// handle does the work of an event. Do slow work after answering (queue it); Nodus waits 15 seconds for a response.\nfunc (r receiver) handle(e event) {\n obj := e.Data.Object\n fmt.Fprintf(r.out, \"%s %s/%s\\n\", e.Type, obj.Kind, obj.Metadata.Name)\n}\n\nfunc main() {\n rcv, err := newReceiver(os.Getenv(\"NODUS WEBHOOK SECRET\"), os.Stdout)\n if err != nil {\n log.Fatal(err)\n }\n srv := &http.Server{Addr: \":8080\", Handler: rcv, ReadHeaderTimeout: 5 time.Second}\n log.Fatal(srv.ListenAndServe())\n}\n\nThis receiver runs as is: NODUS WEBHOOK SECRET=whsec ... go run ./examples/webhooks/receiver . The same check in Python:\n\nfrom standardwebhooks.webhooks import Webhook\n\nwh = Webhook(secret) # the whsec secret\nevent = wh.verify(request body, request headers) # raises if the signature or timestamp is wrong\n\nIf you cannot use a library, compute base64(HMAC-SHA256(key, id + \".\" + timestamp + \".\" + body)) , where key is the secret after the whsec prefix, base64-decoded, and compare it in constant time with each v1, signature in the header.\n\nCaution\n\nAnswer with a 2xx status as soon as you have verified and recorded the delivery, and do the slow work after. Nodus waits 15 seconds for a response, and treats a timeout like a failure.\n\nRetries and delivery order\n\nA 2xx response is success. Anything else, a timeout or a connection error is retried with exponential backoff, starting at 30 seconds and doubling up to 8 hours between attempts, for 3 days after the event. After that the delivery is marked failed. A URL that Nodus is not allowed to connect to, such as a private address, fails at once.\n\nDeliveries can arrive more than once and out of order. Use webhook-id to ignore a duplicate, and the timestamp in the body or the object’s status to see which state is newer.\n\nThe delivery log\n\nNodus keeps the log of an endpoint’s deliveries for 7 days:\n\nTerminal window\n\n$ curl -s \"$NODUS API URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries?limit=20\" \\\n -H \"Authorization: Bearer $NODUS API KEY\"\n\nEach entry shows the event type, the number of attempts, the receiver’s last response code, how long it took, and a phase : Pending while retries remain, then Succeeded or Failed . A delivery to an endpoint that was deleted or disabled is Abandoned . The response includes continue when there is more; pass it back as after .\n\nTo send a delivery again, replay it. The replay has the same body and a new webhook-id :\n\nTerminal window\n\n$ curl -s -X POST \"$NODUS API URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries/$ID/replays\" \\\n -H \"Authorization: Bearer $NODUS API KEY\"\n\nRotate the signing secret\n\nRotate when a secret may have leaked, or when you lost it:\n\nTerminal window\n\n$ curl -s -X POST \"$NODUS API URL/apis/nodus.dev/v1/webhookendpoints/ci/rotations\" \\\n -H \"Authorization: Bearer $NODUS API KEY\" -H \"Content-Type: application/json\" \\\n -H \"Idempotency-Key: $(uuidgen)\" -d '{\"previousSecretTTL\": \"24h\"}' jq -r .status.signingSecret\n\nThe response holds the new secret, once. For previousSecretTTL (24 hours by default, at most 7 days) every delivery carries a signature for the new secret and one for the old, so a receiver using either verifies it. Switch your receiver to the new secret, and the old one stops signing when the time is up. status.secretRotation shows when.\n\nCheck an endpoint’s health\n\nTerminal window\n\n$ nodus get webhookendpoint ci -o yaml\n\n status shows the latest delivery ( lastDelivery ), the deliveries of the last 24 hours that succeeded or failed ( deliveries24h ), and, when the last three deliveries in a row failed, failingSince with the Healthy condition set to False and reason DeliveriesFailing . Nodus also records a webhookendpoint.deliveries failing event when that happens, which you can send to another endpoint.\n\nNote\n\nWebhooks never carry provider names, machine ids or other internal details. Their bodies hold what the API would show you for the same object."},{"id":"docs/guides/workspaces","url":"https://nodus-platform-site.pages.dev/docs/guides/workspaces/","markdown":"https://nodus-platform-site.pages.dev/docs/source/guides/workspaces.md","title":"Workspaces","description":"A GPU or CPU development machine with a saved home directory, reached with SSH, VS Code or JupyterLab.","stage":"GA","headings":[{"depth":2,"slug":"create-a-workspace","text":"Create a Workspace"},{"depth":2,"slug":"connect","text":"Connect"},{"depth":3,"slug":"ssh","text":"SSH"},{"depth":3,"slug":"vs-code-remote-ssh-cursor-and-jetbrains-gateway","text":"VS Code Remote-SSH, Cursor and JetBrains Gateway"},{"depth":3,"slug":"browser-vs-code-and-jupyterlab","text":"Browser VS Code and JupyterLab"},{"depth":2,"slug":"stop-start-and-what-is-saved","text":"Stop, start and what is saved"},{"depth":3,"slug":"schedules","text":"Schedules"},{"depth":2,"slug":"run-a-job-on-your-workspace-files","text":"Run a job on your Workspace files"},{"depth":2,"slug":"sessions-and-billing","text":"Sessions and billing"},{"depth":2,"slug":"shared-access","text":"Shared access"},{"depth":2,"slug":"when-a-workspace-fails","text":"When a Workspace fails"}],"text":"Workspaces\n\n A GPU or CPU development machine with a saved home directory, reached with SSH, VS Code or JupyterLab.\n\nSource: https://nodus-platform-site.pages.dev/docs/guides/workspaces/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nA Workspace is your own development machine: a GPU or CPU machine with VS Code in the browser, JupyterLab and SSH. Its home directory, /home/nodus , lives on a Volume and is saved every time the Workspace stops, so you can stop it at the end of the day and pick up where you left off. While it is stopped you pay only for the saved files.\n\nYou need the nodus CLI to connect over SSH (install it and run nodus login ). Browser VS Code and JupyterLab need only the console.\n\nCreate a Workspace\n\nTerminal window\n\n$ nodus create workspace lab --gpu H100\n$ nodus get workspaces\nNAME PHASE COMPUTE HOME STOP-REASON COST AGE\nlab Running 1 × H100 lab-home $0.09 3m\n\n --gpu H100:2 asks for two GPUs; --cpu 8 --memory 32Gi without --gpu creates a CPU Workspace. The home Volume is lab-home , created on the first start if it does not exist; set spec.volume to use another one. The default image is nodus/workspace-pytorch-cuda (PyTorch and CUDA) on NVIDIA GPUs, nodus/workspace-rocm on AMD GPUs and nodus/workspace-cpu on CPU. Every catalog image has VS Code, JupyterLab, Python and the user nodus .\n\nThe same Workspace as a manifest, applied with nodus apply -f workspace.yaml :\n\nworkspace.yaml\n\napiVersion: nodus.dev/v1\nkind: Workspace\nmetadata:\n name: ws-home\nspec:\n image: nodus/workspace-cpu\n resources:\n cpu: \"2\"\n memory: 8Gi\n volume: ws-home-home\n idleTimeout: 30m\n\nIn the console, Workspaces › New workspace shows the GPU picker, the estimate and the same manifest as YAML, CLI and Python before you create it.\n\nConnect\n\nSSH\n\nTerminal window\n\n$ nodus login\n$ nodus ssh workspace/lab\nnodus@lab:~$ nvidia-smi\n$ nodus ssh workspace/lab -- python train.py --epochs 1\n\nSign in once with nodus login on each computer using an up-to-date CLI and OpenSSH 8.5 or newer. Login creates a dedicated local SSH key for your account, registers only its public key, and configures every existing and future CPU or GPU Workspace in the organizations you selected. New Workspaces need no separate setup. Your private key stays on your computer with owner-only permissions; sign-in does not start or connect to a Workspace.\n\nEach teammate signs in to their own account. Organization roles and project access determine which Workspaces they can enter. Remove a device’s key under Account › SSH keys to revoke it. Removing a teammate’s access also prevents their keys from opening that organization’s Workspaces.\n\n nodus ssh connects through the Nodus gateway. No port is open on the machine. OpenSSH verifies the Workspace’s host key against the authenticated API on every connection, with strict host-key checking. A missing or revoked identity fails closed; it does not trust a host key on first use.\n\nVS Code Remote-SSH, Cursor and JetBrains Gateway\n\nAfter nodus login , open Desktop VS Code or Cursor in the console, or select the Workspace host in your editor’s Remote-SSH connection dialog. The same account key works for all your authorized Workspaces, including ones created later. Signing in to the website alone cannot configure SSH on your computer.\n\nThe .nodus host is an SSH alias, not a public DNS address. Login records the CLI executable and saved context, so editors do not depend on your terminal’s environment. Plain ssh , scp and rsync use the same connection. nodus ssh workspace/lab --config prints the rule; nodus ssh workspace/lab --setup is an optional repair command. Neither command starts or connects to a Workspace.\n\nBrowser VS Code and JupyterLab\n\nOpen the Workspace in the console and choose Browser VS Code or JupyterLab . Both open on a private address that only members of your organization can reach after signing in.\n\nStop, start and what is saved\n\nTerminal window\n\n$ nodus stop workspace/lab\n$ nodus start workspace/lab\n\nWhen a Workspace stops, everything in /home/nodus is saved to its home Volume and the machine is released. Everything else is temporary, including the /workspace directory (the container’s working directory, which has nothing to do with the Workspace kind) and anything installed outside your home. Install Python packages with pip install --user or in a virtual environment under /home/nodus to keep them.\n\nThe next start restores the saved home on a fresh machine. Use ephemeral: true for a Workspace with no home Volume, whose files are deleted on stop.\n\nA Workspace also stops by itself:\n\n Stop reason When How it starts again \n - - - \n Idle No SSH session, editor traffic or process for idleTimeout (default 1h ) Connect, open a tool, or nodus start \n Schedule At schedule.stopAt The next schedule.readyBy , connecting, or nodus start \n InsufficientCredits , BudgetExceeded , MaxCostReached Money ran out or a limit was reached Add credits or raise the limit, then connect or nodus start \n User You ran nodus stop or chose Stop Only nodus start or Start \n\nConnecting to a Workspace that stopped by itself starts it; nodus ssh waits and connects when it is ready. The console asks before a connect starts a stopped Workspace and shows the rate it will bill at.\n\nSchedules\n\nspec:\n schedule:\n readyBy: \"2026-10-01T09:00:00-07:00\"\n stopAt: \"2026-10-01T19:00:00-07:00\"\n\nThe Workspace starts 15 minutes before readyBy so it is ready on time, and stops at stopAt . It records a ScheduledStart event, or ScheduleMissed when it could not start in time.\n\nRun a job on your Workspace files\n\nTerminal window\n\n$ nodus run --from workspace/lab -- python train.py\n\nThe Job runs on a copy of the last saved home, labelled nodus.dev/workspace=lab , so the Workspace’s Jobs from this workspace list shows it.\n\nSessions and billing\n\nEach period a Workspace runs is a session, recorded as an Attempt with its receipt:\n\nTerminal window\n\n$ nodus get attempts -l nodus.dev/workspace=lab\n\n A GPU Workspace bills at the rate shown when it starts, from when its machine is acquired until the machine is released. A session that stops for any reason closes its charge.\n A CPU Workspace bills at the listed CPU rates while it runs.\n Saved files in the home Volume bill as storage after your organization’s included storage.\n\n maxCostUSD stops and saves the Workspace when its cost reaches the limit; raising the limit lets it start again.\n\nShared access\n\nA Workspace belongs to its project. Every member of the organization with access to the project can see it, start it and connect with their own account and SSH key; spec.sshKeys limits SSH to the keys you name.\n\nWhen a Workspace fails\n\nA Workspace in the Failed phase does not restart. Delete it and create it again with the same name (or the same spec.volume ): the home Volume keeps your files.\n\nTerminal window\n\n$ nodus delete workspace/lab\n$ nodus create workspace lab --gpu H100\n\nTo delete the saved files too, delete the home Volume ( nodus delete volume/lab-home ), or empty it with nodus volume clear lab-home ."},{"id":"docs/reference","url":"https://nodus-platform-site.pages.dev/docs/reference/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference.md","title":"Reference","description":"Look up exact CLI commands, Python signatures, API fields, errors, and prices.","stage":"GA","headings":[],"text":"Reference\n\n Look up exact CLI commands, Python signatures, API fields, errors, and prices.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nUse reference pages to check a specific detail. For a walkthrough, find a guide first.\n\n I need… Reference \n - - \n Command syntax and flags CLI \n Python classes, methods, and arguments Python SDK \n HTTP endpoints and request fields HTTP API · OpenAPI JSON \n An error’s meaning and fix Error codes \n Help diagnosing a failure Troubleshooting \n Compute, storage, and transfer prices Pricing \n Training runtime options Training runtimes (Beta) \n Training and evaluation environments Environments \n Server configuration variables Server configuration \n\nCommand, SDK, API, error, catalog, and price details are generated from their product sources."},{"id":"docs/reference/api","url":"https://nodus-platform-site.pages.dev/docs/reference/api/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/api.md","title":"HTTP API","description":"The Nodus REST API, Kubernetes-style resources under /apis/nodus.dev, generated from the Go types.","stage":"GA","headings":[],"text":"HTTP API\n\n The Nodus REST API, Kubernetes-style resources under /apis/nodus.dev, generated from the Go types.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/api/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nNodus serves every resource at a Kubernetes-style path, so kubectl , client-go and the generated clients all work against it:\n\nGET /apis/nodus.dev/v1/namespaces/{project}/jobs\nPOST /apis/nodus.dev/v1/namespaces/{project}/jobs\nGET /apis/nodus.dev/v1/namespaces/{project}/jobs/{name}\nGET /apis/nodus.dev/v1/namespaces/{project}/jobs?watch=true\n\nAuthenticate with an API key as a bearer token ( Authorization: Bearer nodus sk … ). Lists page with limit and continue , writes accept an Idempotency-Key header, and ?dryRun=All validates a create and returns its estimate without charging anything. Errors are described in Error codes.\n\nThe per-operation reference in the sidebar is generated from the OpenAPI document, which is also served at /docs/openapi.json ."},{"id":"docs/reference/cli","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli.md","title":"CLI reference","description":"Every nodus command, flag and example, generated from the CLI's command tree.","stage":"GA","headings":[{"depth":2,"slug":"commands","text":"Commands"}],"text":"CLI reference\n\n Every nodus command, flag and example, generated from the CLI's command tree.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nThe nodus CLI follows kubectl’s verbs ( get , describe , apply , delete , logs , exec , cp ) over every Nodus resource, plus task commands such as run , login and billing . Global flags:\n\n Flag Meaning \n - - \n --context , --org Select a context or org \n -p, --project ( -n accepted), -A, --all-projects Project scope \n -o, --output table , wide , json , yaml , name , jsonpath=… , custom-columns=… , csv , estimate \n -l, --selector , --field-selector Label and field selectors \n -w, --watch Keep watching after listing \n --dry-run=server Validate and return the estimate without creating anything \n\n nodus run exits with your command’s exit code. Nodus and API errors exit 125 and timeouts 124 .\n\nCommands\n\nnodus Run compute Jobs on Nodus\n\nnodus agentrun Message an AgentRun and read its steps and answer\n\nnodus agentrun answer Print an AgentRun's last answer in full\n\nnodus agentrun send Send a message to an AgentRun, which takes it as its next prompt\n\nnodus agentrun steps List an AgentRun's journal steps: what it did, in order\n\nnodus annotate Set or remove annotations\n\nnodus api-resources List the kinds the server offers\n\nnodus apply Create or update resources from manifests (client-side three-way merge)\n\nnodus attach Attach to a running Process of a Job, Sandbox or Workspace\n\nnodus auth Credentials for other tools, and permission checks\n\nnodus auth can-i Check whether the current credential may perform an action\n\nnodus auth token Print an ExecCredential for kubectl and client-go (the kubeconfig exec plugin)\n\nnodus billing Show the org's balance; top up, redeem a code and manage auto-recharge\n\nnodus billing auto-recharge Top up automatically when available credit falls below a threshold\n\nnodus billing deal View partner offers or request an eligibility review\n\nnodus billing portal Open the Stripe Customer Portal: cards, invoices and billing details\n\nnodus billing receipts List top-ups with their Stripe receipt links\n\nnodus billing redeem Redeem a promo code for credit\n\nnodus billing top-up Buy prepaid credit with Stripe Checkout ($5–$1,000)\n\nnodus billing usage Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank\n\nnodus cancel Cancel a run, FunctionCall or Process\n\nnodus cloud Inspect connected cloud accounts\n\nnodus cloud inventory List the instances, GPUs or spend a cloud account's last sync saw\n\nnodus completion Generate the autocompletion script for the specified shell\n\nnodus completion bash Generate the autocompletion script for bash\n\nnodus completion fish Generate the autocompletion script for fish\n\nnodus completion powershell Generate the autocompletion script for powershell\n\nnodus completion zsh Generate the autocompletion script for zsh\n\nnodus config Show and change contexts in \\~/.nodus/config\n\nnodus config current-context Print the current context\n\nnodus config get-contexts List contexts\n\nnodus config set-context Change a context's project or server (NAME defaults to the current context)\n\nnodus config use-context Switch the current context\n\nnodus config view Print the config (keys stay in the keychain)\n\nnodus convert Turn a 0.x nodus.toml or training file into manifests for nodus apply\n\nnodus cp Copy files to and from a running Sandbox, Workspace or Job, or download a Job output\n\nnodus create Create resources from manifests or with a typed generator\n\nnodus create agentgroup Create a agentgroup\n\nnodus create agentrun Create a agentrun\n\nnodus create apikey Create an API key (printed once)\n\nnodus create budget Create a budget\n\nnodus create enrollmenttoken Create a token that enrolls one host into a Pool (installer printed once)\n\nnodus create inferenceendpoint Create a inferenceendpoint\n\nnodus create invite Invite someone to the org by email\n\nnodus create sandbox Create a sandbox\n\nnodus create secret Create a secret\n\nnodus create sshkey Add an SSH public key, or generate a key pair and add it\n\nnodus create token Mint a ServiceAccount API key (printed once)\n\nnodus create volume Create a volume\n\nnodus create workspace Create a workspace\n\nnodus delete Delete resources\n\nnodus deploy Deploy the nodus.App in APP.py as a persistent App\n\nnodus describe Show everything about one resource: attempts, placement, checkpoint, cost, conditions and events\n\nnodus diff Show what apply would change, from a server dry-run (exit 1 when there are differences)\n\nnodus edit Edit a resource in $EDITOR and save it as a merge patch\n\nnodus events List events in the project, or for one object, and watch for new ones\n\nnodus exec Run a command in a running Job, Sandbox or Workspace (recorded as a Process)\n\nnodus explain Show the documentation of a kind or one of its fields\n\nnodus get List or show resources\n\nnodus get transactions List postings on the org's wallet with the balance after each\n\nnodus get usage Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank\n\nnodus inference Call hosted models: list them, send a chat completion and read a request's receipt\n\nnodus inference chat Send one chat completion and print the answer; the request id and charge go to stderr\n\nnodus inference models List the models you can call now\n\nnodus inference receipt Show what one inference request was charged (receipts are kept 30 days)\n\nnodus init Scaffold a manifest or app.py in the current directory\n\nnodus label Set or remove labels\n\nnodus login Sign in through the console and store one API key per chosen org\n\nnodus logout Revoke the context's CLI key and remove the context\n\nnodus logs Print or follow a resource's output\n\nnodus mcp Serve Nodus tools to a local MCP client over stdio\n\nnodus mcp install Add the Nodus MCP server to Claude Code, Cursor or Codex\n\nnodus open Open a resource's console page, or a port of its container, in the browser\n\nnodus patch Update fields of a resource with a merge or JSON patch\n\nnodus pool Pool utilization, forecasts, recommendations and actions\n\nnodus pool approve Approve a proposed PoolAction\n\nnodus pool forecast Show a pool's demand forecast with p50 and p90 bands\n\nnodus pool pause Pause every automated action in a pool\n\nnodus pool recommendations List a pool's recommendations\n\nnodus pool reject Reject a proposed PoolAction\n\nnodus pool resume Resume every automated action in a pool\n\nnodus pool revert Undo an executed PoolAction where the action allows it\n\nnodus pool utilization Show a pool's GPU utilization and device-hour ledger\n\nnodus port-forward Forward local ports to a Job, Sandbox or Workspace\n\nnodus request Request an action: restart, refresh-secrets, reload-sinks, verify, resend, start, reimport, approve-burst, drain\n\nnodus resume Resume a suspended run\n\nnodus rollout Roll out changes\n\nnodus rollout restart Restart without a spec change (alias of \\ request restart\\ )\n\nnodus run Run a command as a Job: estimate, phases, logs and the final cost\n\nnodus secret Create Secrets from literals, env files or registry logins\n\nnodus secret create Create a Secret\n\nnodus serve Run the App in APP.py and redeploy it whenever a file changes\n\nnodus shell Open an interactive shell in a Sandbox, creating it if it is not there\n\nnodus ssh Connect to a Workspace over SSH\n\nnodus start Start a stopped Sandbox, Workspace, Function, Agent or InferenceEndpoint\n\nnodus stop Stop a Sandbox, Workspace, Function, Agent or InferenceEndpoint\n\nnodus suspend Suspend a Job, Pipeline, Sweep, TrainingJob or AgentRun (checkpointing first)\n\nnodus top Show GPU, CPU and memory use, the hourly rate and spend of running Jobs\n\nnodus version Print the client version\n\nnodus volume Put, get, list and remove the files of a Volume\n\nnodus volume clear Commit an empty revision; the Volume and its history stay\n\nnodus volume get Download a file or directory from a Volume\n\nnodus volume ls List a directory of a Volume\n\nnodus volume put Add a local file or directory to a Volume as a new revision\n\nnodus volume rm Remove a file or directory from a Volume as a new revision\n\nnodus wait Wait until a resource meets a condition\n\nnodus whoami Show the signed-in principal, org, role, project, scopes and available credit"},{"id":"docs/reference/cli/nodus","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus.md","title":"nodus","description":"Run compute Jobs on Nodus","stage":"GA","headings":[{"depth":2,"slug":"nodus","text":"nodus"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus\n\n Run compute Jobs on Nodus\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus\n\nRun compute Jobs on Nodus\n\nSynopsis\n\nnodus submits and manages Nodus resources: Jobs, Sandboxes, Functions, Agents and more.\n\nOptions\n\n --context string Context from the config file to use\n -h, --help help for nodus\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus agentrun - Message an AgentRun and read its steps and answer\n nodus annotate - Set or remove annotations\n nodus api-resources - List the kinds the server offers\n nodus apply - Create or update resources from manifests (client-side three-way merge)\n nodus attach - Attach to a running Process of a Job, Sandbox or Workspace\n nodus auth - Credentials for other tools, and permission checks\n nodus billing - Show the org’s balance; top up, redeem a code and manage auto-recharge\n nodus cancel - Cancel a run, FunctionCall or Process\n nodus cloud - Inspect connected cloud accounts\n nodus completion - Generate the autocompletion script for the specified shell\n nodus config - Show and change contexts in \\~/.nodus/config\n nodus convert - Turn a 0.x nodus.toml or training file into manifests for nodus apply\n nodus cp - Copy files to and from a running Sandbox, Workspace or Job, or download a Job output\n nodus create - Create resources from manifests or with a typed generator\n nodus delete - Delete resources\n nodus deploy - Deploy the nodus.App in APP.py as a persistent App\n nodus describe - Show everything about one resource: attempts, placement, checkpoint, cost, conditions and events\n nodus diff - Show what apply would change, from a server dry-run (exit 1 when there are differences)\n nodus edit - Edit a resource in $EDITOR and save it as a merge patch\n nodus events - List events in the project, or for one object, and watch for new ones\n nodus exec - Run a command in a running Job, Sandbox or Workspace (recorded as a Process)\n nodus explain - Show the documentation of a kind or one of its fields\n nodus get - List or show resources\n nodus inference - Call hosted models: list them, send a chat completion and read a request’s receipt\n nodus init - Scaffold a manifest or app.py in the current directory\n nodus label - Set or remove labels\n nodus login - Sign in through the console and store one API key per chosen org\n nodus logout - Revoke the context’s CLI key and remove the context\n nodus logs - Print or follow a resource’s output\n nodus mcp - Serve Nodus tools to a local MCP client over stdio\n nodus open - Open a resource’s console page, or a port of its container, in the browser\n nodus patch - Update fields of a resource with a merge or JSON patch\n nodus pool - Pool utilization, forecasts, recommendations and actions\n nodus port-forward - Forward local ports to a Job, Sandbox or Workspace\n nodus request - Request an action: restart, refresh-secrets, reload-sinks, verify, resend, start, reimport, approve-burst, drain\n nodus resume - Resume a suspended run\n nodus rollout - Roll out changes\n nodus run - Run a command as a Job: estimate, phases, logs and the final cost\n nodus secret - Create Secrets from literals, env files or registry logins\n nodus serve - Run the App in APP.py and redeploy it whenever a file changes\n nodus shell - Open an interactive shell in a Sandbox, creating it if it is not there\n nodus ssh - Connect to a Workspace over SSH\n nodus start - Start a stopped Sandbox, Workspace, Function, Agent or InferenceEndpoint\n nodus stop - Stop a Sandbox, Workspace, Function, Agent or InferenceEndpoint\n nodus suspend - Suspend a Job, Pipeline, Sweep, TrainingJob or AgentRun (checkpointing first)\n nodus top - Show GPU, CPU and memory use, the hourly rate and spend of running Jobs\n nodus version - Print the client version\n nodus volume - Put, get, list and remove the files of a Volume\n nodus wait - Wait until a resource meets a condition\n nodus whoami - Show the signed-in principal, org, role, project, scopes and available credit"},{"id":"docs/reference/cli/nodus_agentrun","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_agentrun/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_agentrun.md","title":"nodus agentrun","description":"Message an AgentRun and read its steps and answer","stage":"GA","headings":[{"depth":2,"slug":"nodus-agentrun","text":"nodus agentrun"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus agentrun\n\n Message an AgentRun and read its steps and answer\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus agentrun/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus agentrun\n\nMessage an AgentRun and read its steps and answer\n\nOptions\n\n -h, --help help for agentrun\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus agentrun answer - Print an AgentRun’s last answer in full\n nodus agentrun send - Send a message to an AgentRun, which takes it as its next prompt\n nodus agentrun steps - List an AgentRun’s journal steps: what it did, in order"},{"id":"docs/reference/cli/nodus_agentrun_answer","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_agentrun_answer/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_agentrun_answer.md","title":"nodus agentrun answer","description":"Print an AgentRun's last answer in full","stage":"GA","headings":[{"depth":2,"slug":"nodus-agentrun-answer","text":"nodus agentrun answer"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus agentrun answer\n\n Print an AgentRun's last answer in full\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus agentrun answer/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus agentrun answer\n\nPrint an AgentRun’s last answer in full\n\nnodus agentrun answer NAME [flags]\n\nExamples\n\n nodus agentrun answer triage\n\nOptions\n\n -h, --help help for answer\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus agentrun - Message an AgentRun and read its steps and answer"},{"id":"docs/reference/cli/nodus_agentrun_send","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_agentrun_send/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_agentrun_send.md","title":"nodus agentrun send","description":"Send a message to an AgentRun, which takes it as its next prompt","stage":"GA","headings":[{"depth":2,"slug":"nodus-agentrun-send","text":"nodus agentrun send"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus agentrun send\n\n Send a message to an AgentRun, which takes it as its next prompt\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus agentrun send/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus agentrun send\n\nSend a message to an AgentRun, which takes it as its next prompt\n\nSynopsis\n\nSend a message to an AgentRun. A run created with spec.keepAlive waits for messages after each turn and wakes for this one; a running run reads it after its current turn. A finished run refuses it.\n\nnodus agentrun send NAME MESSAGE [flags]\n\nExamples\n\n nodus agentrun send triage \"Now check the closed ones too\"\n nodus agentrun send triage \"Retry\" --key retry-1\n\nOptions\n\n -h, --help help for send\n --key string Message key: sending the same key again delivers the message once\n -o, --output string Output format: json prints the receipt\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus agentrun - Message an AgentRun and read its steps and answer"},{"id":"docs/reference/cli/nodus_agentrun_steps","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_agentrun_steps/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_agentrun_steps.md","title":"nodus agentrun steps","description":"List an AgentRun's journal steps: what it did, in order","stage":"GA","headings":[{"depth":2,"slug":"nodus-agentrun-steps","text":"nodus agentrun steps"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus agentrun steps\n\n List an AgentRun's journal steps: what it did, in order\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus agentrun steps/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus agentrun steps\n\nList an AgentRun’s journal steps: what it did, in order\n\nnodus agentrun steps NAME [flags]\n\nExamples\n\n nodus agentrun steps triage\n\nOptions\n\n -h, --help help for steps\n -o, --output string Output format: json\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus agentrun - Message an AgentRun and read its steps and answer"},{"id":"docs/reference/cli/nodus_annotate","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_annotate/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_annotate.md","title":"nodus annotate","description":"Set or remove annotations","stage":"GA","headings":[{"depth":2,"slug":"nodus-annotate","text":"nodus annotate"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus annotate\n\n Set or remove annotations\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus annotate/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus annotate\n\nSet or remove annotations\n\nnodus annotate KIND/NAME KEY=VALUE... [KEY-] [flags]\n\nOptions\n\n -h, --help help for annotate\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus annotate job/train note=hi"},{"id":"docs/reference/cli/nodus_api-resources","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_api-resources/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_api-resources.md","title":"nodus api-resources","description":"List the kinds the server offers","stage":"GA","headings":[{"depth":2,"slug":"nodus-api-resources","text":"nodus api-resources"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus api-resources\n\n List the kinds the server offers\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus api-resources/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus api-resources\n\nList the kinds the server offers\n\nnodus api-resources [flags]\n\nOptions\n\n -h, --help help for api-resources\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus api-resources"},{"id":"docs/reference/cli/nodus_apply","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_apply/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_apply.md","title":"nodus apply","description":"Create or update resources from manifests (client-side three-way merge)","stage":"GA","headings":[{"depth":2,"slug":"nodus-apply","text":"nodus apply"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus apply\n\n Create or update resources from manifests (client-side three-way merge)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus apply/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus apply\n\nCreate or update resources from manifests (client-side three-way merge)\n\nnodus apply -f FILE DIR - [flags]\n\nExamples\n\n nodus apply -f job.yaml\n nodus apply -f job.yaml --dry-run=server -o estimate\n nodus apply -f manifests/ --prune -l app=trainer\n\nOptions\n\n --dry-run string[=\"server\"] server returns the estimate without creating; none, server or client (default \"none\")\n -f, --filename strings Manifest files or directories; - reads stdin\n -h, --help help for apply\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --prune Delete objects matching -l that the manifests no longer contain\n -R, --recursive Walk directories recursively\n -l, --selector string Label selector for --prune\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus apply -f keep.yaml --dry-run=client\nnodus apply -f keep.yaml -f old.yaml -f other.yaml\nnodus apply -f keep.yaml --prune -l app=trainer --dry-run=server\nnodus apply -f keep.yaml --prune -l app=trainer --dry-run=client -o name\nnodus apply -f keep.yaml --prune -l app=trainer\nnodus apply -f jobs.yaml\nnodus apply -f job.yaml\nnodus apply -f job-v2.yaml\nnodus apply -f job.yaml --dry-run=server -o estimate\nnodus apply -f bad.yaml\nnodus apply -f dev.yaml\nnodus apply -f pool.yaml"},{"id":"docs/reference/cli/nodus_attach","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_attach/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_attach.md","title":"nodus attach","description":"Attach to a running Process of a Job, Sandbox or Workspace","stage":"GA","headings":[{"depth":2,"slug":"nodus-attach","text":"nodus attach"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus attach\n\n Attach to a running Process of a Job, Sandbox or Workspace\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus attach/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus attach\n\nAttach to a running Process of a Job, Sandbox or Workspace\n\nnodus attach KIND/NAME --process N [-it] [flags]\n\nExamples\n\n nodus attach sb/dev --process 12 -it\n nodus attach job/train --process 3 --rank 1\n\nOptions\n\n -h, --help help for attach\n --index int Job index (Indexed Jobs) (default -1)\n --process nodus get processes Process sequence number (as nodus get processes shows) or full name\n --rank string Gang member rank (default 0)\n -i, --stdin Pass stdin to the process\n -t, --tty The process has a terminal\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus attach sb/dev --process 1"},{"id":"docs/reference/cli/nodus_auth","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_auth/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_auth.md","title":"nodus auth","description":"Credentials for other tools, and permission checks","stage":"GA","headings":[{"depth":2,"slug":"nodus-auth","text":"nodus auth"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus auth\n\n Credentials for other tools, and permission checks\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus auth/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus auth\n\nCredentials for other tools, and permission checks\n\nOptions\n\n -h, --help help for auth\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus auth can-i - Check whether the current credential may perform an action\n nodus auth token - Print an ExecCredential for kubectl and client-go (the kubeconfig exec plugin)"},{"id":"docs/reference/cli/nodus_auth_can-i","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_auth_can-i/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_auth_can-i.md","title":"nodus auth can-i","description":"Check whether the current credential may perform an action","stage":"GA","headings":[{"depth":2,"slug":"nodus-auth-can-i","text":"nodus auth can-i"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus auth can-i\n\n Check whether the current credential may perform an action\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus auth can-i/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus auth can-i\n\nCheck whether the current credential may perform an action\n\nnodus auth can-i VERB RESOURCE [flags]\n\nExamples\n\n nodus auth can-i create jobs\n nodus auth can-i exec sandboxes -p research\n\nOptions\n\n -h, --help help for can-i\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus auth - Credentials for other tools, and permission checks\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus auth can-i create jobs\nnodus auth can-i get jobs\nnodus auth can-i exec sandboxes"},{"id":"docs/reference/cli/nodus_auth_token","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_auth_token/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_auth_token.md","title":"nodus auth token","description":"Print an ExecCredential for kubectl and client-go (the kubeconfig exec plugin)","stage":"GA","headings":[{"depth":2,"slug":"nodus-auth-token","text":"nodus auth token"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus auth token\n\n Print an ExecCredential for kubectl and client-go (the kubeconfig exec plugin)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus auth token/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus auth token\n\nPrint an ExecCredential for kubectl and client-go (the kubeconfig exec plugin)\n\nnodus auth token [flags]\n\nOptions\n\n -h, --help help for token\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus auth - Credentials for other tools, and permission checks\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus auth token"},{"id":"docs/reference/cli/nodus_billing","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_billing/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_billing.md","title":"nodus billing","description":"Show the org's balance; top up, redeem a code and manage auto-recharge","stage":"GA","headings":[{"depth":2,"slug":"nodus-billing","text":"nodus billing"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus billing\n\n Show the org's balance; top up, redeem a code and manage auto-recharge\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus billing/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus billing\n\nShow the org’s balance; top up, redeem a code and manage auto-recharge\n\nSynopsis\n\nWithout a subcommand, billing summarizes the org’s BillingAccount: what is available to spend, what open holds reserve, spend this month against the fullest budget, and auto-recharge. Billing changes need the billing:write scope (Owners and Admins).\n\nnodus billing [flags]\n\nExamples\n\n nodus billing\n nodus billing top-up 20\n nodus billing redeem LAUNCH25\n nodus billing auto-recharge --threshold 10 --amount 50\n\nOptions\n\n -h, --help help for billing\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus billing auto-recharge - Top up automatically when available credit falls below a threshold\n nodus billing deal - View partner offers or request an eligibility review\n nodus billing portal - Open the Stripe Customer Portal: cards, invoices and billing details\n nodus billing receipts - List top-ups with their Stripe receipt links\n nodus billing redeem - Redeem a promo code for credit\n nodus billing top-up - Buy prepaid credit with Stripe Checkout ($5–$1,000)\n nodus billing usage - Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank"},{"id":"docs/reference/cli/nodus_billing_auto-recharge","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_billing_auto-recharge/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_billing_auto-recharge.md","title":"nodus billing auto-recharge","description":"Top up automatically when available credit falls below a threshold","stage":"GA","headings":[{"depth":2,"slug":"nodus-billing-auto-recharge","text":"nodus billing auto-recharge"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus billing auto-recharge\n\n Top up automatically when available credit falls below a threshold\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus billing auto-recharge/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus billing auto-recharge\n\nTop up automatically when available credit falls below a threshold\n\nSynopsis\n\nauto-recharge charges the saved card for –amount whenever available credit falls below –threshold. With no card on file it opens Stripe’s card setup first and turns auto-recharge on once the card is saved. Three failed charges in a row turn it off and notify the org’s Owners and Admins.\n\nnodus billing auto-recharge --threshold USD --amount USD --off [flags]\n\nExamples\n\n nodus billing auto-recharge --threshold 10 --amount 50\n nodus billing auto-recharge --off\n\nOptions\n\n --amount string Dollars to add each time ($5–$1,000) (default \"20\")\n -h, --help help for auto-recharge\n --no-open Print the card setup link instead of opening the browser\n --no-wait With no card on file, print the setup link and return without turning auto-recharge on\n --off Turn auto-recharge off\n --threshold string Recharge when available credit falls below this many dollars (at least 5) (default \"10\")\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus billing - Show the org’s balance; top up, redeem a code and manage auto-recharge"},{"id":"docs/reference/cli/nodus_billing_deal","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_billing_deal/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_billing_deal.md","title":"nodus billing deal","description":"View partner offers or request an eligibility review","stage":"GA","headings":[{"depth":2,"slug":"nodus-billing-deal","text":"nodus billing deal"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus billing deal\n\n View partner offers or request an eligibility review\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus billing deal/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus billing deal\n\nView partner offers or request an eligibility review\n\nnodus billing deal [CODE] [flags]\n\nExamples\n\n nodus billing deal\n nodus billing deal YC --note 'Company and batch'\n\nOptions\n\n -h, --help help for deal\n --note string Company or batch information for eligibility review\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus billing - Show the org’s balance; top up, redeem a code and manage auto-recharge"},{"id":"docs/reference/cli/nodus_billing_portal","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_billing_portal/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_billing_portal.md","title":"nodus billing portal","description":"Open the Stripe Customer Portal: cards, invoices and billing details","stage":"GA","headings":[{"depth":2,"slug":"nodus-billing-portal","text":"nodus billing portal"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus billing portal\n\n Open the Stripe Customer Portal: cards, invoices and billing details\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus billing portal/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus billing portal\n\nOpen the Stripe Customer Portal: cards, invoices and billing details\n\nnodus billing portal [flags]\n\nOptions\n\n -h, --help help for portal\n --print Print the link instead of opening the browser\n --setup Add a card without buying credits\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus billing - Show the org’s balance; top up, redeem a code and manage auto-recharge"},{"id":"docs/reference/cli/nodus_billing_receipts","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_billing_receipts/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_billing_receipts.md","title":"nodus billing receipts","description":"List top-ups with their Stripe receipt links","stage":"GA","headings":[{"depth":2,"slug":"nodus-billing-receipts","text":"nodus billing receipts"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus billing receipts\n\n List top-ups with their Stripe receipt links\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus billing receipts/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus billing receipts\n\nList top-ups with their Stripe receipt links\n\nnodus billing receipts [flags]\n\nOptions\n\n -h, --help help for receipts\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus billing - Show the org’s balance; top up, redeem a code and manage auto-recharge"},{"id":"docs/reference/cli/nodus_billing_redeem","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_billing_redeem/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_billing_redeem.md","title":"nodus billing redeem","description":"Redeem a promo code for credit","stage":"GA","headings":[{"depth":2,"slug":"nodus-billing-redeem","text":"nodus billing redeem"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus billing redeem\n\n Redeem a promo code for credit\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus billing redeem/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus billing redeem\n\nRedeem a promo code for credit\n\nSynopsis\n\nredeem adds the code’s credit to the org as a CreditGrant. Redeeming the same code again changes nothing.\n\nnodus billing redeem CODE [flags]\n\nExamples\n\n nodus billing redeem LAUNCH25\n\nOptions\n\n -h, --help help for redeem\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus billing - Show the org’s balance; top up, redeem a code and manage auto-recharge"},{"id":"docs/reference/cli/nodus_billing_top-up","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_billing_top-up/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_billing_top-up.md","title":"nodus billing top-up","description":"Buy prepaid credit with Stripe Checkout ($5–$1,000)","stage":"GA","headings":[{"depth":2,"slug":"nodus-billing-top-up","text":"nodus billing top-up"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus billing top-up\n\n Buy prepaid credit with Stripe Checkout ($5–$1,000)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus billing top-up/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus billing top-up\n\nBuy prepaid credit with Stripe Checkout ($5–$1,000)\n\nSynopsis\n\ntop-up creates a TopUp, opens its Stripe Checkout page and waits until the payment is credited. Top-ups settle any arrears before they credit the wallet. Stripe emails the receipt to the billing email.\n\nnodus billing top-up USD [flags]\n\nExamples\n\n nodus billing top-up 20\n nodus billing top-up 50 --save-card\n nodus billing top-up 20 --no-open --no-wait\n\nOptions\n\n -h, --help help for top-up\n --no-open Print the Checkout link instead of opening the browser\n --no-wait Return once the Checkout link exists, without waiting for the payment\n --save-card Save the card for auto-recharge\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus billing - Show the org’s balance; top up, redeem a code and manage auto-recharge"},{"id":"docs/reference/cli/nodus_billing_usage","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_billing_usage/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_billing_usage.md","title":"nodus billing usage","description":"Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank","stage":"GA","headings":[{"depth":2,"slug":"nodus-billing-usage","text":"nodus billing usage"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus billing usage\n\n Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus billing usage/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus billing usage\n\nShow metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank\n\nSynopsis\n\nusage lists UsageRecords, or with –group-by sums them into a UsageSummary. Group keys are project, kind, meter, day, segment (Boot, Running, Restore, Teardown), offering (the capacity it ran on), rank (a gang member), member (the person who started the work, with any of their keys) and label:. Every billed interval of your allocation is on a line, from its billing start to confirmed deletion.\n\nnodus billing usage [--group-by KEY,...] [--since 30d] [--until TIME] [-o csv json yaml wide] [flags]\n\nExamples\n\n nodus get usage --group-by member --since 30d\n nodus get usage --group-by project,label:team --since 30d\n nodus get usage --object job/ddp-2node --group-by rank,segment\n nodus get usage --group-by day,meter --since 7d -o csv usage.csv\n\nOptions\n\n --field-selector string Field selector, such as meter=ComputeSeconds\n --group-by strings Sum by these keys: project, kind, meter, day, segment, offering, rank, member, label: \n -h, --help help for usage\n --no-headers Omit table headers\n --object string Only the usage of one object, as KIND/NAME\n -o, --output string Output format: wide, csv, json or yaml\n --since string Start of the range: a duration back from now (30d, 12h) or a time\n --until string End of the range (default now)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus billing - Show the org’s balance; top up, redeem a code and manage auto-recharge"},{"id":"docs/reference/cli/nodus_cancel","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_cancel/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_cancel.md","title":"nodus cancel","description":"Cancel a run, FunctionCall or Process","stage":"GA","headings":[{"depth":2,"slug":"nodus-cancel","text":"nodus cancel"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus cancel\n\n Cancel a run, FunctionCall or Process\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus cancel/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus cancel\n\nCancel a run, FunctionCall or Process\n\nnodus cancel KIND/NAME [flags]\n\nOptions\n\n -h, --help help for cancel\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus cancel job/train"},{"id":"docs/reference/cli/nodus_cloud","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_cloud/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_cloud.md","title":"nodus cloud","description":"Inspect connected cloud accounts","stage":"GA","headings":[{"depth":2,"slug":"nodus-cloud","text":"nodus cloud"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus cloud\n\n Inspect connected cloud accounts\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus cloud/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus cloud\n\nInspect connected cloud accounts\n\nOptions\n\n -h, --help help for cloud\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus cloud inventory - List the instances, GPUs or spend a cloud account’s last sync saw"},{"id":"docs/reference/cli/nodus_cloud_inventory","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_cloud_inventory/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_cloud_inventory.md","title":"nodus cloud inventory","description":"List the instances, GPUs or spend a cloud account's last sync saw","stage":"GA","headings":[{"depth":2,"slug":"nodus-cloud-inventory","text":"nodus cloud inventory"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus cloud inventory\n\n List the instances, GPUs or spend a cloud account's last sync saw\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus cloud inventory/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus cloud inventory\n\nList the instances, GPUs or spend a cloud account’s last sync saw\n\nnodus cloud inventory ACCOUNT [flags]\n\nExamples\n\n nodus cloud inventory aws-main\n nodus cloud inventory aws-main --resource gpus\n\nOptions\n\n -h, --help help for inventory\n --json Print the response as JSON\n --resource string instances, gpus or spend (default instances)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus cloud - Inspect connected cloud accounts\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus cloud inventory aws-main"},{"id":"docs/reference/cli/nodus_completion","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_completion/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_completion.md","title":"nodus completion","description":"Generate the autocompletion script for the specified shell","stage":"GA","headings":[{"depth":2,"slug":"nodus-completion","text":"nodus completion"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus completion\n\n Generate the autocompletion script for the specified shell\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus completion/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus completion\n\nGenerate the autocompletion script for the specified shell\n\nSynopsis\n\nGenerate the autocompletion script for nodus for the specified shell. See each sub-command’s help for details on how to use the generated script.\n\nOptions\n\n -h, --help help for completion\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus completion bash - Generate the autocompletion script for bash\n nodus completion fish - Generate the autocompletion script for fish\n nodus completion powershell - Generate the autocompletion script for powershell\n nodus completion zsh - Generate the autocompletion script for zsh"},{"id":"docs/reference/cli/nodus_completion_bash","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_completion_bash/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_completion_bash.md","title":"nodus completion bash","description":"Generate the autocompletion script for bash","stage":"GA","headings":[{"depth":2,"slug":"nodus-completion-bash","text":"nodus completion bash"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":4,"slug":"linux","text":"Linux:"},{"depth":4,"slug":"macos","text":"macOS:"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus completion bash\n\n Generate the autocompletion script for bash\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus completion bash/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus completion bash\n\nGenerate the autocompletion script for bash\n\nSynopsis\n\nGenerate the autocompletion script for the bash shell.\n\nThis script depends on the ‘bash-completion’ package. If it is not installed already, you can install it via your OS’s package manager.\n\nTo load completions in your current shell session:\n\nsource <(nodus completion bash)\n\nTo load completions for every new session, execute once:\n\nLinux:\n\nnodus completion bash /etc/bash completion.d/nodus\n\nmacOS:\n\nnodus completion bash $(brew --prefix)/etc/bash completion.d/nodus\n\nYou will need to start a new shell for this setup to take effect.\n\nnodus completion bash\n\nOptions\n\n -h, --help help for bash\n --no-descriptions disable completion descriptions\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus completion - Generate the autocompletion script for the specified shell"},{"id":"docs/reference/cli/nodus_completion_fish","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_completion_fish/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_completion_fish.md","title":"nodus completion fish","description":"Generate the autocompletion script for fish","stage":"GA","headings":[{"depth":2,"slug":"nodus-completion-fish","text":"nodus completion fish"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus completion fish\n\n Generate the autocompletion script for fish\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus completion fish/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus completion fish\n\nGenerate the autocompletion script for fish\n\nSynopsis\n\nGenerate the autocompletion script for the fish shell.\n\nTo load completions in your current shell session:\n\nnodus completion fish source\n\nTo load completions for every new session, execute once:\n\nnodus completion fish ~/.config/fish/completions/nodus.fish\n\nYou will need to start a new shell for this setup to take effect.\n\nnodus completion fish [flags]\n\nOptions\n\n -h, --help help for fish\n --no-descriptions disable completion descriptions\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus completion - Generate the autocompletion script for the specified shell"},{"id":"docs/reference/cli/nodus_completion_powershell","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_completion_powershell/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_completion_powershell.md","title":"nodus completion powershell","description":"Generate the autocompletion script for powershell","stage":"GA","headings":[{"depth":2,"slug":"nodus-completion-powershell","text":"nodus completion powershell"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus completion powershell\n\n Generate the autocompletion script for powershell\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus completion powershell/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus completion powershell\n\nGenerate the autocompletion script for powershell\n\nSynopsis\n\nGenerate the autocompletion script for powershell.\n\nTo load completions in your current shell session:\n\nnodus completion powershell Out-String Invoke-Expression\n\nTo load completions for every new session, add the output of the above command to your powershell profile.\n\nnodus completion powershell [flags]\n\nOptions\n\n -h, --help help for powershell\n --no-descriptions disable completion descriptions\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus completion - Generate the autocompletion script for the specified shell"},{"id":"docs/reference/cli/nodus_completion_zsh","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_completion_zsh/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_completion_zsh.md","title":"nodus completion zsh","description":"Generate the autocompletion script for zsh","stage":"GA","headings":[{"depth":2,"slug":"nodus-completion-zsh","text":"nodus completion zsh"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":4,"slug":"linux","text":"Linux:"},{"depth":4,"slug":"macos","text":"macOS:"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus completion zsh\n\n Generate the autocompletion script for zsh\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus completion zsh/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus completion zsh\n\nGenerate the autocompletion script for zsh\n\nSynopsis\n\nGenerate the autocompletion script for the zsh shell.\n\nIf shell completion is not already enabled in your environment you will need to enable it. You can execute the following once:\n\necho \"autoload -U compinit; compinit\" ~/.zshrc\n\nTo load completions in your current shell session:\n\nsource <(nodus completion zsh)\n\nTo load completions for every new session, execute once:\n\nLinux:\n\nnodus completion zsh \"${fpath[1]}/ nodus\"\n\nmacOS:\n\nnodus completion zsh $(brew --prefix)/share/zsh/site-functions/ nodus\n\nYou will need to start a new shell for this setup to take effect.\n\nnodus completion zsh [flags]\n\nOptions\n\n -h, --help help for zsh\n --no-descriptions disable completion descriptions\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus completion - Generate the autocompletion script for the specified shell\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus completion zsh"},{"id":"docs/reference/cli/nodus_config","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_config/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_config.md","title":"nodus config","description":"Show and change contexts in ~/.nodus/config","stage":"GA","headings":[{"depth":2,"slug":"nodus-config","text":"nodus config"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus config\n\n Show and change contexts in ~/.nodus/config\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus config/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus config\n\nShow and change contexts in \\~/.nodus/config\n\nOptions\n\n -h, --help help for config\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus config current-context - Print the current context\n nodus config get-contexts - List contexts\n nodus config set-context - Change a context’s project or server (NAME defaults to the current context)\n nodus config use-context - Switch the current context\n nodus config view - Print the config (keys stay in the keychain)"},{"id":"docs/reference/cli/nodus_config_current-context","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_config_current-context/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_config_current-context.md","title":"nodus config current-context","description":"Print the current context","stage":"GA","headings":[{"depth":2,"slug":"nodus-config-current-context","text":"nodus config current-context"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus config current-context\n\n Print the current context\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus config current-context/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus config current-context\n\nPrint the current context\n\nnodus config current-context [flags]\n\nOptions\n\n -h, --help help for current-context\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus config - Show and change contexts in \\~/.nodus/config\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus config current-context"},{"id":"docs/reference/cli/nodus_config_get-contexts","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_config_get-contexts/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_config_get-contexts.md","title":"nodus config get-contexts","description":"List contexts","stage":"GA","headings":[{"depth":2,"slug":"nodus-config-get-contexts","text":"nodus config get-contexts"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus config get-contexts\n\n List contexts\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus config get-contexts/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus config get-contexts\n\nList contexts\n\nnodus config get-contexts [flags]\n\nOptions\n\n -h, --help help for get-contexts\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus config - Show and change contexts in \\~/.nodus/config\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus config get-contexts"},{"id":"docs/reference/cli/nodus_config_set-context","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_config_set-context/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_config_set-context.md","title":"nodus config set-context","description":"Change a context's project or server (NAME defaults to the current context)","stage":"GA","headings":[{"depth":2,"slug":"nodus-config-set-context","text":"nodus config set-context"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus config set-context\n\n Change a context's project or server (NAME defaults to the current context)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus config set-context/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus config set-context\n\nChange a context’s project or server (NAME defaults to the current context)\n\nnodus config set-context NAME [--project P] [--server URL] [flags]\n\nOptions\n\n -h, --help help for set-context\n --project string Project for this context\n --server string API server for this context\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus config - Show and change contexts in \\~/.nodus/config\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus config set-context acme --project research"},{"id":"docs/reference/cli/nodus_config_use-context","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_config_use-context/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_config_use-context.md","title":"nodus config use-context","description":"Switch the current context","stage":"GA","headings":[{"depth":2,"slug":"nodus-config-use-context","text":"nodus config use-context"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus config use-context\n\n Switch the current context\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus config use-context/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus config use-context\n\nSwitch the current context\n\nnodus config use-context NAME [flags]\n\nOptions\n\n -h, --help help for use-context\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus config - Show and change contexts in \\~/.nodus/config\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus config use-context nope"},{"id":"docs/reference/cli/nodus_config_view","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_config_view/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_config_view.md","title":"nodus config view","description":"Print the config (keys stay in the keychain)","stage":"GA","headings":[{"depth":2,"slug":"nodus-config-view","text":"nodus config view"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus config view\n\n Print the config (keys stay in the keychain)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus config view/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus config view\n\nPrint the config (keys stay in the keychain)\n\nnodus config view [flags]\n\nOptions\n\n -h, --help help for view\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus config - Show and change contexts in \\~/.nodus/config"},{"id":"docs/reference/cli/nodus_convert","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_convert/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_convert.md","title":"nodus convert","description":"Turn a 0.x nodus.toml or training file into manifests for nodus apply","stage":"GA","headings":[{"depth":2,"slug":"nodus-convert","text":"nodus convert"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus convert\n\n Turn a 0.x nodus.toml or training file into manifests for nodus apply\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus convert/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus convert\n\nTurn a 0.x nodus.toml or training file into manifests for nodus apply\n\nSynopsis\n\nconvert maps a 0.x nodus.toml or nodus.json onto a Job (a Pipeline when it has stages), and a nodus.training.json onto a TrainingRuntime and a TrainingJob. Every field that does not carry over is named on stderr; nothing is sent to the API.\n\nnodus convert nodus.toml nodus.json nodus.training.json [flags]\n\nExamples\n\n nodus convert nodus.toml job.yaml && nodus apply -f job.yaml\n nodus convert nodus.training.json --image ghcr.io/acme/trainer@sha256:... nodus apply -f -\n\nOptions\n\n -h, --help help for convert\n --image string Trainer image for a nodus.training.json runtime\n --name string metadata.name for the manifests (default: the file's directory name)\n -o, --output string Output format: yaml or json (default \"yaml\")\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus convert nodus.toml\nnodus convert --name 'Train Eval' nodus.toml\nnodus convert -o json payload/nodus.json\nnodus convert --image ghcr.io/acme/trainer@sha256:0123 training/nodus.training.json"},{"id":"docs/reference/cli/nodus_cp","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_cp/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_cp.md","title":"nodus cp","description":"Copy files to and from a running Sandbox, Workspace or Job, or download a Job output","stage":"GA","headings":[{"depth":2,"slug":"nodus-cp","text":"nodus cp"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus cp\n\n Copy files to and from a running Sandbox, Workspace or Job, or download a Job output\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus cp/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus cp\n\nCopy files to and from a running Sandbox, Workspace or Job, or download a Job output\n\nSynopsis\n\nEither side is KIND/NAME:/absolute/path inside the running container; the other side is a local path. Directories copy recursively. KIND/NAME:outputs/OUTPUT downloads a declared output of a Job, verified against its SHA-256 and written atomically; KIND/NAME:outputs/ or outputs/PREFIX/ downloads every output under it into a directory, so a TrainingJob’s adapter comes down in one command.\n\nnodus cp SRC DST [flags]\n\nExamples\n\n nodus cp ./data sb/dev:/workspace/data\n nodus cp ws/lab:/home/nodus/results ./results\n nodus cp job/train:outputs/model ./model\n nodus cp trainingjob/tune:outputs/ ./outputs\n nodus cp job/ddp:outputs/metrics . --rank 1\n\nOptions\n\n -h, --help help for cp\n --index int Job index (Indexed Jobs) (default -1)\n --rank string Gang member rank (default 0)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus cp job/train:outputs/model model.bin\nnodus cp job/train:outputs/model out\nnodus cp want-model.bin sb/dev:/workspace/\nnodus cp src sb/dev:/workspace/src\nnodus cp sb/dev:/workspace back\nnodus cp sb/dev:/workspace/src/a.bin single.bin\nnodus cp sb/dev:/workspace/missing gone.bin\nnodus cp want-model.bin sb/dev:/hostile/ok.bin\nnodus cp sb/dev:/hostile got"},{"id":"docs/reference/cli/nodus_create","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create.md","title":"nodus create","description":"Create resources from manifests or with a typed generator","stage":"GA","headings":[{"depth":2,"slug":"nodus-create","text":"nodus create"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus create\n\n Create resources from manifests or with a typed generator\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create\n\nCreate resources from manifests or with a typed generator\n\nnodus create -f FILE create KIND NAME [flags]\n\nOptions\n\n --dry-run string[=\"server\"] server returns the estimate without creating; none, server or client (default \"none\")\n -f, --filename strings Manifest files or directories; - reads stdin\n -h, --help help for create\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n -R, --recursive Walk directories recursively\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus create agentgroup - Create a agentgroup\n nodus create agentrun - Create a agentrun\n nodus create apikey - Create an API key (printed once)\n nodus create budget - Create a budget\n nodus create enrollmenttoken - Create a token that enrolls one host into a Pool (installer printed once)\n nodus create inferenceendpoint - Create a inferenceendpoint\n nodus create invite - Invite someone to the org by email\n nodus create sandbox - Create a sandbox\n nodus create secret - Create a secret\n nodus create sshkey - Add an SSH public key, or generate a key pair and add it\n nodus create token - Mint a ServiceAccount API key (printed once)\n nodus create volume - Create a volume\n nodus create workspace - Create a workspace\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus create -f generated.yaml -o name\nnodus create -f stray.yaml"},{"id":"docs/reference/cli/nodus_create_agentgroup","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_agentgroup/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_agentgroup.md","title":"nodus create agentgroup","description":"Create a agentgroup","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-agentgroup","text":"nodus create agentgroup"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus create agentgroup\n\n Create a agentgroup\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create agentgroup/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create agentgroup\n\nCreate a agentgroup\n\nnodus create agentgroup NAME [flags]\n\nExamples\n\n nodus create agentgroup triage --agent researcher --max-active 8 --max-cost 30\n\nOptions\n\n --agent string Agent every run of the group runs\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n -h, --help help for agentgroup\n --max-active int Runs that run at once\n --max-cost string Spend limit in USD shared by every run\n --max-pending int Unfinished runs the group admits\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator"},{"id":"docs/reference/cli/nodus_create_agentrun","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_agentrun/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_agentrun.md","title":"nodus create agentrun","description":"Create a agentrun","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-agentrun","text":"nodus create agentrun"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus create agentrun\n\n Create a agentrun\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create agentrun/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create agentrun\n\nCreate a agentrun\n\nnodus create agentrun NAME [flags]\n\nExamples\n\n nodus create agentrun triage --agent researcher --prompt \"Summarize the open issues\"\n\nOptions\n\n --agent string Agent to run: one in the project, or a template of project nodus\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n -h, --help help for agentrun\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --prompt string The first prompt\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator"},{"id":"docs/reference/cli/nodus_create_apikey","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_apikey/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_apikey.md","title":"nodus create apikey","description":"Create an API key (printed once)","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-apikey","text":"nodus create apikey"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus create apikey\n\n Create an API key (printed once)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create apikey/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create apikey\n\nCreate an API key (printed once)\n\nnodus create apikey NAME [flags]\n\nExamples\n\n nodus create apikey ci --scopes jobs:write,volumes:read --projects research --expires 720h\n\nOptions\n\n --description string What the key is for (at most 256 characters)\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n --expires string Lifetime of the key, such as 720h (default: it does not expire)\n -h, --help help for apikey\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --projects strings Projects the key may reach, comma-separated or repeated (default: all)\n --scopes strings Scopes such as jobs:write, comma-separated or repeated (default: whatever your role allows)\n --service-account string Bind the key to this ServiceAccount of the project (-p)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus create apikey ci --scopes jobs:write,volumes:read --projects research --expires 720h\nnodus create apikey ci2 -o name\nnodus create apikey bot --service-account deploy --dry-run=client"},{"id":"docs/reference/cli/nodus_create_budget","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_budget/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_budget.md","title":"nodus create budget","description":"Create a budget","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-budget","text":"nodus create budget"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus create budget\n\n Create a budget\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create budget/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create budget\n\nCreate a budget\n\nnodus create budget NAME [flags]\n\nExamples\n\n nodus create budget research --limit 2000 --period Monthly --scope-project research\n nodus create budget ada-monthly --limit 300 --period Monthly --scope-member ada@example.com\n\nOptions\n\n --action string Block or Notify\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n -h, --help help for budget\n --limit string Limit in USD\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --period string Monthly or Total\n --scope-member string Limit spend of one member, by email, across every key they use\n --scope-project string Limit spend of one project\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus create budget research --limit 2000 --period Monthly --scope-project research"},{"id":"docs/reference/cli/nodus_create_enrollmenttoken","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_enrollmenttoken/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_enrollmenttoken.md","title":"nodus create enrollmenttoken","description":"Create a token that enrolls one host into a Pool (installer printed once)","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-enrollmenttoken","text":"nodus create enrollmenttoken"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus create enrollmenttoken\n\n Create a token that enrolls one host into a Pool (installer printed once)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create enrollmenttoken/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create enrollmenttoken\n\nCreate a token that enrolls one host into a Pool (installer printed once)\n\nnodus create enrollmenttoken [NAME] [flags]\n\nExamples\n\n nodus create enrollmenttoken --pool lab-a100 --ttl 2h\n\nOptions\n\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n -h, --help help for enrollmenttoken\n --label stringArray KEY=VALUE label copied onto the Node when it enrolls (repeatable)\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --pool string Pool the host joins\n --ttl string How long the token works, such as 2h (at most 24h; default 24h)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus create enrollmenttoken --pool lab --ttl 2h --label rack=a"},{"id":"docs/reference/cli/nodus_create_inferenceendpoint","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_inferenceendpoint/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_inferenceendpoint.md","title":"nodus create inferenceendpoint","description":"Create a inferenceendpoint","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-inferenceendpoint","text":"nodus create inferenceendpoint"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus create inferenceendpoint\n\n Create a inferenceendpoint\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create inferenceendpoint/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create inferenceendpoint\n\nCreate a inferenceendpoint\n\nnodus create inferenceendpoint NAME [flags]\n\nExamples\n\n nodus create inferenceendpoint support-bot --model nodus/gpt-oss-120b --rpm 120 --max-cost 5\n\nOptions\n\n --allowed-key stringArray API key name allowed to call it (repeatable)\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n -h, --help help for inferenceendpoint\n --max-concurrent int Open requests\n --max-cost string Spend limit in USD\n --model string Catalog model, as 'nodus inference models' lists them\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --rpm int Requests per minute\n --tpm int Tokens per minute\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator"},{"id":"docs/reference/cli/nodus_create_invite","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_invite/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_invite.md","title":"nodus create invite","description":"Invite someone to the org by email","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-invite","text":"nodus create invite"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus create invite\n\n Invite someone to the org by email\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create invite/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create invite\n\nInvite someone to the org by email\n\nnodus create invite --email EMAIL [--role ROLE] [--projects P,...] [flags]\n\nExamples\n\n nodus create invite --email ada@example.com --role Admin\n nodus create invite --email lin@example.com --projects research,evals\n\nOptions\n\n --email string The invitee's email address; they accept by signing in with it\n -h, --help help for invite\n --projects strings Projects the member may reach, comma-separated (default: all)\n --role string Owner, Admin, Member or Viewer (default \"Member\")\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator"},{"id":"docs/reference/cli/nodus_create_sandbox","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_sandbox/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_sandbox.md","title":"nodus create sandbox","description":"Create a sandbox","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-sandbox","text":"nodus create sandbox"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus create sandbox\n\n Create a sandbox\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create sandbox/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create sandbox\n\nCreate a sandbox\n\nnodus create sandbox NAME [flags]\n\nExamples\n\n nodus create sandbox dev --cpu 2 --memory 4Gi --idle-timeout 10m\n\nOptions\n\n --cpu string vCPU floor\n --disk string Disk floor, such as 50Gi\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n --gpu string Accelerator TYPE[:COUNT], such as L4 or H100:8\n -h, --help help for sandbox\n --idle-timeout string Stop after this idle time\n --image string Container image (default \"nodus/agent-tools\")\n --max-cost string Spend limit in USD\n --memory string Memory floor, such as 16Gi\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --secret stringArray Secret to mount (repeatable)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus create sandbox dev --cpu 2 --secret gh-token --dry-run=client\nnodus create sandbox dev --cpu 2"},{"id":"docs/reference/cli/nodus_create_secret","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_secret/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_secret.md","title":"nodus create secret","description":"Create a secret","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-secret","text":"nodus create secret"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus create secret\n\n Create a secret\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create secret/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create secret\n\nCreate a secret\n\nnodus create secret NAME [flags]\n\nExamples\n\n nodus create secret hf --from-literal HF TOKEN=hf xxx\n\nOptions\n\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n --from-env-file stringArray File of KEY=VALUE lines\n --from-literal stringArray KEY=VALUE (repeatable)\n -h, --help help for secret\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --type string Opaque or Registry\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus create secret hf --from-literal HF TOKEN=hf test --from-env-file wandb.env -o jsonpath={.spec.stringData}"},{"id":"docs/reference/cli/nodus_create_sshkey","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_sshkey/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_sshkey.md","title":"nodus create sshkey","description":"Add an SSH public key, or generate a key pair and add it","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-sshkey","text":"nodus create sshkey"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus create sshkey\n\n Add an SSH public key, or generate a key pair and add it\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create sshkey/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create sshkey\n\nAdd an SSH public key, or generate a key pair and add it\n\nnodus create sshkey [NAME] [flags]\n\nExamples\n\n nodus create sshkey laptop --from-file ~/.ssh/id ed25519.pub\n nodus create sshkey --generate\n\nOptions\n\n --display-name string A label for the console\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n --from-file string OpenSSH public key file to add, or - for standard input\n --generate Create a new Ed25519 key pair here and add its public key\n -h, --help help for sshkey\n --key-file string With --generate: where to write the private key (the public key goes next to it, with .pub) (default \"~/.ssh/id ed25519\")\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus create sshkey laptop --from-file laptop.pub\nnodus create sshkey fresh --generate --key-file id fresh"},{"id":"docs/reference/cli/nodus_create_token","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_token/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_token.md","title":"nodus create token","description":"Mint a ServiceAccount API key (printed once)","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-token","text":"nodus create token"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus create token\n\n Mint a ServiceAccount API key (printed once)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create token/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create token\n\nMint a ServiceAccount API key (printed once)\n\nnodus create token sa/NAME [flags]\n\nExamples\n\n nodus create token sa/ci --duration 720h --scope jobs:write\n\nOptions\n\n --duration duration Lifetime of the key, such as 720h (default the server's 90 days; at most 8760h)\n -h, --help help for token\n --scope stringArray Scope such as jobs:write (repeatable; default the ServiceAccount's)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus create token sa/ci --duration 720h --scope jobs:write\nnodus create token sa/ci"},{"id":"docs/reference/cli/nodus_create_volume","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_volume/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_volume.md","title":"nodus create volume","description":"Create a volume","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-volume","text":"nodus create volume"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus create volume\n\n Create a volume\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create volume/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create volume\n\nCreate a volume\n\nnodus create volume NAME [flags]\n\nExamples\n\n nodus create volume weights --size 200Gi --access-mode ReadOnlyMany\n\nOptions\n\n --access-mode string ReadWriteOnce, ReadOnlyMany or ReadWriteMany\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n -h, --help help for volume\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --size string Capacity, such as 100Gi\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator"},{"id":"docs/reference/cli/nodus_create_workspace","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_create_workspace/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_create_workspace.md","title":"nodus create workspace","description":"Create a workspace","stage":"GA","headings":[{"depth":2,"slug":"nodus-create-workspace","text":"nodus create workspace"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus create workspace\n\n Create a workspace\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus create workspace/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus create workspace\n\nCreate a workspace\n\nnodus create workspace NAME [flags]\n\nExamples\n\n nodus create workspace lab --gpu L4 --image nodus/workspace-pytorch-cuda\n\nOptions\n\n --cpu string vCPU floor\n --disk string Disk floor, such as 50Gi\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n --gpu string Accelerator TYPE[:COUNT], such as L4 or H100:8\n -h, --help help for workspace\n --image string Container image\n --memory string Memory floor, such as 16Gi\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --secret stringArray Secret to mount (repeatable)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus create - Create resources from manifests or with a typed generator"},{"id":"docs/reference/cli/nodus_delete","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_delete/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_delete.md","title":"nodus delete","description":"Delete resources","stage":"GA","headings":[{"depth":2,"slug":"nodus-delete","text":"nodus delete"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus delete\n\n Delete resources\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus delete/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus delete\n\nDelete resources\n\nnodus delete KIND/NAME... -f FILE [flags]\n\nExamples\n\n nodus delete job/finetune --wait\n nodus delete -f job.yaml\n\nOptions\n\n --cascade string background, foreground or orphan (default \"background\")\n -f, --filename strings Manifests naming the resources to delete\n -h, --help help for delete\n --ignore-not-found Succeed when a resource is already gone\n -l, --selector string Label selector\n --timeout duration How long --wait waits (default 10m0s)\n --wait Wait until the resources are gone (finalizers cleared)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus delete job/train\nnodus delete job/train --ignore-not-found\nnodus delete job/train --wait\nnodus delete sshkey laptop"},{"id":"docs/reference/cli/nodus_deploy","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_deploy/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_deploy.md","title":"nodus deploy","description":"Deploy the nodus.App in APP.py as a persistent App","stage":"GA","headings":[{"depth":2,"slug":"nodus-deploy","text":"nodus deploy"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus deploy\n\n Deploy the nodus.App in APP.py as a persistent App\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus deploy/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus deploy\n\nDeploy the nodus.App in APP.py as a persistent App\n\nnodus deploy APP.py [flags]\n\nExamples\n\n nodus deploy app.py\n\nOptions\n\n -h, --help help for deploy\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus deploy app.py"},{"id":"docs/reference/cli/nodus_describe","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_describe/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_describe.md","title":"nodus describe","description":"Show everything about one resource: attempts, placement, checkpoint, cost, conditions and events","stage":"GA","headings":[{"depth":2,"slug":"nodus-describe","text":"nodus describe"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus describe\n\n Show everything about one resource: attempts, placement, checkpoint, cost, conditions and events\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus describe/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus describe\n\nShow everything about one resource: attempts, placement, checkpoint, cost, conditions and events\n\nnodus describe KIND/NAME [flags]\n\nOptions\n\n -h, --help help for describe\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus describe job/train\nnodus describe job/late"},{"id":"docs/reference/cli/nodus_diff","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_diff/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_diff.md","title":"nodus diff","description":"Show what apply would change, from a server dry-run (exit 1 when there are differences)","stage":"GA","headings":[{"depth":2,"slug":"nodus-diff","text":"nodus diff"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus diff\n\n Show what apply would change, from a server dry-run (exit 1 when there are differences)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus diff/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus diff\n\nShow what apply would change, from a server dry-run (exit 1 when there are differences)\n\nnodus diff -f FILE [flags]\n\nOptions\n\n -f, --filename strings Manifest file, directory or - for stdin\n -h, --help help for diff\n -R, --recursive Read directories recursively\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus diff -f job.yaml\nnodus diff -f job-v2.yaml"},{"id":"docs/reference/cli/nodus_edit","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_edit/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_edit.md","title":"nodus edit","description":"Edit a resource in $EDITOR and save it as a merge patch","stage":"GA","headings":[{"depth":2,"slug":"nodus-edit","text":"nodus edit"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus edit\n\n Edit a resource in $EDITOR and save it as a merge patch\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus edit/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus edit\n\nEdit a resource in $EDITOR and save it as a merge patch\n\nnodus edit KIND/NAME [flags]\n\nExamples\n\n nodus edit job/finetune\n EDITOR=nano nodus edit budget/research-monthly\n\nOptions\n\n -h, --help help for edit\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus edit job/train"},{"id":"docs/reference/cli/nodus_events","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_events/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_events.md","title":"nodus events","description":"List events in the project, or for one object, and watch for new ones","stage":"GA","headings":[{"depth":2,"slug":"nodus-events","text":"nodus events"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus events\n\n List events in the project, or for one object, and watch for new ones\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus events/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus events\n\nList events in the project, or for one object, and watch for new ones\n\nnodus events [--for KIND/NAME] [-w] [flags]\n\nExamples\n\n nodus events\n nodus events --for job/finetune -w\n nodus events -w -o json\n\nOptions\n\n --for string Only events about KIND/NAME\n -h, --help help for events\n -o, --output string Output format: json, yaml, name, jsonpath=..., go-template=... or custom-columns=...\n -w, --watch Print the events, then each new one as it is recorded\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus events --for job/train"},{"id":"docs/reference/cli/nodus_exec","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_exec/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_exec.md","title":"nodus exec","description":"Run a command in a running Job, Sandbox or Workspace (recorded as a Process)","stage":"GA","headings":[{"depth":2,"slug":"nodus-exec","text":"nodus exec"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus exec\n\n Run a command in a running Job, Sandbox or Workspace (recorded as a Process)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus exec/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus exec\n\nRun a command in a running Job, Sandbox or Workspace (recorded as a Process)\n\nnodus exec [-it] KIND/NAME -- COMMAND [ARGS...] [flags]\n\nExamples\n\n nodus exec -it job/finetune -- bash\n nodus exec sb/dev -- nvidia-smi\n\nOptions\n\n -h, --help help for exec\n --index int Job index (Indexed Jobs) (default -1)\n --rank string Gang member rank (default 0)\n -i, --stdin Pass stdin to the command\n -t, --tty Allocate a terminal\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus exec sb/dev -- python -c 'raise SystemExit(3)'"},{"id":"docs/reference/cli/nodus_explain","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_explain/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_explain.md","title":"nodus explain","description":"Show the documentation of a kind or one of its fields","stage":"GA","headings":[{"depth":2,"slug":"nodus-explain","text":"nodus explain"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus explain\n\n Show the documentation of a kind or one of its fields\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus explain/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus explain\n\nShow the documentation of a kind or one of its fields\n\nnodus explain KIND[.FIELD...] [flags]\n\nExamples\n\n nodus explain jobs\n nodus explain job.spec.resources.gpu\n nodus explain sandbox.spec --recursive\n\nOptions\n\n -h, --help help for explain\n -o, --output string Output format: plaintext or plaintext-openapiv2 (default \"plaintext\")\n --recursive Print every nested field\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus explain jobs\nnodus explain job.spec.maxCostUSD"},{"id":"docs/reference/cli/nodus_get","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_get/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_get.md","title":"nodus get","description":"List or show resources","stage":"GA","headings":[{"depth":2,"slug":"nodus-get","text":"nodus get"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus get\n\n List or show resources\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus get/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus get\n\nList or show resources\n\nnodus get KIND[/NAME]... [flags]\n\nExamples\n\n nodus get jobs\n nodus get job/finetune -o yaml\n nodus get jobs -l team=nlp -w\n nodus get gpus --gpu H100 --count 8 --region us\n\nOptions\n\n -A, --all-projects List across every project\n --count int get gpus: accelerators per offering\n --field-selector string Field selector\n -f, --filename strings Files naming the resources to get\n --gpu string get gpus: accelerator family, such as H100\n -h, --help help for get\n --interruptible get gpus: interruptible capacity only\n --mine Only resources I created\n --no-headers Omit table headers\n -o, --output string Output format: wide, json, yaml, name, jsonpath=..., go-template=..., custom-columns=HEADER:.path,...\n --region string get gpus: region class, such as us or eu\n -l, --selector string Label selector\n --show-labels Show labels as the last column\n -w, --watch Watch for changes after listing\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus get transactions - List postings on the org’s wallet with the balance after each\n nodus get usage - Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus get job/keep\nnodus get job/old -o name\nnodus get job/old\nnodus get jobs -o name\nnodus get jobs\nnodus get jobs -o wide\nnodus get job/train -o jsonpath={.metadata.labels.team}\nnodus get job train -o yaml\nnodus get jobs -l team=nlp\nnodus get job/train -o jsonpath={.spec.maxCostUSD}\nnodus get jobs -w -o custom-columns=NAME:.metadata.name,PHASE:.status.phase\nnodus get job/train"},{"id":"docs/reference/cli/nodus_get_transactions","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_get_transactions/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_get_transactions.md","title":"nodus get transactions","description":"List postings on the org's wallet with the balance after each","stage":"GA","headings":[{"depth":2,"slug":"nodus-get-transactions","text":"nodus get transactions"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus get transactions\n\n List postings on the org's wallet with the balance after each\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus get transactions/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus get transactions\n\nList postings on the org’s wallet with the balance after each\n\nnodus get transactions [NAME] [--since 7d] [--type TYPE] [-o csv json yaml wide] [flags]\n\nExamples\n\n nodus get transactions --since 7d\n nodus get transactions --type Capture -o csv\n\nOptions\n\n --field-selector string Field selector, such as meter=ComputeSeconds\n -h, --help help for transactions\n --no-headers Omit table headers\n -o, --output string Output format: wide, csv, json or yaml\n --since string Start of the range: a duration back from now (30d, 12h) or a time\n --type string Only this type, such as TopUp, Capture, Grant or Refund\n --until string End of the range (default now)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus get - List or show resources"},{"id":"docs/reference/cli/nodus_get_usage","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_get_usage/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_get_usage.md","title":"nodus get usage","description":"Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank","stage":"GA","headings":[{"depth":2,"slug":"nodus-get-usage","text":"nodus get usage"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus get usage\n\n Show metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus get usage/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus get usage\n\nShow metered usage, itemized or summed by project, kind, member, label, meter, day, segment, offering or rank\n\nSynopsis\n\nusage lists UsageRecords, or with –group-by sums them into a UsageSummary. Group keys are project, kind, meter, day, segment (Boot, Running, Restore, Teardown), offering (the capacity it ran on), rank (a gang member), member (the person who started the work, with any of their keys) and label:. Every billed interval of your allocation is on a line, from its billing start to confirmed deletion.\n\nnodus get usage [--group-by KEY,...] [--since 30d] [--until TIME] [-o csv json yaml wide] [flags]\n\nExamples\n\n nodus get usage --group-by member --since 30d\n nodus get usage --group-by project,label:team --since 30d\n nodus get usage --object job/ddp-2node --group-by rank,segment\n nodus get usage --group-by day,meter --since 7d -o csv usage.csv\n\nOptions\n\n --field-selector string Field selector, such as meter=ComputeSeconds\n --group-by strings Sum by these keys: project, kind, meter, day, segment, offering, rank, member, label: \n -h, --help help for usage\n --no-headers Omit table headers\n --object string Only the usage of one object, as KIND/NAME\n -o, --output string Output format: wide, csv, json or yaml\n --since string Start of the range: a duration back from now (30d, 12h) or a time\n --until string End of the range (default now)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus get - List or show resources"},{"id":"docs/reference/cli/nodus_inference","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_inference/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_inference.md","title":"nodus inference","description":"Call hosted models: list them, send a chat completion and read a request's receipt","stage":"GA","headings":[{"depth":2,"slug":"nodus-inference","text":"nodus inference"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus inference\n\n Call hosted models: list them, send a chat completion and read a request's receipt\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus inference/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus inference\n\nCall hosted models: list them, send a chat completion and read a request’s receipt\n\nSynopsis\n\ninference talks to the OpenAI-compatible data plane of the current context with its credential: the inference host of the API (inference. for api., else the API host itself), or –base-url.\n\nOptions\n\n --base-url string Inference origin (default: derived from the API URL)\n -h, --help help for inference\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus inference chat - Send one chat completion and print the answer; the request id and charge go to stderr\n nodus inference models - List the models you can call now\n nodus inference receipt - Show what one inference request was charged (receipts are kept 30 days)"},{"id":"docs/reference/cli/nodus_inference_chat","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_inference_chat/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_inference_chat.md","title":"nodus inference chat","description":"Send one chat completion and print the answer; the request id and charge go to stderr","stage":"GA","headings":[{"depth":2,"slug":"nodus-inference-chat","text":"nodus inference chat"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus inference chat\n\n Send one chat completion and print the answer; the request id and charge go to stderr\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus inference chat/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus inference chat\n\nSend one chat completion and print the answer; the request id and charge go to stderr\n\nSynopsis\n\nchat sends PROMPT (or standard input, with - or no argument) as one user message and prints the answer. The request is held at its maximum cost (the input plus –max-tokens of output at the model’s rates) and charged for the tokens it used; the request id, the model that answered, the tokens and the charge are printed to stderr.\n\nnodus inference chat [PROMPT -] [--model MODEL --endpoint NAME] [--max-tokens N] [flags]\n\nExamples\n\n nodus inference chat \"What is the capital of France?\"\n nodus inference chat --model nodus/gpt-oss-20b --max-tokens 64 \"Say hello\"\n nodus inference chat --endpoint support-bot < question.txt\n nodus inference chat -o json \"Hi\" jq .usage\n\nOptions\n\n --endpoint string Send through this InferenceEndpoint of the project instead of --model\n -h, --help help for chat\n --idempotency-key string Retry safely: a repeat with the same key and prompt is answered once and charged once\n --max-tokens int Most output tokens; the request's hold is sized from it (default 512)\n --model string Catalog model, as 'nodus inference models' lists them; nodus/indra (Indra) routes each request (default \"nodus/indra\")\n -o, --output string Output format: json or yaml (the whole completion)\n --system string System message\n\nOptions inherited from parent commands\n\n --base-url string Inference origin (default: derived from the API URL)\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus inference - Call hosted models: list them, send a chat completion and read a request’s receipt"},{"id":"docs/reference/cli/nodus_inference_models","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_inference_models/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_inference_models.md","title":"nodus inference models","description":"List the models you can call now","stage":"GA","headings":[{"depth":2,"slug":"nodus-inference-models","text":"nodus inference models"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus inference models\n\n List the models you can call now\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus inference models/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus inference models\n\nList the models you can call now\n\nnodus inference models [-o json] [flags]\n\nExamples\n\n nodus inference models\n\nOptions\n\n -h, --help help for models\n -o, --output string Output format: json or yaml\n\nOptions inherited from parent commands\n\n --base-url string Inference origin (default: derived from the API URL)\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus inference - Call hosted models: list them, send a chat completion and read a request’s receipt"},{"id":"docs/reference/cli/nodus_inference_receipt","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_inference_receipt/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_inference_receipt.md","title":"nodus inference receipt","description":"Show what one inference request was charged (receipts are kept 30 days)","stage":"GA","headings":[{"depth":2,"slug":"nodus-inference-receipt","text":"nodus inference receipt"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus inference receipt\n\n Show what one inference request was charged (receipts are kept 30 days)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus inference receipt/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus inference receipt\n\nShow what one inference request was charged (receipts are kept 30 days)\n\nSynopsis\n\nreceipt shows a request’s state (Running, Settled, Released or Unknown), its usage and its charge. Every response of the data plane names its request in the Nodus-Request-Id header. A Released or Unknown request is never charged.\n\nnodus inference receipt REQUEST ID [-o json yaml] [flags]\n\nExamples\n\n nodus inference receipt ireq 01k6d2x7q9fvjt3y8m0c4r5n2e\n\nOptions\n\n -h, --help help for receipt\n -o, --output string Output format: json or yaml\n\nOptions inherited from parent commands\n\n --base-url string Inference origin (default: derived from the API URL)\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus inference - Call hosted models: list them, send a chat completion and read a request’s receipt"},{"id":"docs/reference/cli/nodus_init","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_init/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_init.md","title":"nodus init","description":"Scaffold a manifest or app.py in the current directory","stage":"GA","headings":[{"depth":2,"slug":"nodus-init","text":"nodus init"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus init\n\n Scaffold a manifest or app.py in the current directory\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus init/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus init\n\nScaffold a manifest or app.py in the current directory\n\nnodus init [job app agent sandbox] [flags]\n\nExamples\n\n nodus init # job.yaml\n nodus init app # app.py for nodus run app.py\n\nOptions\n\n --force Replace an existing file\n -h, --help help for init\n --name string Resource or App name (default: the directory name)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus init\nnodus init app --name demo"},{"id":"docs/reference/cli/nodus_label","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_label/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_label.md","title":"nodus label","description":"Set or remove labels","stage":"GA","headings":[{"depth":2,"slug":"nodus-label","text":"nodus label"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus label\n\n Set or remove labels\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus label/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus label\n\nSet or remove labels\n\nnodus label KIND/NAME KEY=VALUE... [KEY-] [flags]\n\nOptions\n\n -h, --help help for label\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus label job/train tier=gold"},{"id":"docs/reference/cli/nodus_login","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_login/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_login.md","title":"nodus login","description":"Sign in through the console and store one API key per chosen org","stage":"GA","headings":[{"depth":2,"slug":"nodus-login","text":"nodus login"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus login\n\n Sign in through the console and store one API key per chosen org\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus login/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus login\n\nSign in through the console and store one API key per chosen org\n\nSynopsis\n\nOpens the console, where you sign in (or sign up) and choose orgs; nodus stores one 90-day API key per org in the OS keychain and writes one context per org. Interactive sign-in also creates or reuses your local SSH key and configures access to all current and future Workspaces in those orgs. –device prints a code to enter on another device; –with-token reads an API key from stdin (CI).\n\nnodus login [--device --with-token] [flags]\n\nExamples\n\n nodus login\n nodus login --device\n echo \"$NODUS API KEY\" nodus login --with-token\n\nOptions\n\n --api-url string API server (default https://api.nodus-compute.ai)\n --console-url string Console origin (default: what the API server names)\n --device Sign in from another device with a one-time code (RFC 8628)\n -h, --help help for login\n --with-token Read an API key from stdin\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus login --with-token\nnodus login --device\nnodus login"},{"id":"docs/reference/cli/nodus_logout","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_logout/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_logout.md","title":"nodus logout","description":"Revoke the context's CLI key and remove the context","stage":"GA","headings":[{"depth":2,"slug":"nodus-logout","text":"nodus logout"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus logout\n\n Revoke the context's CLI key and remove the context\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus logout/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus logout\n\nRevoke the context’s CLI key and remove the context\n\nnodus logout [flags]\n\nOptions\n\n -h, --help help for logout\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus logout"},{"id":"docs/reference/cli/nodus_logs","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_logs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_logs.md","title":"nodus logs","description":"Print or follow a resource's output","stage":"GA","headings":[{"depth":2,"slug":"nodus-logs","text":"nodus logs"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus logs\n\n Print or follow a resource's output\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus logs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus logs\n\nPrint or follow a resource’s output\n\nnodus logs KIND/NAME [-f] [flags]\n\nExamples\n\n nodus logs -f job/finetune\n nodus logs job/ddp --all-ranks\n nodus logs sb/dev --process 12\n\nOptions\n\n --all-indexes Every index of an Indexed Job\n --all-ranks Every gang member, each line prefixed with its rank\n --attempt int Attempt number (default: the current one)\n -f, --follow Stream new output\n -h, --help help for logs\n --index int Job index (Indexed Jobs) (default -1)\n --process string A Process of a Sandbox or Workspace, by sequence number\n --rank string Gang member rank (default 0)\n --since duration Only output newer than this, such as 10m\n --tail int Lines of backlog to show (default -1)\n --timestamps Prefix each line with its time\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus logs job/train\nnodus logs job/train --rank 1\nnodus logs job/train --all-ranks\nnodus logs job/shards --all-indexes"},{"id":"docs/reference/cli/nodus_mcp","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_mcp/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_mcp.md","title":"nodus mcp","description":"Serve Nodus tools to a local MCP client over stdio","stage":"GA","headings":[{"depth":2,"slug":"nodus-mcp","text":"nodus mcp"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus mcp\n\n Serve Nodus tools to a local MCP client over stdio\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus mcp/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus mcp\n\nServe Nodus tools to a local MCP client over stdio\n\nSynopsis\n\nRuns the Nodus MCP server on standard input and output, for clients that launch it as a subprocess. It acts with this machine’s login (nodus login or NODUS API KEY), so it can do what that credential can, and its writes return a dry-run that the client’s user confirms before anything is created. On top of the hosted tools it can upload the current directory as source, download outputs and copy files to and from a Sandbox.\n\nTo connect a client, run nodus mcp install CLIENT, or add the hosted server from .\n\nnodus mcp [flags]\n\nOptions\n\n -h, --help help for mcp\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus mcp install - Add the Nodus MCP server to Claude Code, Cursor or Codex"},{"id":"docs/reference/cli/nodus_mcp_install","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_mcp_install/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_mcp_install.md","title":"nodus mcp install","description":"Add the Nodus MCP server to Claude Code, Cursor or Codex","stage":"GA","headings":[{"depth":2,"slug":"nodus-mcp-install","text":"nodus mcp install"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus mcp install\n\n Add the Nodus MCP server to Claude Code, Cursor or Codex\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus mcp install/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus mcp install\n\nAdd the Nodus MCP server to Claude Code, Cursor or Codex\n\nSynopsis\n\nWrites the nodus server into the client’s user-level configuration file and leaves everything else in it as it was. CLIENT is claude, cursor or codex. By default the client launches nodus mcp (local, acts with this machine’s login); with –hosted it connects to the hosted server and signs in with your browser instead. The original file is kept beside it as FILE.nodus-backup. A different existing nodus entry is not replaced unless you pass –force.\n\nnodus mcp install CLIENT [flags]\n\nExamples\n\n nodus mcp install claude\n nodus mcp install cursor --hosted\n nodus mcp install codex --dry-run\n\nOptions\n\n --dry-run Show what would change without writing\n --force Replace a different existing nodus entry\n -h, --help help for install\n --hosted Connect to the hosted server instead of launching nodus mcp\n --url string Hosted server URL (default: this context's server plus /mcp)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus mcp - Serve Nodus tools to a local MCP client over stdio"},{"id":"docs/reference/cli/nodus_open","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_open/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_open.md","title":"nodus open","description":"Open a resource's console page, or a port of its container, in the browser","stage":"GA","headings":[{"depth":2,"slug":"nodus-open","text":"nodus open"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus open\n\n Open a resource's console page, or a port of its container, in the browser\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus open/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus open\n\nOpen a resource’s console page, or a port of its container, in the browser\n\nSynopsis\n\nWithout a port flag, open shows the resource’s console page. With –vscode, –jupyter or –port, it mints a preview of that container port (a single-use link valid for 60 seconds that signs the browser in for an hour) and opens it.\n\nnodus open KIND/NAME [--vscode --jupyter --port N] [flags]\n\nExamples\n\n nodus open job/train\n nodus open ws/lab --vscode\n nodus open sb/dev --port 3000 --print\n\nOptions\n\n -h, --help help for open\n --jupyter Open Jupyter (port 8888)\n --port int Open this container port\n --print Print the link instead of opening a browser\n --vscode Open VS Code in the browser (port 8080)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus open job/train --print\nnodus open job/train\nnodus open sb/dev --vscode --print\nnodus open sb/dev --jupyter --print\nnodus open sb/dev --port 3000\nnodus open job/train --port 6006"},{"id":"docs/reference/cli/nodus_patch","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_patch/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_patch.md","title":"nodus patch","description":"Update fields of a resource with a merge or JSON patch","stage":"GA","headings":[{"depth":2,"slug":"nodus-patch","text":"nodus patch"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus patch\n\n Update fields of a resource with a merge or JSON patch\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus patch/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus patch\n\nUpdate fields of a resource with a merge or JSON patch\n\nnodus patch KIND/NAME --patch PATCH [--type merge json] [flags]\n\nExamples\n\n nodus patch job/x --patch '{\"spec\":{\"maxCostUSD\":\"50.00\"}}'\n\nOptions\n\n -h, --help help for patch\n --patch string The patch document (JSON)\n --type string merge or json (default \"merge\")\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus patch job/train --patch '{\"spec\":{\"maxCostUSD\":\"50.00\"}}'"},{"id":"docs/reference/cli/nodus_pool","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_pool/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_pool.md","title":"nodus pool","description":"Pool utilization, forecasts, recommendations and actions","stage":"GA","headings":[{"depth":2,"slug":"nodus-pool","text":"nodus pool"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus pool\n\n Pool utilization, forecasts, recommendations and actions\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus pool/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus pool\n\nPool utilization, forecasts, recommendations and actions\n\nOptions\n\n -h, --help help for pool\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus pool approve - Approve a proposed PoolAction\n nodus pool forecast - Show a pool’s demand forecast with p50 and p90 bands\n nodus pool pause - Pause every automated action in a pool\n nodus pool recommendations - List a pool’s recommendations\n nodus pool reject - Reject a proposed PoolAction\n nodus pool resume - Resume every automated action in a pool\n nodus pool revert - Undo an executed PoolAction where the action allows it\n nodus pool utilization - Show a pool’s GPU utilization and device-hour ledger"},{"id":"docs/reference/cli/nodus_pool_approve","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_pool_approve/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_pool_approve.md","title":"nodus pool approve","description":"Approve a proposed PoolAction","stage":"GA","headings":[{"depth":2,"slug":"nodus-pool-approve","text":"nodus pool approve"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus pool approve\n\n Approve a proposed PoolAction\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus pool approve/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus pool approve\n\nApprove a proposed PoolAction\n\nnodus pool approve ACTION [flags]\n\nOptions\n\n -h, --help help for approve\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus pool - Pool utilization, forecasts, recommendations and actions\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus pool approve lab-idle-1"},{"id":"docs/reference/cli/nodus_pool_forecast","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_pool_forecast/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_pool_forecast.md","title":"nodus pool forecast","description":"Show a pool's demand forecast with p50 and p90 bands","stage":"GA","headings":[{"depth":2,"slug":"nodus-pool-forecast","text":"nodus pool forecast"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus pool forecast\n\n Show a pool's demand forecast with p50 and p90 bands\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus pool forecast/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus pool forecast\n\nShow a pool’s demand forecast with p50 and p90 bands\n\nnodus pool forecast POOL [flags]\n\nOptions\n\n -h, --help help for forecast\n --horizon string Forecast horizon: 24h or 7d (default 7d)\n --json Print the response as JSON\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus pool - Pool utilization, forecasts, recommendations and actions\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus pool forecast lab --horizon 7d\nnodus pool forecast missing"},{"id":"docs/reference/cli/nodus_pool_pause","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_pool_pause/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_pool_pause.md","title":"nodus pool pause","description":"Pause every automated action in a pool","stage":"GA","headings":[{"depth":2,"slug":"nodus-pool-pause","text":"nodus pool pause"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus pool pause\n\n Pause every automated action in a pool\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus pool pause/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus pool pause\n\nPause every automated action in a pool\n\nnodus pool pause POOL [flags]\n\nOptions\n\n -h, --help help for pause\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus pool - Pool utilization, forecasts, recommendations and actions\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus pool pause lab"},{"id":"docs/reference/cli/nodus_pool_recommendations","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_pool_recommendations/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_pool_recommendations.md","title":"nodus pool recommendations","description":"List a pool's recommendations","stage":"GA","headings":[{"depth":2,"slug":"nodus-pool-recommendations","text":"nodus pool recommendations"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus pool recommendations\n\n List a pool's recommendations\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus pool recommendations/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus pool recommendations\n\nList a pool’s recommendations\n\nnodus pool recommendations POOL [flags]\n\nOptions\n\n -h, --help help for recommendations\n --json Print the response as JSON\n --state string Only Open, Dismissed, Actioned or Expired recommendations\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus pool - Pool utilization, forecasts, recommendations and actions\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus pool recommendations lab\nnodus pool recommendations lab --json"},{"id":"docs/reference/cli/nodus_pool_reject","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_pool_reject/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_pool_reject.md","title":"nodus pool reject","description":"Reject a proposed PoolAction","stage":"GA","headings":[{"depth":2,"slug":"nodus-pool-reject","text":"nodus pool reject"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus pool reject\n\n Reject a proposed PoolAction\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus pool reject/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus pool reject\n\nReject a proposed PoolAction\n\nnodus pool reject ACTION [flags]\n\nOptions\n\n -h, --help help for reject\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus pool - Pool utilization, forecasts, recommendations and actions"},{"id":"docs/reference/cli/nodus_pool_resume","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_pool_resume/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_pool_resume.md","title":"nodus pool resume","description":"Resume every automated action in a pool","stage":"GA","headings":[{"depth":2,"slug":"nodus-pool-resume","text":"nodus pool resume"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus pool resume\n\n Resume every automated action in a pool\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus pool resume/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus pool resume\n\nResume every automated action in a pool\n\nnodus pool resume POOL [flags]\n\nOptions\n\n -h, --help help for resume\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus pool - Pool utilization, forecasts, recommendations and actions\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus pool resume lab"},{"id":"docs/reference/cli/nodus_pool_revert","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_pool_revert/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_pool_revert.md","title":"nodus pool revert","description":"Undo an executed PoolAction where the action allows it","stage":"GA","headings":[{"depth":2,"slug":"nodus-pool-revert","text":"nodus pool revert"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus pool revert\n\n Undo an executed PoolAction where the action allows it\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus pool revert/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus pool revert\n\nUndo an executed PoolAction where the action allows it\n\nnodus pool revert ACTION [flags]\n\nOptions\n\n -h, --help help for revert\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus pool - Pool utilization, forecasts, recommendations and actions\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus pool revert pact/lab-idle-1"},{"id":"docs/reference/cli/nodus_pool_utilization","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_pool_utilization/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_pool_utilization.md","title":"nodus pool utilization","description":"Show a pool's GPU utilization and device-hour ledger","stage":"GA","headings":[{"depth":2,"slug":"nodus-pool-utilization","text":"nodus pool utilization"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus pool utilization\n\n Show a pool's GPU utilization and device-hour ledger\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus pool utilization/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus pool utilization\n\nShow a pool’s GPU utilization and device-hour ledger\n\nnodus pool utilization POOL [flags]\n\nOptions\n\n -h, --help help for utilization\n --json Print the response as JSON\n --window string Window to summarize, up to 30d (default 24h)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus pool - Pool utilization, forecasts, recommendations and actions\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus pool utilization lab"},{"id":"docs/reference/cli/nodus_port-forward","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_port-forward/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_port-forward.md","title":"nodus port-forward","description":"Forward local ports to a Job, Sandbox or Workspace","stage":"GA","headings":[{"depth":2,"slug":"nodus-port-forward","text":"nodus port-forward"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus port-forward\n\n Forward local ports to a Job, Sandbox or Workspace\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus port-forward/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus port-forward\n\nForward local ports to a Job, Sandbox or Workspace\n\nnodus port-forward KIND/NAME LOCAL[:REMOTE]... [flags]\n\nExamples\n\n nodus port-forward job/train 6006\n nodus port-forward sb/dev 8080:80\n\nOptions\n\n --address string Local address to listen on (default \"127.0.0.1\")\n -h, --help help for port-forward\n --index int Job index (Indexed Jobs) (default -1)\n --rank string Gang member rank (default 0)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus"},{"id":"docs/reference/cli/nodus_request","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_request/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_request.md","title":"nodus request","description":"Request an action: restart, refresh-secrets, reload-sinks, verify, resend, start, reimport, approve-burst, drain","stage":"GA","headings":[{"depth":2,"slug":"nodus-request","text":"nodus request"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus request\n\n Request an action: restart, refresh-secrets, reload-sinks, verify, resend, start, reimport, approve-burst, drain\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus request/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus request\n\nRequest an action: restart, refresh-secrets, reload-sinks, verify, resend, start, reimport, approve-burst, drain\n\nnodus request ACTION KIND/NAME [flags]\n\nOptions\n\n -h, --help help for request\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus request restart job/train"},{"id":"docs/reference/cli/nodus_resume","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_resume/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_resume.md","title":"nodus resume","description":"Resume a suspended run","stage":"GA","headings":[{"depth":2,"slug":"nodus-resume","text":"nodus resume"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus resume\n\n Resume a suspended run\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus resume/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus resume\n\nResume a suspended run\n\nnodus resume KIND/NAME [flags]\n\nOptions\n\n -h, --help help for resume\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus resume job/train"},{"id":"docs/reference/cli/nodus_rollout","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_rollout/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_rollout.md","title":"nodus rollout","description":"Roll out changes","stage":"GA","headings":[{"depth":2,"slug":"nodus-rollout","text":"nodus rollout"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus rollout\n\n Roll out changes\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus rollout/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus rollout\n\nRoll out changes\n\nOptions\n\n -h, --help help for rollout\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus rollout restart - Restart without a spec change (alias of request restart )"},{"id":"docs/reference/cli/nodus_rollout_restart","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_rollout_restart/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_rollout_restart.md","title":"nodus rollout restart","description":"Restart without a spec change (alias of `request restart`)","stage":"GA","headings":[{"depth":2,"slug":"nodus-rollout-restart","text":"nodus rollout restart"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus rollout restart\n\n Restart without a spec change (alias of request restart )\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus rollout restart/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus rollout restart\n\nRestart without a spec change (alias of request restart )\n\nnodus rollout restart KIND/NAME [flags]\n\nOptions\n\n -h, --help help for restart\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus rollout - Roll out changes\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus rollout restart job/train"},{"id":"docs/reference/cli/nodus_run","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_run/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_run.md","title":"nodus run","description":"Run a command as a Job: estimate, phases, logs and the final cost","stage":"GA","headings":[{"depth":2,"slug":"nodus-run","text":"nodus run"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus run\n\n Run a command as a Job: estimate, phases, logs and the final cost\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus run/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus run\n\nRun a command as a Job: estimate, phases, logs and the final cost\n\nSynopsis\n\nUploads the current directory (unless –no-source), creates a Job, prints its estimate and phases, streams its logs and ends with the cost line. The command’s exit code becomes nodus’s; 125 means a Nodus or API error, 124 a timeout and 130 an interrupt. On a terminal Ctrl+C asks before canceling; -d detaches.\n\nnodus run [flags] -- COMMAND [ARGS...] run APP.py[::FUNCTION] [ARGS...]\n\nExamples\n\n nodus run --gpu L4 --image nodus/pytorch -- python train.py\n nodus run --gpu H100:8 --max-cost 40 -d -- torchrun train.py\n nodus run app.py::main\n\nOptions\n\n --allow-large-source Upload source above 500 MiB compressed\n --checkpoint string Checkpoint path (default /nodus/state)\n --complete-by string Finish-by time (RFC 3339) for placement\n --completions int Indexed completions\n --continuity string Recovery: Checkpointed, Restartable or Ephemeral\n --cpu string vCPUs\n -d, --detach Return once the Job is created\n --disk string Disk, such as 200Gi\n --dry-run Print the estimate only\n --env stringArray Environment variable K=V\n --expected-duration duration Expected runtime, for the estimate\n --gpu string Accelerator TYPE[:COUNT], such as L4 or H100:8\n --gpus-per-node int Beta: GPUs per gang node\n -h, --help help for run\n --idempotency-key string Override the generated idempotency key\n --image string Image, such as nodus/pytorch\n --interruptible string[=\"prefer\"] Interruptible capacity: allow, prefer or never\n --keep Keep the Job after it finishes (no 30-day TTL)\n -l, --label strings Labels k=v\n --launcher string Beta: plain, torchrun, ray or verl\n --max-cost string Spend cap in USD\n --memory string Memory, such as 64Gi\n --name string Job name (default: generated run-xxxxx)\n --network string Beta: colocated, regional or global\n --no-source Do not upload the current directory\n --no-warm Release capacity at once instead of keeping it warm for 60 s\n --nodes int Beta: gang nodes\n --output stringArray Output NAME=/path\n --parallelism int Indexes running at once\n --profile string Placement profile: Balanced, Cost or Speed\n --region string Region class, such as us or eu\n --secret stringArray Secret to expose as environment variables\n --startup-timeout duration Beta: gang assembly budget\n --timeout duration Wall-clock limit, such as 2h\n --total-gpus int Beta: total GPUs across the gang\n --transport string Beta: direct or auto\n --volume stringArray Volume NAME:/path[:ro]\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus run app.py::main --epochs 3\nnodus run fail.py\nnodus run --gpu L4 --image nodus/pytorch -- python train.py\nnodus run --gpu L4 --no-source -- sh -c 'exit 3'\nnodus run --gpu L4 --no-source --name late -- python deadline.py\nnodus run -d --no-source --gpu L4 --keep --name detached -- python train.py\nnodus run --dry-run --no-source --gpu L4 -- python train.py\nnodus run -d --no-source --name gang --gpu H100 --gpus-per-node 8 --nodes 2 --launcher torchrun --network regional --max-cost 40 --env A=1 --interruptible -- torchrun train.py\nnodus run --dry-run --no-source --gpu H100 --gpus-per-node 8 --nodes 2 -- torchrun train.py\nnodus run --no-source --gpu none -- python train.py"},{"id":"docs/reference/cli/nodus_secret","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_secret/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_secret.md","title":"nodus secret","description":"Create Secrets from literals, env files or registry logins","stage":"GA","headings":[{"depth":2,"slug":"nodus-secret","text":"nodus secret"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus secret\n\n Create Secrets from literals, env files or registry logins\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus secret/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus secret\n\nCreate Secrets from literals, env files or registry logins\n\nOptions\n\n -h, --help help for secret\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus secret create - Create a Secret"},{"id":"docs/reference/cli/nodus_secret_create","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_secret_create/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_secret_create.md","title":"nodus secret create","description":"Create a Secret","stage":"GA","headings":[{"depth":2,"slug":"nodus-secret-create","text":"nodus secret create"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus secret create\n\n Create a Secret\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus secret create/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus secret create\n\nCreate a Secret\n\nnodus secret create NAME [flags]\n\nExamples\n\n nodus secret create hf-token --from-literal HF TOKEN=hf xxx\n nodus secret create app-env --from-env-file .env\n nodus secret create registry --type Registry --from-literal server=ghcr.io --from-literal username=ada --from-literal password=…\n\nOptions\n\n --dry-run string[=\"server\"] server returns the estimate without creating; client prints the manifest (default \"none\")\n --from-env-file stringArray File of KEY=VALUE lines\n --from-literal stringArray KEY=VALUE (repeatable)\n -h, --help help for create\n -o, --output string Output format: name, json, yaml, or estimate (with --dry-run=server)\n --type string Opaque or Registry\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus secret - Create Secrets from literals, env files or registry logins"},{"id":"docs/reference/cli/nodus_serve","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_serve/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_serve.md","title":"nodus serve","description":"Run the App in APP.py and redeploy it whenever a file changes","stage":"GA","headings":[{"depth":2,"slug":"nodus-serve","text":"nodus serve"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus serve\n\n Run the App in APP.py and redeploy it whenever a file changes\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus serve/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus serve\n\nRun the App in APP.py and redeploy it whenever a file changes\n\nnodus serve APP.py [flags]\n\nExamples\n\n nodus serve app.py\n\nOptions\n\n -h, --help help for serve\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus serve app.py"},{"id":"docs/reference/cli/nodus_shell","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_shell/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_shell.md","title":"nodus shell","description":"Open an interactive shell in a Sandbox, creating it if it is not there","stage":"GA","headings":[{"depth":2,"slug":"nodus-shell","text":"nodus shell"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus shell\n\n Open an interactive shell in a Sandbox, creating it if it is not there\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus shell/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus shell\n\nOpen an interactive shell in a Sandbox, creating it if it is not there\n\nSynopsis\n\nshell opens a terminal in a Sandbox. A name that does not exist yet is created from the same flags as nodus create sandbox (image nodus/agent-tools unless –image says otherwise), and the Sandbox stays after the shell exits, so the next nodus shell finds /workspace as you left it; –rm deletes a Sandbox this command created when the shell exits. Without a name the shell gets a Sandbox of its own, which is deleted when it exits unless –keep is set. The command after – replaces bash.\n\nnodus shell [sandbox/NAME] [flags] [-- COMMAND...]\n\nExamples\n\n nodus shell sandbox/dev\n nodus shell sandbox/dev --cpu 4 --memory 8Gi\n nodus shell sandbox/dev --rm\n nodus shell\n nodus shell sandbox/dev -- zsh\n\nOptions\n\n --cpu string vCPU floor\n --disk string Disk floor, such as 50Gi\n --gpu string Accelerator TYPE[:COUNT], such as L4 or H100:8\n -h, --help help for shell\n --idle-timeout string Stop after this idle time\n --image string Container image (default \"nodus/agent-tools\")\n --keep Without a name: keep the Sandbox after the shell exits\n --max-cost string Spend limit in USD\n --memory string Memory floor, such as 16Gi\n --rm Delete the Sandbox when the shell exits (only one this command created)\n --secret stringArray Secret to mount (repeatable)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus shell sandbox/lab\nnodus shell sandbox/lab -- zsh\nnodus shell sandbox/scratch --rm\nnodus shell"},{"id":"docs/reference/cli/nodus_ssh","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_ssh/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_ssh.md","title":"nodus ssh","description":"Connect to a Workspace over SSH","stage":"GA","headings":[{"depth":2,"slug":"nodus-ssh","text":"nodus ssh"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus ssh\n\n Connect to a Workspace over SSH\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus ssh/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus ssh\n\nConnect to a Workspace over SSH\n\nSynopsis\n\nssh uses the account-bound connection rule installed by nodus login for ...nodus (so plain ssh, scp, rsync and VS Code Remote-SSH work too) and connects. A Workspace stopped by idle, schedule or credits starts when you connect; one you stopped needs nodus start . Sign in once per computer with nodus login to create a local SSH key and register only its public key. Use –setup to repair SSH enrollment without starting or connecting to any Workspace.\n\nnodus ssh workspace/NAME [-- COMMAND...] [flags]\n\nExamples\n\n nodus login\n nodus ssh workspace/lab\n nodus ssh workspace/lab -- nvidia-smi\n\nOptions\n\n --config Print the ~/.ssh/config entry instead of connecting\n -h, --help help for ssh\n --setup Repair this account's SSH setup for all Workspaces without starting or connecting\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus"},{"id":"docs/reference/cli/nodus_start","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_start/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_start.md","title":"nodus start","description":"Start a stopped Sandbox, Workspace, Function, Agent or InferenceEndpoint","stage":"GA","headings":[{"depth":2,"slug":"nodus-start","text":"nodus start"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus start\n\n Start a stopped Sandbox, Workspace, Function, Agent or InferenceEndpoint\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus start/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus start\n\nStart a stopped Sandbox, Workspace, Function, Agent or InferenceEndpoint\n\nnodus start KIND/NAME [flags]\n\nOptions\n\n -h, --help help for start\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus start sb/dev"},{"id":"docs/reference/cli/nodus_stop","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_stop/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_stop.md","title":"nodus stop","description":"Stop a Sandbox, Workspace, Function, Agent or InferenceEndpoint","stage":"GA","headings":[{"depth":2,"slug":"nodus-stop","text":"nodus stop"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus stop\n\n Stop a Sandbox, Workspace, Function, Agent or InferenceEndpoint\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus stop/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus stop\n\nStop a Sandbox, Workspace, Function, Agent or InferenceEndpoint\n\nnodus stop KIND/NAME [flags]\n\nOptions\n\n -h, --help help for stop\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus stop sb/dev\nnodus stop sb/missing"},{"id":"docs/reference/cli/nodus_suspend","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_suspend/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_suspend.md","title":"nodus suspend","description":"Suspend a Job, Pipeline, Sweep, TrainingJob or AgentRun (checkpointing first)","stage":"GA","headings":[{"depth":2,"slug":"nodus-suspend","text":"nodus suspend"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus suspend\n\n Suspend a Job, Pipeline, Sweep, TrainingJob or AgentRun (checkpointing first)\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus suspend/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus suspend\n\nSuspend a Job, Pipeline, Sweep, TrainingJob or AgentRun (checkpointing first)\n\nnodus suspend KIND/NAME [flags]\n\nOptions\n\n -h, --help help for suspend\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus suspend job/train"},{"id":"docs/reference/cli/nodus_top","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_top/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_top.md","title":"nodus top","description":"Show GPU, CPU and memory use, the hourly rate and spend of running Jobs","stage":"GA","headings":[{"depth":2,"slug":"nodus-top","text":"nodus top"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus top\n\n Show GPU, CPU and memory use, the hourly rate and spend of running Jobs\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus top/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus top\n\nShow GPU, CPU and memory use, the hourly rate and spend of running Jobs\n\nnodus top jobs sandboxes workspaces functions [NAME] [flags]\n\nExamples\n\n nodus top jobs\n nodus top sandbox dev\n\nOptions\n\n -h, --help help for top\n --no-headers Omit table headers\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus top jobs\nnodus top job idle\nnodus top jobs --no-headers\nnodus top sandboxes\nnodus top job missing"},{"id":"docs/reference/cli/nodus_version","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_version/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_version.md","title":"nodus version","description":"Print the client version","stage":"GA","headings":[{"depth":2,"slug":"nodus-version","text":"nodus version"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus version\n\n Print the client version\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus version/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus version\n\nPrint the client version\n\nnodus version [flags]\n\nOptions\n\n -h, --help help for version\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus version"},{"id":"docs/reference/cli/nodus_volume","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_volume/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_volume.md","title":"nodus volume","description":"Put, get, list and remove the files of a Volume","stage":"GA","headings":[{"depth":2,"slug":"nodus-volume","text":"nodus volume"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus volume\n\n Put, get, list and remove the files of a Volume\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus volume/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus volume\n\nPut, get, list and remove the files of a Volume\n\nOptions\n\n -h, --help help for volume\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n nodus volume clear - Commit an empty revision; the Volume and its history stay\n nodus volume get - Download a file or directory from a Volume\n nodus volume ls - List a directory of a Volume\n nodus volume put - Add a local file or directory to a Volume as a new revision\n nodus volume rm - Remove a file or directory from a Volume as a new revision"},{"id":"docs/reference/cli/nodus_volume_clear","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_volume_clear/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_volume_clear.md","title":"nodus volume clear","description":"Commit an empty revision; the Volume and its history stay","stage":"GA","headings":[{"depth":2,"slug":"nodus-volume-clear","text":"nodus volume clear"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus volume clear\n\n Commit an empty revision; the Volume and its history stay\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus volume clear/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus volume clear\n\nCommit an empty revision; the Volume and its history stay\n\nnodus volume clear NAME [flags]\n\nExamples\n\n nodus volume clear home\n\nOptions\n\n -h, --help help for clear\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus volume - Put, get, list and remove the files of a Volume"},{"id":"docs/reference/cli/nodus_volume_get","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_volume_get/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_volume_get.md","title":"nodus volume get","description":"Download a file or directory from a Volume","stage":"GA","headings":[{"depth":2,"slug":"nodus-volume-get","text":"nodus volume get"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus volume get\n\n Download a file or directory from a Volume\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus volume get/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus volume get\n\nDownload a file or directory from a Volume\n\nSynopsis\n\nDownloads REMOTE from the latest revision (or –revision) into LOCAL (default .). LOCAL - writes a file to standard output. Existing local files are kept unless –force.\n\nnodus volume get NAME REMOTE [LOCAL] [flags]\n\nExamples\n\n nodus volume get data /train/part-0.parquet ./\n nodus volume get data /train ./train --revision 3\n nodus volume get data /config.yaml -\n\nOptions\n\n --force Overwrite existing local files\n -h, --help help for get\n --revision int32 Revision to read (default the latest)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus volume - Put, get, list and remove the files of a Volume"},{"id":"docs/reference/cli/nodus_volume_ls","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_volume_ls/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_volume_ls.md","title":"nodus volume ls","description":"List a directory of a Volume","stage":"GA","headings":[{"depth":2,"slug":"nodus-volume-ls","text":"nodus volume ls"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus volume ls\n\n List a directory of a Volume\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus volume ls/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus volume ls\n\nList a directory of a Volume\n\nnodus volume ls NAME [PATH] [flags]\n\nExamples\n\n nodus volume ls data /train\n nodus volume ls data --revision 2\n\nOptions\n\n -h, --help help for ls\n --revision int32 Revision to list (default the latest)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus volume - Put, get, list and remove the files of a Volume"},{"id":"docs/reference/cli/nodus_volume_put","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_volume_put/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_volume_put.md","title":"nodus volume put","description":"Add a local file or directory to a Volume as a new revision","stage":"GA","headings":[{"depth":2,"slug":"nodus-volume-put","text":"nodus volume put"},{"depth":3,"slug":"synopsis","text":"Synopsis"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus volume put\n\n Add a local file or directory to a Volume as a new revision\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus volume put/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus volume put\n\nAdd a local file or directory to a Volume as a new revision\n\nSynopsis\n\nAdds LOCAL at REMOTE (default /) on top of the latest revision; everything else in the Volume stays. A directory’s contents merge into REMOTE; a file lands at REMOTE, or inside it when REMOTE is a directory or ends in /. –extract expands a .zip, .tar, .tar.gz, .tgz or .tar.zst archive into REMOTE instead.\n\nnodus volume put NAME LOCAL [REMOTE] [flags]\n\nExamples\n\n nodus volume put data ./dataset /train\n nodus volume put data ./config.yaml /etc/\n nodus volume put data ./archive.tar.gz /raw --extract\n\nOptions\n\n --extract Expand the LOCAL archive into REMOTE\n -h, --help help for put\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus volume - Put, get, list and remove the files of a Volume"},{"id":"docs/reference/cli/nodus_volume_rm","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_volume_rm/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_volume_rm.md","title":"nodus volume rm","description":"Remove a file or directory from a Volume as a new revision","stage":"GA","headings":[{"depth":2,"slug":"nodus-volume-rm","text":"nodus volume rm"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"}],"text":"nodus volume rm\n\n Remove a file or directory from a Volume as a new revision\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus volume rm/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus volume rm\n\nRemove a file or directory from a Volume as a new revision\n\nnodus volume rm NAME PATH [flags]\n\nExamples\n\n nodus volume rm data /train/stale.csv\n\nOptions\n\n -h, --help help for rm\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus volume - Put, get, list and remove the files of a Volume"},{"id":"docs/reference/cli/nodus_wait","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_wait/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_wait.md","title":"nodus wait","description":"Wait until a resource meets a condition","stage":"GA","headings":[{"depth":2,"slug":"nodus-wait","text":"nodus wait"},{"depth":3,"slug":"examples","text":"Examples"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus wait\n\n Wait until a resource meets a condition\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus wait/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus wait\n\nWait until a resource meets a condition\n\nnodus wait KIND/NAME --for=condition=C jsonpath='{...}'=VALUE delete [flags]\n\nExamples\n\n nodus wait job/x --for=jsonpath='{.status.phase}'=Succeeded --timeout 1h\n\nOptions\n\n --for string condition=NAME[=STATUS], jsonpath='{.path}'=VALUE, or delete\n -h, --help help for wait\n --timeout duration Give up after this long (exit 124) (default 30m0s)\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus wait job/train --for=jsonpath={.status.phase}=Succeeded --timeout 1m\nnodus wait job/train --for=condition=Scheduled\nnodus wait job/train --for=jsonpath={.status.phase}=Running"},{"id":"docs/reference/cli/nodus_whoami","url":"https://nodus-platform-site.pages.dev/docs/reference/cli/nodus_whoami/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/cli/nodus_whoami.md","title":"nodus whoami","description":"Show the signed-in principal, org, role, project, scopes and available credit","stage":"GA","headings":[{"depth":2,"slug":"nodus-whoami","text":"nodus whoami"},{"depth":3,"slug":"options","text":"Options"},{"depth":3,"slug":"options-inherited-from-parent-commands","text":"Options inherited from parent commands"},{"depth":3,"slug":"see-also","text":"SEE ALSO"},{"depth":3,"slug":"tested-examples","text":"Tested examples"}],"text":"nodus whoami\n\n Show the signed-in principal, org, role, project, scopes and available credit\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/cli/nodus whoami/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nnodus whoami\n\nShow the signed-in principal, org, role, project, scopes and available credit\n\nnodus whoami [flags]\n\nOptions\n\n -h, --help help for whoami\n -o, --output string json for machine-readable output\n\nOptions inherited from parent commands\n\n --context string Context from the config file to use\n --org string Organization (selects the context for that org)\n -p, --project string Project to work in\n -v, --verbose Log each API request (never credentials)\n\nSEE ALSO\n\n nodus - Run compute Jobs on Nodus\n\nTested examples\n\nThese invocations run in CI against the CLI’s fake API.\n\nTerminal window\n\nnodus whoami"},{"id":"docs/reference/environments","url":"https://nodus-platform-site.pages.dev/docs/reference/environments/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/environments.md","title":"Environments","description":"Every catalog Environment, its category, reward and readiness.","stage":"GA","headings":[],"text":"Environments\n\n Every catalog Environment, its category, reward and readiness.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/environments/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nThe catalog Environments in project nodus ; reference one as spec.environment.name: nodus/ @ on a TrainingJob or an agent eval.\n\n Environment Category Reward Readiness Summary \n - - - - - \n arithmetic-v2 Math Binary Stable Evaluate short integer expressions with standard precedence and give the exact result. \n graph-coloring Reasoning Binary Stable Color a small graph with three colors so that no edge joins two nodes of the same color. \n gsm8k Math Binary Research Eight thousand grade-school maths word problems, each with a worked solution and a single numeric answer. \n letter-counting-legacy-eval Reasoning Scalar Research Legacy letter-counting evaluation protocol: first 64 generated tasks, raw completion scoring. \n letter-counting-legacy-rl Reasoning Scalar Research Legacy letter-counting RL protocol: first 64 generated tasks train, next 64 held out. \n python-functions Code Binary Stable Write solve(values: list\\[int]) - int from a one-line specification; private cases decide the verdict. \n reasoning-gym Reasoning Scalar Research Procedural reasoning generators with deterministic algorithmic scorers; only reviewed families are served."},{"id":"docs/reference/environments/arithmetic-v2","url":"https://nodus-platform-site.pages.dev/docs/reference/environments/arithmetic-v2/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/environments/arithmetic-v2.md","title":"arithmetic-v2","description":"Evaluate short integer expressions with standard precedence and give the exact result.","stage":"GA","headings":[{"depth":2,"slug":"grading","text":"Grading"},{"depth":2,"slug":"sample-task","text":"Sample task"},{"depth":2,"slug":"examples","text":"Examples"}],"text":"arithmetic-v2\n\n Evaluate short integer expressions with standard precedence and give the exact result.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/environments/arithmetic-v2/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nEvaluate short integer expressions with standard precedence and give the exact result.\n\n Field Value \n - - \n Reference nodus/arithmetic-v2@2.0.0 \n Image nodus/env-arithmetic-v2:2.0.0 \n Publisher Nodus \n Category Math \n Readiness Stable \n Modes Train, Evaluate \n Reward Binary \n Held-out measures TrainedTask \n Splits 500 train, 128 test (disjoint by canonical identity) \n Licenses code Apache-2.0, data Apache-2.0 \n Source \n\nGrading\n\nCompletions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.\n\n Grader Kind What it checks \n - - - \n exact-integer Program The last N equals the expression’s value \n\nSample task\n\nOne line of nodus-env tasks --split test --seed 0 ; tasks never carry answers.\n\n{\n \"metadata\": {\n \"terms\": 3\n },\n \"prompt\": \"Compute 41 - 38 33. Give only the final integer, like \\u003canswer\\u003e42\\u003c/answer\\u003e.\",\n \"taskId\": \"arithmetic-v2:test:0:0\"\n}\n\nExamples\n\nEach example is a TrainingJob template. A baseline is shown only where the example was measured by running it.\n\n Example Mode Runtime Model Tasks Baseline Trained Measured on \n - - - - - - - - \n arithmetic-grpo Train nodus/grpo-lora Qwen/Qwen3-0.6B @ c1899de 64 not measured not measured not measured \n\nRun one with a server dry-run first:\n\nTerminal window\n\n$ nodus create trainingjob my-run --from-example nodus/arithmetic-v2:arithmetic-grpo --dry-run=server -o estimate\n\nimport nodus\n\njob = nodus.recipes.TrainingJob.from example(\"nodus/arithmetic-v2:arithmetic-grpo\")\nplan = job.preview()\nrun = plan.run(max cost=5)\nprint(run.wait().summary)"},{"id":"docs/reference/environments/graph-coloring","url":"https://nodus-platform-site.pages.dev/docs/reference/environments/graph-coloring/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/environments/graph-coloring.md","title":"graph-coloring","description":"Color a small graph with three colors so that no edge joins two nodes of the same color.","stage":"GA","headings":[{"depth":2,"slug":"grading","text":"Grading"},{"depth":2,"slug":"sample-task","text":"Sample task"},{"depth":2,"slug":"examples","text":"Examples"}],"text":"graph-coloring\n\n Color a small graph with three colors so that no edge joins two nodes of the same color.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/environments/graph-coloring/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nColor a small graph with three colors so that no edge joins two nodes of the same color.\n\n Field Value \n - - \n Reference nodus/graph-coloring@1.0.0 \n Image nodus/env-graph-coloring:1.0.0 \n Publisher Nodus \n Category Reasoning \n Readiness Stable \n Modes Train, Evaluate \n Reward Binary \n Held-out measures TrainedTask \n Splits 200 train, 64 test (disjoint by canonical identity) \n Licenses code Apache-2.0, data Apache-2.0 \n Source \n\nGrading\n\nCompletions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.\n\n Grader Kind What it checks \n - - - \n proper-coloring Program Checks every edge joins two different colors; any proper coloring is correct \n\nSample task\n\nOne line of nodus-env tasks --split test --seed 0 ; tasks never carry answers.\n\n{\n \"metadata\": {\n \"edges\": 12,\n \"nodes\": 7\n },\n \"prompt\": \"A graph has 7 nodes numbered 0 to 6 and these edges: 0-1, 0-2, 0-4, 1-3, 1-6, 2-3, 2-4, 2-5, 3-4, 3-5, 3-6, 4-6.\\nAssign each node one of the colors 0, 1 or 2 so that no edge connects two nodes of the same color.\\nAnswer with the colors of nodes 0 to 6 in order, like \\u003canswer\\u003e[0, 1, 2, 0]\\u003c/answer\\u003e.\",\n \"taskId\": \"graph-coloring:test:0:0\"\n}\n\nExamples\n\nEach example is a TrainingJob template. A baseline is shown only where the example was measured by running it.\n\n Example Mode Runtime Model Tasks Baseline Trained Measured on \n - - - - - - - - \n graph-coloring-grpo Train nodus/grpo-lora Qwen/Qwen3-0.6B @ c1899de 64 not measured not measured not measured \n\nRun one with a server dry-run first:\n\nTerminal window\n\n$ nodus create trainingjob my-run --from-example nodus/graph-coloring:graph-coloring-grpo --dry-run=server -o estimate\n\nimport nodus\n\njob = nodus.recipes.TrainingJob.from example(\"nodus/graph-coloring:graph-coloring-grpo\")\nplan = job.preview()\nrun = plan.run(max cost=5)\nprint(run.wait().summary)"},{"id":"docs/reference/environments/gsm8k","url":"https://nodus-platform-site.pages.dev/docs/reference/environments/gsm8k/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/environments/gsm8k.md","title":"gsm8k","description":"Eight thousand grade-school maths word problems, each with a worked solution and a single numeric answer.","stage":"GA","headings":[{"depth":2,"slug":"grading","text":"Grading"},{"depth":2,"slug":"sample-task","text":"Sample task"},{"depth":2,"slug":"examples","text":"Examples"}],"text":"gsm8k\n\n Eight thousand grade-school maths word problems, each with a worked solution and a single numeric answer.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/environments/gsm8k/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nEight thousand grade-school maths word problems, each with a worked solution and a single numeric answer.\n\n Field Value \n - - \n Reference nodus/gsm8k@1.0.0 \n Image nodus/env-gsm8k:1.0.0 \n Publisher OpenAI \n Category Math \n Readiness Research \n Modes Train, Evaluate \n Reward Binary \n Held-out measures TrainedTask \n Splits 7473 train, 1319 test (disjoint by canonical identity) \n Licenses code MIT, data MIT \n Source \n\nGrading\n\nCompletions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.\n\n Grader Kind What it checks \n - - - \n final-number Program The last number, else the last number, equals the value after #### \n\nSample task\n\nOne line of nodus-env tasks --split test --seed 0 ; tasks never carry answers.\n\n{\n \"metadata\": {\n \"source\": \"openai/gsm8k\"\n },\n \"prompt\": \"A fruit vendor bought 50 watermelons for $80. He sold all of them at a profit of 25%. How much was each watermelon sold?\\n\\nWork through the problem step by step, then give the final number inside answer tags, like \\u003canswer\\u003e42\\u003c/answer\\u003e.\",\n \"taskId\": \"gsm8k:test:0:0\"\n}\n\nExamples\n\nEach example is a TrainingJob template. A baseline is shown only where the example was measured by running it.\n\n Example Mode Runtime Model Tasks Baseline Trained Measured on \n - - - - - - - - \n gsm8k-trained Train nodus/grpo-lora Qwen/Qwen3-1.7B @ 70d244c 64 79.7 % 87.5 % a40-48g-x1, 2026-09-24 \n\nRun one with a server dry-run first:\n\nTerminal window\n\n$ nodus create trainingjob my-run --from-example nodus/gsm8k:gsm8k-trained --dry-run=server -o estimate\n\nimport nodus\n\njob = nodus.recipes.TrainingJob.from example(\"nodus/gsm8k:gsm8k-trained\")\nplan = job.preview()\nrun = plan.run(max cost=5)\nprint(run.wait().summary)"},{"id":"docs/reference/environments/letter-counting-legacy-eval","url":"https://nodus-platform-site.pages.dev/docs/reference/environments/letter-counting-legacy-eval/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/environments/letter-counting-legacy-eval.md","title":"letter-counting-legacy-eval","description":"Legacy letter-counting evaluation protocol: first 64 generated tasks, raw completion scoring.","stage":"GA","headings":[{"depth":2,"slug":"grading","text":"Grading"},{"depth":2,"slug":"sample-task","text":"Sample task"},{"depth":2,"slug":"examples","text":"Examples"}],"text":"letter-counting-legacy-eval\n\n Legacy letter-counting evaluation protocol: first 64 generated tasks, raw completion scoring.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/environments/letter-counting-legacy-eval/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nLegacy letter-counting evaluation protocol: first 64 generated tasks, raw completion scoring.\n\n Field Value \n - - \n Reference nodus/letter-counting-legacy-eval@1.0.0 \n Image nodus/env-letter-counting-legacy-eval:1.0.0 \n Publisher Open-Thought \n Category Reasoning \n Readiness Research \n Modes Evaluate \n Reward Scalar \n Held-out measures TrainedTask \n Splits 0 train, 64 test (disjoint by canonical identity) \n Licenses code Apache-2.0, data Apache-2.0 \n Source \n\nGrading\n\nCompletions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.\n\n Grader Kind What it checks \n - - - \n score-answer Program Raw completion scored by pinned Reasoning Gym; only full credit is Correct \n\nSample task\n\nOne line of nodus-env tasks --split test --seed 0 ; tasks never carry answers.\n\n{\n \"metadata\": {\n \"family\": \"letter counting\",\n \"systemPrompt\": \"Reply with the final answer only. Do not restate the question, do not show working, and do not add any words, labels or punctuation around it.\"\n },\n \"prompt\": \"How many times does the letter \\\"b\\\" appear in the text: \\\"phrase Project Gutenberg associated with or appearing on the work you\\\"?\",\n \"taskId\": \"letter-counting-legacy-eval:test:0:0\"\n}\n\nExamples\n\nEach example is a TrainingJob template. A baseline is shown only where the example was measured by running it.\n\n Example Mode Runtime Model Tasks Baseline Trained Measured on \n - - - - - - - - \n letter-counting Evaluate nodus/evaluate Qwen/Qwen3-1.7B @ 70d244c 64 not measured n/a (evaluation) not measured \n\nRun one with a server dry-run first:\n\nTerminal window\n\n$ nodus create trainingjob my-run --from-example nodus/letter-counting-legacy-eval:letter-counting --dry-run=server -o estimate\n\nimport nodus\n\njob = nodus.recipes.TrainingJob.from example(\"nodus/letter-counting-legacy-eval:letter-counting\")\nplan = job.preview()\nrun = plan.run(max cost=5)\nprint(run.wait().summary)"},{"id":"docs/reference/environments/letter-counting-legacy-rl","url":"https://nodus-platform-site.pages.dev/docs/reference/environments/letter-counting-legacy-rl/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/environments/letter-counting-legacy-rl.md","title":"letter-counting-legacy-rl","description":"Legacy letter-counting RL protocol: first 64 generated tasks train, next 64 held out.","stage":"GA","headings":[{"depth":2,"slug":"grading","text":"Grading"},{"depth":2,"slug":"sample-task","text":"Sample task"},{"depth":2,"slug":"examples","text":"Examples"}],"text":"letter-counting-legacy-rl\n\n Legacy letter-counting RL protocol: first 64 generated tasks train, next 64 held out.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/environments/letter-counting-legacy-rl/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nLegacy letter-counting RL protocol: first 64 generated tasks train, next 64 held out.\n\n Field Value \n - - \n Reference nodus/letter-counting-legacy-rl@1.0.0 \n Image nodus/env-letter-counting-legacy-rl:1.0.0 \n Publisher Open-Thought \n Category Reasoning \n Readiness Research \n Modes Train, Evaluate \n Reward Scalar \n Held-out measures TrainedTask \n Splits 64 train, 64 test (disjoint by canonical identity) \n Licenses code Apache-2.0, data Apache-2.0 \n Source \n\nGrading\n\nCompletions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.\n\n Grader Kind What it checks \n - - - \n score-answer Program Raw completion scored by pinned Reasoning Gym; only full credit is Correct \n\nSample task\n\nOne line of nodus-env tasks --split test --seed 0 ; tasks never carry answers.\n\n{\n \"metadata\": {\n \"family\": \"letter counting\",\n \"systemPrompt\": \"Reply with the final answer only. Do not restate the question, do not show working, and do not add any words, labels or punctuation around it.\"\n },\n \"prompt\": \"How many times does the letter \\\"t\\\" appear in the text: \\\"are far more complex than all that In real life every act\\\"?\",\n \"taskId\": \"letter-counting-legacy-rl:test:0:0\"\n}\n\nExamples\n\nEach example is a TrainingJob template. A baseline is shown only where the example was measured by running it.\n\n Example Mode Runtime Model Tasks Baseline Trained Measured on \n - - - - - - - - \n letter-counting Train nodus/grpo-lora Qwen/Qwen3-1.7B @ 70d244c 64 not measured not measured not measured \n\nRun one with a server dry-run first:\n\nTerminal window\n\n$ nodus create trainingjob my-run --from-example nodus/letter-counting-legacy-rl:letter-counting --dry-run=server -o estimate\n\nimport nodus\n\njob = nodus.recipes.TrainingJob.from example(\"nodus/letter-counting-legacy-rl:letter-counting\")\nplan = job.preview()\nrun = plan.run(max cost=5)\nprint(run.wait().summary)"},{"id":"docs/reference/environments/python-functions","url":"https://nodus-platform-site.pages.dev/docs/reference/environments/python-functions/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/environments/python-functions.md","title":"python-functions","description":"Write solve(values: list[int]) -> int from a one-line specification; private cases decide the verdict.","stage":"GA","headings":[{"depth":2,"slug":"grading","text":"Grading"},{"depth":2,"slug":"sample-task","text":"Sample task"},{"depth":2,"slug":"examples","text":"Examples"}],"text":"python-functions\n\n Write solve(values: list[int]) - int from a one-line specification; private cases decide the verdict.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/environments/python-functions/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nWrite solve(values: list\\[int]) - int from a one-line specification; private cases decide the verdict.\n\n Field Value \n - - \n Reference nodus/python-functions@1.0.0 \n Image nodus/env-python-functions:1.0.0 \n Publisher Nodus \n Category Code \n Readiness Stable \n Modes Train, Evaluate \n Reward Binary \n Held-out measures TrainedTask \n Splits 20 train, 16 test (disjoint by canonical identity) \n Licenses code Apache-2.0, data Apache-2.0 \n Source \n\nGrading\n\nCompletions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.\n\n Grader Kind What it checks \n - - - \n private-cases Program Runs the candidate as an unprivileged uid on public inputs; every private expected output must match \n\nGrading runs in a sandbox ( nodus/env-python-functions:1.0.0 , 1 CPU, 1Gi memory, 30s timeout, no network); candidate code runs as an unprivileged user that cannot read the expected answers.\n\nSample task\n\nOne line of nodus-env tasks --split test --seed 0 ; tasks never carry answers.\n\n{\n \"metadata\": {\n \"function\": \"count-divisible-by-last\"\n },\n \"prompt\": \"Write Python source defining solve(values: list[int]) -\\u003e int. Return how many values are divisible by the last value, or 0 when it is zero. Return 0 for an empty list. Output only Python source, without Markdown fences.\",\n \"taskId\": \"python-functions:test:0:0\"\n}\n\nExamples\n\nEach example is a TrainingJob template. A baseline is shown only where the example was measured by running it.\n\n Example Mode Runtime Model Tasks Baseline Trained Measured on \n - - - - - - - - \n python-functions-grpo Train nodus/grpo-lora Qwen/Qwen3-0.6B @ c1899de 16 not measured not measured not measured \n\nRun one with a server dry-run first:\n\nTerminal window\n\n$ nodus create trainingjob my-run --from-example nodus/python-functions:python-functions-grpo --dry-run=server -o estimate\n\nimport nodus\n\njob = nodus.recipes.TrainingJob.from example(\"nodus/python-functions:python-functions-grpo\")\nplan = job.preview()\nrun = plan.run(max cost=5)\nprint(run.wait().summary)"},{"id":"docs/reference/environments/reasoning-gym","url":"https://nodus-platform-site.pages.dev/docs/reference/environments/reasoning-gym/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/environments/reasoning-gym.md","title":"reasoning-gym","description":"Procedural reasoning generators with deterministic algorithmic scorers; only reviewed families are served.","stage":"GA","headings":[{"depth":2,"slug":"grading","text":"Grading"},{"depth":2,"slug":"sample-task","text":"Sample task"},{"depth":2,"slug":"examples","text":"Examples"}],"text":"reasoning-gym\n\n Procedural reasoning generators with deterministic algorithmic scorers; only reviewed families are served.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/environments/reasoning-gym/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nProcedural reasoning generators with deterministic algorithmic scorers; only reviewed families are served.\n\n Field Value \n - - \n Reference nodus/reasoning-gym@1.0.0 \n Image nodus/env-reasoning-gym:1.0.0 \n Publisher Open-Thought \n Category Reasoning \n Readiness Research \n Modes Train, Evaluate \n Reward Scalar \n Held-out measures TrainedTask \n Splits 2000 train, 320 test (disjoint by canonical identity) \n Licenses code Apache-2.0, data Apache-2.0 \n Source \n\nGrading\n\nCompletions are graded by the platform, never by the trainer: the trainer submits {taskId, completion} batches and the verdicts come back as task events.\n\n Grader Kind What it checks \n - - - \n score-answer Program The family’s score answer; only a full score is Correct, the score is the reward \n\nSample task\n\nOne line of nodus-env tasks --split test --seed 0 ; tasks never carry answers.\n\n{\n \"metadata\": {\n \"family\": \"basic arithmetic\"\n },\n \"prompt\": \"Calculate -3 - 8 / ( -8 + 1 + 8 ).\\n\\nReply with the final answer only. Do not restate the question, do not show working, and do not add any words, labels or punctuation around it.\",\n \"taskId\": \"reasoning-gym:test:0:0\"\n}\n\nExamples\n\nEach example is a TrainingJob template. A baseline is shown only where the example was measured by running it.\n\n Example Mode Runtime Model Tasks Baseline Trained Measured on \n - - - - - - - - \n chain-sum Evaluate nodus/evaluate Qwen/Qwen3-1.7B @ 70d244c 64 not measured n/a (evaluation) not measured \n number-format Evaluate nodus/evaluate Qwen/Qwen3-1.7B @ 70d244c 64 not measured n/a (evaluation) not measured \n basic-arithmetic Evaluate nodus/evaluate Qwen/Qwen3-1.7B @ 70d244c 64 not measured n/a (evaluation) not measured \n products Evaluate nodus/evaluate Qwen/Qwen3-1.7B @ 70d244c 64 not measured n/a (evaluation) not measured \n letter-counting Evaluate nodus/evaluate Qwen/Qwen3-1.7B @ 70d244c 64 not measured n/a (evaluation) not measured \n\nRun one with a server dry-run first:\n\nTerminal window\n\n$ nodus create trainingjob my-run --from-example nodus/reasoning-gym:chain-sum --dry-run=server -o estimate\n\nimport nodus\n\njob = nodus.recipes.TrainingJob.from example(\"nodus/reasoning-gym:chain-sum\")\nplan = job.preview()\nrun = plan.run(max cost=5)\nprint(run.wait().summary)"},{"id":"docs/reference/errors","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors.md","title":"Error codes","description":"One page per API error code, with its HTTP status, what it means and how to fix it.","stage":"GA","headings":[],"text":"Error codes\n\n One page per API error code, with its HTTP status, what it means and how to fix it.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nEvery API error is a Kubernetes-style Status object. reason is a stable code, code is the HTTP status, and three extra fields tell you what to do next:\n\n{\n \"kind\": \"Status\",\n \"apiVersion\": \"v1\",\n \"status\": \"Failure\",\n \"reason\": \"InsufficientCredits\",\n \"code\": 402,\n \"message\": \"job \\\"train-a\\\" needs a $3.20 hold to start (released when it ends); available $1.10\",\n \"fix\": \"nodus billing top-up 20, or lower spec.maxCostUSD\",\n \"docs\": \"https://nodus-compute.ai/docs/reference/errors/insufficient-credits\",\n \"requestId\": \"req 01j9…\"\n}\n\nInclude the requestId when you contact support. The docs link is the code’s page below, named after the code in lowercase words ( InsufficientCredits is insufficient-credits ).\n\nAdmissionsPaused Operators paused admissions for this kind; existing objects keep running.\n\nAgentGroupClosed A sealed, canceled, finished or deleting AgentGroup admits no new member runs.\n\nAgentGroupFull An AgentGroup admits at most spec.limits.maxPending unfinished member runs.\n\nAgentRunFinished A run in a terminal phase (Succeeded, Failed or canceled) no longer reads its inbox.\n\nAgentStopped An Agent in the Stopped state refuses new runs; runs already created are not affected.\n\nAlreadyExists A create by name found an object with that name whose spec differs; an identical create returns the object.\n\nArrearsOutstanding Charges that the org's credits could not cover are unpaid; they block new work until they are settled.\n\nBadRequest The request is malformed: a query parameter, selector or body could not be parsed.\n\nBudgetExceeded A Block budget or a maxCostUSD cap that applies to the object has less room than the action's hold.\n\nCapacityUnavailable No offering that meets the requirements has capacity at the moment.\n\nCloudAccountNotVerified Nodus reads inventory only after it has verified the account's read-only access.\n\nConflict The object changed after the caller read it (the resourceVersion is stale).\n\nConnectionNotReady A Connection is usable once a verification passed.\n\nEncryptionFailed The org's encryption key was unavailable; nothing was stored.\n\nEnrollmentTokenInvalid An enrollment token works once, for 24 hours, for the pool it was created for.\n\nExpired The watch position or list continue token is older than the 24-hour change history.\n\nFieldImmutable The update changes a spec field that is immutable after create; the causes list each changed field.\n\nFileTooLarge A file request carries at most 64 MiB; larger data does not go through the files API.\n\nForbidden The credential is valid but lacks the scope or project access the operation needs.\n\nFunctionCallBatchTooLarge A batch create takes at most 1,000 calls and 8 MiB.\n\nFunctionNotFound The Function does not exist in the project, or it was deleted since the call was prepared.\n\nIdempotencyKeyReused The Idempotency-Key was used in the last 24 hours with a different request body.\n\nImageBuildUnavailable This deployment cannot build image steps, Dockerfiles or Sandbox snapshots; prebuilt registry images remain supported.\n\nImageNotFound The tag or digest does not exist in the registry, or the pull credentials cannot see it.\n\nImageNotReady An imageRef runs the digest the Image built; the Image has none yet.\n\nImagePullFailed The registry refused the credentials or was unavailable while pinning the image to a digest.\n\nInsufficientCredits The org's available credits do not cover the hold that the action reserves before it starts.\n\nInternalError The server hit an unexpected failure; the request id finds it in the logs.\n\nInvalid The object fails strict decoding or validation; \\ details.causes\\ lists every field and reason.\n\nMethodNotAllowed The resource does not support this verb (for example, a read-only kind).\n\nNoSSHKey The Workspace accepts only the SSHKeys named in spec.sshKeys, or every project member's keys when it names none.\n\nNotFound The object, project or path does not exist in this org.\n\nOutputIndexRequired Several indexes of an Indexed Job committed the output; the download must pick one.\n\nOutputNotFound The Job did not commit an output by that name, or not at that index.\n\nPaymentDisputed A card payment on the org is disputed; new work cannot start until the dispute closes.\n\nPaymentMethodRequired New work requires a verified card, including work funded by promotional credit. Adding a card does not charge it.\n\nPaymentsNotConfigured The deployment holds no Stripe key, so it takes no card payments; balances, grants and usage still work.\n\nPaymentVerificationUnavailable Stripe could not confirm the payment method. Existing running work retains its funded holds.\n\nPoolActionNotPending Only a queued, undecided action waits for a decision; it may have expired or been decided by someone else.\n\nPoolActionNotReversible Only an executed drain, undrain or routing change can be reverted, once.\n\nPoolActPaused The pool's kill switch stops every action until it is resumed.\n\nPoolInUse A Pool is deleted only once no Node is enrolled in it and no object names it in spec.placement.pool.\n\nPreconditionFailed The If-Match value no longer matches the compiled spec and price book of a fresh dry-run.\n\nPriceConsentRequired Pool routing and Predict are billed, so turning either on records consent to its current price.\n\nPromoCodeExpired The promo code exists but can no longer be redeemed.\n\nPromoCodeUsed The org has already redeemed the code as often as it may, or the code has no redemptions left.\n\nQuotaExceeded The org or project already uses everything its Quota allows of this counter.\n\nRequestEntityTooLarge The request body is larger than the route accepts (1 MiB for JSON).\n\nRequestInProgress The original request with this Idempotency-Key has not finished yet.\n\nSandboxFailed The Sandbox is in the Failed phase, which is final: status.reason says why.\n\nSandboxNotRunning The Sandbox is being deleted, so no command, file or port request can reach it and none starts it.\n\nSandboxStarting The request needs the Sandbox's container, which is not running yet. A stopped Sandbox is started by the request itself; the answer carries Retry-After.\n\nSecretValueInEnv Secret values reach a Job or Sandbox only through secrets or valueFrom.secretKeyRef, never as literals.\n\nSHA256Mismatch A write with expectedSHA256 found different content than the writer started from, so nothing was written.\n\nStdinBackpressure The process's stdin buffer is full, so the write was refused without losing earlier input.\n\nTooManyRequests The credential sent too many requests of this class; the RateLimit headers show the budget.\n\nTopUpLimitReached An org buys at most $5,000 of credit through Checkout per UTC day.\n\nUnauthorized The request carries no valid credential, or the credential expired.\n\nUnavailable A dependency of the API is temporarily unavailable.\n\nUnsupported The request is well formed but asks for a feature this API does not offer.\n\nUnsupportedMediaType The request body's Content-Type is not accepted; strategic merge patch is not supported.\n\nVolumeBusy A ReadWriteOnce Volume has one writer; another attempt holds its lease.\n\nWorkerNotAllowed Only the Function's own worker attempts, through their ServiceAccount token, can claim its calls.\n\nWorkspaceStarting The Workspace has no running session yet: it is starting, or the request woke it from a stop.\n\nWorkspaceToolUnavailable A browser tool has a preview URL only while the Workspace is running and lists the tool in spec.tools."},{"id":"docs/reference/errors/admissions-paused","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/admissions-paused/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/admissions-paused.md","title":"AdmissionsPaused","description":"Operators paused admissions for this kind; existing objects keep running.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"AdmissionsPaused\n\n Operators paused admissions for this kind; existing objects keep running.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/admissions-paused/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 503 · Retryable: yes\n\nOperators paused admissions for this kind; existing objects keep running.\n\nMessage\n\nnew {resource} are not being admitted right now\n\nWhat to do\n\nretry later; shows the incident\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: AdmissionsPaused and code: 503 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.AdmissionsPaused ."},{"id":"docs/reference/errors/agent-group-closed","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/agent-group-closed/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/agent-group-closed.md","title":"AgentGroupClosed","description":"A sealed, canceled, finished or deleting AgentGroup admits no new member runs.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"AgentGroupClosed\n\n A sealed, canceled, finished or deleting AgentGroup admits no new member runs.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/agent-group-closed/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nA sealed, canceled, finished or deleting AgentGroup admits no new member runs.\n\nMessage\n\nagent group \"{group}\" is {state} and takes no new runs\n\nWhat to do\n\ncreate the runs in a new group ( nodus create agentgroup --agent )\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: AgentGroupClosed and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.AgentGroupClosed ."},{"id":"docs/reference/errors/agent-group-full","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/agent-group-full/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/agent-group-full.md","title":"AgentGroupFull","description":"An AgentGroup admits at most spec.limits.maxPending unfinished member runs.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"AgentGroupFull\n\n An AgentGroup admits at most spec.limits.maxPending unfinished member runs.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/agent-group-full/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: yes\n\nAn AgentGroup admits at most spec.limits.maxPending unfinished member runs.\n\nMessage\n\nagent group \"{group}\" already has {limit} unfinished runs\n\nWhat to do\n\nwait for runs to finish, or raise the limit ( nodus patch ag/{group} -p '{\"spec\":{\"limits\":{\"maxPending\": }}}' )\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: AgentGroupFull and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.AgentGroupFull ."},{"id":"docs/reference/errors/agent-run-finished","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/agent-run-finished/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/agent-run-finished.md","title":"AgentRunFinished","description":"A run in a terminal phase (Succeeded, Failed or canceled) no longer reads its inbox.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"AgentRunFinished\n\n A run in a terminal phase (Succeeded, Failed or canceled) no longer reads its inbox.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/agent-run-finished/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nA run in a terminal phase (Succeeded, Failed or canceled) no longer reads its inbox.\n\nMessage\n\nagent run \"{name}\" has finished and takes no more messages\n\nWhat to do\n\nstart a new run with nodus create agentrun --agent \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: AgentRunFinished and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.AgentRunFinished ."},{"id":"docs/reference/errors/agent-stopped","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/agent-stopped/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/agent-stopped.md","title":"AgentStopped","description":"An Agent in the Stopped state refuses new runs; runs already created are not affected.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"AgentStopped\n\n An Agent in the Stopped state refuses new runs; runs already created are not affected.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/agent-stopped/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nAn Agent in the Stopped state refuses new runs; runs already created are not affected.\n\nMessage\n\nagent \"{agent}\" is stopped and takes no new runs\n\nWhat to do\n\nset the agent’s state to Running ( nodus patch agent/{agent} -p '{\"spec\":{\"state\":\"Running\"}}' ) and create the run again\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: AgentStopped and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.AgentStopped ."},{"id":"docs/reference/errors/already-exists","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/already-exists/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/already-exists.md","title":"AlreadyExists","description":"A create by name found an object with that name whose spec differs; an identical create returns the object.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"AlreadyExists\n\n A create by name found an object with that name whose spec differs; an identical create returns the object.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/already-exists/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nA create by name found an object with that name whose spec differs; an identical create returns the object.\n\nMessage\n\n{resource} \"{name}\" already exists with a different spec: {fields} differ\n\nWhat to do\n\nchoose another name, delete the existing object, or apply the change with nodus apply \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: AlreadyExists and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.AlreadyExists .\n\n details also carries:\n\n diff , which the Python SDK exposes as diff"},{"id":"docs/reference/errors/arrears-outstanding","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/arrears-outstanding/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/arrears-outstanding.md","title":"ArrearsOutstanding","description":"Charges that the org's credits could not cover are unpaid; they block new work until they are settled.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"ArrearsOutstanding\n\n Charges that the org's credits could not cover are unpaid; they block new work until they are settled.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/arrears-outstanding/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 402 · Retryable: no\n\nCharges that the org’s credits could not cover are unpaid; they block new work until they are settled.\n\nMessage\n\nthis org has {arrears} of unpaid charges; new work is paused until they are paid\n\nWhat to do\n\nadd credits with nodus billing top-up {topUp} ; a top-up pays the unpaid charges first\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: ArrearsOutstanding and code: 402 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.ArrearsOutstanding .\n\n details also carries:\n\n arrearsUSD , which the Python SDK exposes as arrears usd"},{"id":"docs/reference/errors/bad-request","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/bad-request/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/bad-request.md","title":"BadRequest","description":"The request is malformed: a query parameter, selector or body could not be parsed.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"BadRequest\n\n The request is malformed: a query parameter, selector or body could not be parsed.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/bad-request/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 400 · Retryable: no\n\nThe request is malformed: a query parameter, selector or body could not be parsed.\n\nMessage\n\n{reason}\n\nWhat to do\n\ncorrect the request; nodus explain shows the accepted fields\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: BadRequest and code: 400 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.BadRequest ."},{"id":"docs/reference/errors/budget-exceeded","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/budget-exceeded/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/budget-exceeded.md","title":"BudgetExceeded","description":"A Block budget or a maxCostUSD cap that applies to the object has less room than the action's hold.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"BudgetExceeded\n\n A Block budget or a maxCostUSD cap that applies to the object has less room than the action's hold.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/budget-exceeded/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 402 · Retryable: no\n\nA Block budget or a maxCostUSD cap that applies to the object has less room than the action’s hold.\n\nMessage\n\n{limit} has {remaining} left of {limitUSD} {period}; {subject} needs {need}\n\nWhat to do\n\nraise the limit ({field} on {limitRef}), or wait until {resetsAt}\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: BudgetExceeded and code: 402 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.BudgetExceeded .\n\n details also carries:\n\n budget , which the Python SDK exposes as budget \n neededUSD , which the Python SDK exposes as needed usd \n remainingUSD , which the Python SDK exposes as remaining usd"},{"id":"docs/reference/errors/capacity-unavailable","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/capacity-unavailable/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/capacity-unavailable.md","title":"CapacityUnavailable","description":"No offering that meets the requirements has capacity at the moment.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"CapacityUnavailable\n\n No offering that meets the requirements has capacity at the moment.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/capacity-unavailable/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 503 · Retryable: no\n\nNo offering that meets the requirements has capacity at the moment.\n\nMessage\n\nno matching capacity is available right now{where}\n\nWhat to do\n\nwait, or widen placement.regions or resources.gpu.type\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: CapacityUnavailable and code: 503 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.CapacityUnavailable .\n\n details also carries:\n\n eta , which the Python SDK exposes as eta"},{"id":"docs/reference/errors/cloud-account-not-verified","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/cloud-account-not-verified/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/cloud-account-not-verified.md","title":"CloudAccountNotVerified","description":"Nodus reads inventory only after it has verified the account's read-only access.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"CloudAccountNotVerified\n\n Nodus reads inventory only after it has verified the account's read-only access.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/cloud-account-not-verified/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nNodus reads inventory only after it has verified the account’s read-only access.\n\nMessage\n\ncloud account \"{name}\" is not verified: {reason}\n\nWhat to do\n\nfinish onboarding with nodus get cloudaccount/{name} --subresource onboarding \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: CloudAccountNotVerified and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.CloudAccountNotVerified ."},{"id":"docs/reference/errors/conflict","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/conflict/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/conflict.md","title":"Conflict","description":"The object changed after the caller read it (the resourceVersion is stale).","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"Conflict\n\n The object changed after the caller read it (the resourceVersion is stale).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/conflict/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nThe object changed after the caller read it (the resourceVersion is stale).\n\nMessage\n\n{resource} \"{name}\" was modified since it was read; the update was not applied\n\nWhat to do\n\nread the object again and reapply the change\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: Conflict and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.Conflict ."},{"id":"docs/reference/errors/connection-not-ready","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/connection-not-ready/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/connection-not-ready.md","title":"ConnectionNotReady","description":"A Connection is usable once a verification passed.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"ConnectionNotReady\n\n A Connection is usable once a verification passed.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/connection-not-ready/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: yes\n\nA Connection is usable once a verification passed.\n\nMessage\n\nconnection \"{name}\" is not verified: {reason}\n\nWhat to do\n\nfix the credentials Secret and run nodus describe connection/{name} to see the last check\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: ConnectionNotReady and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.ConnectionNotReady ."},{"id":"docs/reference/errors/encryption-failed","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/encryption-failed/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/encryption-failed.md","title":"EncryptionFailed","description":"The org's encryption key was unavailable; nothing was stored.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"EncryptionFailed\n\n The org's encryption key was unavailable; nothing was stored.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/encryption-failed/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 503 · Retryable: yes\n\nThe org’s encryption key was unavailable; nothing was stored.\n\nMessage\n\nthe values of secret \"{name}\" could not be encrypted\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: EncryptionFailed and code: 503 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.EncryptionFailed ."},{"id":"docs/reference/errors/enrollment-token-invalid","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/enrollment-token-invalid/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/enrollment-token-invalid.md","title":"EnrollmentTokenInvalid","description":"An enrollment token works once, for 24 hours, for the pool it was created for.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"EnrollmentTokenInvalid\n\n An enrollment token works once, for 24 hours, for the pool it was created for.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/enrollment-token-invalid/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 401 · Retryable: no\n\nAn enrollment token works once, for 24 hours, for the pool it was created for.\n\nMessage\n\nthe enrollment token is invalid, expired, revoked or already used\n\nWhat to do\n\ncreate a new token with nodus create enrollmenttoken --pool {pool} and run the installer it prints\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: EnrollmentTokenInvalid and code: 401 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.EnrollmentTokenInvalid ."},{"id":"docs/reference/errors/expired","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/expired/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/expired.md","title":"Expired","description":"The watch position or list continue token is older than the 24-hour change history.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"Expired\n\n The watch position or list continue token is older than the 24-hour change history.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/expired/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 410 · Retryable: no\n\nThe watch position or list continue token is older than the 24-hour change history.\n\nMessage\n\nresourceVersion {resourceVersion} is older than the retained history\n\nWhat to do\n\nlist again and watch from the new resourceVersion\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: Expired and code: 410 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.Expired ."},{"id":"docs/reference/errors/field-immutable","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/field-immutable/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/field-immutable.md","title":"FieldImmutable","description":"The update changes a spec field that is immutable after create; the causes list each changed field.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"FieldImmutable\n\n The update changes a spec field that is immutable after create; the causes list each changed field.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/field-immutable/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nThe update changes a spec field that is immutable after create; the causes list each changed field.\n\nMessage\n\n{resource} \"{name}\": {fields} cannot be changed after create\n\nWhat to do\n\ncreate a new object with the new values, or revert those fields\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: FieldImmutable and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.FieldImmutable .\n\n details also carries:\n\n diff , which the Python SDK exposes as diff"},{"id":"docs/reference/errors/file-too-large","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/file-too-large/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/file-too-large.md","title":"FileTooLarge","description":"A file request carries at most 64 MiB; larger data does not go through the files API.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"FileTooLarge\n\n A file request carries at most 64 MiB; larger data does not go through the files API.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/file-too-large/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 413 · Retryable: no\n\nA file request carries at most 64 MiB; larger data does not go through the files API.\n\nMessage\n\n{path} is larger than {limit}\n\nWhat to do\n\nput large data on a Volume ( nodus volume put ) and mount it in the Sandbox\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: FileTooLarge and code: 413 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.FileTooLarge ."},{"id":"docs/reference/errors/forbidden","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/forbidden/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/forbidden.md","title":"Forbidden","description":"The credential is valid but lacks the scope or project access the operation needs.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"Forbidden\n\n The credential is valid but lacks the scope or project access the operation needs.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/forbidden/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 403 · Retryable: no\n\nThe credential is valid but lacks the scope or project access the operation needs.\n\nMessage\n\n{principal} cannot {verb} {resource}{where}: missing scope {scope}\n\nWhat to do\n\nask an org admin for the {scope} scope, or use a key that has it ( nodus auth can-i {verb} {resource} )\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: Forbidden and code: 403 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.Forbidden .\n\n details also carries:\n\n missingScope , which the Python SDK exposes as missing scope"},{"id":"docs/reference/errors/function-call-batch-too-large","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/function-call-batch-too-large/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/function-call-batch-too-large.md","title":"FunctionCallBatchTooLarge","description":"A batch create takes at most 1,000 calls and 8 MiB.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"FunctionCallBatchTooLarge\n\n A batch create takes at most 1,000 calls and 8 MiB.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/function-call-batch-too-large/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 413 · Retryable: no\n\nA batch create takes at most 1,000 calls and 8 MiB.\n\nMessage\n\na FunctionCallList holds at most {max} calls, got {count}\n\nWhat to do\n\nsend .map() inputs in batches of at most 1,000 calls; the SDK does this for you\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: FunctionCallBatchTooLarge and code: 413 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.FunctionCallBatchTooLarge ."},{"id":"docs/reference/errors/function-not-found","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/function-not-found/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/function-not-found.md","title":"FunctionNotFound","description":"The Function does not exist in the project, or it was deleted since the call was prepared.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"FunctionNotFound\n\n The Function does not exist in the project, or it was deleted since the call was prepared.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/function-not-found/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 404 · Retryable: no\n\nThe Function does not exist in the project, or it was deleted since the call was prepared.\n\nMessage\n\nproject \"{project}\" has no function \"{name}\"\n\nWhat to do\n\nlist the project’s Functions with nodus get functions , or deploy the App with nodus deploy app.py \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: FunctionNotFound and code: 404 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.FunctionNotFound ."},{"id":"docs/reference/errors/idempotency-key-reused","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/idempotency-key-reused/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/idempotency-key-reused.md","title":"IdempotencyKeyReused","description":"The Idempotency-Key was used in the last 24 hours with a different request body.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"IdempotencyKeyReused\n\n The Idempotency-Key was used in the last 24 hours with a different request body.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/idempotency-key-reused/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nThe Idempotency-Key was used in the last 24 hours with a different request body.\n\nMessage\n\nIdempotency-Key {key} was already used for a different request\n\nWhat to do\n\nsend a new Idempotency-Key for a different request\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: IdempotencyKeyReused and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.IdempotencyKeyReused ."},{"id":"docs/reference/errors/image-build-unavailable","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/image-build-unavailable/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/image-build-unavailable.md","title":"ImageBuildUnavailable","description":"This deployment cannot build image steps, Dockerfiles or Sandbox snapshots; prebuilt registry images remain supported.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"ImageBuildUnavailable\n\n This deployment cannot build image steps, Dockerfiles or Sandbox snapshots; prebuilt registry images remain supported.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/image-build-unavailable/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 503 · Retryable: no\n\nThis deployment cannot build image steps, Dockerfiles or Sandbox snapshots; prebuilt registry images remain supported.\n\nMessage\n\nimage builds are unavailable on this deployment\n\nWhat to do\n\npublish the image to a registry and use image or a base-only Image without steps; for an Environment use package.image instead of package.pip\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: ImageBuildUnavailable and code: 503 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.ImageBuildUnavailable ."},{"id":"docs/reference/errors/image-not-found","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/image-not-found/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/image-not-found.md","title":"ImageNotFound","description":"The tag or digest does not exist in the registry, or the pull credentials cannot see it.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"ImageNotFound\n\n The tag or digest does not exist in the registry, or the pull credentials cannot see it.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/image-not-found/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 422 · Retryable: no\n\nThe tag or digest does not exist in the registry, or the pull credentials cannot see it.\n\nMessage\n\nimage \"{image}\" was not found\n\nWhat to do\n\ncheck the reference, or add a Registry Secret under imagePullSecrets\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: ImageNotFound and code: 422 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.ImageNotFound ."},{"id":"docs/reference/errors/image-not-ready","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/image-not-ready/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/image-not-ready.md","title":"ImageNotReady","description":"An imageRef runs the digest the Image built; the Image has none yet.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"ImageNotReady\n\n An imageRef runs the digest the Image built; the Image has none yet.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/image-not-ready/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: yes\n\nAn imageRef runs the digest the Image built; the Image has none yet.\n\nMessage\n\nimage \"{name}\" is {phase}\n\nWhat to do\n\nwait with nodus wait image/{name} --for=condition=Ready , then retry\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: ImageNotReady and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.ImageNotReady ."},{"id":"docs/reference/errors/image-pull-failed","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/image-pull-failed/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/image-pull-failed.md","title":"ImagePullFailed","description":"The registry refused the credentials or was unavailable while pinning the image to a digest.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"ImagePullFailed\n\n The registry refused the credentials or was unavailable while pinning the image to a digest.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/image-pull-failed/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 422 · Retryable: yes\n\nThe registry refused the credentials or was unavailable while pinning the image to a digest.\n\nMessage\n\nimage \"{image}\" could not be resolved: {reason}\n\nWhat to do\n\ncheck the imagePullSecrets, or retry\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: ImagePullFailed and code: 422 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.ImagePullFailed ."},{"id":"docs/reference/errors/insufficient-credits","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/insufficient-credits/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/insufficient-credits.md","title":"InsufficientCredits","description":"The org's available credits do not cover the hold that the action reserves before it starts.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"InsufficientCredits\n\n The org's available credits do not cover the hold that the action reserves before it starts.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/insufficient-credits/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 402 · Retryable: no\n\nThe org’s available credits do not cover the hold that the action reserves before it starts.\n\nMessage\n\n{subject} needs a {need} hold to start (released when it ends); available {available}\n\nWhat to do\n\nadd credits with nodus billing top-up {topUp} , or lower spec.maxCostUSD\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: InsufficientCredits and code: 402 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.InsufficientCredits .\n\n details also carries:\n\n availableUSD , which the Python SDK exposes as available usd \n neededUSD , which the Python SDK exposes as needed usd"},{"id":"docs/reference/errors/internal-error","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/internal-error/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/internal-error.md","title":"InternalError","description":"The server hit an unexpected failure; the request id finds it in the logs.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"InternalError\n\n The server hit an unexpected failure; the request id finds it in the logs.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/internal-error/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 500 · Retryable: yes\n\nThe server hit an unexpected failure; the request id finds it in the logs.\n\nMessage\n\ninternal error\n\nWhat to do\n\nretry; if it persists, contact support with the requestId\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: InternalError and code: 500 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.InternalError ."},{"id":"docs/reference/errors/invalid","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/invalid/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/invalid.md","title":"Invalid","description":"The object fails strict decoding or validation; `details.causes` lists every field and reason.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"Invalid\n\n The object fails strict decoding or validation; details.causes lists every field and reason.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/invalid/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 422 · Retryable: no\n\nThe object fails strict decoding or validation; details.causes lists every field and reason.\n\nMessage\n\n{object} is invalid: {errors}\n\nWhat to do\n\nfix the listed fields; nodus explain . documents each one\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: Invalid and code: 422 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.Invalid ."},{"id":"docs/reference/errors/method-not-allowed","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/method-not-allowed/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/method-not-allowed.md","title":"MethodNotAllowed","description":"The resource does not support this verb (for example, a read-only kind).","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"MethodNotAllowed\n\n The resource does not support this verb (for example, a read-only kind).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/method-not-allowed/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 405 · Retryable: no\n\nThe resource does not support this verb (for example, a read-only kind).\n\nMessage\n\n{verb} is not allowed on {resource}\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: MethodNotAllowed and code: 405 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.MethodNotAllowed ."},{"id":"docs/reference/errors/no-ssh-key","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/no-ssh-key/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/no-ssh-key.md","title":"NoSSHKey","description":"The Workspace accepts only the SSHKeys named in spec.sshKeys, or every project member's keys when it names none.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"NoSSHKey\n\n The Workspace accepts only the SSHKeys named in spec.sshKeys, or every project member's keys when it names none.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/no-ssh-key/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 403 · Retryable: no\n\nThe Workspace accepts only the SSHKeys named in spec.sshKeys, or every project member’s keys when it names none.\n\nMessage\n\nno SSH key of yours may log in to workspace \"{name}\"\n\nWhat to do\n\nadd a key with nodus create sshkey --from-file ~/.ssh/id ed25519.pub , or ask the owner to list yours in spec.sshKeys\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: NoSSHKey and code: 403 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.NoSSHKey ."},{"id":"docs/reference/errors/not-found","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/not-found/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/not-found.md","title":"NotFound","description":"The object, project or path does not exist in this org.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"NotFound\n\n The object, project or path does not exist in this org.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/not-found/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 404 · Retryable: no\n\nThe object, project or path does not exist in this org.\n\nMessage\n\n{resource} \"{name}\" not found\n\nWhat to do\n\ncheck the name and project ( nodus get {resource} -p )\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: NotFound and code: 404 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.NotFound ."},{"id":"docs/reference/errors/output-index-required","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/output-index-required/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/output-index-required.md","title":"OutputIndexRequired","description":"Several indexes of an Indexed Job committed the output; the download must pick one.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"OutputIndexRequired\n\n Several indexes of an Indexed Job committed the output; the download must pick one.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/output-index-required/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 400 · Retryable: no\n\nSeveral indexes of an Indexed Job committed the output; the download must pick one.\n\nMessage\n\noutput \"{output}\" of job \"{name}\" was produced by {count} indexes\n\nWhat to do\n\npass ?index= &#x20;( nodus cp job/{name}:{output} . --index ) \n\n ## On the wireThe response is a Kubernetes Status with reason: OutputIndexRequired and code: 400 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.OutputIndexRequired ."},{"id":"docs/reference/errors/output-not-found","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/output-not-found/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/output-not-found.md","title":"OutputNotFound","description":"The Job did not commit an output by that name, or not at that index.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"OutputNotFound\n\n The Job did not commit an output by that name, or not at that index.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/output-not-found/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 404 · Retryable: no\n\nThe Job did not commit an output by that name, or not at that index.\n\nMessage\n\njob \"{name}\" has no committed output \"{output}\"{where}\n\nWhat to do\n\nlist the committed outputs with nodus get job/{name} -o jsonpath='{.status.outputs}' \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: OutputNotFound and code: 404 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.OutputNotFound ."},{"id":"docs/reference/errors/payment-disputed","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/payment-disputed/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/payment-disputed.md","title":"PaymentDisputed","description":"A card payment on the org is disputed; new work cannot start until the dispute closes.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PaymentDisputed\n\n A card payment on the org is disputed; new work cannot start until the dispute closes.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/payment-disputed/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 402 · Retryable: no\n\nA card payment on the org is disputed; new work cannot start until the dispute closes.\n\nMessage\n\na payment on this org is disputed; new work is paused until the dispute closes\n\nWhat to do\n\ncontact support to resolve the dispute\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PaymentDisputed and code: 402 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PaymentDisputed ."},{"id":"docs/reference/errors/payment-method-required","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/payment-method-required/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/payment-method-required.md","title":"PaymentMethodRequired","description":"New work requires a verified card, including work funded by promotional credit. Adding a card does not charge it.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PaymentMethodRequired\n\n New work requires a verified card, including work funded by promotional credit. Adding a card does not charge it.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/payment-method-required/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 402 · Retryable: no\n\nNew work requires a verified card, including work funded by promotional credit. Adding a card does not charge it.\n\nMessage\n\nadd a valid payment method before starting work\n\nWhat to do\n\nopen Billing and choose Add a card, or run nodus billing portal --setup ; if you just added it, wait for verification and retry\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PaymentMethodRequired and code: 402 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PaymentMethodRequired ."},{"id":"docs/reference/errors/payment-verification-unavailable","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/payment-verification-unavailable/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/payment-verification-unavailable.md","title":"PaymentVerificationUnavailable","description":"Stripe could not confirm the payment method. Existing running work retains its funded holds.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PaymentVerificationUnavailable\n\n Stripe could not confirm the payment method. Existing running work retains its funded holds.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/payment-verification-unavailable/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 503 · Retryable: yes\n\nStripe could not confirm the payment method. Existing running work retains its funded holds.\n\nMessage\n\npayment method verification is temporarily unavailable\n\nWhat to do\n\nretry shortly; contact support if verification remains unavailable\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PaymentVerificationUnavailable and code: 503 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PaymentVerificationUnavailable ."},{"id":"docs/reference/errors/payments-not-configured","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/payments-not-configured/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/payments-not-configured.md","title":"PaymentsNotConfigured","description":"The deployment holds no Stripe key, so it takes no card payments; balances, grants and usage still work.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PaymentsNotConfigured\n\n The deployment holds no Stripe key, so it takes no card payments; balances, grants and usage still work.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/payments-not-configured/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 503 · Retryable: no\n\nThe deployment holds no Stripe key, so it takes no card payments; balances, grants and usage still work.\n\nMessage\n\npayments are not configured on this deployment\n\nWhat to do\n\ncontact support to restore payment setup and verification\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PaymentsNotConfigured and code: 503 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PaymentsNotConfigured ."},{"id":"docs/reference/errors/pool-act-paused","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/pool-act-paused/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/pool-act-paused.md","title":"PoolActPaused","description":"The pool's kill switch stops every action until it is resumed.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PoolActPaused\n\n The pool's kill switch stops every action until it is resumed.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/pool-act-paused/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nThe pool’s kill switch stops every action until it is resumed.\n\nMessage\n\nactions on pool \"{pool}\" are paused\n\nWhat to do\n\nresume actions with nodus pool resume {pool} \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PoolActPaused and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PoolActPaused ."},{"id":"docs/reference/errors/pool-action-not-pending","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/pool-action-not-pending/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/pool-action-not-pending.md","title":"PoolActionNotPending","description":"Only a queued, undecided action waits for a decision; it may have expired or been decided by someone else.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PoolActionNotPending\n\n Only a queued, undecided action waits for a decision; it may have expired or been decided by someone else.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/pool-action-not-pending/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nOnly a queued, undecided action waits for a decision; it may have expired or been decided by someone else.\n\nMessage\n\npool action \"{name}\" is {phase} and no longer waits for a decision\n\nWhat to do\n\ncheck it with nodus get poolaction/{name} \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PoolActionNotPending and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PoolActionNotPending ."},{"id":"docs/reference/errors/pool-action-not-reversible","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/pool-action-not-reversible/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/pool-action-not-reversible.md","title":"PoolActionNotReversible","description":"Only an executed drain, undrain or routing change can be reverted, once.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PoolActionNotReversible\n\n Only an executed drain, undrain or routing change can be reverted, once.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/pool-action-not-reversible/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nOnly an executed drain, undrain or routing change can be reverted, once.\n\nMessage\n\npool action \"{name}\" cannot be reverted: {reason}\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PoolActionNotReversible and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PoolActionNotReversible ."},{"id":"docs/reference/errors/pool-in-use","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/pool-in-use/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/pool-in-use.md","title":"PoolInUse","description":"A Pool is deleted only once no Node is enrolled in it and no object names it in spec.placement.pool.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PoolInUse\n\n A Pool is deleted only once no Node is enrolled in it and no object names it in spec.placement.pool.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/pool-in-use/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nA Pool is deleted only once no Node is enrolled in it and no object names it in spec.placement.pool.\n\nMessage\n\npool \"{pool}\" is in use by {users}\n\nWhat to do\n\ndelete its nodes with nodus delete node and move or delete the objects that place on it\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PoolInUse and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PoolInUse ."},{"id":"docs/reference/errors/precondition-failed","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/precondition-failed/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/precondition-failed.md","title":"PreconditionFailed","description":"The If-Match value no longer matches the compiled spec and price book of a fresh dry-run.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PreconditionFailed\n\n The If-Match value no longer matches the compiled spec and price book of a fresh dry-run.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/precondition-failed/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 412 · Retryable: no\n\nThe If-Match value no longer matches the compiled spec and price book of a fresh dry-run.\n\nMessage\n\nthe object or its price changed since the dry-run it was reviewed with{reason}\n\nWhat to do\n\nrun the dry-run again ( --dry-run=server ), review the new estimate and create with its ETag\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PreconditionFailed and code: 412 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PreconditionFailed ."},{"id":"docs/reference/errors/price-consent-required","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/price-consent-required/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/price-consent-required.md","title":"PriceConsentRequired","description":"Pool routing and Predict are billed, so turning either on records consent to its current price.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PriceConsentRequired\n\n Pool routing and Predict are billed, so turning either on records consent to its current price.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/price-consent-required/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 422 · Retryable: no\n\nPool routing and Predict are billed, so turning either on records consent to its current price.\n\nMessage\n\n{feature} costs ${price}; agree to the price to turn it on\n\nWhat to do\n\nset the annotation {annotation}: “{price}” on the Pool, or agree in the console\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PriceConsentRequired and code: 422 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PriceConsentRequired .\n\n details also carries:\n\n annotation , which the Python SDK exposes as annotation \n feature , which the Python SDK exposes as feature \n price , which the Python SDK exposes as price"},{"id":"docs/reference/errors/promo-code-expired","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/promo-code-expired/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/promo-code-expired.md","title":"PromoCodeExpired","description":"The promo code exists but can no longer be redeemed.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PromoCodeExpired\n\n The promo code exists but can no longer be redeemed.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/promo-code-expired/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 410 · Retryable: no\n\nThe promo code exists but can no longer be redeemed.\n\nMessage\n\npromo code {code} has expired\n\nWhat to do\n\nask whoever gave you the code for a current one, or add credits with nodus billing top-up \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PromoCodeExpired and code: 410 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PromoCodeExpired ."},{"id":"docs/reference/errors/promo-code-used","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/promo-code-used/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/promo-code-used.md","title":"PromoCodeUsed","description":"The org has already redeemed the code as often as it may, or the code has no redemptions left.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"PromoCodeUsed\n\n The org has already redeemed the code as often as it may, or the code has no redemptions left.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/promo-code-used/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nThe org has already redeemed the code as often as it may, or the code has no redemptions left.\n\nMessage\n\npromo code {code} {why}\n\nWhat to do\n\nsee the credit it added with nodus get creditgrants \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: PromoCodeUsed and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.PromoCodeUsed ."},{"id":"docs/reference/errors/quota-exceeded","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/quota-exceeded/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/quota-exceeded.md","title":"QuotaExceeded","description":"The org or project already uses everything its Quota allows of this counter.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"QuotaExceeded\n\n The org or project already uses everything its Quota allows of this counter.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/quota-exceeded/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 429 · Retryable: no\n\nThe org or project already uses everything its Quota allows of this counter.\n\nMessage\n\nquota {quota} exceeded\n\nWhat to do\n\ndelete objects you no longer need, or ask for a higher limit ( nodus get quota default )\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: QuotaExceeded and code: 429 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.QuotaExceeded .\n\n details also carries:\n\n quota , which the Python SDK exposes as quota"},{"id":"docs/reference/errors/request-entity-too-large","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/request-entity-too-large/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/request-entity-too-large.md","title":"RequestEntityTooLarge","description":"The request body is larger than the route accepts (1 MiB for JSON).","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"RequestEntityTooLarge\n\n The request body is larger than the route accepts (1 MiB for JSON).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/request-entity-too-large/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 413 · Retryable: no\n\nThe request body is larger than the route accepts (1 MiB for JSON).\n\nMessage\n\nthe request body exceeds {limit}\n\nWhat to do\n\nupload large content as a blob ( /blobs/v1 ) and reference it by sha256\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: RequestEntityTooLarge and code: 413 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.RequestEntityTooLarge ."},{"id":"docs/reference/errors/request-in-progress","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/request-in-progress/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/request-in-progress.md","title":"RequestInProgress","description":"The original request with this Idempotency-Key has not finished yet.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"RequestInProgress\n\n The original request with this Idempotency-Key has not finished yet.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/request-in-progress/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: yes\n\nThe original request with this Idempotency-Key has not finished yet.\n\nMessage\n\na request with Idempotency-Key {key} is still in progress\n\nWhat to do\n\nretry after the Retry-After interval with the same key\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: RequestInProgress and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.RequestInProgress ."},{"id":"docs/reference/errors/sandbox-failed","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/sandbox-failed/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/sandbox-failed.md","title":"SandboxFailed","description":"The Sandbox is in the Failed phase, which is final: status.reason says why.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"SandboxFailed\n\n The Sandbox is in the Failed phase, which is final: status.reason says why.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/sandbox-failed/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nThe Sandbox is in the Failed phase, which is final: status.reason says why.\n\nMessage\n\nsandbox \"{name}\" failed: {cause}\n\nWhat to do\n\nread the cause with nodus describe sandbox/{name} , then delete the Sandbox and create it again\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: SandboxFailed and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.SandboxFailed ."},{"id":"docs/reference/errors/sandbox-not-running","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/sandbox-not-running/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/sandbox-not-running.md","title":"SandboxNotRunning","description":"The Sandbox is being deleted, so no command, file or port request can reach it and none starts it.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"SandboxNotRunning\n\n The Sandbox is being deleted, so no command, file or port request can reach it and none starts it.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/sandbox-not-running/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nThe Sandbox is being deleted, so no command, file or port request can reach it and none starts it.\n\nMessage\n\nsandbox \"{name}\" is {state} and takes no more requests\n\nWhat to do\n\ncreate a new Sandbox once the old one is gone ( nodus get sandbox/{name} shows when)\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: SandboxNotRunning and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.SandboxNotRunning ."},{"id":"docs/reference/errors/sandbox-starting","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/sandbox-starting/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/sandbox-starting.md","title":"SandboxStarting","description":"The request needs the Sandbox's container, which is not running yet. A stopped Sandbox is started by the request itself; the answer carries Retry-After.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"SandboxStarting\n\n The request needs the Sandbox's container, which is not running yet. A stopped Sandbox is started by the request itself; the answer carries Retry-After.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/sandbox-starting/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 503 · Retryable: yes\n\nThe request needs the Sandbox’s container, which is not running yet. A stopped Sandbox is started by the request itself; the answer carries Retry-After.\n\nMessage\n\nsandbox \"{name}\" is starting\n\nWhat to do\n\nretry after the Retry-After interval, or run nodus wait sandbox/{name} --for=condition=Ready \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: SandboxStarting and code: 503 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.SandboxStarting ."},{"id":"docs/reference/errors/secret-value-in-env","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/secret-value-in-env/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/secret-value-in-env.md","title":"SecretValueInEnv","description":"Secret values reach a Job or Sandbox only through secrets or valueFrom.secretKeyRef, never as literals.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"SecretValueInEnv\n\n Secret values reach a Job or Sandbox only through secrets or valueFrom.secretKeyRef, never as literals.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/secret-value-in-env/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 422 · Retryable: no\n\nSecret values reach a Job or Sandbox only through secrets or valueFrom.secretKeyRef, never as literals.\n\nMessage\n\nenv contains the value of a Secret{where}\n\nWhat to do\n\nreference the Secret with valueFrom.secretKeyRef or list it under secrets\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: SecretValueInEnv and code: 422 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.SecretValueInEnv ."},{"id":"docs/reference/errors/sha256-mismatch","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/sha256-mismatch/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/sha256-mismatch.md","title":"SHA256Mismatch","description":"A write with expectedSHA256 found different content than the writer started from, so nothing was written.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"SHA256Mismatch\n\n A write with expectedSHA256 found different content than the writer started from, so nothing was written.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/sha256-mismatch/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 412 · Retryable: no\n\nA write with expectedSHA256 found different content than the writer started from, so nothing was written.\n\nMessage\n\nthe file {path} changed: its SHA-256 does not match expectedSHA256\n\nWhat to do\n\nread the file again, merge your change, and write it with the new digest as expectedSHA256\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: SHA256Mismatch and code: 412 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.SHA256Mismatch ."},{"id":"docs/reference/errors/stdin-backpressure","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/stdin-backpressure/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/stdin-backpressure.md","title":"StdinBackpressure","description":"The process's stdin buffer is full, so the write was refused without losing earlier input.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"StdinBackpressure\n\n The process's stdin buffer is full, so the write was refused without losing earlier input.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/stdin-backpressure/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 429 · Retryable: yes\n\nThe process’s stdin buffer is full, so the write was refused without losing earlier input.\n\nMessage\n\nthe process is not reading stdin as fast as it is sent\n\nWhat to do\n\nretry after a short backoff; the SDK and CLI do this automatically\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: StdinBackpressure and code: 429 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.StdinBackpressure ."},{"id":"docs/reference/errors/too-many-requests","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/too-many-requests/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/too-many-requests.md","title":"TooManyRequests","description":"The credential sent too many requests of this class; the RateLimit headers show the budget.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"TooManyRequests\n\n The credential sent too many requests of this class; the RateLimit headers show the budget.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/too-many-requests/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 429 · Retryable: yes\n\nThe credential sent too many requests of this class; the RateLimit headers show the budget.\n\nMessage\n\nrate limit exceeded for {class} requests\n\nWhat to do\n\nretry after the Retry-After interval\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: TooManyRequests and code: 429 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.TooManyRequests ."},{"id":"docs/reference/errors/top-up-limit-reached","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/top-up-limit-reached/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/top-up-limit-reached.md","title":"TopUpLimitReached","description":"An org buys at most $5,000 of credit through Checkout per UTC day.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"TopUpLimitReached\n\n An org buys at most $5,000 of credit through Checkout per UTC day.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/top-up-limit-reached/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 429 · Retryable: no\n\nAn org buys at most $5,000 of credit through Checkout per UTC day.\n\nMessage\n\ntop-ups of this org would exceed {limit} today (UTC)\n\nWhat to do\n\nbuy the rest after midnight UTC, or contact support for a larger purchase\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: TopUpLimitReached and code: 429 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.TopUpLimitReached ."},{"id":"docs/reference/errors/unauthorized","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/unauthorized/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/unauthorized.md","title":"Unauthorized","description":"The request carries no valid credential, or the credential expired.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"Unauthorized\n\n The request carries no valid credential, or the credential expired.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/unauthorized/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 401 · Retryable: no\n\nThe request carries no valid credential, or the credential expired.\n\nMessage\n\nauthentication required{reason}\n\nWhat to do\n\nrun nodus login , or send a valid API key as Authorization: Bearer \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: Unauthorized and code: 401 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.Unauthorized ."},{"id":"docs/reference/errors/unavailable","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/unavailable/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/unavailable.md","title":"Unavailable","description":"A dependency of the API is temporarily unavailable.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"Unavailable\n\n A dependency of the API is temporarily unavailable.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/unavailable/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 503 · Retryable: yes\n\nA dependency of the API is temporarily unavailable.\n\nMessage\n\nthe service is temporarily unavailable{reason}\n\nWhat to do\n\nretry with backoff\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: Unavailable and code: 503 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.Unavailable ."},{"id":"docs/reference/errors/unsupported","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/unsupported/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/unsupported.md","title":"Unsupported","description":"The request is well formed but asks for a feature this API does not offer.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"Unsupported\n\n The request is well formed but asks for a feature this API does not offer.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/unsupported/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 422 · Retryable: no\n\nThe request is well formed but asks for a feature this API does not offer.\n\nMessage\n\n{what} is not supported\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: Unsupported and code: 422 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.Unsupported ."},{"id":"docs/reference/errors/unsupported-media-type","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/unsupported-media-type/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/unsupported-media-type.md","title":"UnsupportedMediaType","description":"The request body's Content-Type is not accepted; strategic merge patch is not supported.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"UnsupportedMediaType\n\n The request body's Content-Type is not accepted; strategic merge patch is not supported.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/unsupported-media-type/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 415 · Retryable: no\n\nThe request body’s Content-Type is not accepted; strategic merge patch is not supported.\n\nMessage\n\ncontent type {contentType} is not supported{hint}\n\nWhat to do\n\nsend application/json, application/merge-patch+json, application/json-patch+json or application/apply-patch+yaml\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: UnsupportedMediaType and code: 415 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.UnsupportedMediaType ."},{"id":"docs/reference/errors/volume-busy","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/volume-busy/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/volume-busy.md","title":"VolumeBusy","description":"A ReadWriteOnce Volume has one writer; another attempt holds its lease.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"VolumeBusy\n\n A ReadWriteOnce Volume has one writer; another attempt holds its lease.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/volume-busy/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: yes\n\nA ReadWriteOnce Volume has one writer; another attempt holds its lease.\n\nMessage\n\nvolume \"{name}\" is held by {holder}\n\nWhat to do\n\nwait for {holder} to finish, or mount the Volume with readOnly: true\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: VolumeBusy and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.VolumeBusy .\n\n details also carries:\n\n holder , which the Python SDK exposes as holder"},{"id":"docs/reference/errors/worker-not-allowed","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/worker-not-allowed/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/worker-not-allowed.md","title":"WorkerNotAllowed","description":"Only the Function's own worker attempts, through their ServiceAccount token, can claim its calls.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"WorkerNotAllowed\n\n Only the Function's own worker attempts, through their ServiceAccount token, can claim its calls.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/worker-not-allowed/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 403 · Retryable: no\n\nOnly the Function’s own worker attempts, through their ServiceAccount token, can claim its calls.\n\nMessage\n\nthis credential is not a worker of function \"{name}\"\n\nWhat to do\n\ncall the Function with .remote() , .spawn() or .map() instead of claiming calls yourself\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: WorkerNotAllowed and code: 403 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.WorkerNotAllowed ."},{"id":"docs/reference/errors/workspace-starting","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/workspace-starting/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/workspace-starting.md","title":"WorkspaceStarting","description":"The Workspace has no running session yet: it is starting, or the request woke it from a stop.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"WorkspaceStarting\n\n The Workspace has no running session yet: it is starting, or the request woke it from a stop.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/workspace-starting/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 503 · Retryable: yes\n\nThe Workspace has no running session yet: it is starting, or the request woke it from a stop.\n\nMessage\n\nworkspace \"{name}\" is starting; its saved files are being restored\n\nWhat to do\n\nretry after the Retry-After interval, or run nodus wait workspace/{name} --for=condition=Ready \n\nOn the wire\n\nThe response is a Kubernetes Status with reason: WorkspaceStarting and code: 503 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.WorkspaceStarting ."},{"id":"docs/reference/errors/workspace-tool-unavailable","url":"https://nodus-platform-site.pages.dev/docs/reference/errors/workspace-tool-unavailable/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/errors/workspace-tool-unavailable.md","title":"WorkspaceToolUnavailable","description":"A browser tool has a preview URL only while the Workspace is running and lists the tool in spec.tools.","stage":"GA","headings":[{"depth":2,"slug":"message","text":"Message"},{"depth":2,"slug":"what-to-do","text":"What to do"},{"depth":2,"slug":"on-the-wire","text":"On the wire"}],"text":"WorkspaceToolUnavailable\n\n A browser tool has a preview URL only while the Workspace is running and lists the tool in spec.tools.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/errors/workspace-tool-unavailable/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n HTTP status: 409 · Retryable: no\n\nA browser tool has a preview URL only while the Workspace is running and lists the tool in spec.tools.\n\nMessage\n\nworkspace \"{name}\" serves no {tool} right now\n\nWhat to do\n\nrun nodus start workspace/{name} and wait for Ready, or add the tool to spec.tools\n\nOn the wire\n\nThe response is a Kubernetes Status with reason: WorkspaceToolUnavailable and code: 409 , plus fix , docs (this page) and requestId . details.causes lists the fields involved, when there are any. The CLI prints the fix and the request id; the Python SDK raises nodus.errors.WorkspaceToolUnavailable ."},{"id":"docs/reference/pricing","url":"https://nodus-platform-site.pages.dev/docs/reference/pricing/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/pricing.md","title":"Pricing","description":"Every published Nodus rate, from GPU and CPU list prices to Nodus nodes, storage, egress and models.","stage":"GA","headings":[{"depth":2,"slug":"how-you-are-billed","text":"How you are billed"},{"depth":2,"slug":"gpus","text":"GPUs"},{"depth":2,"slug":"cpu-machines","text":"CPU machines"},{"depth":2,"slug":"nodus-nodes","text":"Nodus nodes"},{"depth":2,"slug":"storage-and-egress","text":"Storage and egress"},{"depth":2,"slug":"bring-your-own-machines","text":"Bring your own machines"},{"depth":2,"slug":"models","text":"Models"},{"depth":2,"slug":"plans","text":"Plans"},{"depth":2,"slug":"what-you-pay-for","text":"What you pay for"},{"depth":2,"slug":"worked-examples","text":"Worked examples"},{"depth":2,"slug":"starter-grant","text":"Starter grant"}],"text":"Pricing\n\n Every published Nodus rate, from GPU and CPU list prices to Nodus nodes, storage, egress and models.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/pricing/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nPricebook version 2026.10.8 , effective 2026-10-01. All prices are in US dollars.\n\nHow you are billed\n\nNever above list; your rate is shown before launch and frozen for the run. You pay what the provider bills for your machine, from the moment it is created until it is deleted, at provider cost ÷ 0.875 (Nodus keeps 12.5 % of what you pay). Each charge is itemized as Boot , Running , Restore or Teardown . Nodus-operated capacity bills at the node rates below. Credit is prepaid and every paid action reserves a hold first.\n\nGPUs\n\nList price per hour for the whole machine, by GPU count.\n\n GPU Memory ×1 ×2 ×4 ×8 \n - - - - - - \n A10 24 GB $0.89 $1.78 $3.56 $7.12 \n A100 40G 40 GB $1.49 $2.98 $5.96 $11.92 \n A100 40G PCIE 40 GB $1.39 $2.78 $5.56 $11.12 \n A100 80G 80 GB $1.99 $3.98 $7.96 $15.92 \n A100 80G PCIE 80 GB $1.89 $3.78 $7.56 $15.12 \n B200 180 GB $6.49 $12.98 $25.96 $51.92 \n H100 PCIE 80 GB $2.99 $5.98 $11.96 $23.92 \n H100 SXM 80 GB $3.29 $6.58 $13.16 $26.32 \n H200 141 GB $4.29 $8.58 $17.16 $34.32 \n L4 24 GB $0.89 $1.78 $3.56 $7.12 \n L40S 48 GB $1.29 $2.58 $5.16 $10.32 \n RTX 3090 24 GB $0.59 $1.18 $2.36 $4.72 \n RTX 4090 24 GB $0.69 $1.38 $2.76 $5.52 \n RTX 6000 ADA 48 GB $1.09 $2.18 $4.36 $8.72 \n RTX A6000 48 GB $0.79 $1.58 $3.16 $6.32 \n\nCPU machines\n\n Shape vCPU Memory Disk Per hour \n - - - - - \n cpu-2x4 2 4 GiB — $0.16 \n cpu-2x8 2 8 GiB — $0.12 \n cpu-4x8 4 8 GiB — $0.31 \n cpu-4x16 4 16 GiB — $0.23 \n cpu-8x16 8 16 GiB — $0.62 \n cpu-8x32 8 32 GiB — $0.45 \n cpu-16x32 16 32 GiB — $0.79 \n cpu-16x64 16 64 GiB — $0.89 \n cpu-32x64 32 64 GiB — $1.58 \n cpu-32x128 32 128 GiB — $1.78 \n cpu-64x256 64 256 GiB — $3.56 \n\nNodus nodes\n\nSandboxes, CPU Functions, CPU Workspaces and warm pools run on Nodus-operated nodes, billed per second from these rates.\n\n Resource Per hour \n - - \n vCPU $0.048 \n GiB of memory $0.006 \n GiB of disk above 10 GiB per vCPU $0.00014 \n\n Example Per hour \n - - \n Smallest Sandbox $0.015 \n 2 vCPU / 4 GiB Sandbox $0.12 \n 8 vCPU / 32 GiB CPU Job $0.576 \n\nStorage and egress\n\n Meter Included Price above the inclusion \n - - - \n Storage 10 GB per org $0.017143 per GB-month \n Egress 10 GiB per org per day $0.09 per GiB \n Relayed gang traffic — $0.011429 per GiB \n\nBring your own machines\n\n Item Price \n - - \n Route: Nodus-scheduled work on your pool GPUs $0.02 per device-hour \n Predict $99.00 per pool-month \n\nModels\n\nHosted inference is billed per request from its usage.\n\n Model Operation Price \n - - - \n baai/bge-m3 input $0.013685 per 1M tokens \n canopylabs/orpheus-v1-english speech $23.157895 per 1M characters \n deepseek-ai/deepseek-v4-flash input $0.052632 per 1M tokens \n deepseek-ai/deepseek-v4-flash cacheRead $0.046211 per 1M tokens \n deepseek-ai/deepseek-v4-flash output $0.305264 per 1M tokens \n deepseek-ai/deepseek-v4-pro input $0.410527 per 1M tokens \n deepseek-ai/deepseek-v4-pro cacheRead $0.332527 per 1M tokens \n deepseek-ai/deepseek-v4-pro output $3.673685 per 1M tokens \n deepseek-ai/deepseek-v4.1-flash input $0.105264 per 1M tokens \n deepseek-ai/deepseek-v4.1-flash cacheRead $0.047369 per 1M tokens \n deepseek-ai/deepseek-v4.1-flash output $0.463158 per 1M tokens \n moonshotai/kimi-k3 input $1.51579 per 1M tokens \n moonshotai/kimi-k3 cacheRead $0.31579 per 1M tokens \n moonshotai/kimi-k3 output $9.473685 per 1M tokens \n openai/gpt-oss-120b input $0.157895 per 1M tokens \n openai/gpt-oss-120b cacheRead $0.078948 per 1M tokens \n openai/gpt-oss-120b output $0.631579 per 1M tokens \n openai/gpt-oss-20b input $0.052632 per 1M tokens \n openai/gpt-oss-20b cacheRead $0.005264 per 1M tokens \n openai/gpt-oss-20b output $0.210527 per 1M tokens \n openai/whisper-large-v3 audio $0.001948 per audio minute \n openai/whisper-large-v3-turbo audio $0.000702 per audio minute \n qwen/qwen3.8-27b input $0.442106 per 1M tokens \n qwen/qwen3.8-27b cacheRead $0.089474 per 1M tokens \n qwen/qwen3.8-27b output $3.157895 per 1M tokens \n zai-org/glm-5.2 input $0.242106 per 1M tokens \n zai-org/glm-5.2 cacheRead $0.196948 per 1M tokens \n zai-org/glm-5.2 output $4.631579 per 1M tokens \n zai-org/glm-5.3 input $0.231579 per 1M tokens \n zai-org/glm-5.3 cacheRead $0.186843 per 1M tokens \n zai-org/glm-5.3 output $3.568422 per 1M tokens \n zai-org/glm-5.3-flash input $0.205264 per 1M tokens \n zai-org/glm-5.3-flash cacheRead $0.041053 per 1M tokens \n zai-org/glm-5.3-flash output $0.684211 per 1M tokens \n nodus/indra routing input $0.044211 per 1M tokens \n nodus/indra routing output $0.00 per 1M tokens \n\nPlans\n\nA plan’s allowance pays for its models at the prices above; it expires at the end of each month, and usage beyond it draws from your credits.\n\n Plan Per month Allowance Models \n - - - - \n composer $20.00 $20.00 nodus/indra \n\nWhat you pay for\n\n Item Who pays \n - - \n Your rented machine for every second the provider bills it, from creation to confirmed deletion: boot, image pull, restore, running, teardown and the provider’s rounding You, at provider cost ÷ 0.875, never above list \n Sandboxes, Functions, builds and CPU work on Nodus nodes, from placement to release You, at the published node rates \n Model tokens, including the routing call of nodus/auto You, at the cheapest available model cost ÷ 0.95 \n Agent runs on Nodus-managed models You, at model cost ÷ 0.875, from credits \n Storage above the included amount and egress above the daily inclusion You, at the published rates \n Spare machines Nodus starts to finish sooner, and failures Nodus causes Nodus \n Requests whose outcome Nodus cannot confirm Nodus \n Logs, and egress within the daily inclusion Nodus \n\nWorked examples\n\n Example Cost \n - - \n One H100 for one hour, at most $3.29 \n A 2 vCPU / 4 GiB Sandbox for one hour $0.12 \n 1M output tokens on openai/gpt-oss-120b $0.631579 \n\nStarter grant\n\nThe first org a verified user creates receives $30 of credit that expires 30 days after it is granted. Grant credit is spent before purchased credit."},{"id":"docs/reference/python","url":"https://nodus-platform-site.pages.dev/docs/reference/python/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python.md","title":"Python SDK reference","description":"Every public module, class and function of the nodus-compute package, generated from its docstrings.","stage":"GA","headings":[],"text":"Python SDK reference\n\n Every public module, class and function of the nodus-compute package, generated from its docstrings.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nInstall the SDK with pip install nodus-compute and import it as nodus (Python 3.10 or newer). The pages below are generated from the package’s docstrings and type annotations.\n\nnodus Nodus Python SDK (\\ pip install nodus-compute\\ ).\n\nnodus.agent Agents: durable runs on the journal engine (resources.md §4.1 to §4.3, ADR-048, ADR-101).\n\nnodus.api Layer 1 of the SDK: the generic resource client over every kind (resources.md §10.3).\n\nnodus.api.models Pydantic v2 models for every kind, generated from \\ api/openapi\\ by \\ make gen\\ (datamodel-code-generator).\n\nnodus.app \\ App\\ : the group of Functions defined in one Python file (resources.md §3.9, ADR-047).\n\nnodus.checkpoint The checkpoint handshake and gang checkpoints (resources.md §8, distributed-training.md §10.3).\n\nnodus.cluster \\ nodus.cluster.info()\\ : the gang env contract of distributed-training.md §3.6, as one object.\n\nnodus.envs Helpers for building your own Environment and datasets: \\ nodus.envs.split\\ .\n\nnodus.errors Errors raised by the Nodus SDK.\n\nnodus.examples Runnable examples that ship with the SDK: \\ python -m nodus.examples.rl\\ trains a model with RL.\n\nnodus.examples.rl Reinforcement learning in one file: teach a small model to spell words backwards.\n\nnodus.functions Functions, FunctionCalls and \\ @app.cls\\ classes (resources.md §3.10, §3.11, ADR-047, ADR-110).\n\nnodus.image The Modal-shaped \\ Image\\ builder (resources.md §5.3).\n\nnodus.integrations Framework callbacks that report training progress to Nodus from inside a container.\n\nnodus.integrations.lightning \\ NodusCallback\\ for a Lightning \\ Trainer\\ (ADR-103).\n\nnodus.integrations.transformers \\ NodusCallback\\ for the Hugging Face \\ Trainer\\ and every TRL trainer built on it (ADR-103).\n\nnodus.job \\ Job\\ : a run-to-completion container with checkpoints, recovery and outputs (resources.md §3.1).\n\nnodus.llm OpenAI- and Anthropic-compatible inference (resources.md §4.5, §4.6, ADR-049, ADR-100).\n\nnodus.log Structured telemetry from inside a container: metrics, RL task outcomes, work units and phases.\n\nnodus.outputs \\ nodus.outputs.verify()\\ : fail fast inside a Job when its model output would not load (resources.md §10.5).\n\nnodus.process Processes and files inside a running Job, Sandbox or Workspace (resources.md §3.6).\n\nnodus.recipes Training recipes: TrainingJob builders for fine-tuning, pretraining, distillation, preference training and RL.\n\nnodus.recipes.finetune Fine-tuning, preference training, distillation and pretraining recipes on the catalog runtimes.\n\nnodus.recipes.rl Reinforcement learning and evaluation on catalog Environments or your own.\n\nnodus.runtime In-container helpers that speak nodusd's sockets (resources.md §8); each is a no-op outside Nodus.\n\nnodus.sandbox \\ Sandbox\\ : an isolated long-running container driven by exec, files and tunnels (resources.md §3.5, ADR-044).\n\nnodus.secret \\ Secret\\ references (resources.md §5.2). Values are write-only: nothing the API returns contains them.\n\nnodus.sweep \\ Sweep\\ : one Job template run across a matrix of GPUs, regions and parameters (resources.md §3.3, ADR-051).\n\nnodus.volume \\ Volume\\ : named storage with Modal's commit and reload semantics for \\ ReadWriteMany\\ (resources.md §5.1, ADR-091).\n\nnodus.workspace \\ Workspace\\ : a development machine with SSH, VS Code, JupyterLab and a persistent home (resources.md §3.8)."},{"id":"docs/reference/python/nodus","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus.md","title":"nodus","description":"Nodus Python SDK (`pip install nodus-compute`).","stage":"GA","headings":[{"depth":2,"slug":"exports","text":"Exports"}],"text":"nodus\n\n Nodus Python SDK ( pip install nodus-compute ).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nNodus Python SDK ( pip install nodus-compute ).\n\nLayer 2, the Modal-shaped layer, is what most code uses: App , @app.function with .remote() , .spawn() and .map() , @app.cls , Image , Volume , Secret , Sandbox , Job , Sweep , Workspace , Agent , nodus.llm and @nodus.clustered . Layer 1, nodus.api , is the generic resource client under it. Every blocking call also has an .aio form for asyncio code ( await f.remote.aio(x) ). Errors are in nodus.errors .\n\nSubmodules load on first attribute access, so import nodus stays fast and in-container helpers such as nodus.progress pull in nothing they do not need.\n\nExports\n\n Agent : defined in nodus.agent \n AgentGroup : defined in nodus.agent \n AgentRun : defined in nodus.agent \n App : defined in nodus.app \n ClaudeAgent : defined in nodus. runtime.agents.claude \n Client : defined in nodus.api \n Cls : defined in nodus.functions \n Distributed : defined in nodus.job \n Egress : defined in nodus. spec \n Function : defined in nodus.functions \n FunctionCall : defined in nodus.functions \n GPU : defined in nodus. spec \n Image : defined in nodus.image \n InferenceEndpoint : defined in nodus.llm \n Init : defined in nodus.sandbox \n Job : defined in nodus.job \n Process : defined in nodus.process \n Retries : defined in nodus. spec \n RunContext : defined in nodus.agent \n Sandbox : defined in nodus.sandbox \n Secret : defined in nodus.secret \n Service : defined in nodus.sandbox \n Sweep : defined in nodus.sweep \n TrainingJob : defined in nodus.recipes \n Tunnel : defined in nodus.sandbox \n Volume : defined in nodus.volume \n Workspace : defined in nodus.workspace \n clustered : defined in nodus.functions \n enter : defined in nodus.functions \n exit : defined in nodus.functions \n method : defined in nodus.functions \n progress : defined in nodus.runtime \n restored : defined in nodus.runtime \n self : defined in nodus.runtime \n state dir : defined in nodus.runtime"},{"id":"docs/reference/python/nodus-agent","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-agent/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-agent.md","title":"nodus.agent","description":"Agents: durable runs on the journal engine (resources.md §4.1 to §4.3, ADR-048, ADR-101).","stage":"GA","headings":[{"depth":2,"slug":"agent","text":"Agent"},{"depth":3,"slug":"agentdeploy","text":"Agent.deploy"},{"depth":3,"slug":"agententrypoint","text":"Agent.entrypoint"},{"depth":3,"slug":"agentfrom_name","text":"Agent.from_name"},{"depth":3,"slug":"agentlocal","text":"Agent.local"},{"depth":3,"slug":"agentmap","text":"Agent.map"},{"depth":3,"slug":"agentname","text":"Agent.name"},{"depth":3,"slug":"agentremote","text":"Agent.remote"},{"depth":3,"slug":"agentspawn","text":"Agent.spawn"},{"depth":3,"slug":"agentstep","text":"Agent.step"},{"depth":3,"slug":"agentsubmit","text":"Agent.submit"},{"depth":2,"slug":"agentgroup","text":"AgentGroup"},{"depth":3,"slug":"agentgroupcancel","text":"AgentGroup.cancel"},{"depth":3,"slug":"agentgroupcreate","text":"AgentGroup.create"},{"depth":3,"slug":"agentgroupdelete","text":"AgentGroup.delete"},{"depth":3,"slug":"agentgroupfrom_name","text":"AgentGroup.from_name"},{"depth":3,"slug":"agentgroupname","text":"AgentGroup.name"},{"depth":3,"slug":"agentgroupresults","text":"AgentGroup.results"},{"depth":3,"slug":"agentgroupruns","text":"AgentGroup.runs"},{"depth":3,"slug":"agentgroupseal","text":"AgentGroup.seal"},{"depth":3,"slug":"agentgroupsubmit_many","text":"AgentGroup.submit_many"},{"depth":3,"slug":"agentgroupwait","text":"AgentGroup.wait"},{"depth":2,"slug":"agentrun","text":"AgentRun"},{"depth":3,"slug":"agentrunanswer","text":"AgentRun.answer"},{"depth":3,"slug":"agentruncancel","text":"AgentRun.cancel"},{"depth":3,"slug":"agentrunchildren","text":"AgentRun.children"},{"depth":3,"slug":"agentrunfrom_name","text":"AgentRun.from_name"},{"depth":3,"slug":"agentrunlogs","text":"AgentRun.logs"},{"depth":3,"slug":"agentrunname","text":"AgentRun.name"},{"depth":3,"slug":"agentrunoutputs","text":"AgentRun.outputs"},{"depth":3,"slug":"agentrunresolve","text":"AgentRun.resolve"},{"depth":3,"slug":"agentrunresult","text":"AgentRun.result"},{"depth":3,"slug":"agentrunresume","text":"AgentRun.resume"},{"depth":3,"slug":"agentrunretry","text":"AgentRun.retry"},{"depth":3,"slug":"agentrunsend","text":"AgentRun.send"},{"depth":3,"slug":"agentrunsteps","text":"AgentRun.steps"},{"depth":3,"slug":"agentrunsuspend","text":"AgentRun.suspend"},{"depth":3,"slug":"agentrunwait","text":"AgentRun.wait"},{"depth":2,"slug":"runcontext","text":"RunContext"},{"depth":3,"slug":"runcontextchild_output","text":"RunContext.child_output"},{"depth":3,"slug":"runcontextcontinue_as_new","text":"RunContext.continue_as_new"},{"depth":3,"slug":"runcontextgather","text":"RunContext.gather"},{"depth":3,"slug":"runcontextidempotency_key","text":"RunContext.idempotency_key"},{"depth":3,"slug":"runcontextmap","text":"RunContext.map"},{"depth":3,"slug":"runcontextsave_output","text":"RunContext.save_output"},{"depth":3,"slug":"runcontextsend","text":"RunContext.send"},{"depth":3,"slug":"runcontextsleep","text":"RunContext.sleep"},{"depth":3,"slug":"runcontextsleep_until","text":"RunContext.sleep_until"},{"depth":3,"slug":"runcontextspawn","text":"RunContext.spawn"},{"depth":3,"slug":"runcontextstate_dir","text":"RunContext.state_dir"},{"depth":3,"slug":"runcontextstep","text":"RunContext.step"},{"depth":3,"slug":"runcontextwait_for_message","text":"RunContext.wait_for_message"}],"text":"nodus.agent\n\n Agents: durable runs on the journal engine (resources.md §4.1 to §4.3, ADR-048, ADR-101).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-agent/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nAgents: durable runs on the journal engine (resources.md §4.1 to §4.3, ADR-048, ADR-101).\n\nThis module is the client side: defining and deploying an Agent, submitting AgentRuns, waiting for results, and fanning out through AgentGroups. Inside a run, the agent runtime ( nodus. runtime.agents ) executes the entrypoint with a RunContext and journals every @agent.step ; outside a run a step is a plain function call.\n\nThe API serves an Agent’s image and perRunMaxCostUSD , an AgentRun’s input , deadline and group membership ( group , taskKey , dependsOn ), a run’s answer , steps , messages and cancel , and an AgentGroup’s agent , limits , maxCostUSD , sealed , state and bounded Environment evaluation (ADR-119 defers the rest). Anything else a call names raises errors.Unsupported before anything is sent, because the API rejects an unknown field outright; nodus.ClaudeAgent defines an agent on the served kind.\n\n Agent \n\nclass Agent(name: str, , image: Any = None, source: str Path dict[str, Any] None = None, entrypoint: str None = None, setup: str None = None, secrets: list[Any] None = None, env: dict[str, str] None = None, network: spec.Egress None = None, models: list[str] None = None, min workers: int None = None, max workers: int None = None, scaledown window: Any = None, per run max cost: Any = None, max cost: Any = None, cpu: Any = None, memory: Any = None, app: Any = None, project: str None = None) - None\n\nA durable agent definition; deploy it, then submit runs with .remote() , .spawn() or .map() .\n\n Agent.deploy \n\ndeploy() - View\n\nCreate or update the Agent; every accepted change is a new revision, and new runs pin it.\n\n Agent.entrypoint \n\nentrypoint(fn: Callable[..., Any]) - Callable[..., Any]\n\nMark fn(ctx, input) as the run entrypoint; its module is uploaded as the Agent’s source.\n\n Agent.from name \n\nfrom name(name: str, project: str None = None) - Agent\n\nA deployed Agent, submitted to without redeploying it.\n\n Agent.local \n\nlocal(input: Any = None) - Any\n\nRun the entrypoint in this process with an in-memory journal (needs the agent runtime).\n\n Agent.map \n\nmap(inputs: Iterable[Any], , max active: int None = None, order outputs: bool = True, return exceptions: bool = False) - AsyncIterator[Any]\n\nOne run per input in an ephemeral AgentGroup; yields each run’s answer text as the run finishes.\n\nAnswers come in input order unless order outputs=False . A run that does not succeed raises AgentRunFailed , or is yielded as that exception with return exceptions=True .\n\n Agent.name \n\nType: str \n\n Agent.remote \n\nremote(input: Any = None, kwargs: Any) - Any\n\nSubmit, wait and return the run’s answer text; raises AgentRunFailed .\n\n Agent.spawn \n\nspawn(input: Any = None, kwargs: Any) - AgentRun\n\nThe same as submit .\n\n Agent.step \n\nstep( fn: Callable[..., Any] None = None, , effect: str = 'pure', name: str None = None) - Any\n\nJournal a function as a step: pure , idempotent or external (an unknown outcome needs resolution).\n\n Agent.submit \n\nsubmit(input: Any = None, , idempotency key: str None = None, session key: str None = None, deadline: str None = None, group: str None = None, hold worker: str None = None, name: str None = None) - AgentRun\n\nStart a run and return its handle. A name derived from an event makes redeliveries idempotent.\n\n group is refused (submit group runs through AgentGroup.submit many ), and so are session key and a hold worker other than \"auto\" , which the API does not serve yet.\n\n AgentGroup \n\nclass AgentGroup(obj: Obj) - None\n\nRuns of one Agent under a shared concurrency limit and cost cap, with task dependencies and one cancel.\n\n AgentGroup.cancel \n\ncancel() - None\n\nCancel every unfinished run in the group.\n\n AgentGroup.create \n\ncreate(name: str, agent: Agent str, , max active: int None = None, max pending: int None = None, max held: int None = None, max cost: Any = None, evaluation: dict[str, Any] None = None, project: str None = None) - AgentGroup\n\nCreate the group; max active runs go at once and runs are released while max cost has room for them.\n\n max cost caps the runs’ Claude usage (model and routing calls) together; their sandboxes are billed apart.\n\n evaluation={\"environment\": \"nodus/arithmetic-v2@2.0.0\", \"tasks\": 10, \"seed\": 42} generates and scores a fixed batch automatically; split defaults to test, repetitions to 1, and tasks times repetitions is at most 100. timeout defaults to “30m” (1m to 24h) from group creation, including queued time. Evaluation and agent Sandbox compute is billed separately from the model-only max cost ; use a project Budget to cap total spend. max held remains unsupported.\n\n AgentGroup.delete \n\ndelete() - None\n\nDelete the group and, with it, its runs.\n\n AgentGroup.from name \n\nfrom name(name: str, project: str None = None) - AgentGroup\n\nA group that exists already, such as one created from YAML.\n\n AgentGroup.name \n\nType: str \n\n AgentGroup.results \n\nresults() - list[View]\n\nPer-case evaluation outcomes and grader evidence; no hidden answers or task payloads.\n\n AgentGroup.runs \n\nruns() - list[ AgentRun]\n\nThe group’s member runs.\n\n AgentGroup.seal \n\nseal() - None\n\nClose the group to new runs; it finishes once every run is terminal.\n\n AgentGroup.submit many \n\nsubmit many(tasks: list[dict[str, Any]]) - list[ AgentRun]\n\nCreate one run per task {key, input, depends on?} and return them in the caller’s order.\n\nA task starts after the tasks in depends on , which name tasks of this call or of an earlier one. Runs go out in AgentRunList batches (each batch is all-or-nothing); a cycle or a repeated key raises Invalid .\n\n AgentGroup.wait \n\nwait(timeout: float None = None) - View\n\nReturn the terminal status; submitted batches must be sealed, while evaluations close automatically.\n\n AgentRun \n\nclass AgentRun(obj: Obj) - None\n\nOne AgentRun: wait() , result() (the answer text), answer() , steps() , send() and cancel() .\n\n AgentRun.answer \n\nanswer() - str\n\nThe run’s full answer text.\n\n AgentRun.cancel \n\ncancel() - None\n\n AgentRun.children \n\nchildren() - list[ AgentRun]\n\n AgentRun.from name \n\nfrom name(name: str, project: str None = None) - AgentRun\n\n AgentRun.logs \n\nlogs(follow: bool = False) - AsyncIterator[str]\n\n AgentRun.name \n\nType: str \n\n AgentRun.outputs \n\nType: Outputs \n\n AgentRun.resolve \n\nresolve(step id: str, decision: str, evidence: dict[str, str] None = None, result: Any = None, checkpoint seq: int None = None, expected revision: int None = None) - View\n\nResolve an external step with an unknown outcome: Completed , NoEffect or Cancelled .\n\n AgentRun.result \n\nresult() - str\n\nThe run’s answer text (waits for the run first); raises AgentRunFailed unless it succeeded.\n\n AgentRun.resume \n\nresume() - None\n\n AgentRun.retry \n\nretry() - None\n\n AgentRun.send \n\nsend(name: str, payload: Any, message key: str None = None) - View\n\n AgentRun.steps \n\nsteps() - list[View]\n\n AgentRun.suspend \n\nsuspend() - None\n\n AgentRun.wait \n\nwait(timeout: float None = None) - View\n\nBlock until the run is terminal; raises AgentRunFailed unless it succeeded.\n\n RunContext \n\nclass RunContext(Protocol)\n\nWhat an agent entrypoint receives as ctx ; the agent runtime provides the implementation.\n\n RunContext.child output \n\nchild output(child: Any, name: str) - Any\n\n RunContext.continue as new \n\ncontinue as new(input: Any) - None\n\n RunContext.gather \n\ngather(children: list[Any], return exceptions: bool = False) - list[Any]\n\n RunContext.idempotency key \n\nType: str \n\n RunContext.map \n\nmap(inputs: Iterable[Any]) - list[Any]\n\n RunContext.save output \n\nsave output(name: str, path: str) - None\n\n RunContext.send \n\nsend(target: str, name: str, payload: Any) - None\n\n RunContext.sleep \n\nsleep(duration: str float) - None\n\n RunContext.sleep until \n\nsleep until(time: str) - None\n\n RunContext.spawn \n\nspawn(input: Any, key: str, permissions: Any = None, deadline: str None = None) - Any\n\n RunContext.state dir \n\nType: Path \n\n RunContext.step \n\nstep(name: str, fn: Callable[[], Any], effect: str = 'pure') - Any\n\n RunContext.wait for message \n\nwait for message(name: str, timeout: str float None = None) - Any"},{"id":"docs/reference/python/nodus-api","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-api/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-api.md","title":"nodus.api","description":"Layer 1 of the SDK: the generic resource client over every kind (resources.md §10.3).","stage":"GA","headings":[{"depth":2,"slug":"exports","text":"Exports"}],"text":"nodus.api\n\n Layer 1 of the SDK: the generic resource client over every kind (resources.md §10.3).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-api/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nLayer 1 of the SDK: the generic resource client over every kind (resources.md §10.3).\n\n nodus.api.get , list , watch , apply , create , patch , delete , estimate , logs , exec , files , healthz and readyz use the default client; nodus.Client(...) makes another one. Every call blocks and has an .aio form. Objects are JSON dicts shaped like the kinds in resources.md; nodus.api.models holds the generated pydantic models.\n\nExports\n\n Config : defined in nodus.api. config \n KINDS : defined in nodus.api. kinds \n Resource : defined in nodus.api. kinds \n default client : defined in nodus.api. client \n resolve : defined in nodus.api. config"},{"id":"docs/reference/python/nodus-api-models","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-api-models/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-api-models.md","title":"nodus.api.models","description":"Pydantic v2 models for every kind, generated from `api/openapi` by `make gen` (datamodel-code-generator).","stage":"GA","headings":[],"text":"nodus.api.models\n\n Pydantic v2 models for every kind, generated from api/openapi by make gen (datamodel-code-generator).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-api-models/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nPydantic v2 models for every kind, generated from api/openapi by make gen (datamodel-code-generator).\n\nThe nodus.dev/v1 kinds and their parts are at the top level ( nodus.api.models.Job ); v1beta1 holds the nodus.dev/v1beta1 ones ( nodus.api.models.v1beta1.TrainingJob ). The SDK’s calls exchange plain JSON dicts, so a model is opt-in: Job.model validate(obj) reads an object and job.model dump(mode=\"json\", by alias=True, exclude none=True) writes it back. Fields a newer server adds are kept. Every module here but this one is generated; never edit them by hand."},{"id":"docs/reference/python/nodus-app","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-app/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-app.md","title":"nodus.app","description":"`App`: the group of Functions defined in one Python file (resources.md §3.9, ADR-047).","stage":"GA","headings":[{"depth":2,"slug":"app","text":"App"},{"depth":3,"slug":"appcls","text":"App.cls"},{"depth":3,"slug":"appdeploy","text":"App.deploy"},{"depth":3,"slug":"appfunction","text":"App.function"},{"depth":3,"slug":"applocal_entrypoint","text":"App.local_entrypoint"},{"depth":3,"slug":"appname","text":"App.name"},{"depth":3,"slug":"appregistered_entrypoints","text":"App.registered_entrypoints"},{"depth":3,"slug":"appregistered_functions","text":"App.registered_functions"},{"depth":3,"slug":"apprun","text":"App.run"}],"text":"nodus.app\n\n App : the group of Functions defined in one Python file (resources.md §3.9, ADR-047).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-app/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n App : the group of Functions defined in one Python file (resources.md §3.9, ADR-047).\n\n with app.run(): creates an ephemeral App (renewed every 30 s, deleted on exit, so a vanished client leaves nothing running); app.deploy() creates or updates the persistent App and prunes members removed from the file. Code is uploaded once per content hash. Decorating never touches the network, so the worker can import the same file.\n\n App \n\nclass App(name: str, , project: str None = None, image: Any = None, secrets: list[Any] None = None, volumes: dict[str, Any] None = None, description: str None = None) - None\n\n App.cls \n\ncls( cls: type None = None, options: Any) - Any\n\nRegister a class as one Function; @nodus.method() s are callable remotely, hooks run once per worker.\n\n App.deploy \n\ndeploy() - View\n\nCreate or update the persistent App and its members; members no longer in the file are deleted.\n\n App.function \n\nfunction( fn: Callable[..., Any] None = None, , name: str None = None, gpu: Any = None, cpu: Any = None, memory: Any = None, ephemeral disk: Any = None, image: Any = None, secrets: list[Any] None = None, volumes: dict[str, Any] None = None, env: dict[str, str] None = None, network: Any = None, timeout: Any = None, retries: Any = None, checkpoint: Any = None, interruptible: Any = None, region: Any = None, profile: Any = None, max cost: Any = None, min workers: int None = None, max workers: int None = None, scaledown window: Any = None, target concurrency: int None = None, min containers: int None = None, max containers: int None = None) - Any\n\nRegister a Function; the arguments map onto Function spec fields (resources.md §10.4).\n\n App.local entrypoint \n\nlocal entrypoint( fn: Callable[..., Any] None = None, , name: str None = None) - Any\n\nMark the function nodus run app.py calls inside an ephemeral run of this App.\n\n App.name \n\nType: str \n\n App.registered entrypoints \n\nType: dict[str, Callable[..., Any]] \n\n App.registered functions \n\nType: dict[str, Function] \n\n App.run \n\nrun( , show logs: bool = True) - AsyncIterator[ App]\n\nRun the App ephemerally for the duration of the with block; it is deleted on exit."},{"id":"docs/reference/python/nodus-checkpoint","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-checkpoint/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-checkpoint.md","title":"nodus.checkpoint","description":"The checkpoint handshake and gang checkpoints (resources.md §8, distributed-training.md §10.3).","stage":"GA","headings":[{"depth":2,"slug":"ack","text":"ack"},{"depth":2,"slug":"handle","text":"handle"},{"depth":2,"slug":"on_request","text":"on_request"},{"depth":2,"slug":"requested","text":"requested"},{"depth":2,"slug":"subscribe","text":"subscribe"}],"text":"nodus.checkpoint\n\n The checkpoint handshake and gang checkpoints (resources.md §8, distributed-training.md §10.3).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-checkpoint/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nThe checkpoint handshake and gang checkpoints (resources.md §8, distributed-training.md §10.3).\n\nNodus owns checkpoint cadence and storage; programs write their own files into nodus.state dir() . Before a snapshot the node sends checkpoint.request : on request(callback) saves, then ack() tells the node the files are consistent. Restoring files never restores process memory: a restarted program loads what it saved.\n\n ack \n\nack(seq: str None = None) - None\n\nTell the node the checkpoint files are complete for request seq (the latest one by default).\n\n handle \n\nhandle(message: dict[str, Any]) - None\n\nProcess one message from the events socket.\n\n on request \n\non request(callback: Callable[[], None]) - None\n\nCall callback (then ack() ) whenever the node asks for a checkpoint; a no-op outside Nodus.\n\n requested \n\nrequested() - bool\n\nTrue while a checkpoint request is pending; poll it at safe points in a training loop.\n\n subscribe \n\nsubscribe() - None\n\nListen for checkpoint requests without a callback: poll requested() and call ack() after saving."},{"id":"docs/reference/python/nodus-cluster","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-cluster/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-cluster.md","title":"nodus.cluster","description":"`nodus.cluster.info()`: the gang env contract of distributed-training.md §3.6, as one object.","stage":"GA","headings":[{"depth":2,"slug":"clusterinfo","text":"ClusterInfo"},{"depth":3,"slug":"clusterinfoepoch","text":"ClusterInfo.epoch"},{"depth":3,"slug":"clusterinfogpus_per_node","text":"ClusterInfo.gpus_per_node"},{"depth":3,"slug":"clusterinfoips","text":"ClusterInfo.ips"},{"depth":3,"slug":"clusterinfois_leader","text":"ClusterInfo.is_leader"},{"depth":3,"slug":"clusterinfomaster_addr","text":"ClusterInfo.master_addr"},{"depth":3,"slug":"clusterinfomaster_port","text":"ClusterInfo.master_port"},{"depth":3,"slug":"clusterinforank","text":"ClusterInfo.rank"},{"depth":3,"slug":"clusterinfosize","text":"ClusterInfo.size"},{"depth":3,"slug":"clusterinfotransport","text":"ClusterInfo.transport"},{"depth":2,"slug":"info","text":"info"}],"text":"nodus.cluster\n\n nodus.cluster.info() : the gang env contract of distributed-training.md §3.6, as one object.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-cluster/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n nodus.cluster.info() : the gang env contract of distributed-training.md §3.6, as one object.\n\nIt works in @nodus.clustered Functions and in plain gang Jobs alike; outside a gang it describes a gang of one.\n\n ClusterInfo \n\nclass ClusterInfo(rank: int, size: int, ips: list[str], master addr: str, master port: int, epoch: int, transport: str None, gpus per node: int) - None\n\n ClusterInfo.epoch \n\nType: int \n\n ClusterInfo.gpus per node \n\nType: int \n\n ClusterInfo.ips \n\nType: list[str] \n\n ClusterInfo.is leader \n\nType: bool \n\nRank 0 hosts the rendezvous, and its return value is the clustered call’s result.\n\n ClusterInfo.master addr \n\nType: str \n\n ClusterInfo.master port \n\nType: int \n\n ClusterInfo.rank \n\nType: int \n\n ClusterInfo.size \n\nType: int \n\n ClusterInfo.transport \n\nType: str None \n\n info \n\ninfo(env: Mapping[str, str] None = None) - ClusterInfo"},{"id":"docs/reference/python/nodus-envs","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-envs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-envs.md","title":"nodus.envs","description":"Helpers for building your own Environment and datasets: `nodus.envs.split`.","stage":"GA","headings":[{"depth":2,"slug":"canonical","text":"canonical"},{"depth":2,"slug":"split","text":"split"}],"text":"nodus.envs\n\n Helpers for building your own Environment and datasets: nodus.envs.split .\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-envs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nHelpers for building your own Environment and datasets: nodus.envs.split .\n\n split keeps train and test disjoint under a canonical identity, the rule every catalog Environment follows: two items with the same identity (the same question, the same graph up to relabeling, the same expression up to commutativity) always land in the same split, so the held-out score measures generalisation, not recall.\n\ntrain, test = nodus.envs.split(items, identity=lambda x: x[\"question\"], test size=0.2, seed=7)\n\n canonical \n\ncanonical(value: Any) - str\n\nA stable identity for any JSON-like value: sha256 of its sorted, compact JSON (strings are whitespace- and case-normalised, so trivially reformatted duplicates collide).\n\n split \n\nsplit(items: Iterable[T], identity: Callable[[T], Any] None = None, test size: float int = 0.2, seed: int = 0) - tuple[list[T], list[T]]\n\nSeeded train and test lists with no identity in both; each keeps the input order.\n\n test size is a fraction of the identity groups (0 < f < 1) or a number of groups. Groups, not items, are assigned, so every duplicate of a held-out item is held out too."},{"id":"docs/reference/python/nodus-errors","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-errors/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-errors.md","title":"nodus.errors","description":"Errors raised by the Nodus SDK.","stage":"GA","headings":[{"depth":2,"slug":"exports","text":"Exports"},{"depth":2,"slug":"from_status","text":"from_status"}],"text":"nodus.errors\n\n Errors raised by the Nodus SDK.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-errors/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nErrors raised by the Nodus SDK.\n\nEvery error derives from NodusError . API errors map one-to-one onto the error-code registry (ADR-028): the class is chosen by the metav1.Status reason, so except nodus.errors.InsufficientCredits works the same for the CLI, the SDK and raw HTTP. Client-side failures are APIConnectionError and APITimeoutError ; outcome errors are JobFailed , AgentRunFailed , ImageBuildFailed , FunctionCallFailed and RemoteError .\n\nExports\n\n APIConnectionError : defined in nodus.errors. base \n APITimeoutError : defined in nodus.errors. base \n AdmissionsPaused : defined in nodus.errors. codes \n AgentRunFailed : defined in nodus.errors. base \n AlreadyExists : defined in nodus.errors. codes \n ArrearsOutstanding : defined in nodus.errors. codes \n BY STATUS : defined in nodus.errors. codes \n BadRequest : defined in nodus.errors. codes \n BudgetExceeded : defined in nodus.errors. codes \n CapacityUnavailable : defined in nodus.errors. codes \n Conflict : defined in nodus.errors. codes \n Expired : defined in nodus.errors. codes \n FieldImmutable : defined in nodus.errors. codes \n Forbidden : defined in nodus.errors. codes \n FunctionCallFailed : defined in nodus.errors. base \n IdempotencyKeyReused : defined in nodus.errors. codes \n ImageBuildFailed : defined in nodus.errors. base \n InsufficientCredits : defined in nodus.errors. codes \n Invalid : defined in nodus.errors. codes \n JobFailed : defined in nodus.errors. base \n NodusError : defined in nodus.errors. base \n NotFound : defined in nodus.errors. codes \n PaymentDisputed : defined in nodus.errors. codes \n PaymentMethodRequired : defined in nodus.errors. codes \n PaymentVerificationUnavailable : defined in nodus.errors. codes \n PreconditionFailed : defined in nodus.errors. codes \n QuotaExceeded : defined in nodus.errors. codes \n REGISTRY : defined in nodus.errors. codes \n RemoteError : defined in nodus.errors. base \n RequestInProgress : defined in nodus.errors. codes \n SandboxStarting : defined in nodus.errors. codes \n StdinBackpressure : defined in nodus.errors. codes \n TooManyRequests : defined in nodus.errors. codes \n Unauthorized : defined in nodus.errors. codes \n Unavailable : defined in nodus.errors. codes \n Unsupported : defined in nodus.errors. codes \n UnsupportedMediaType : defined in nodus.errors. codes \n VolumeBusy : defined in nodus.errors. codes \n\n from status \n\nfrom status(http status: int, body: Any, , request id: str None = None, retry after: float None = None, idempotency key: str None = None) - NodusError\n\nBuild the registry error for an HTTP error response whose body is a metav1.Status (or anything else)."},{"id":"docs/reference/python/nodus-examples","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-examples/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-examples.md","title":"nodus.examples","description":"Runnable examples that ship with the SDK: `python -m nodus.examples.rl` trains a model with RL.","stage":"GA","headings":[],"text":"nodus.examples\n\n Runnable examples that ship with the SDK: python -m nodus.examples.rl trains a model with RL.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-examples/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nRunnable examples that ship with the SDK: python -m nodus.examples.rl trains a model with RL."},{"id":"docs/reference/python/nodus-examples-rl","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-examples-rl/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-examples-rl.md","title":"nodus.examples.rl","description":"Reinforcement learning in one file: teach a small model to spell words backwards.","stage":"GA","headings":[{"depth":2,"slug":"main","text":"main"},{"depth":2,"slug":"reward","text":"reward"}],"text":"nodus.examples.rl\n\n Reinforcement learning in one file: teach a small model to spell words backwards.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-examples-rl/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nReinforcement learning in one file: teach a small model to spell words backwards.\n\nRun it with python -m nodus.examples.rl . To train on your own task, copy this file and change tasks and reward : a task is a prompt and the answer only reward sees, and reward scores one completion against that answer.\n\n main \n\nmain() - None\n\n reward \n\nreward(completion, answer)\n\n1.0 for the exact reversed word, 0.0 for a wrong one, None when the completion holds no answer."},{"id":"docs/reference/python/nodus-functions","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-functions/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-functions.md","title":"nodus.functions","description":"Functions, FunctionCalls and `@app.cls` classes (resources.md §3.10, §3.11, ADR-047, ADR-110).","stage":"GA","headings":[{"depth":2,"slug":"cls","text":"Cls"},{"depth":3,"slug":"clsfrom_name","text":"Cls.from_name"},{"depth":2,"slug":"clsobject","text":"ClsObject"},{"depth":2,"slug":"function","text":"Function"},{"depth":3,"slug":"functionestimate","text":"Function.estimate"},{"depth":3,"slug":"functionfor_each","text":"Function.for_each"},{"depth":3,"slug":"functionfrom_name","text":"Function.from_name"},{"depth":3,"slug":"functionget_raw_f","text":"Function.get_raw_f"},{"depth":3,"slug":"functionlocal","text":"Function.local"},{"depth":3,"slug":"functionlookup","text":"Function.lookup"},{"depth":3,"slug":"functionmap","text":"Function.map"},{"depth":3,"slug":"functionmember","text":"Function.member"},{"depth":3,"slug":"functionremote","text":"Function.remote"},{"depth":3,"slug":"functionspawn","text":"Function.spawn"},{"depth":3,"slug":"functionstarmap","text":"Function.starmap"},{"depth":2,"slug":"functioncall","text":"FunctionCall"},{"depth":3,"slug":"functioncallcancel","text":"FunctionCall.cancel"},{"depth":3,"slug":"functioncallfrom_name","text":"FunctionCall.from_name"},{"depth":3,"slug":"functioncallfunction","text":"FunctionCall.function"},{"depth":3,"slug":"functioncallget","text":"FunctionCall.get"},{"depth":3,"slug":"functioncalllogs","text":"FunctionCall.logs"},{"depth":3,"slug":"functioncallname","text":"FunctionCall.name"},{"depth":3,"slug":"functioncallstatus","text":"FunctionCall.status"},{"depth":2,"slug":"functionoptions","text":"FunctionOptions"},{"depth":3,"slug":"functionoptionscheckpoint","text":"FunctionOptions.checkpoint"},{"depth":3,"slug":"functionoptionscpu","text":"FunctionOptions.cpu"},{"depth":3,"slug":"functionoptionsenv","text":"FunctionOptions.env"},{"depth":3,"slug":"functionoptionsephemeral_disk","text":"FunctionOptions.ephemeral_disk"},{"depth":3,"slug":"functionoptionsgpu","text":"FunctionOptions.gpu"},{"depth":3,"slug":"functionoptionsimage","text":"FunctionOptions.image"},{"depth":3,"slug":"functionoptionsinterruptible","text":"FunctionOptions.interruptible"},{"depth":3,"slug":"functionoptionsmax_cost","text":"FunctionOptions.max_cost"},{"depth":3,"slug":"functionoptionsmax_workers","text":"FunctionOptions.max_workers"},{"depth":3,"slug":"functionoptionsmemory","text":"FunctionOptions.memory"},{"depth":3,"slug":"functionoptionsmin_workers","text":"FunctionOptions.min_workers"},{"depth":3,"slug":"functionoptionsname","text":"FunctionOptions.name"},{"depth":3,"slug":"functionoptionsnetwork","text":"FunctionOptions.network"},{"depth":3,"slug":"functionoptionsprofile","text":"FunctionOptions.profile"},{"depth":3,"slug":"functionoptionsregion","text":"FunctionOptions.region"},{"depth":3,"slug":"functionoptionsretries","text":"FunctionOptions.retries"},{"depth":3,"slug":"functionoptionsscaledown_window","text":"FunctionOptions.scaledown_window"},{"depth":3,"slug":"functionoptionssecrets","text":"FunctionOptions.secrets"},{"depth":3,"slug":"functionoptionstarget_concurrency","text":"FunctionOptions.target_concurrency"},{"depth":3,"slug":"functionoptionstimeout","text":"FunctionOptions.timeout"},{"depth":3,"slug":"functionoptionsvolumes","text":"FunctionOptions.volumes"},{"depth":2,"slug":"clustered","text":"clustered"}],"text":"nodus.functions\n\n Functions, FunctionCalls and @app.cls classes (resources.md §3.10, §3.11, ADR-047, ADR-110).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-functions/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nFunctions, FunctionCalls and @app.cls classes (resources.md §3.10, §3.11, ADR-047, ADR-110).\n\n .remote() , .spawn() and .map() all create FunctionCall objects, the one fan-out mechanism of R3. Arguments are cloudpickled ( nodus. serialize ), results come back through a watch on the call, and a remote exception is re-raised as its own type when that type is importable here.\n\n Cls \n\nclass Cls(fn: Function) - None\n\nAn @app.cls class: Embedder() gives an object whose @nodus.method() s have .remote , .map , .spawn .\n\n Cls.from name \n\nfrom name(app: str, name: str, project: str None = None) - Cls\n\n ClsObject \n\nclass ClsObject(fn: Function) - None\n\n Function \n\nclass Function( , raw: Callable[..., Any] None = None, app: App None = None, options: FunctionOptions None = None, user cls: type None = None, method: str None = None, deployed: str None = None, deployed app: str None = None, project: str None = None, instance: Obj None = None) - None\n\nA Function defined with @app.function , a method of an @app.cls class, or a deployed one by name.\n\n Function.estimate \n\nestimate( args: Any, kwargs: Any) - View\n\nDry-run one call: cold and warm start ETAs and the rate; .etag binds a reviewed launch.\n\n Function.for each \n\nfor each( iterables: Iterable[Any], ignore exceptions: bool = False) - None\n\nRun the function over the inputs and discard the results.\n\n Function.from name \n\nfrom name(app: str, name: str, project: str None = None) - Function\n\nA deployed Function: - in the project ( Function.lookup is an alias).\n\n Function.get raw f \n\nget raw f() - Callable[..., Any]\n\nThe undecorated function (for a class method, the class).\n\n Function.local \n\nlocal( args: Any, kwargs: Any) - Any\n\nRun in this process, without Nodus.\n\n Function.lookup \n\n Function.map \n\nmap( iterables: Iterable[Any], kwargs: dict[str, Any] None = None, order outputs: bool = True, return exceptions: bool = False) - AsyncIterator[Any]\n\nOne FunctionCall per input, created in batches (≤1,000 calls, ≤8 MiB); results stream back as they finish.\n\n Function.member \n\n Function.remote \n\nremote( args: Any, kwargs: Any) - Any\n\nOne FunctionCall: block until it finishes and return its result, or raise its exception.\n\n Function.spawn \n\nspawn( args: Any, kwargs: Any) - FunctionCall\n\nStart one FunctionCall and return its handle without waiting.\n\n Function.starmap \n\nstarmap(iterable: Iterable[Iterable[Any]], order outputs: bool = True, return exceptions: bool = False) - AsyncIterator[Any]\n\n FunctionCall \n\nclass FunctionCall(name: str, function: str None = None, project: str None = None) - None\n\nA handle on one FunctionCall: .get(timeout=None) , .cancel() , .status() , .logs() .\n\n FunctionCall.cancel \n\ncancel() - None\n\n FunctionCall.from name \n\nfrom name(name: str, project: str None = None) - FunctionCall\n\n FunctionCall.function \n\nType: str None \n\n FunctionCall.get \n\nget(timeout: float None = None) - Any\n\nWait for the result; raises the remote exception, or TimeoutError after timeout seconds.\n\n FunctionCall.logs \n\nlogs(follow: bool = False) - AsyncIterator[str]\n\nLog lines of the Function’s workers.\n\n FunctionCall.name \n\nType: str \n\n FunctionCall.status \n\nstatus() - str None\n\nThe call’s phase: Queued , Running , Recovering , Succeeded , Failed , Cancelling or Cancelled .\n\n FunctionOptions \n\nclass FunctionOptions(name: str None = None, gpu: Any = None, cpu: Any = None, memory: Any = None, ephemeral disk: Any = None, image: Any = None, secrets: list[Any] = list(), volumes: dict[str, Any] = dict(), env: dict[str, str] None = None, network: Any = None, timeout: Any = None, retries: Any = None, checkpoint: Any = None, interruptible: Any = None, region: Any = None, profile: Any = None, max cost: Any = None, min workers: int None = None, max workers: int None = None, scaledown window: Any = None, target concurrency: int None = None) - None\n\nThe @app.function / @app.cls arguments, mapped to Function spec fields by spec() .\n\n FunctionOptions.checkpoint \n\nType: Any \n\n FunctionOptions.cpu \n\nType: Any \n\n FunctionOptions.env \n\nType: dict[str, str] None \n\n FunctionOptions.ephemeral disk \n\nType: Any \n\n FunctionOptions.gpu \n\nType: Any \n\n FunctionOptions.image \n\nType: Any \n\n FunctionOptions.interruptible \n\nType: Any \n\n FunctionOptions.max cost \n\nType: Any \n\n FunctionOptions.max workers \n\nType: int None \n\n FunctionOptions.memory \n\nType: Any \n\n FunctionOptions.min workers \n\nType: int None \n\n FunctionOptions.name \n\nType: str None \n\n FunctionOptions.network \n\nType: Any \n\n FunctionOptions.profile \n\nType: Any \n\n FunctionOptions.region \n\nType: Any \n\n FunctionOptions.retries \n\nType: Any \n\n FunctionOptions.scaledown window \n\nType: Any \n\n FunctionOptions.secrets \n\nType: list[Any] \n\n FunctionOptions.target concurrency \n\nType: int None \n\n FunctionOptions.timeout \n\nType: Any \n\n FunctionOptions.volumes \n\nType: dict[str, Any] \n\n clustered \n\nclustered(size: int, launcher: str = 'Plain', network: str = 'Colocated', transport: str = 'Direct') - Callable[[Callable[..., Any]], Callable[..., Any]]\n\nRun each .remote() or .spawn() as one gang of size nodes; rank 0’s return value is the result (Beta).\n\nApply it below @app.function . nodus.cluster.info() gives each member its rank and the gang’s addresses."},{"id":"docs/reference/python/nodus-image","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-image/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-image.md","title":"nodus.image","description":"The Modal-shaped `Image` builder (resources.md §5.3).","stage":"GA","headings":[{"depth":2,"slug":"image","text":"Image"},{"depth":3,"slug":"imageadd_local_dir","text":"Image.add_local_dir"},{"depth":3,"slug":"imageadd_local_file","text":"Image.add_local_file"},{"depth":3,"slug":"imageadd_local_python_source","text":"Image.add_local_python_source"},{"depth":3,"slug":"imageapt_install","text":"Image.apt_install"},{"depth":3,"slug":"imagebuild","text":"Image.build"},{"depth":3,"slug":"imagedebian_slim","text":"Image.debian_slim"},{"depth":3,"slug":"imagedigest","text":"Image.digest"},{"depth":3,"slug":"imageenv","text":"Image.env"},{"depth":3,"slug":"imagefrom_dockerfile","text":"Image.from_dockerfile"},{"depth":3,"slug":"imagefrom_name","text":"Image.from_name"},{"depth":3,"slug":"imagefrom_registry","text":"Image.from_registry"},{"depth":3,"slug":"imagename","text":"Image.name"},{"depth":3,"slug":"imagepip_install","text":"Image.pip_install"},{"depth":3,"slug":"imagepython","text":"Image.python"},{"depth":3,"slug":"imagerun_commands","text":"Image.run_commands"},{"depth":3,"slug":"imageuv_pip_install","text":"Image.uv_pip_install"},{"depth":3,"slug":"imageuv_sync","text":"Image.uv_sync"},{"depth":3,"slug":"imageworkdir","text":"Image.workdir"}],"text":"nodus.image\n\n The Modal-shaped Image builder (resources.md §5.3).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-image/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nThe Modal-shaped Image builder (resources.md §5.3).\n\nA chain of immutable builder calls describes an Image spec. Nothing touches the network until the image is used by a published App or .build() is called; the Image object is named by the hash of its spec, so an identical chain reuses the built digest.\n\n Image \n\nclass Image( , base: str None = None, dockerfile: Obj None = None, python: str None = None, steps: tuple[tuple[Any, ...], ...] = , existing: str None = None, pull secrets: tuple[str, ...] = ) - None\n\n Image.add local dir \n\nadd local dir(local path: str Path, remote path: str) - Image\n\n Image.add local file \n\nadd local file(local path: str Path, remote path: str) - Image\n\n Image.add local python source \n\nadd local python source( modules: str) - Image\n\nCopy importable local modules or packages to /workspace/ , where the worker imports from.\n\n Image.apt install \n\napt install( packages: str) - Image\n\n Image.build \n\nbuild( , show logs: bool = True) - Image\n\nBuild now (otherwise the first use builds it), streaming the build log; raises ImageBuildFailed .\n\n Image.debian slim \n\ndebian slim(python version: str None = None) - Image\n\nThe catalog nodus/python: image; defaults to the caller’s Python minor version.\n\n Image.digest \n\nType: str None \n\nThe built digest, known after .build() or after an App that uses the image was published.\n\n Image.env \n\nenv(values: dict[str, str]) - Image\n\n Image.from dockerfile \n\nfrom dockerfile(path: str Path, context: str Path = '.') - Image\n\n Image.from name \n\nfrom name(name: str) - Image\n\nAn Image object that already exists in the project.\n\n Image.from registry \n\nfrom registry(ref: str, secret: Any = None) - Image\n\nStart from an OCI reference or a catalog image ( nodus/pytorch:2.8-cuda12.8 ).\n\n Image.name \n\nType: str None \n\n Image.pip install \n\npip install( packages: str, index url: str None = None, extra index urls: list[str] None = None) - Image\n\n Image.python \n\nType: str None \n\n Image.run commands \n\nrun commands( commands: str) - Image\n\n Image.uv pip install \n\nuv pip install( packages: str, index url: str None = None, extra index urls: list[str] None = None) - Image\n\n Image.uv sync \n\nuv sync(project dir: str Path = '.', frozen: bool = True) - Image\n\n Image.workdir \n\nworkdir(path: str) - Image"},{"id":"docs/reference/python/nodus-integrations","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-integrations/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-integrations.md","title":"nodus.integrations","description":"Framework callbacks that report training progress to Nodus from inside a container.","stage":"GA","headings":[],"text":"nodus.integrations\n\n Framework callbacks that report training progress to Nodus from inside a container.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-integrations/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nFramework callbacks that report training progress to Nodus from inside a container.\n\n nodus.integrations.transformers.NodusCallback plugs into a Hugging Face Trainer (and every TRL trainer); nodus.integrations.lightning.NodusCallback into a Lightning Trainer . Both send metrics to nodus.log.metrics , report progress, save when the node asks for a checkpoint, and do nothing outside Nodus. Only the global rank 0 process reports, so a multi-node gang reports once."},{"id":"docs/reference/python/nodus-integrations-lightning","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-integrations-lightning/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-integrations-lightning.md","title":"nodus.integrations.lightning","description":"`NodusCallback` for a Lightning `Trainer` (ADR-103).","stage":"GA","headings":[{"depth":2,"slug":"noduscallback","text":"NodusCallback"},{"depth":3,"slug":"noduscallbacklast_checkpoint","text":"NodusCallback.last_checkpoint"},{"depth":3,"slug":"noduscallbackon_train_batch_end","text":"NodusCallback.on_train_batch_end"},{"depth":3,"slug":"noduscallbackon_train_start","text":"NodusCallback.on_train_start"}],"text":"nodus.integrations.lightning\n\n NodusCallback for a Lightning Trainer (ADR-103).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-integrations-lightning/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n NodusCallback for a Lightning Trainer (ADR-103).\n\nfrom nodus.integrations.lightning import NodusCallback\n\ntrainer = L.Trainer(callbacks=[NodusCallback()], default root dir=nodus.state dir())\ntrainer.fit(model, ckpt path=NodusCallback.last checkpoint())\n\n NodusCallback \n\nclass NodusCallback(every n steps: int = 10, phase: str = 'Training') - None\n\nMetrics from trainer.callback metrics , progress, phases and the checkpoint handshake.\n\n NodusCallback.last checkpoint \n\nlast checkpoint(directory: str os.PathLike[str] None = None) - str None\n\nThe checkpoint this callback saved in an earlier attempt, for trainer.fit(ckpt path=...) .\n\n NodusCallback.on train batch end \n\non train batch end(trainer: Any, pl module: Any, outputs: Any, batch: Any, batch idx: int) - None\n\n NodusCallback.on train start \n\non train start(trainer: Any, pl module: Any) - None"},{"id":"docs/reference/python/nodus-integrations-transformers","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-integrations-transformers/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-integrations-transformers.md","title":"nodus.integrations.transformers","description":"`NodusCallback` for the Hugging Face `Trainer` and every TRL trainer built on it (ADR-103).","stage":"GA","headings":[{"depth":2,"slug":"noduscallback","text":"NodusCallback"},{"depth":3,"slug":"noduscallbackon_log","text":"NodusCallback.on_log"},{"depth":3,"slug":"noduscallbackon_save","text":"NodusCallback.on_save"},{"depth":3,"slug":"noduscallbackon_step_end","text":"NodusCallback.on_step_end"},{"depth":3,"slug":"noduscallbackon_train_begin","text":"NodusCallback.on_train_begin"},{"depth":2,"slug":"last_checkpoint","text":"last_checkpoint"}],"text":"nodus.integrations.transformers\n\n NodusCallback for the Hugging Face Trainer and every TRL trainer built on it (ADR-103).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-integrations-transformers/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n NodusCallback for the Hugging Face Trainer and every TRL trainer built on it (ADR-103).\n\nfrom nodus.integrations.transformers import NodusCallback\n\ntrainer = SFTTrainer(..., callbacks=[NodusCallback()])\ntrainer.train(resume from checkpoint=nodus.integrations.transformers.last checkpoint())\n\nPoint output dir at nodus.state dir() so the Trainer’s checkpoint- folders are the recovery state.\n\n NodusCallback \n\nclass NodusCallback(phase: str = 'Training') - None\n\nMetrics, progress, phases and the checkpoint handshake for a transformers.Trainer .\n\n NodusCallback.on log \n\non log(args: Any, state: Any, control: Any, logs: dict[str, Any] None = None, kwargs: Any) - None\n\n NodusCallback.on save \n\non save(args: Any, state: Any, control: Any, kwargs: Any) - None\n\n NodusCallback.on step end \n\non step end(args: Any, state: Any, control: Any, kwargs: Any) - Any\n\n NodusCallback.on train begin \n\non train begin(args: Any, state: Any, control: Any, kwargs: Any) - None\n\n last checkpoint \n\nlast checkpoint(directory: str os.PathLike[str] None = None) - str None\n\nThe newest complete checkpoint- under directory (default the state dir), for resume from checkpoint .\n\nA folder counts only when the Trainer finished writing it ( trainer state.json is present), so an attempt stopped mid-save resumes from the checkpoint before it."},{"id":"docs/reference/python/nodus-job","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-job/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-job.md","title":"nodus.job","description":"`Job`: a run-to-completion container with checkpoints, recovery and outputs (resources.md §3.1).","stage":"GA","headings":[{"depth":2,"slug":"distributed","text":"Distributed"},{"depth":3,"slug":"distributedgpus_per_node","text":"Distributed.gpus_per_node"},{"depth":3,"slug":"distributedlauncher","text":"Distributed.launcher"},{"depth":3,"slug":"distributednetwork","text":"Distributed.network"},{"depth":3,"slug":"distributednodes","text":"Distributed.nodes"},{"depth":3,"slug":"distributedspec","text":"Distributed.spec"},{"depth":3,"slug":"distributedstartup_timeout","text":"Distributed.startup_timeout"},{"depth":3,"slug":"distributedtotal_gpus","text":"Distributed.total_gpus"},{"depth":3,"slug":"distributedtransport","text":"Distributed.transport"},{"depth":2,"slug":"job","text":"Job"},{"depth":3,"slug":"jobattempts","text":"Job.attempts"},{"depth":3,"slug":"jobcancel","text":"Job.cancel"},{"depth":3,"slug":"jobcreate","text":"Job.create"},{"depth":3,"slug":"jobestimate","text":"Job.estimate"},{"depth":3,"slug":"jobexec","text":"Job.exec"},{"depth":3,"slug":"jobfiles","text":"Job.files"},{"depth":3,"slug":"jobfrom_name","text":"Job.from_name"},{"depth":3,"slug":"joblogs","text":"Job.logs"},{"depth":3,"slug":"jobname","text":"Job.name"},{"depth":3,"slug":"joboutputs","text":"Job.outputs"},{"depth":3,"slug":"jobresume","text":"Job.resume"},{"depth":3,"slug":"jobrun","text":"Job.run"},{"depth":3,"slug":"jobstatus","text":"Job.status"},{"depth":3,"slug":"jobsuspend","text":"Job.suspend"},{"depth":3,"slug":"jobwait","text":"Job.wait"},{"depth":2,"slug":"output","text":"Output"},{"depth":3,"slug":"outputdownload","text":"Output.download"},{"depth":3,"slug":"outputname","text":"Output.name"},{"depth":2,"slug":"outputs","text":"Outputs"},{"depth":3,"slug":"outputsdownload","text":"Outputs.download"},{"depth":3,"slug":"outputslist","text":"Outputs.list"}],"text":"nodus.job\n\n Job : a run-to-completion container with checkpoints, recovery and outputs (resources.md §3.1).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-job/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n Job : a run-to-completion container with checkpoints, recovery and outputs (resources.md §3.1).\n\n distributed=nodus.Distributed(...) makes it a multi-node gang (Beta, ADR-112); logs and exec then take a rank .\n\n Distributed \n\nclass Distributed(nodes: int None = None, total gpus: int None = None, gpus per node: int None = None, launcher: str None = None, network: str None = None, transport: str None = None, startup timeout: Any = None) - None\n\nA gang: nodes or total gpus ; launcher Plain , Torchrun , Ray , Verl ; network and transport .\n\n Distributed.gpus per node \n\nType: int None \n\n Distributed.launcher \n\nType: str None \n\n Distributed.network \n\nType: str None \n\n Distributed.nodes \n\nType: int None \n\n Distributed.spec \n\nspec() - Obj\n\n Distributed.startup timeout \n\nType: Any \n\n Distributed.total gpus \n\nType: int None \n\n Distributed.transport \n\nType: str None \n\n Job \n\nclass Job(obj: Obj) - None\n\n Job.attempts \n\nattempts() - list[View]\n\n Job.cancel \n\ncancel() - None\n\n Job.create \n\ncreate( , name: str None = None, image: Any = None, command: list[str] None = None, args: list[str] None = None, source: str os.PathLike[str] Mapping[str, Any] None = None, gpu: Any = None, cpu: Any = None, memory: Any = None, disk: Any = None, env: dict[str, str] None = None, secrets: list[Any] None = None, volumes: Mapping[str, Any] None = None, workdir: str None = None, network: spec.Egress None = None, checkpoint: Any = None, timeout: Any = None, expected duration: Any = None, interruptible: Any = None, region: Any = None, profile: Any = None, max cost: Any = None, completions: int None = None, parallelism: int None = None, distributed: Distributed None = None, outputs: Mapping[str, str] None = None, labels: dict[str, str] None = None, allow large source: bool = False, project: str None = None) - Job\n\n Job.estimate \n\nestimate() - View\n\nThe Job’s status.estimate : cost p50/p90, start ETA, hold (and topology and gang hold for gangs).\n\n Job.exec \n\nexec( command: str, pty: bool = False, rank: int None = None, index: int None = None) - Process\n\n Job.files \n\nType: Files \n\n Job.from name \n\nfrom name(name: str, project: str None = None) - Job\n\n Job.logs \n\nlogs(follow: bool = False, rank: int str None = None, params: Any) - AsyncIterator[str]\n\nLog lines; for a gang, rank=n reads one rank and rank=\"all\" merges them with [r ] prefixes.\n\n Job.name \n\nType: str \n\n Job.outputs \n\nType: Outputs \n\n job.outputs[\"model\"].download(\"./model\") .\n\n Job.resume \n\nresume() - None\n\n Job.run \n\nrun( kwargs: Any) - Job\n\nCreate the Job and return its handle (the same as create ); .wait() blocks until it finishes.\n\n Job.status \n\nstatus() - View\n\n Job.suspend \n\nsuspend() - None\n\n Job.wait \n\nwait(timeout: float None = None) - View\n\nBlock until the Job finishes; raises JobFailed with the exit code and log tail when it fails.\n\n Output \n\nclass Output(parent: Any, name: str, kind: str = 'Job') - None\n\n Output.download \n\ndownload(path: str os.PathLike[str], index: int None = None) - Path\n\nSave the output at path : the sha256 from X-Nodus-SHA256 is verified, then the file is renamed in.\n\n Output.name \n\nType: str \n\n Outputs \n\nclass Outputs(parent: Any, kind: str = 'Job') - None\n\nDeclared and collected outputs of a Job, TrainingJob or AgentRun: outputs[\"name\"].download(path) for one, outputs.download(dir) for all of them, outputs.list() .\n\n Outputs.download \n\ndownload(path: str os.PathLike[str], prefix: str = '') - Path\n\nSave every output named under prefix (all by default) into the directory path , each verified, at its name below prefix : outputs.download(\"./adapter\", prefix=\"adapter/\") .\n\n Outputs.list \n\nlist() - list[Obj]"},{"id":"docs/reference/python/nodus-llm","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-llm/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-llm.md","title":"nodus.llm","description":"OpenAI- and Anthropic-compatible inference (resources.md §4.5, §4.6, ADR-049, ADR-100).","stage":"GA","headings":[{"depth":2,"slug":"inferenceendpoint","text":"InferenceEndpoint"},{"depth":3,"slug":"inferenceendpointanthropic","text":"InferenceEndpoint.anthropic"},{"depth":3,"slug":"inferenceendpointbase_url","text":"InferenceEndpoint.base_url"},{"depth":3,"slug":"inferenceendpointcreate","text":"InferenceEndpoint.create"},{"depth":3,"slug":"inferenceendpointfrom_name","text":"InferenceEndpoint.from_name"},{"depth":3,"slug":"inferenceendpointname","text":"InferenceEndpoint.name"},{"depth":3,"slug":"inferenceendpointopenai","text":"InferenceEndpoint.openai"},{"depth":3,"slug":"inferenceendpointusage","text":"InferenceEndpoint.usage"},{"depth":2,"slug":"anthropic","text":"anthropic"},{"depth":2,"slug":"async_anthropic","text":"async_anthropic"},{"depth":2,"slug":"async_openai","text":"async_openai"},{"depth":2,"slug":"openai","text":"openai"}],"text":"nodus.llm\n\n OpenAI- and Anthropic-compatible inference (resources.md §4.5, §4.6, ADR-049, ADR-100).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-llm/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nOpenAI- and Anthropic-compatible inference (resources.md §4.5, §4.6, ADR-049, ADR-100).\n\n nodus.llm.openai() and nodus.llm.anthropic() return the official clients pointed at the Nodus inference data plane with your key, so every stock SDK feature works. Inside a Nodus container they use OPENAI BASE URL and ANTHROPIC BASE URL , the loopback proxy that adds the attempt’s token and bills the owning run; name an endpoint there with model=\"endpoint/ \" . Requires the openai or anthropic extra: pip install \"nodus-compute[openai]\" .\n\n InferenceEndpoint \n\nclass InferenceEndpoint(obj: Obj) - None\n\nA named access policy (limits, allowed keys, cost cap) over a catalog Model, with its own base URL.\n\n InferenceEndpoint.anthropic \n\nanthropic( kwargs: Any) - Any\n\n InferenceEndpoint.base url \n\nType: str None \n\n InferenceEndpoint.create \n\ncreate(name: str, model: str, , rpm: int None = None, tpm: int None = None, max concurrent: int None = None, allowed keys: list[str] None = None, max cost: Any = None, project: str None = None) - InferenceEndpoint\n\n InferenceEndpoint.from name \n\nfrom name(name: str, project: str None = None) - InferenceEndpoint\n\n InferenceEndpoint.name \n\nType: str \n\n InferenceEndpoint.openai \n\nopenai( kwargs: Any) - Any\n\n InferenceEndpoint.usage \n\nusage() - View\n\nRequests, errors, tokens and cost over the last 24 hours (refreshed every minute).\n\n anthropic \n\nanthropic(endpoint: str None = None, project: str None = None, kwargs: Any) - Any\n\nAn anthropic.Anthropic client for /v1/messages on the inference data plane.\n\n async anthropic \n\nasync anthropic(endpoint: str None = None, project: str None = None, kwargs: Any) - Any\n\n async openai \n\nasync openai(endpoint: str None = None, project: str None = None, kwargs: Any) - Any\n\n openai \n\nopenai(endpoint: str None = None, project: str None = None, kwargs: Any) - Any\n\nAn openai.OpenAI client for the inference data plane (or a named InferenceEndpoint)."},{"id":"docs/reference/python/nodus-log","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-log/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-log.md","title":"nodus.log","description":"Structured telemetry from inside a container: metrics, RL task outcomes, work units and phases.","stage":"GA","headings":[{"depth":2,"slug":"metrics","text":"metrics"},{"depth":2,"slug":"phase","text":"phase"},{"depth":2,"slug":"task","text":"task"},{"depth":2,"slug":"task_started","text":"task_started"},{"depth":2,"slug":"unit","text":"unit"}],"text":"nodus.log\n\n Structured telemetry from inside a container: metrics, RL task outcomes, work units and phases.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-log/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nStructured telemetry from inside a container: metrics, RL task outcomes, work units and phases.\n\n metrics \n\nmetrics(step: int None = None, values: float) - None\n\nNumeric metrics ( loss=0.4 ); they appear in status.progress.metrics and the console charts.\n\n phase \n\nphase(name: str) - None\n\n task \n\ntask(phase: str, task id: str, outcome: str None = None, reward: float None = None, attempt: int = 1, trace: Any = None, kind: str = 'completed') - None\n\nAn RL or evaluation task event; kind=\"started\" marks the start.\n\n task started \n\ntask started(phase: str, task id: str) - None\n\n unit \n\nunit(id: str, ms: float) - None\n\nOne finished unit of work and its duration, for per-unit cost and latency."},{"id":"docs/reference/python/nodus-outputs","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-outputs/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-outputs.md","title":"nodus.outputs","description":"`nodus.outputs.verify()`: fail fast inside a Job when its model output would not load (resources.md §10.5).","stage":"GA","headings":[{"depth":2,"slug":"verify","text":"verify"}],"text":"nodus.outputs\n\n nodus.outputs.verify() : fail fast inside a Job when its model output would not load (resources.md §10.5).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-outputs/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n nodus.outputs.verify() : fail fast inside a Job when its model output would not load (resources.md §10.5).\n\n verify \n\nverify(kind: str = 'huggingface', path: str os.PathLike[str] None = None) - Path\n\nCheck that path (default NODUS OUTPUT DIR ) holds a loadable peft adapter or huggingface model."},{"id":"docs/reference/python/nodus-process","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-process/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-process.md","title":"nodus.process","description":"Processes and files inside a running Job, Sandbox or Workspace (resources.md §3.6).","stage":"GA","headings":[{"depth":2,"slug":"attach","text":"Attach"},{"depth":3,"slug":"attachclose_stdin","text":"Attach.close_stdin"},{"depth":3,"slug":"attachexit_code","text":"Attach.exit_code"},{"depth":3,"slug":"attachresize","text":"Attach.resize"},{"depth":3,"slug":"attachwrite","text":"Attach.write"},{"depth":2,"slug":"file","text":"File"},{"depth":3,"slug":"fileclose","text":"File.close"},{"depth":3,"slug":"fileread","text":"File.read"},{"depth":3,"slug":"filewrite","text":"File.write"},{"depth":2,"slug":"files","text":"Files"},{"depth":3,"slug":"fileslist","text":"Files.list"},{"depth":3,"slug":"filesopen","text":"Files.open"},{"depth":3,"slug":"filesread","text":"Files.read"},{"depth":3,"slug":"filesremove","text":"Files.remove"},{"depth":3,"slug":"fileswrite","text":"Files.write"},{"depth":2,"slug":"process","text":"Process"},{"depth":3,"slug":"processattach","text":"Process.attach"},{"depth":3,"slug":"processcancel","text":"Process.cancel"},{"depth":3,"slug":"processname","text":"Process.name"},{"depth":3,"slug":"processresize","text":"Process.resize"},{"depth":3,"slug":"processreturncode","text":"Process.returncode"},{"depth":3,"slug":"processsignal","text":"Process.signal"},{"depth":3,"slug":"processstart","text":"Process.start"},{"depth":3,"slug":"processstderr","text":"Process.stderr"},{"depth":3,"slug":"processstdin","text":"Process.stdin"},{"depth":3,"slug":"processstdout","text":"Process.stdout"},{"depth":3,"slug":"processwait","text":"Process.wait"},{"depth":3,"slug":"processwrite","text":"Process.write"},{"depth":2,"slug":"stdinwriter","text":"StdinWriter"},{"depth":3,"slug":"stdinwriterdrain","text":"StdinWriter.drain"},{"depth":3,"slug":"stdinwriterwrite","text":"StdinWriter.write"},{"depth":3,"slug":"stdinwriterwrite_eof","text":"StdinWriter.write_eof"},{"depth":2,"slug":"streamreader","text":"StreamReader"},{"depth":3,"slug":"streamreaderread","text":"StreamReader.read"},{"depth":3,"slug":"streamreaderread_bytes","text":"StreamReader.read_bytes"}],"text":"nodus.process\n\n Processes and files inside a running Job, Sandbox or Workspace (resources.md §3.6).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-process/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nProcesses and files inside a running Job, Sandbox or Workspace (resources.md §3.6).\n\n stdout , stderr and stdin use the long-poll protocol ( output , stdin , signals ), which needs no WebSocket and works through every proxy; a Process’s output is retained, so a reader that reconnects resumes from its offset. p.attach() opens the interactive remotecommand v5 WebSocket session instead, for terminals.\n\n Attach \n\nclass Attach(ws: Any) - None\n\nFrames of a remotecommand v5 session: iterate for (stream, bytes) ; write , resize , close stdin .\n\n Attach.close stdin \n\nclose stdin() - None\n\n Attach.exit code \n\nType: int None \n\nSet once the session reports the process’s exit status.\n\n Attach.resize \n\nresize(rows: int, cols: int) - None\n\n Attach.write \n\nwrite(data: bytes str) - None\n\n File \n\nclass File(files: Files, path: str, mode: str) - None\n\nA file opened with sb.open(path, mode) : reads come from one download, writes upload on close() .\n\n File.close \n\nclose() - None\n\n File.read \n\nread(n: int = 1) - str bytes\n\n File.write \n\nwrite(data: str bytes) - int\n\n Files \n\nclass Files(client: Client, kind: str, name: str, project: str None = None) - None\n\nThe files subresource of a Job, Sandbox or Workspace.\n\n Files.list \n\nlist(path: str = '/') - list[Obj]\n\n Files.open \n\nopen(path: str, mode: str = 'r') - File\n\n Files.read \n\nread(path: str) - bytes\n\n Files.remove \n\nremove(path: str) - None\n\n Files.write \n\nwrite(path: str, data: bytes str, expected sha256: str None = None) - None\n\n Process \n\nclass Process(client: Client, obj: Obj) - None\n\nOne command executing in a running attempt: stdout , stderr , stdin , wait() , returncode .\n\n Process.attach \n\nattach(stdin: bool = True, tty: bool = False) - AsyncIterator[ Attach]\n\nAn interactive session over the remotecommand v5 WebSocket (one stdin writer; many readers).\n\n Process.cancel \n\ncancel() - None\n\n Process.name \n\nType: str \n\n Process.resize \n\nresize(rows: int, cols: int) - None\n\n Process.returncode \n\nType: int None \n\nThe exit code once the process has finished, else None.\n\n Process.signal \n\nsignal(sig: str) - None\n\nSend SIGINT , SIGTERM , SIGKILL , SIGHUP , SIGQUIT , SIGUSR1 or SIGUSR2 .\n\n Process.start \n\nstart(client: Client, kind: str, parent: str, command: list[str], , tty: bool = False, stdin: bool = False, env: dict[str, str] None = None, workdir: str None = None, timeout: str None = None, project: str None = None, params: Any) - Process\n\n Process.stderr \n\nType: StreamReader \n\n Process.stdin \n\nType: StdinWriter \n\n Process.stdout \n\nType: StreamReader \n\n Process.wait \n\nwait(timeout: float None = None) - int\n\nBlock until the process exits and return its exit code; TimeoutError after timeout seconds.\n\n Process.write \n\nwrite(data: bytes str) - None\n\n StdinWriter \n\nclass StdinWriter(proc: Process) - None\n\n StdinWriter.drain \n\ndrain() - None\n\nWrites are sent as they are made, so there is nothing buffered to flush.\n\n StdinWriter.write \n\nwrite(data: bytes str) - None\n\n StdinWriter.write eof \n\nwrite eof() - None\n\n StreamReader \n\nclass StreamReader(proc: Process, stream: str) - None\n\nstdout or stderr of a Process: read() returns everything; iterating yields text as it arrives.\n\n StreamReader.read \n\nread() - str\n\n StreamReader.read bytes \n\nread bytes() - bytes"},{"id":"docs/reference/python/nodus-recipes","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-recipes/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-recipes.md","title":"nodus.recipes","description":"Training recipes: TrainingJob builders for fine-tuning, pretraining, distillation, preference training and RL.","stage":"GA","headings":[{"depth":2,"slug":"exports","text":"Exports"}],"text":"nodus.recipes\n\n Training recipes: TrainingJob builders for fine-tuning, pretraining, distillation, preference training and RL.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-recipes/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nTraining recipes: TrainingJob builders for fine-tuning, pretraining, distillation, preference training and RL.\n\n finetune and rl build TrainingJob s on the catalog runtimes (resources.md §4.8, ADR-103). A builder previews with a server dry-run ( .preview() returns a Plan with the estimate, the compiled Job, blocking reasons and the ETag) and runs bound to that ETag ( .run(gpu=, nodes=, max cost=, idempotency key=) creates with If-Match ). nodes 1 asks for a gang: the TrainingJob compiles to a Job with spec.distributed and the runtime launches one worker per GPU on every node with torchrun, Accelerate or DeepSpeed.\n\nfrom nodus.recipes import TrainingJob, finetune, rl\n\njob = TrainingJob.from example(\"nodus/gsm8k:gsm8k-trained\")\nprint(job.preview().estimate)\n\nExports\n\n Data : defined in nodus.recipes. job \n LoRA : defined in nodus.recipes. job \n Plan : defined in nodus.recipes. job \n Run : defined in nodus.recipes. job \n TrainingJob : defined in nodus.recipes. job"},{"id":"docs/reference/python/nodus-recipes-finetune","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-recipes-finetune/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-recipes-finetune.md","title":"nodus.recipes.finetune","description":"Fine-tuning, preference training, distillation and pretraining recipes on the catalog runtimes.","stage":"GA","headings":[{"depth":2,"slug":"exports","text":"Exports"},{"depth":2,"slug":"distill","text":"distill"},{"depth":2,"slug":"dpo","text":"dpo"},{"depth":2,"slug":"kto","text":"kto"},{"depth":2,"slug":"orpo","text":"orpo"},{"depth":2,"slug":"pretrain","text":"pretrain"},{"depth":2,"slug":"reward","text":"reward"},{"depth":2,"slug":"sft","text":"sft"}],"text":"nodus.recipes.finetune\n\n Fine-tuning, preference training, distillation and pretraining recipes on the catalog runtimes.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-recipes-finetune/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nFine-tuning, preference training, distillation and pretraining recipes on the catalog runtimes.\n\nEach function returns a TrainingJob builder; nothing runs until .preview() or .run(gpu=…, max cost=…) . Keyword arguments not named here are runtime parameters in snake case ( learning rate=1e-5 is parameters.learningRate ), validated by the server against the runtime’s schema.\n\nsft = finetune.sft(model=\"Qwen/Qwen3-1.7B\", data=nodus.Volume.from name(\"support-chats\"),\n lora=finetune.LoRA(r=16), max steps=2000).run(gpu=\"H100\", max cost=20)\npt = finetune.pretrain(model=\"Qwen/Qwen3-0.6B\", data=corpus, initialization=\"Scratch\")\nbig = finetune.sft(model=m, data=d, distributed=\"ZeRO3\").run(gpu=\"H100:8\", nodes=2, max cost=200)\n\nExports\n\n Data : defined in nodus.recipes. job \n LoRA : defined in nodus.recipes. job \n\n distill \n\ndistill( , student: Any, teacher: Any, data: Any, temperature: float None = None, alpha: float None = None, lmbda: float None = None, params: Any) - Any\n\nKnowledge distillation from a teacher model ( distill ); lmbda=0 is offline KD on the dataset’s text.\n\n dpo \n\ndpo( , model: Any, data: Any, beta: float None = None, params: Any) - Any\n\nDirect preference optimisation on prompt, chosen and rejected rows ( dpo ).\n\n kto \n\nkto( , model: Any, data: Any, beta: float None = None, params: Any) - Any\n\nKTO on prompt, completion and a boolean label ( kto ).\n\n orpo \n\norpo( , model: Any, data: Any, beta: float None = None, params: Any) - Any\n\nOdds-ratio preference optimisation, no reference model ( orpo ).\n\n pretrain \n\npretrain( , model: Any, data: Any, initialization: str = 'Continued', packing: bool = True, params: Any) - Any\n\nCausal-LM pretraining on raw text ( sft , task: Pretrain ).\n\n initialization=\"Continued\" keeps the model’s weights; \"Scratch\" uses only its config and tokenizer (optionally resized by architecture={...} ) and starts from random weights.\n\n reward \n\nreward( , model: Any, data: Any, params: Any) - Any\n\nA reward model from chosen and rejected pairs ( reward-model ).\n\n sft \n\nsft( , model: Any, data: Any, lora: LoRA None = None, max steps: int None = None, params: Any) - Any\n\nSupervised fine-tuning on prompt and completion (or chat messages ) rows ( sft )."},{"id":"docs/reference/python/nodus-recipes-rl","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-recipes-rl/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-recipes-rl.md","title":"nodus.recipes.rl","description":"Reinforcement learning and evaluation on catalog Environments or your own.","stage":"GA","headings":[{"depth":2,"slug":"evaluate","text":"evaluate"},{"depth":2,"slug":"grpo_lora","text":"grpo_lora"}],"text":"nodus.recipes.rl\n\n Reinforcement learning and evaluation on catalog Environments or your own.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-recipes-rl/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nReinforcement learning and evaluation on catalog Environments or your own.\n\nFrom a task list and a reward function to a training run in one call. python -m nodus.examples.rl runs a whole example ( nodus/examples/rl.py ), a template to copy:\n\ndef reward(completion, answer):\n return 1.0 if answer in completion else 0.0\n\nrun = rl.train([(\"Spell 'cat' backwards.\", \"tac\"), ...], reward, max cost=2) # model= picks another base model\nrun.watch() # each stage, then reward, loss and KL per step\nrun.outputs.download(\"./outputs\") # the LoRA adapter and the before/after comparison\n\n train packages the reward with the statements of its file it uses as an Environment in your project, then runs GRPO on it, learning from groups of replies. rl.train(\"nodus/gsm8k@1.0.0\") trains on a catalog Environment the same way. To set every GRPO parameter yourself, on a catalog Environment or one you build ( examples/training/custom-reward ):\n\njob = rl.grpo lora(model=\"Qwen/Qwen3-0.6B\", environment=\"nodus/graph-coloring@1.0.0\",\n steps=50, held out=64, seed=42)\nplan = job.preview() # estimate, compiled Job, blocking reasons, ETag\nrun = plan.run(gpu=\"RTX-4090\", max cost=5, idempotency key=\"gc-001\")\nprint(run.wait().summary.comparison) # the numbers the console shows\n\nThe trainer never grades itself: completions go to the platform grader and come back as task events, and the comparison is computed server-side on the same held-out tasks before and after training.\n\n evaluate \n\nevaluate( , model: Any, environment: str None = None, tasks: Iterable[str] None = None, held out: int = 64, seed: int = 0, task filter: Mapping[str, str] None = None, params: Any) - Any\n\n mode: Evaluate on evaluate : an Environment’s held-out tasks, or benchmark tasks such as arc easy.\n\n grpo lora \n\ngrpo lora( , model: Any, environment: str, steps: int = 50, train tasks: int = 256, held out: int = 64, seed: int = 0, lora: LoRA None = None, task filter: Mapping[str, str] None = None, min comparison tasks: int = 16, params: Any) - Any\n\nGRPO with a LoRA adapter ( grpo-lora ): baseline, train, then the final eval on the same held-out tasks."},{"id":"docs/reference/python/nodus-runtime","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-runtime/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-runtime.md","title":"nodus.runtime","description":"In-container helpers that speak nodusd's sockets (resources.md §8); each is a no-op outside Nodus.","stage":"GA","headings":[{"depth":2,"slug":"emit","text":"emit"},{"depth":2,"slug":"in_nodus","text":"in_nodus"},{"depth":2,"slug":"progress","text":"progress"},{"depth":2,"slug":"restored","text":"restored"},{"depth":2,"slug":"self","text":"self"},{"depth":2,"slug":"state_dir","text":"state_dir"}],"text":"nodus.runtime\n\n In-container helpers that speak nodusd's sockets (resources.md §8); each is a no-op outside Nodus.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-runtime/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nIn-container helpers that speak nodusd’s sockets (resources.md §8); each is a no-op outside Nodus.\n\nTelemetry goes to /run/nodus/events.sock as JSON lines. Without the socket, inside a Nodus container, events are printed as nodus.event {json} lines, which nodusd also parses; outside Nodus nothing is emitted.\n\n emit \n\nemit(event: dict[str, Any]) - None\n\nSend one telemetry event; ids are unique per process so the node’s per-attempt dedup keeps every event.\n\n in nodus \n\nin nodus() - bool\n\n progress \n\nprogress(completed: int float, total: int float None = None) - None\n\nReport progress; it appears in status.progress and the console, and drives Restartable cursors.\n\n restored \n\nrestored() - bool\n\nTrue when this attempt started from a restored checkpoint.\n\n self \n\nself() - dict[str, Any]\n\n {kind, name, project, attempt, epoch, index} of the object this container runs for, from api.sock .\n\n state dir \n\nstate dir() - Path\n\nThe checkpointed directory: write model, optimizer and progress files here."},{"id":"docs/reference/python/nodus-sandbox","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-sandbox/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-sandbox.md","title":"nodus.sandbox","description":"`Sandbox`: an isolated long-running container driven by exec, files and tunnels (resources.md §3.5, ADR-044).","stage":"GA","headings":[{"depth":2,"slug":"init","text":"Init"},{"depth":3,"slug":"initgit","text":"Init.git"},{"depth":3,"slug":"initpath","text":"Init.path"},{"depth":3,"slug":"initproject","text":"Init.project"},{"depth":3,"slug":"initref","text":"Init.ref"},{"depth":3,"slug":"initsetup","text":"Init.setup"},{"depth":2,"slug":"sandbox","text":"Sandbox"},{"depth":3,"slug":"sandboxcreate","text":"Sandbox.create"},{"depth":3,"slug":"sandboxexec","text":"Sandbox.exec"},{"depth":3,"slug":"sandboxfiles","text":"Sandbox.files"},{"depth":3,"slug":"sandboxfrom_id","text":"Sandbox.from_id"},{"depth":3,"slug":"sandboxfrom_name","text":"Sandbox.from_name"},{"depth":3,"slug":"sandboxlist","text":"Sandbox.list"},{"depth":3,"slug":"sandboxlogs","text":"Sandbox.logs"},{"depth":3,"slug":"sandboxname","text":"Sandbox.name"},{"depth":3,"slug":"sandboxobject_id","text":"Sandbox.object_id"},{"depth":3,"slug":"sandboxopen","text":"Sandbox.open"},{"depth":3,"slug":"sandboxrefresh_secrets","text":"Sandbox.refresh_secrets"},{"depth":3,"slug":"sandboxsnapshot_filesystem","text":"Sandbox.snapshot_filesystem"},{"depth":3,"slug":"sandboxstart","text":"Sandbox.start"},{"depth":3,"slug":"sandboxstatus","text":"Sandbox.status"},{"depth":3,"slug":"sandboxstop","text":"Sandbox.stop"},{"depth":3,"slug":"sandboxterminate","text":"Sandbox.terminate"},{"depth":3,"slug":"sandboxtunnels","text":"Sandbox.tunnels"},{"depth":3,"slug":"sandboxwait","text":"Sandbox.wait"},{"depth":2,"slug":"service","text":"Service"},{"depth":3,"slug":"servicecommand","text":"Service.command"},{"depth":3,"slug":"servicehealth_path","text":"Service.health_path"},{"depth":3,"slug":"serviceport","text":"Service.port"},{"depth":2,"slug":"tunnel","text":"Tunnel"},{"depth":3,"slug":"tunnelport","text":"Tunnel.port"},{"depth":3,"slug":"tunnelpublic","text":"Tunnel.public"},{"depth":3,"slug":"tunnelurl","text":"Tunnel.url"},{"depth":2,"slug":"tunnels","text":"Tunnels"},{"depth":3,"slug":"tunnelsopen","text":"Tunnels.open"}],"text":"nodus.sandbox\n\n Sandbox : an isolated long-running container driven by exec, files and tunnels (resources.md §3.5, ADR-044).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-sandbox/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n Sandbox : an isolated long-running container driven by exec, files and tunnels (resources.md §3.5, ADR-044).\n\n Sandbox.create(name=...) is create-by-name: an identical create reconnects to (and wakes) the existing Sandbox, a different spec is AlreadyExists with a diff. The image defaults to nodus/agent-tools client-side.\n\nSpec fields the API does not serve yet ( volumes , ports , init , service , egress allow-lists) raise errors.Unsupported before anything is sent, because the API rejects an unknown field outright. secrets is sent when set, and a server that does not take it yet answers Unsupported as well.\n\n Init \n\nclass Init(git: str None = None, ref: str None = None, path: str None = None, project: str Path None = None, setup: str None = None) - None\n\nOne-time setup: clone git (or upload the local project directory), then run setup as Process setup .\n\n Init.git \n\nType: str None \n\n Init.path \n\nType: str None \n\n Init.project \n\nType: str Path None \n\n Init.ref \n\nType: str None \n\n Init.setup \n\nType: str None \n\n Sandbox \n\nclass Sandbox(obj: Obj) - None\n\n Sandbox.create \n\ncreate( command: str, name: str None = None, image: Any = None, cpu: Any = None, memory: Any = None, gpu: Any = None, disk: Any = None, timeout: Any = None, idle timeout: Any = None, on idle: str = 'stop', secrets: list[Any] None = None, volumes: Mapping[str, Any] None = None, env: dict[str, str] None = None, workdir: str None = None, network: spec.Egress None = None, ports: list[int] None = None, init: Init None = None, service: Service None = None, max cost: Any = None, labels: dict[str, str] None = None, idempotency key: str None = None, project: str None = None) - Sandbox\n\nCreate (or reconnect to, by name ) a Sandbox; returns without waiting for it to start.\n\n Sandbox.exec \n\nexec( command: str, pty: bool = False, stdin: bool = False, env: dict[str, str] None = None, workdir: str None = None, timeout: Any = None) - Process\n\nStart a recorded Process; a single string with spaces runs under /bin/sh -c . Wakes a stopped Sandbox.\n\n Sandbox.files \n\nType: Files \n\n Sandbox.from id \n\nfrom id(object id: str, project: str None = None) - Sandbox\n\n Sandbox.from name \n\nfrom name(name: str, project: str None = None) - Sandbox\n\nReconnect by name; a Sandbox the user stopped is started again.\n\n Sandbox.list \n\nlist(labels: Mapping[str, str] None = None, project: str None = None) - list[ Sandbox]\n\n Sandbox.logs \n\nlogs(follow: bool = False) - AsyncIterator[str]\n\n Sandbox.name \n\nType: str \n\n Sandbox.object id \n\nType: str None \n\n Sandbox.open \n\nopen(path: str, mode: str = 'r') - File\n\n Sandbox.refresh secrets \n\nrefresh secrets() - None\n\n Sandbox.snapshot filesystem \n\nsnapshot filesystem(name: str None = None) - Any\n\nBuild an Image from this Sandbox’s filesystem (Beta); use it as image= for new Sandboxes.\n\n Sandbox.start \n\nstart() - None\n\n Sandbox.status \n\nstatus() - View\n\nThe Sandbox’s status (phase, activity, endpoints, stop reason, cost).\n\n Sandbox.stop \n\nstop() - None\n\n Sandbox.terminate \n\nterminate() - None\n\n Sandbox.tunnels \n\nType: Tunnels \n\n Sandbox.wait \n\nwait() - int None\n\nWait for the main command to exit and return its exit code (None when the Sandbox stopped without one).\n\n Service \n\nclass Service(command: str list[str], port: int, health path: str None = None) - None\n\nA supervised main service: restarted on exit, health-checked on health path .\n\n Service.command \n\nType: str list[str] \n\n Service.health path \n\nType: str None \n\n Service.port \n\nType: int \n\n Tunnel \n\nclass Tunnel(port: int, url: str, public: bool = False) - None\n\n Tunnel.port \n\nType: int \n\n Tunnel.public \n\nType: bool \n\n Tunnel.url \n\nType: str \n\n Tunnels \n\nclass Tunnels(sb: Sandbox) - None\n\n sb.tunnels.open(port) adds an ingress port and returns its URL; sb.tunnels() lists them.\n\n Tunnels.open \n\nopen(port: int, public: bool = False) - Tunnel"},{"id":"docs/reference/python/nodus-secret","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-secret/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-secret.md","title":"nodus.secret","description":"`Secret` references (resources.md §5.2). Values are write-only: nothing the API returns contains them.","stage":"GA","headings":[{"depth":2,"slug":"secret","text":"Secret"},{"depth":3,"slug":"secretcreate","text":"Secret.create"},{"depth":3,"slug":"secretfrom_dict","text":"Secret.from_dict"},{"depth":3,"slug":"secretfrom_dotenv","text":"Secret.from_dotenv"},{"depth":3,"slug":"secretfrom_name","text":"Secret.from_name"},{"depth":3,"slug":"secretname","text":"Secret.name"}],"text":"nodus.secret\n\n Secret references (resources.md §5.2). Values are write-only: nothing the API returns contains them.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-secret/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n Secret references (resources.md §5.2). Values are write-only: nothing the API returns contains them.\n\n Secret \n\nclass Secret(name: str None = None, , data: dict[str, str] None = None, type: str = 'Opaque') - None\n\nA named Secret, or inline values that become one when first used.\n\n from name never touches the network. from dict and from dotenv are stored as - each time an App using them is published (owned by that App object and removed with it), or as sdk- outside an App.\n\n Secret.create \n\ncreate(name: str, data: dict[str, str], type: str = 'Opaque', project: str None = None) - Secret\n\nCreate the Secret, or write a new version when it exists (running attempts keep their pinned version).\n\n Secret.from dict \n\nfrom dict(data: dict[str, str]) - Secret\n\n Secret.from dotenv \n\nfrom dotenv(path: str Path = '.env') - Secret\n\n Secret.from name \n\nfrom name(name: str) - Secret\n\n Secret.name \n\nType: str None"},{"id":"docs/reference/python/nodus-sweep","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-sweep/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-sweep.md","title":"nodus.sweep","description":"`Sweep`: one Job template run across a matrix of GPUs, regions and parameters (resources.md §3.3, ADR-051).","stage":"GA","headings":[{"depth":2,"slug":"sweep","text":"Sweep"},{"depth":3,"slug":"sweepcancel","text":"Sweep.cancel"},{"depth":3,"slug":"sweepcells","text":"Sweep.cells"},{"depth":3,"slug":"sweepestimate","text":"Sweep.estimate"},{"depth":3,"slug":"sweepfrom_name","text":"Sweep.from_name"},{"depth":3,"slug":"sweepjobs","text":"Sweep.jobs"},{"depth":3,"slug":"sweepname","text":"Sweep.name"},{"depth":3,"slug":"sweepresume","text":"Sweep.resume"},{"depth":3,"slug":"sweeprun","text":"Sweep.run"},{"depth":3,"slug":"sweepstatus","text":"Sweep.status"},{"depth":3,"slug":"sweepsuspend","text":"Sweep.suspend"},{"depth":3,"slug":"sweepwait","text":"Sweep.wait"}],"text":"nodus.sweep\n\n Sweep : one Job template run across a matrix of GPUs, regions and parameters (resources.md §3.3, ADR-051).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-sweep/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n Sweep : one Job template run across a matrix of GPUs, regions and parameters (resources.md §3.3, ADR-051).\n\n Sweep(job spec, grid={\"gpu\": [\"H100\", \"L4\"], \"BATCH SIZE\": [8, 16]}) describes the matrix and run() creates it; each cell is a Job named - and the report lists cost, wall time and throughput per cell. The API serves Job templates, so a Function or a recipe as the target raises errors.Unsupported .\n\n Sweep \n\nclass Sweep(target: Mapping[str, Any] None = None, , grid: Mapping[str, Iterable[Any] str] None = None, repetitions: int None = None, max parallel: int None = None, max cost: Any = None, name: str None = None, labels: dict[str, str] None = None, project: str None = None) - None\n\nA Sweep: describe it, run() it, then read the report with wait() , cells() and status() .\n\n Sweep.cancel \n\ncancel() - None\n\n Sweep.cells \n\ncells() - list[View]\n\nOne entry per cell, in index order: GPU, region, parameters, phase, cost, wall time and throughput.\n\n Sweep.estimate \n\nestimate() - View\n\nThe Sweep’s status.estimate : the total cost estimate over every cell.\n\n Sweep.from name \n\nfrom name(name: str, project: str None = None) - Sweep\n\nA Sweep that exists already, to wait on or read.\n\n Sweep.jobs \n\njobs() - list[View]\n\nThe cell Jobs the Sweep created.\n\n Sweep.name \n\nType: str None \n\nThe Sweep’s name; a Sweep without one gets it from run() .\n\n Sweep.resume \n\nresume() - None\n\n Sweep.run \n\nrun(idempotency key: str None = None) - Sweep\n\nCreate the Sweep and return it; .wait() blocks until every cell is terminal.\n\n Sweep.status \n\nstatus() - View\n\nThe Sweep’s status : phase, cells, best, cost and estimate.\n\n Sweep.suspend \n\nsuspend() - None\n\n Sweep.wait \n\nwait(timeout: float None = None) - View\n\nBlock until the Sweep is terminal and return its report (the status ).\n\nA Sweep with failed cells ends Failed with reason CellsFailed and still has its report, so only a timeout or a missing Sweep raises; read phase and best on the result."},{"id":"docs/reference/python/nodus-volume","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-volume/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-volume.md","title":"nodus.volume","description":"`Volume`: named storage with Modal's commit and reload semantics for `ReadWriteMany` (resources.md §5.1, ADR-091).","stage":"GA","headings":[{"depth":2,"slug":"batchupload","text":"BatchUpload"},{"depth":3,"slug":"batchuploadput_directory","text":"BatchUpload.put_directory"},{"depth":3,"slug":"batchuploadput_file","text":"BatchUpload.put_file"},{"depth":2,"slug":"volume","text":"Volume"},{"depth":3,"slug":"volumebatch_upload","text":"Volume.batch_upload"},{"depth":3,"slug":"volumeclear","text":"Volume.clear"},{"depth":3,"slug":"volumecommit","text":"Volume.commit"},{"depth":3,"slug":"volumefrom_name","text":"Volume.from_name"},{"depth":3,"slug":"volumeimport_from","text":"Volume.import_from"},{"depth":3,"slug":"volumelistdir","text":"Volume.listdir"},{"depth":3,"slug":"volumename","text":"Volume.name"},{"depth":3,"slug":"volumeput_directory","text":"Volume.put_directory"},{"depth":3,"slug":"volumeput_file","text":"Volume.put_file"},{"depth":3,"slug":"volumeread_file","text":"Volume.read_file"},{"depth":3,"slug":"volumereload","text":"Volume.reload"},{"depth":3,"slug":"volumeremove_file","text":"Volume.remove_file"},{"depth":3,"slug":"volumerevisions","text":"Volume.revisions"}],"text":"nodus.volume\n\n Volume : named storage with Modal's commit and reload semantics for ReadWriteMany (resources.md §5.1, ADR-091).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-volume/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n Volume : named storage with Modal’s commit and reload semantics for ReadWriteMany (resources.md §5.1, ADR-091).\n\nFile transfers go through the kopia client inside the nodus CLI with a storage grant from the Volume’s uploads subresource, so they run at object-store speed and never proxy bytes through the API.\n\n BatchUpload \n\nclass BatchUpload(volume: Volume) - None\n\n BatchUpload.put directory \n\nput directory(local path: str os.PathLike[str], remote path: str) - None\n\n BatchUpload.put file \n\nput file(local path: str os.PathLike[str], remote path: str) - None\n\n Volume \n\nclass Volume(name: str, , create if missing: bool = False, access mode: str None = None, size: str = '50Gi', project: str None = None, source: Obj None = None) - None\n\n Volume.batch upload \n\nbatch upload() - BatchUpload\n\n with vol.batch upload() as up: up.put file(...); up.put directory(...) uploads as one revision.\n\n Volume.clear \n\nclear() - None\n\nCommit an empty revision; the Volume and its history stay.\n\n Volume.commit \n\ncommit() - None\n\nInside a container that mounts this Volume: publish this attempt’s changed paths as a new revision.\n\n Volume.from name \n\nfrom name(name: str, create if missing: bool = False, access mode: str None = None, size: str = '50Gi', project: str None = None) - Volume\n\nA reference resolved on first use; create if missing=True creates a ReadWriteMany Volume (as Modal).\n\n Volume.import from \n\nimport from(name: str, , huggingface: str None = None, revision: str None = None, git: str None = None, url: str None = None, sha256: str None = None, extract: str = 'auto', s3: str None = None, connection: str None = None, query: str None = None, secret: str None = None, size: str = '50Gi', project: str None = None) - Volume\n\nCreate a ReadOnlyMany Volume imported from Hugging Face, git, a URL, S3 or a Connection query.\n\n Volume.listdir \n\nlistdir(path: str = '/') - list[str]\n\n Volume.name \n\nType: str \n\n Volume.put directory \n\nput directory(local path: str os.PathLike[str], remote path: str) - None\n\n Volume.put file \n\nput file(local path: str os.PathLike[str], remote path: str) - None\n\n Volume.read file \n\nread file(remote path: str) - bytes\n\n Volume.reload \n\nreload() - None\n\nInside a container: remount the latest revision (close open files first).\n\n Volume.remove file \n\nremove file(path: str) - None\n\n Volume.revisions \n\nrevisions() - list[Obj]"},{"id":"docs/reference/python/nodus-workspace","url":"https://nodus-platform-site.pages.dev/docs/reference/python/nodus-workspace/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/python/nodus-workspace.md","title":"nodus.workspace","description":"`Workspace`: a development machine with SSH, VS Code, JupyterLab and a persistent home (resources.md §3.8).","stage":"GA","headings":[{"depth":2,"slug":"workspace","text":"Workspace"},{"depth":3,"slug":"workspacecreate","text":"Workspace.create"},{"depth":3,"slug":"workspaceexec","text":"Workspace.exec"},{"depth":3,"slug":"workspacefiles","text":"Workspace.files"},{"depth":3,"slug":"workspacefrom_name","text":"Workspace.from_name"},{"depth":3,"slug":"workspacelogs","text":"Workspace.logs"},{"depth":3,"slug":"workspacename","text":"Workspace.name"},{"depth":3,"slug":"workspaceopen","text":"Workspace.open"},{"depth":3,"slug":"workspaceschedule","text":"Workspace.schedule"},{"depth":3,"slug":"workspacesessions","text":"Workspace.sessions"},{"depth":3,"slug":"workspacessh","text":"Workspace.ssh"},{"depth":3,"slug":"workspacestart","text":"Workspace.start"},{"depth":3,"slug":"workspacestatus","text":"Workspace.status"},{"depth":3,"slug":"workspacestop","text":"Workspace.stop"},{"depth":3,"slug":"workspaceunschedule","text":"Workspace.unschedule"},{"depth":3,"slug":"workspacewait_ready","text":"Workspace.wait_ready"}],"text":"nodus.workspace\n\n Workspace : a development machine with SSH, VS Code, JupyterLab and a persistent home (resources.md §3.8).\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-workspace/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\n Workspace : a development machine with SSH, VS Code, JupyterLab and a persistent home (resources.md §3.8).\n\n Workspace \n\nclass Workspace(obj: Obj) - None\n\n Workspace.create \n\ncreate(name: str, , gpu: Any = None, cpu: Any = None, memory: Any = None, disk: Any = None, image: str None = None, volume: str Volume None = None, ephemeral: bool = False, tools: list[str] None = None, idle timeout: Any = None, secrets: list[Any] None = None, env: dict[str, str] None = None, max cost: Any = None, project: str None = None) - Workspace\n\nCreate a Workspace; unless ephemeral=True its home is volume (default -home ), made if missing.\n\n Workspace.exec \n\nexec( command: str, pty: bool = False) - Process\n\n Workspace.files \n\nType: Files \n\n Workspace.from name \n\nfrom name(name: str, project: str None = None) - Workspace\n\n Workspace.logs \n\nlogs(follow: bool = False) - AsyncIterator[str]\n\n Workspace.name \n\nType: str \n\n Workspace.open \n\nopen(tool: str = 'vscode', browser: bool = True) - str\n\nMint a preview URL for VS Code or JupyterLab in the browser, open it, and return it.\n\n Workspace.schedule \n\nschedule(ready by: str datetime, stop at: str datetime None = None) - None\n\nBe ready by ready by (it starts 15 minutes earlier) and stop at stop at .\n\n Workspace.sessions \n\nsessions() - list[View]\n\nEvery running period of this Workspace, each an Attempt whose status is the session’s receipt.\n\n Workspace.ssh \n\nssh() - None\n\nOpen an interactive SSH session through the CLI’s ssh-proxy (the CLI writes ~/.ssh/config ).\n\n Workspace.start \n\nstart() - None\n\n Workspace.status \n\nstatus() - View\n\n Workspace.stop \n\nstop() - None\n\n Workspace.unschedule \n\nunschedule() - None\n\n Workspace.wait ready \n\nwait ready(timeout: float None = None) - View\n\nBlock until the tools answer (phase Running ); raises when it fails or stops instead."},{"id":"docs/reference/runtimes","url":"https://nodus-platform-site.pages.dev/docs/reference/runtimes/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/runtimes.md","title":"Training runtimes","description":"Every catalog TrainingRuntime, its task and modes.","stage":"GA","headings":[],"text":"Training runtimes\n\n Every catalog TrainingRuntime, its task and modes.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/runtimes/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nBeta\n\nTrainingRuntimes are Beta.\n\nThe catalog runtimes in project nodus ; reference one as spec.runtime: nodus/ on a TrainingJob.\n\n Runtime Task Modes Summary \n - - - - \n distill Distill Train Knowledge distillation from a teacher model; lmbda 0 is offline distillation \n dpo DPO Train Direct preference optimization on prompt, chosen and rejected rows \n evaluate Evaluate Evaluate Benchmark evaluation on standard tasks, or over an Environment’s held-out tasks \n grpo-lora GRPO Train, Evaluate GRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform \n kto KTO Train Kahneman-Tversky optimization on unpaired desirable and undesirable completions \n orpo ORPO Train Odds-ratio preference optimization: SFT and preference alignment in one pass, no reference model \n reward-model Reward Train Reward-model training: a one-logit head scored on chosen versus rejected \n sft SFT Train Supervised fine-tuning (full, LoRA, QLoRA) and pretraining with task: Pretrain"},{"id":"docs/reference/runtimes/distill","url":"https://nodus-platform-site.pages.dev/docs/reference/runtimes/distill/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/runtimes/distill.md","title":"distill","description":"Knowledge distillation from a teacher model; lmbda 0 is offline distillation","stage":"GA","headings":[{"depth":2,"slug":"parameters","text":"Parameters"},{"depth":2,"slug":"outputs","text":"Outputs"},{"depth":2,"slug":"presets","text":"Presets"},{"depth":2,"slug":"estimates","text":"Estimates"},{"depth":2,"slug":"example","text":"Example"}],"text":"distill\n\n Knowledge distillation from a teacher model; lmbda 0 is offline distillation\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/runtimes/distill/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nBeta\n\nTrainingJob and TrainingRuntime are Beta: fields can still change before GA.\n\nKnowledge distillation from a teacher model; lmbda 0 is offline distillation.\n\n Field Value \n - - \n Reference nodus/distill \n Image nodus/distill:2.0.0 \n Task Distill \n Modes Train \n Launcher Torchrun \n Requires model yes, data yes, environment no \n Checkpoints /nodus/state (HFTrainer) \n\nParameters\n\nSet under spec.parameters ; the server validates them against this schema, and unset parameters take the runtime default.\n\n Parameter Type Default Allowed Description \n - - - - - \n alpha number 0.5 0 to 1 Generalized JSD interpolation: 0 forward KL, 1 reverse KL \n attention string none eager , sdpa , flash attention 2 Attention kernel \n batchSize integer 4 1 to 1024 Examples per GPU per step \n distributed.findUnusedParameters boolean none DDP: tolerate parameters without gradients \n distributed.strategy string Auto Auto , DDP , FSDP , ZeRO1 , ZeRO2 , ZeRO3 Auto and DDP replicate the model; FSDP and ZeRO shard it \n epochs number 1 0.01 to 100 Passes over the data when steps is -1 \n evalSteps integer none 1 to 100000 Steps between validation passes; unset evaluates at the end only \n gradientAccumulation integer 4 1 to 1024 Steps whose gradients are summed before an update \n gradientCheckpointing boolean true Recompute activations to save GPU memory \n learningRate number 0.00001 0 to 1 Peak learning rate \n lmbda number 0 0 to 1 Fraction of batches whose completions the student samples itself; 0 is offline distillation \n loggingSteps integer 10 1 to 10000 Steps between metric reports \n lora.alpha integer 32 1 to 4096 Adapter scaling numerator \n lora.dropout number 0.05 0 to 0.9 Adapter dropout \n lora.r integer 16 1 to 1024 Adapter rank \n lora.targetModules any none Module names, or all-linear \n lrScheduler string cosine cosine , linear , constant , constant with warmup Learning-rate schedule \n maxGradNorm number none 0 to 100 Gradient clipping norm \n maxLength integer 1024 16 to 131072 Tokens per example after truncation (packed block size for pretraining) \n maxNewTokens integer 128 1 to 8192 Most new tokens per student-sampled completion \n method string LoRA Full , LoRA , QLoRA Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) \n precision string Auto Auto , bf16 , fp16 , fp32 Numeric precision; Auto is bf16 where the GPU supports it, else fp16 \n saveSteps integer 200 1 to 100000 Steps between recovery checkpoints \n saveTotalLimit integer none 1 to 10 Recovery checkpoints kept in the state directory \n seed integer 42 0 to 2147483647 Seed for data order, splits and initialization \n sequenceKD boolean none Train on teacher-generated sequences (sequence-level KD) \n steps integer -1 -1 to 1000000 Optimizer steps; -1 trains for epochs instead \n temperature number 2 0.05 to 20 Softmax temperature of teacher and student \n warmupRatio number 0.03 0 to 0.5 Fraction of steps spent warming up the learning rate \n weightDecay number none 0 to 1 AdamW weight decay \n\nOutputs\n\n Output Path \n - - \n outputs /nodus/outputs \n\nEvery run also writes results.json , provenance.json and a sha256 manifest.json . Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/ :outputs/results.json ./results.json .\n\nPresets\n\nWithout spec.resources , the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.\n\n Model Method Quantization GPUs \n - - - - \n Qwen/Qwen3-0.6B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-0.6B Full none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-1.7B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-1.7B Full none 1 × A100-80G or H100 \n Qwen/Qwen3-4B QLoRA nf4 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-4B LoRA none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B QLoRA nf4 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B Full none 8 × H100 \n\nDefault resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.\n\nEstimates\n\nThe dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.\n\n Accelerator Seconds Per \n - - - \n A100-80G 0.7 secondsPerStep \n H100 0.45 secondsPerStep \n L40S 1 secondsPerStep \n RTX-4090 1.2 secondsPerStep \n RTX-A6000 1.3 secondsPerStep \n\nExample\n\napiVersion: nodus.dev/v1beta1\nkind: TrainingJob\nmetadata: {name: my-distill}\nspec:\n runtime: nodus/distill\n model: {uri: \"hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca\"}\n data: {volume: my-dataset, format: JSONL}\n maxCostUSD: \"5.00\""},{"id":"docs/reference/runtimes/dpo","url":"https://nodus-platform-site.pages.dev/docs/reference/runtimes/dpo/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/runtimes/dpo.md","title":"dpo","description":"Direct preference optimization on prompt, chosen and rejected rows","stage":"GA","headings":[{"depth":2,"slug":"parameters","text":"Parameters"},{"depth":2,"slug":"outputs","text":"Outputs"},{"depth":2,"slug":"presets","text":"Presets"},{"depth":2,"slug":"estimates","text":"Estimates"},{"depth":2,"slug":"example","text":"Example"}],"text":"dpo\n\n Direct preference optimization on prompt, chosen and rejected rows\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/runtimes/dpo/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nBeta\n\nTrainingJob and TrainingRuntime are Beta: fields can still change before GA.\n\nDirect preference optimization on prompt, chosen and rejected rows.\n\n Field Value \n - - \n Reference nodus/dpo \n Image nodus/dpo:2.0.0 \n Task DPO \n Modes Train \n Launcher Torchrun \n Requires model yes, data yes, environment no \n Checkpoints /nodus/state (HFTrainer) \n\nParameters\n\nSet under spec.parameters ; the server validates them against this schema, and unset parameters take the runtime default.\n\n Parameter Type Default Allowed Description \n - - - - - \n attention string none eager , sdpa , flash attention 2 Attention kernel \n batchSize integer 4 1 to 1024 Examples per GPU per step \n beta number 0.1 0 to 10 Strength of the KL penalty toward the reference model \n distributed.findUnusedParameters boolean none DDP: tolerate parameters without gradients \n distributed.strategy string Auto Auto , DDP , FSDP , ZeRO1 , ZeRO2 , ZeRO3 Auto and DDP replicate the model; FSDP and ZeRO shard it \n epochs number 1 0.01 to 100 Passes over the data when steps is -1 \n evalSteps integer none 1 to 100000 Steps between validation passes; unset evaluates at the end only \n gradientAccumulation integer 4 1 to 1024 Steps whose gradients are summed before an update \n gradientCheckpointing boolean true Recompute activations to save GPU memory \n learningRate number 0.000005 0 to 1 Peak learning rate \n loggingSteps integer 10 1 to 10000 Steps between metric reports \n lora.alpha integer 32 1 to 4096 Adapter scaling numerator \n lora.dropout number 0.05 0 to 0.9 Adapter dropout \n lora.r integer 16 1 to 1024 Adapter rank \n lora.targetModules any none Module names, or all-linear \n lossType string sigmoid sigmoid , ipo , hinge , robust , sppo hard , nca pair , apo zero , apo down DPO loss \n lrScheduler string cosine cosine , linear , constant , constant with warmup Learning-rate schedule \n maxGradNorm number none 0 to 100 Gradient clipping norm \n maxLength integer 1024 16 to 131072 Tokens per example after truncation (packed block size for pretraining) \n method string LoRA Full , LoRA , QLoRA Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) \n precision string Auto Auto , bf16 , fp16 , fp32 Numeric precision; Auto is bf16 where the GPU supports it, else fp16 \n saveSteps integer 200 1 to 100000 Steps between recovery checkpoints \n saveTotalLimit integer none 1 to 10 Recovery checkpoints kept in the state directory \n seed integer 42 0 to 2147483647 Seed for data order, splits and initialization \n steps integer -1 -1 to 1000000 Optimizer steps; -1 trains for epochs instead \n warmupRatio number 0.03 0 to 0.5 Fraction of steps spent warming up the learning rate \n weightDecay number none 0 to 1 AdamW weight decay \n\nOutputs\n\n Output Path \n - - \n outputs /nodus/outputs \n\nEvery run also writes results.json , provenance.json and a sha256 manifest.json . Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/ :outputs/results.json ./results.json .\n\nPresets\n\nWithout spec.resources , the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.\n\n Model Method Quantization GPUs \n - - - - \n Qwen/Qwen3-0.6B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-0.6B Full none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-1.7B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-1.7B Full none 1 × A100-80G or H100 \n Qwen/Qwen3-4B QLoRA nf4 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-4B LoRA none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B QLoRA nf4 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B Full none 8 × H100 \n\nDefault resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.\n\nEstimates\n\nThe dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.\n\n Accelerator Seconds Per \n - - - \n A100-80G 0.7 secondsPerStep \n H100 0.45 secondsPerStep \n L40S 1 secondsPerStep \n RTX-4090 1.2 secondsPerStep \n RTX-A6000 1.3 secondsPerStep \n\nExample\n\napiVersion: nodus.dev/v1beta1\nkind: TrainingJob\nmetadata: {name: my-dpo}\nspec:\n runtime: nodus/dpo\n model: {uri: \"hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca\"}\n data: {volume: my-dataset, format: JSONL}\n maxCostUSD: \"5.00\""},{"id":"docs/reference/runtimes/evaluate","url":"https://nodus-platform-site.pages.dev/docs/reference/runtimes/evaluate/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/runtimes/evaluate.md","title":"evaluate","description":"Benchmark evaluation on standard tasks, or over an Environment's held-out tasks","stage":"GA","headings":[{"depth":2,"slug":"parameters","text":"Parameters"},{"depth":2,"slug":"outputs","text":"Outputs"},{"depth":2,"slug":"estimates","text":"Estimates"},{"depth":2,"slug":"example","text":"Example"}],"text":"evaluate\n\n Benchmark evaluation on standard tasks, or over an Environment's held-out tasks\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/runtimes/evaluate/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nBeta\n\nTrainingJob and TrainingRuntime are Beta: fields can still change before GA.\n\nBenchmark evaluation on standard tasks, or over an Environment’s held-out tasks.\n\n Field Value \n - - \n Reference nodus/evaluate \n Image nodus/evaluate:2.0.0 \n Task Evaluate \n Modes Evaluate \n Launcher Torchrun \n Requires model yes, data no, environment no \n Checkpoints /nodus/state (HFTrainer) \n\nParameters\n\nSet under spec.parameters ; the server validates them against this schema, and unset parameters take the runtime default.\n\n Parameter Type Default Allowed Description \n - - - - - \n batchSize integer 8 1 to 1024 Examples per forward pass \n evaluationBatchSize integer none 1 to 1024 Prompts generated together during Environment baseline and final evaluation \n limit integer none 1 to 1000000 Samples per task (unset scores the whole task) \n maxCompletionLength integer none 8 to 32768 Environment evaluation: tokens per completion \n numFewshot integer none 0 to 64 Few-shot examples per prompt \n seed integer 0 0 to 2147483647 Seed for sampling and few-shot selection \n taskFilter object none at most 8 keys Keep only tasks whose metadata has these values (such as family on reasoning-gym), before trainTasks and heldOutTasks count them \n tasks array of string none at most 64 items Benchmark task names, such as arc easy or hellaswag, whose datasets the run downloads from the Hugging Face Hub before it scores them offline; set either tasks or spec.environment, which scores an Environment’s held-out tasks instead \n\nOutputs\n\n Output Path \n - - \n outputs /nodus/outputs \n\nEvery run also writes results.json , provenance.json and a sha256 manifest.json . Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/ :outputs/results.json ./results.json .\n\nDefault resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.\n\nEstimates\n\nThe dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.\n\n Accelerator Seconds Per \n - - - \n A100-80G 0.17 secondsPerTask \n H100 0.11 secondsPerTask \n L40S 0.25 secondsPerTask \n RTX-4090 0.3 secondsPerTask \n RTX-A6000 0.33 secondsPerTask \n\nExample\n\napiVersion: nodus.dev/v1beta1\nkind: TrainingJob\nmetadata: {name: my-evaluate}\nspec:\n runtime: nodus/evaluate\n mode: Evaluate\n model: {uri: \"hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca\"}\n parameters: {tasks: [arc easy], limit: 200}\n maxCostUSD: \"5.00\""},{"id":"docs/reference/runtimes/grpo-lora","url":"https://nodus-platform-site.pages.dev/docs/reference/runtimes/grpo-lora/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/runtimes/grpo-lora.md","title":"grpo-lora","description":"GRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform","stage":"GA","headings":[{"depth":2,"slug":"parameters","text":"Parameters"},{"depth":2,"slug":"outputs","text":"Outputs"},{"depth":2,"slug":"presets","text":"Presets"},{"depth":2,"slug":"estimates","text":"Estimates"},{"depth":2,"slug":"example","text":"Example"}],"text":"grpo-lora\n\n GRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/runtimes/grpo-lora/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nBeta\n\nTrainingJob and TrainingRuntime are Beta: fields can still change before GA.\n\nGRPO reinforcement learning with a LoRA adapter over an Environment, graded by the platform.\n\n Field Value \n - - \n Reference nodus/grpo-lora \n Image nodus/grpo-lora:2.0.0 \n Task GRPO \n Modes Train, Evaluate \n Launcher Torchrun \n Requires model yes, data no, environment yes \n Checkpoints /nodus/state (HFTrainer) \n\nParameters\n\nSet under spec.parameters ; the server validates them against this schema, and unset parameters take the runtime default.\n\n Parameter Type Default Allowed Description \n - - - - - \n attention string none eager , sdpa , flash attention 2 Attention kernel \n batchSize integer 1 1 to 1024 Examples per GPU per step \n beta number 0 0 to 10 KL penalty toward the base model; 0 disables the reference model \n distributed.findUnusedParameters boolean none DDP: tolerate parameters without gradients \n distributed.strategy string Auto Auto , DDP , FSDP , ZeRO1 , ZeRO2 , ZeRO3 Auto and DDP replicate the model; FSDP and ZeRO shard it \n epochs number 1 0.01 to 100 Passes over the data when steps is -1 \n evalSteps integer none 1 to 100000 Steps between validation passes; unset evaluates at the end only \n evaluationBatchSize integer none 1 to 1024 Prompts generated together during Environment baseline and final evaluation \n generationBatchSize integer none 2 to 65536 Total sampled completions per generation batch; must divide evenly across GPUs and numGenerations \n gradientAccumulation integer 4 1 to 1024 Steps whose gradients are summed before an update \n gradientCheckpointing boolean true Recompute activations to save GPU memory \n learningRate number 0.000001 0 to 1 Peak learning rate \n loggingSteps integer 10 1 to 10000 Steps between metric reports \n lora.alpha integer 16 1 to 4096 Adapter scaling numerator \n lora.dropout number 0.05 0 to 0.9 Adapter dropout \n lora.r integer 8 1 to 1024 Adapter rank \n lora.targetModules any none Module names, or all-linear \n lrScheduler string cosine cosine , linear , constant , constant with warmup Learning-rate schedule \n maxCompletionLength integer 64 8 to 32768 Tokens per completion; in a multi-turn episode, tokens per reply \n maxEpisodeTokens integer none 8 to 131072 Tokens of a whole multi-turn episode after the prompt, the model’s and the Environment’s \n maxGradNorm number none 0 to 100 Gradient clipping norm \n maxLength integer 2048 16 to 131072 Tokens per example after truncation (packed block size for pretraining) \n maxTurns integer none 1 to 31 Model turns per episode of a multi-turn Environment; 1 grades the first reply alone \n method string LoRA Full , LoRA , QLoRA Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) \n numGenerations integer 4 2 to 64 Completions sampled per prompt (the GRPO group) \n precision string Auto Auto , bf16 , fp16 , fp32 Numeric precision; Auto is bf16 where the GPU supports it, else fp16 \n sampling.temperature number 0.7 0 to 5 Sampling temperature \n sampling.topK integer none 0 to 1000000 Keep the highest-probability K tokens; 0 disables top-k filtering \n sampling.topP number 1 0 to 1 Nucleus sampling mass \n saveSteps integer 25 1 to 100000 Steps between recovery checkpoints \n saveTotalLimit integer none 1 to 10 Recovery checkpoints kept in the state directory \n seed integer 42 0 to 2147483647 Seed for data order, splits and initialization \n steps integer 50 -1 to 1000000 Optimizer steps; -1 trains for epochs instead \n taskFilter object none at most 8 keys Keep only tasks whose metadata has these values (such as family on reasoning-gym), before trainTasks and heldOutTasks count them \n warmupRatio number 0.03 0 to 0.5 Fraction of steps spent warming up the learning rate \n weightDecay number none 0 to 1 AdamW weight decay \n\nOutputs\n\n Output Path \n - - \n outputs /nodus/outputs \n\nEvery run also writes results.json , provenance.json and a sha256 manifest.json . Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/ :outputs/results.json ./results.json .\n\nPresets\n\nWithout spec.resources , the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.\n\n Model Method Quantization GPUs \n - - - - \n Qwen/Qwen3-0.6B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-1.7B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-4B QLoRA nf4 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-4B LoRA none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B QLoRA nf4 1 × L40S or RTX-A6000 or A100-80G \n\nDefault resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.\n\nEstimates\n\nThe dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.\n\n Accelerator Seconds Per \n - - - \n A100-80G 0.7 secondsPerStep \n H100 0.45 secondsPerStep \n L40S 1 secondsPerStep \n RTX-4090 1.2 secondsPerStep \n RTX-A6000 1.3 secondsPerStep \n\nExample\n\napiVersion: nodus.dev/v1beta1\nkind: TrainingJob\nmetadata: {name: my-grpo-lora}\nspec:\n runtime: nodus/grpo-lora\n model: {uri: \"hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca\"}\n environment: {name: nodus/graph-coloring@1.0.0, trainTasks: 200, heldOutTasks: 64, seed: 42}\n maxCostUSD: \"5.00\""},{"id":"docs/reference/runtimes/kto","url":"https://nodus-platform-site.pages.dev/docs/reference/runtimes/kto/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/runtimes/kto.md","title":"kto","description":"Kahneman-Tversky optimization on unpaired desirable and undesirable completions","stage":"GA","headings":[{"depth":2,"slug":"parameters","text":"Parameters"},{"depth":2,"slug":"outputs","text":"Outputs"},{"depth":2,"slug":"presets","text":"Presets"},{"depth":2,"slug":"estimates","text":"Estimates"},{"depth":2,"slug":"example","text":"Example"}],"text":"kto\n\n Kahneman-Tversky optimization on unpaired desirable and undesirable completions\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/runtimes/kto/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nBeta\n\nTrainingJob and TrainingRuntime are Beta: fields can still change before GA.\n\nKahneman-Tversky optimization on unpaired desirable and undesirable completions.\n\n Field Value \n - - \n Reference nodus/kto \n Image nodus/kto:2.0.0 \n Task KTO \n Modes Train \n Launcher Torchrun \n Requires model yes, data yes, environment no \n Checkpoints /nodus/state (HFTrainer) \n\nParameters\n\nSet under spec.parameters ; the server validates them against this schema, and unset parameters take the runtime default.\n\n Parameter Type Default Allowed Description \n - - - - - \n attention string none eager , sdpa , flash attention 2 Attention kernel \n batchSize integer 4 1 to 1024 Examples per GPU per step \n beta number 0.1 0 to 10 Strength of the KL penalty toward the reference model \n desirableWeight number 1 0 to 100 Loss weight of desirable examples \n distributed.findUnusedParameters boolean none DDP: tolerate parameters without gradients \n distributed.strategy string Auto Auto , DDP , FSDP , ZeRO1 , ZeRO2 , ZeRO3 Auto and DDP replicate the model; FSDP and ZeRO shard it \n epochs number 1 0.01 to 100 Passes over the data when steps is -1 \n evalSteps integer none 1 to 100000 Steps between validation passes; unset evaluates at the end only \n gradientAccumulation integer 4 1 to 1024 Steps whose gradients are summed before an update \n gradientCheckpointing boolean true Recompute activations to save GPU memory \n learningRate number 0.000005 0 to 1 Peak learning rate \n loggingSteps integer 10 1 to 10000 Steps between metric reports \n lora.alpha integer 32 1 to 4096 Adapter scaling numerator \n lora.dropout number 0.05 0 to 0.9 Adapter dropout \n lora.r integer 16 1 to 1024 Adapter rank \n lora.targetModules any none Module names, or all-linear \n lrScheduler string cosine cosine , linear , constant , constant with warmup Learning-rate schedule \n maxGradNorm number none 0 to 100 Gradient clipping norm \n maxLength integer 1024 16 to 131072 Tokens per example after truncation (packed block size for pretraining) \n method string LoRA Full , LoRA , QLoRA Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) \n precision string Auto Auto , bf16 , fp16 , fp32 Numeric precision; Auto is bf16 where the GPU supports it, else fp16 \n saveSteps integer 200 1 to 100000 Steps between recovery checkpoints \n saveTotalLimit integer none 1 to 10 Recovery checkpoints kept in the state directory \n seed integer 42 0 to 2147483647 Seed for data order, splits and initialization \n steps integer -1 -1 to 1000000 Optimizer steps; -1 trains for epochs instead \n undesirableWeight number 1 0 to 100 Loss weight of undesirable examples \n warmupRatio number 0.03 0 to 0.5 Fraction of steps spent warming up the learning rate \n weightDecay number none 0 to 1 AdamW weight decay \n\nOutputs\n\n Output Path \n - - \n outputs /nodus/outputs \n\nEvery run also writes results.json , provenance.json and a sha256 manifest.json . Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/ :outputs/results.json ./results.json .\n\nPresets\n\nWithout spec.resources , the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.\n\n Model Method Quantization GPUs \n - - - - \n Qwen/Qwen3-0.6B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-0.6B Full none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-1.7B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-1.7B Full none 1 × A100-80G or H100 \n Qwen/Qwen3-4B QLoRA nf4 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-4B LoRA none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B QLoRA nf4 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B Full none 8 × H100 \n\nDefault resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.\n\nEstimates\n\nThe dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.\n\n Accelerator Seconds Per \n - - - \n A100-80G 0.7 secondsPerStep \n H100 0.45 secondsPerStep \n L40S 1 secondsPerStep \n RTX-4090 1.2 secondsPerStep \n RTX-A6000 1.3 secondsPerStep \n\nExample\n\napiVersion: nodus.dev/v1beta1\nkind: TrainingJob\nmetadata: {name: my-kto}\nspec:\n runtime: nodus/kto\n model: {uri: \"hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca\"}\n data: {volume: my-dataset, format: JSONL}\n maxCostUSD: \"5.00\""},{"id":"docs/reference/runtimes/orpo","url":"https://nodus-platform-site.pages.dev/docs/reference/runtimes/orpo/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/runtimes/orpo.md","title":"orpo","description":"Odds-ratio preference optimization: SFT and preference alignment in one pass, no reference model","stage":"GA","headings":[{"depth":2,"slug":"parameters","text":"Parameters"},{"depth":2,"slug":"outputs","text":"Outputs"},{"depth":2,"slug":"presets","text":"Presets"},{"depth":2,"slug":"estimates","text":"Estimates"},{"depth":2,"slug":"example","text":"Example"}],"text":"orpo\n\n Odds-ratio preference optimization: SFT and preference alignment in one pass, no reference model\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/runtimes/orpo/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nBeta\n\nTrainingJob and TrainingRuntime are Beta: fields can still change before GA.\n\nOdds-ratio preference optimization: SFT and preference alignment in one pass, no reference model.\n\n Field Value \n - - \n Reference nodus/orpo \n Image nodus/orpo:2.0.0 \n Task ORPO \n Modes Train \n Launcher Torchrun \n Requires model yes, data yes, environment no \n Checkpoints /nodus/state (HFTrainer) \n\nParameters\n\nSet under spec.parameters ; the server validates them against this schema, and unset parameters take the runtime default.\n\n Parameter Type Default Allowed Description \n - - - - - \n attention string none eager , sdpa , flash attention 2 Attention kernel \n batchSize integer 4 1 to 1024 Examples per GPU per step \n beta number 0.1 0 to 10 Weight of the odds-ratio term (lambda in the ORPO paper) \n distributed.findUnusedParameters boolean none DDP: tolerate parameters without gradients \n distributed.strategy string Auto Auto , DDP , FSDP , ZeRO1 , ZeRO2 , ZeRO3 Auto and DDP replicate the model; FSDP and ZeRO shard it \n epochs number 1 0.01 to 100 Passes over the data when steps is -1 \n evalSteps integer none 1 to 100000 Steps between validation passes; unset evaluates at the end only \n gradientAccumulation integer 4 1 to 1024 Steps whose gradients are summed before an update \n gradientCheckpointing boolean true Recompute activations to save GPU memory \n learningRate number 0.000008 0 to 1 Peak learning rate \n loggingSteps integer 10 1 to 10000 Steps between metric reports \n lora.alpha integer 32 1 to 4096 Adapter scaling numerator \n lora.dropout number 0.05 0 to 0.9 Adapter dropout \n lora.r integer 16 1 to 1024 Adapter rank \n lora.targetModules any none Module names, or all-linear \n lrScheduler string cosine cosine , linear , constant , constant with warmup Learning-rate schedule \n maxGradNorm number none 0 to 100 Gradient clipping norm \n maxLength integer 1024 16 to 131072 Tokens per example after truncation (packed block size for pretraining) \n method string LoRA Full , LoRA , QLoRA Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) \n precision string Auto Auto , bf16 , fp16 , fp32 Numeric precision; Auto is bf16 where the GPU supports it, else fp16 \n saveSteps integer 200 1 to 100000 Steps between recovery checkpoints \n saveTotalLimit integer none 1 to 10 Recovery checkpoints kept in the state directory \n seed integer 42 0 to 2147483647 Seed for data order, splits and initialization \n steps integer -1 -1 to 1000000 Optimizer steps; -1 trains for epochs instead \n warmupRatio number 0.03 0 to 0.5 Fraction of steps spent warming up the learning rate \n weightDecay number none 0 to 1 AdamW weight decay \n\nOutputs\n\n Output Path \n - - \n outputs /nodus/outputs \n\nEvery run also writes results.json , provenance.json and a sha256 manifest.json . Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/ :outputs/results.json ./results.json .\n\nPresets\n\nWithout spec.resources , the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.\n\n Model Method Quantization GPUs \n - - - - \n Qwen/Qwen3-0.6B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-0.6B Full none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-1.7B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-1.7B Full none 1 × A100-80G or H100 \n Qwen/Qwen3-4B QLoRA nf4 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-4B LoRA none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B QLoRA nf4 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B Full none 8 × H100 \n\nDefault resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.\n\nEstimates\n\nThe dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.\n\n Accelerator Seconds Per \n - - - \n A100-80G 0.7 secondsPerStep \n H100 0.45 secondsPerStep \n L40S 1 secondsPerStep \n RTX-4090 1.2 secondsPerStep \n RTX-A6000 1.3 secondsPerStep \n\nExample\n\napiVersion: nodus.dev/v1beta1\nkind: TrainingJob\nmetadata: {name: my-orpo}\nspec:\n runtime: nodus/orpo\n model: {uri: \"hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca\"}\n data: {volume: my-dataset, format: JSONL}\n maxCostUSD: \"5.00\""},{"id":"docs/reference/runtimes/reward-model","url":"https://nodus-platform-site.pages.dev/docs/reference/runtimes/reward-model/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/runtimes/reward-model.md","title":"reward-model","description":"Reward-model training: a one-logit head scored on chosen versus rejected","stage":"GA","headings":[{"depth":2,"slug":"parameters","text":"Parameters"},{"depth":2,"slug":"outputs","text":"Outputs"},{"depth":2,"slug":"presets","text":"Presets"},{"depth":2,"slug":"estimates","text":"Estimates"},{"depth":2,"slug":"example","text":"Example"}],"text":"reward-model\n\n Reward-model training: a one-logit head scored on chosen versus rejected\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/runtimes/reward-model/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nBeta\n\nTrainingJob and TrainingRuntime are Beta: fields can still change before GA.\n\nReward-model training: a one-logit head scored on chosen versus rejected.\n\n Field Value \n - - \n Reference nodus/reward-model \n Image nodus/reward-model:2.0.0 \n Task Reward \n Modes Train \n Launcher Torchrun \n Requires model yes, data yes, environment no \n Checkpoints /nodus/state (HFTrainer) \n\nParameters\n\nSet under spec.parameters ; the server validates them against this schema, and unset parameters take the runtime default.\n\n Parameter Type Default Allowed Description \n - - - - - \n attention string none eager , sdpa , flash attention 2 Attention kernel \n batchSize integer 4 1 to 1024 Examples per GPU per step \n distributed.findUnusedParameters boolean none DDP: tolerate parameters without gradients \n distributed.strategy string Auto Auto , DDP , FSDP , ZeRO1 , ZeRO2 , ZeRO3 Auto and DDP replicate the model; FSDP and ZeRO shard it \n epochs number 1 0.01 to 100 Passes over the data when steps is -1 \n evalSteps integer none 1 to 100000 Steps between validation passes; unset evaluates at the end only \n gradientAccumulation integer 4 1 to 1024 Steps whose gradients are summed before an update \n gradientCheckpointing boolean true Recompute activations to save GPU memory \n learningRate number 0.00001 0 to 1 Peak learning rate \n loggingSteps integer 10 1 to 10000 Steps between metric reports \n lora.alpha integer 32 1 to 4096 Adapter scaling numerator \n lora.dropout number 0.05 0 to 0.9 Adapter dropout \n lora.r integer 16 1 to 1024 Adapter rank \n lora.targetModules any none Module names, or all-linear \n lrScheduler string cosine cosine , linear , constant , constant with warmup Learning-rate schedule \n maxGradNorm number none 0 to 100 Gradient clipping norm \n maxLength integer 1024 16 to 131072 Tokens per example after truncation (packed block size for pretraining) \n method string LoRA Full , LoRA , QLoRA Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) \n precision string Auto Auto , bf16 , fp16 , fp32 Numeric precision; Auto is bf16 where the GPU supports it, else fp16 \n saveSteps integer 200 1 to 100000 Steps between recovery checkpoints \n saveTotalLimit integer none 1 to 10 Recovery checkpoints kept in the state directory \n seed integer 42 0 to 2147483647 Seed for data order, splits and initialization \n steps integer -1 -1 to 1000000 Optimizer steps; -1 trains for epochs instead \n warmupRatio number 0.03 0 to 0.5 Fraction of steps spent warming up the learning rate \n weightDecay number none 0 to 1 AdamW weight decay \n\nOutputs\n\n Output Path \n - - \n outputs /nodus/outputs \n\nEvery run also writes results.json , provenance.json and a sha256 manifest.json . Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/ :outputs/results.json ./results.json .\n\nPresets\n\nWithout spec.resources , the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.\n\n Model Method Quantization GPUs \n - - - - \n Qwen/Qwen3-0.6B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-0.6B Full none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-1.7B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-1.7B Full none 1 × A100-80G or H100 \n Qwen/Qwen3-4B QLoRA nf4 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-4B LoRA none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B QLoRA nf4 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B Full none 8 × H100 \n\nDefault resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.\n\nEstimates\n\nThe dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.\n\n Accelerator Seconds Per \n - - - \n A100-80G 0.7 secondsPerStep \n H100 0.45 secondsPerStep \n L40S 1 secondsPerStep \n RTX-4090 1.2 secondsPerStep \n RTX-A6000 1.3 secondsPerStep \n\nExample\n\napiVersion: nodus.dev/v1beta1\nkind: TrainingJob\nmetadata: {name: my-reward-model}\nspec:\n runtime: nodus/reward-model\n model: {uri: \"hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca\"}\n data: {volume: my-dataset, format: JSONL}\n maxCostUSD: \"5.00\""},{"id":"docs/reference/runtimes/sft","url":"https://nodus-platform-site.pages.dev/docs/reference/runtimes/sft/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/runtimes/sft.md","title":"sft","description":"Supervised fine-tuning (full, LoRA, QLoRA) and pretraining with task: Pretrain","stage":"GA","headings":[{"depth":2,"slug":"parameters","text":"Parameters"},{"depth":2,"slug":"outputs","text":"Outputs"},{"depth":2,"slug":"presets","text":"Presets"},{"depth":2,"slug":"estimates","text":"Estimates"},{"depth":2,"slug":"example","text":"Example"}],"text":"sft\n\n Supervised fine-tuning (full, LoRA, QLoRA) and pretraining with task: Pretrain\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/runtimes/sft/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nBeta\n\nTrainingJob and TrainingRuntime are Beta: fields can still change before GA.\n\nSupervised fine-tuning (full, LoRA, QLoRA) and pretraining with task: Pretrain.\n\n Field Value \n - - \n Reference nodus/sft \n Image nodus/sft:2.0.0 \n Task SFT \n Modes Train \n Launcher Torchrun \n Requires model yes, data yes, environment no \n Checkpoints /nodus/state (HFTrainer) \n\nParameters\n\nSet under spec.parameters ; the server validates them against this schema, and unset parameters take the runtime default.\n\n Parameter Type Default Allowed Description \n - - - - - \n architecture object none Scratch only: config overrides such as num hidden layers or hidden size \n assistantOnlyLoss boolean none Conversational data: train on assistant turns only \n attention string none eager , sdpa , flash attention 2 Attention kernel \n batchSize integer 4 1 to 1024 Examples per GPU per step \n completionOnlyLoss boolean none Prompt/completion data: train on the completion tokens only \n dataProcesses integer none 1 to 64 Processes for tokenizing the data \n distributed.findUnusedParameters boolean none DDP: tolerate parameters without gradients \n distributed.strategy string Auto Auto , DDP , FSDP , ZeRO1 , ZeRO2 , ZeRO3 Auto and DDP replicate the model; FSDP and ZeRO shard it \n epochs number 1 0.01 to 100 Passes over the data when steps is -1 \n evalSteps integer none 1 to 100000 Steps between validation passes; unset evaluates at the end only \n gradientAccumulation integer 4 1 to 1024 Steps whose gradients are summed before an update \n gradientCheckpointing boolean true Recompute activations to save GPU memory \n initialization string Continued Continued , Scratch Pretraining: continue from the model’s weights or start from a fresh initialization \n learningRate number 0.0002 0 to 1 Peak learning rate \n loggingSteps integer 10 1 to 10000 Steps between metric reports \n lora.alpha integer 32 1 to 4096 Adapter scaling numerator \n lora.dropout number 0.05 0 to 0.9 Adapter dropout \n lora.r integer 16 1 to 1024 Adapter rank \n lora.targetModules any none Module names, or all-linear \n lrScheduler string cosine cosine , linear , constant , constant with warmup Learning-rate schedule \n maxGradNorm number none 0 to 100 Gradient clipping norm \n maxLength integer 2048 16 to 131072 Tokens per example after truncation (packed block size for pretraining) \n method string LoRA Full , LoRA , QLoRA Full fine-tuning, a LoRA adapter, or LoRA on a 4-bit quantized base (GPU only) \n packing boolean none Pack short examples into full-length blocks (default on for pretraining) \n precision string Auto Auto , bf16 , fp16 , fp32 Numeric precision; Auto is bf16 where the GPU supports it, else fp16 \n saveSteps integer 200 1 to 100000 Steps between recovery checkpoints \n saveTotalLimit integer none 1 to 10 Recovery checkpoints kept in the state directory \n seed integer 42 0 to 2147483647 Seed for data order, splits and initialization \n steps integer -1 -1 to 1000000 Optimizer steps; -1 trains for epochs instead \n warmupRatio number 0.03 0 to 0.5 Fraction of steps spent warming up the learning rate \n weightDecay number none 0 to 1 AdamW weight decay \n\nOutputs\n\n Output Path \n - - \n outputs /nodus/outputs \n\nEvery run also writes results.json , provenance.json and a sha256 manifest.json . Each file under /nodus/outputs is an output of the TrainingJob, named by its path: download one with nodus cp tj/ :outputs/results.json ./results.json .\n\nPresets\n\nWithout spec.resources , the TrainingJob takes the resources of the preset matching its model, method and quantization, else the runtime default.\n\n Model Method Quantization GPUs \n - - - - \n Qwen/Qwen3-0.6B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-0.6B Full none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-1.7B LoRA none 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-1.7B Full none 1 × A100-80G or H100 \n Qwen/Qwen3-4B QLoRA nf4 1 × RTX-4090 or L40S or RTX-A6000 \n Qwen/Qwen3-4B LoRA none 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B QLoRA nf4 1 × L40S or RTX-A6000 or A100-80G \n Qwen/Qwen3-8B Full none 8 × H100 \n\nDefault resources: 1 × RTX-4090 or L40S or RTX-A6000 or A100-80G.\n\nEstimates\n\nThe dry-run estimate multiplies these by the run’s steps or tasks; the nightly runs re-measure them.\n\n Accelerator Seconds Per \n - - - \n A100-80G 0.7 secondsPerStep \n H100 0.45 secondsPerStep \n L40S 1 secondsPerStep \n RTX-4090 1.2 secondsPerStep \n RTX-A6000 1.3 secondsPerStep \n\nExample\n\napiVersion: nodus.dev/v1beta1\nkind: TrainingJob\nmetadata: {name: my-sft}\nspec:\n runtime: nodus/sft\n model: {uri: \"hf://Qwen/Qwen3-0.6B@c1899de289a04d12100db370d81485cdf75e47ca\"}\n data: {volume: my-dataset, format: JSONL}\n maxCostUSD: \"5.00\""},{"id":"docs/reference/server-configuration","url":"https://nodus-platform-site.pages.dev/docs/reference/server-configuration/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/server-configuration.md","title":"Server configuration","description":"Every environment variable nodus-server, the node agent and the CLI read.","stage":"GA","headings":[{"depth":2,"slug":"nodus-server-roles-api-controller-gateway-inference","text":"nodus-server (roles api, controller, gateway, inference)"},{"depth":2,"slug":"nodusd-node-agent","text":"nodusd (node agent)"},{"depth":2,"slug":"nodus-cli-and-sdk-client-environment-adr-069","text":"nodus CLI and SDK (client environment, ADR-069)"}],"text":"Server configuration\n\n Every environment variable nodus-server, the node agent and the CLI read.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/server-configuration/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\n \n\nNodus processes read NODUS -prefixed environment variables, validated at startup: a missing or invalid value stops the process with an error that names the variable. Secrets come from the environment only.\n\nnodus-server (roles api, controller, gateway, inference)\n\n Variable Default Required Description \n - - - - \n NODUS ASSISTANT DOCS INDEX URL https://nodus-compute.ai/docs/index.json no The site’s runtime docs index, which the assistant searches and cites; the api role caches it for 15 minutes. Roles: api. \n NODUS AUTH ISSUER — no Expected iss claim of user access tokens (https\\://.supabase.co/auth/v1); required outside NODUS ENV=local, where empty skips the issuer check. Roles: api, gateway. \n NODUS AUTH JWKS URL — no JWKS URL of the Auth signing keys; empty derives \\ /.well-known/jwks.json. Roles: api, gateway. \n NODUS AUTH URL — no Supabase Auth base URL (https\\://.supabase.co/auth/v1; on compose); empty, which only NODUS ENV=local allows, accepts only API keys and ServiceAccount tokens. Roles: api, gateway. \n NODUS CORS ORIGINS — no Comma-separated browser origins allowed to call the API: the console and admin apps (ADR-107). Roles: api, gateway. \n NODUS DATABASE MAX CONNS 0 no Pool size override; 0 uses the role budget of ADR-070 (api 8, controller 12, gateway 6, inference 4). \n NODUS DATABASE URL — yes Postgres URL: Supavisor session mode (port 5432) in the cloud, the compose postgres locally. \n NODUS ENV local no Deployment: local (compose, CI, tests), staging or prod. \n NODUS GITHUB APP — no GitHub App as JSON: app id, slug, private key (PEM), webhook secret, client id, client secret; empty disables GitHub Connections. Roles: api, controller, gateway. \n NODUS HTTP ADDR :8080 no Public HTTP listener: api, gateway client streams, inference and health. \n NODUS INFERENCE CONSOLE ORIGINS — no Comma-separated browser origins allowed to call the data plane with a user token (the console playground). Roles: inference. \n NODUS INFERENCE REPLICAS 1 no Fleet size of the inference role; the per-replica rate limits divide by it. Roles: inference. \n NODUS INFERENCE UPSTREAMS — no Inference upstreams as JSON: groq keys (list, at most 8), groq base url, wafer key, wafer base url, openrelay key, openrelay base url, openrelay catalog url, anthropic key, anthropic base url, jev key, jev endpoint, jev model; every field optional, URLs https; empty serves no upstream unless its dedicated key is configured. Roles: controller, inference. \n NODUS INTERNAL ADDR :9090 no Internal stream-bridge listener. Roles: gateway. \n NODUS KMS KEY URI — no Key-encryption key: aws-kms\\://arn:aws:kms:… in the cloud, local://\\ locally; empty uses a fixed development key when NODUS ENV=local. \n NODUS LOCAL DOCKER HOST — no Docker endpoint the local provider creates node containers on, such as unix:///Users/me/.colima/default/docker.sock; empty keeps providers.yaml’s or unix:///var/run/docker.sock. Roles: controller. \n NODUS LOCAL NODE BINDS — no Comma-separated Docker binds of every local node container, such as /path/to/worktree:/src. Roles: controller. \n NODUS LOCAL NODE COMMAND — no Shell command a local node container runs (sh -c); empty keeps providers.yaml’s or the image’s default. Roles: controller. \n NODUS LOCAL NODE ENV — no Comma-separated KEY=value pairs added to every local node container, such as NODUS GATEWAY URL= . Roles: controller. \n NODUS LOCAL NODE IMAGE — no Image of the local provider’s node containers; empty keeps providers.yaml’s or the compose-built nodus-platform-nodusd. Roles: controller. \n NODUS LOCAL NODE NETWORK — no Docker network the local provider’s node containers join; empty keeps providers.yaml’s or nodus-platform default. Roles: controller. \n NODUS LOCAL NODE OBJSTORE ENDPOINT — no Object store endpoint written into local nodes’ storage grants when they reach the store at another address than NODUS OBJSTORE ENDPOINT, such as for a host-run server; empty keeps NODUS OBJSTORE ENDPOINT. Roles: controller. \n NODUS LOG LEVEL info no Minimum log level: debug, info, warn or error. \n NODUS METRICS ADDR :9100 no Prometheus metrics listener. \n NODUS NODE ADDR :8443 no Node protocol listener. Roles: gateway. \n NODUS NODUSD VERSION — no nodusd release new nodes fetch, such as git-0123456789ab: the server reads nodusd///nodusd.sha256 from the releases bucket and presigns the binary per create. Required outside local. Roles: controller. \n NODUS NOTIFY CONSOLE URL http://localhost:5173 no Console base URL that links in emails point at. Roles: api, controller. \n NODUS NOTIFY FROM — no Sender of every notice as an RFC 5322 address, such as Nodus \\ ; required with NODUS RESEND API KEY. Roles: api, controller. \n NODUS OBJSTORE ACCESS KEY ID — no Access key id; empty uses the AWS default credential chain. \n NODUS OBJSTORE BUCKET PREFIX — yes Bucket name prefix; buckets are -data, -logs, -ephemeral, -registry and -releases. \n NODUS OBJSTORE ENDPOINT — no S3 API endpoint; empty uses AWS S3. R2: https\\://.r2.cloudflarestorage.com. \n NODUS OBJSTORE R2 ACCOUNT ID — no Cloudflare account id; set to issue R2 temporary credentials instead of STS. \n NODUS OBJSTORE R2 API TOKEN — no Cloudflare API token allowed to create R2 temporary credentials. \n NODUS OBJSTORE REGION auto no S3 signing region; R2 uses auto. \n NODUS OBJSTORE SECRET ACCESS KEY — no Secret access key. \n NODUS OPENRELAY API KEY — no Dedicated OpenRelay inference key. When nonempty, overrides only openrelay key in NODUS INFERENCE UPSTREAMS; empty preserves the composite key and all other upstream settings. Roles: controller, inference. \n NODUS OPS ACCESS AUDIENCE — no Deprecated; accepted for rollout compatibility and unused by admin authentication. Roles: api. \n NODUS OPS ACCESS TEAM DOMAIN — no Deprecated; accepted for rollout compatibility and unused by admin authentication. Roles: api. \n NODUS OPS ADDR :8081 no Ops API listener. Roles: api. \n NODUS OPS CONSOLE ORIGIN — no Admin console origin ( .) allowed to call the ops API from the browser; empty allows no cross-origin caller. Roles: api. \n NODUS OPS SUPABASE ISSUER — no Supabase Auth issuer (https\\://.supabase.co/auth/v1) whose verified sessions are checked against the fixed admin email allowlist; empty refuses every ops request. Roles: api. \n NODUS POOLS — no BYOC pools and cloud accounts as JSON: install script url ( ./install/nodusd.sh), node gateway url (wss\\://nodes., the ADR-123 carrier), api url and inference url (workload proxy origins; default to deployed release origins), agent url (ARCH becomes amd64 or arm64), agent sha256 ({arch: sha256}), agent signer identity (cosign certificate identity regexp), aws observer role arn, aws template url, google client id, google client secret, google redirect url, consent key (32+ characters), console url. The installer keys replace the one the API otherwise serves from the deployed nodusd release (NODUS NODUSD VERSION); without either, EnrollmentTokens are refused, and without a cloud’s keys its onboarding is off. Roles: controller. \n NODUS POSTHOG READ API KEY — no Project-scoped PostHog query:read credential for internal aggregate reports; empty leaves analytics unavailable. Roles: api. \n NODUS PROVIDER SECRETS — no Where provider account secrets live: ssm:/nodus-platform//app/providers/ (SSM SecureString parameters, ADR-122) or secretsmanager: (a bare prefix also means Secrets Manager); each account’s secretRef is appended. Required outside local. Roles: controller. \n NODUS PUBLIC API URL — no Public origin of the API ( .), which hosted MCP’s protected-resource metadata and 401 challenge name (RFC 9728); empty uses : under NODUS ENV=local and leaves /mcp unmounted elsewhere. Roles: api, gateway. \n NODUS REGISTRY HOST — no Host of the Nodus registry, registry. (localhost:5000 in compose); empty disables image builds, mirrors and registry tokens. Roles: api, controller. \n NODUS REGISTRY SIGNING KEY — no P-256 private key, PEM (SEC 1 or PKCS #8), that signs registry bearer tokens; the registry trusts its public key. Roles: api, controller. \n NODUS RESEND API KEY — no Resend API key; empty logs each email’s template and recipient count instead of sending, which only NODUS ENV=local allows. Roles: api, controller. \n NODUS SENTRY DSN — no Sentry DSN; empty disables error reporting. \n NODUS STRIPE API BASE — no Stripe API base URL override; the compose stack points it at stripe-mock. Empty uses api.stripe.com. \n NODUS STRIPE LIVE false no Whether this deployment takes live-mode payments; events whose livemode differs are rejected. \n NODUS STRIPE SECRET KEY — no Stripe secret or restricted key (sk test\\ /rk test\\ outside prod), agreeing with NODUS STRIPE LIVE; empty turns payments off. \n NODUS STRIPE WEBHOOK SECRET — no Signing secret (whsec\\ …) of the POST /webhooks/stripe endpoint; empty refuses every delivery. \n NODUS STRIPE WEBHOOK URL — no Public URL of POST /webhooks/stripe that nodus-server seed registers as the Stripe webhook endpoint; empty leaves the endpoint alone. \n NODUS TOKEN PEPPERS — no Token HMAC peppers as :, comma-separated, KMS-encrypted in the cloud; the highest version signs new tokens. Empty uses a fixed development pepper when NODUS ENV=local. \n NODUS USERCONTENT DOMAIN — no Registrable domain of preview URLs and Workspace browser tools, https\\://-. (ADR-020); never the API’s domain, and one Nodus owns, since it receives each exchange token. Unset outside NODUS ENV=local turns the browser tools and previews off; NODUS ENV=local defaults it to nodus-usercontent.net. Roles: controller, gateway. \n NODUS WATCH DATABASE URL — no Postgres URL of the watch tailer’s one connection, logged in as nodus watch (data-model §12); empty uses NODUS DATABASE URL, whose login must then bypass row-level security (compose). Roles: api, gateway. \n NODUS WEBHOOKS ALLOW PREFIXES — no Address prefixes the SSRF dialer would refuse but deliveries may reach, comma-separated; only NODUS ENV=local allows any (compose receivers). Roles: controller. \n NODUS WORKSPACE EDGE — no Optional JSON browser ingress: domain (account.workers.dev), backend origin (dedicated HTTPS gateway origin), key (32 random bytes, standard base64). Shared only by controller and gateway. Roles: controller, gateway. \n NODUS WORKSPACE EDGE ACCOUNT — no Optional JSON browser endpoint provisioning: account id and api token with Workers Scripts edit. Controller only; required with WORKSPACE EDGE. Roles: controller. \n\nnodusd (node agent)\n\n Variable Default Required Description \n - - - - \n NODUS GATEWAY URL — yes Node protocol endpoint the agent dials, such as .. \n NODUS STATE DIR /var/lib/nodus no Agent state directory: credentials, ring buffers, re-adoption records. \n NODUS CONTAINERD ADDRESS /run/containerd/containerd.sock no containerd socket. \n NODUS LOG LEVEL info no Minimum log level: debug, info, warn or error. \n NODUS SENTRY DSN — no Sentry DSN; empty disables error reporting. \n\nnodus CLI and SDK (client environment, ADR-069)\n\n Variable Default Required Description \n - - - - \n NODUS API KEY — no API key (nodus sk\\ …); overrides the stored login. \n NODUS API URL — no API base URL, such as .. \n NODUS BASE URL — no Deprecated alias of NODUS API URL, accepted with a warning for one major version. \n NODUS ORG — no Organization name or id. \n NODUS PROJECT — no Project name. \n NODUS CONTEXT — no Named context from the config file. \n NODUS CONFIG — no Config file path."},{"id":"docs/reference/troubleshooting","url":"https://nodus-platform-site.pages.dev/docs/reference/troubleshooting/","markdown":"https://nodus-platform-site.pages.dev/docs/source/reference/troubleshooting.md","title":"Troubleshooting","description":"What to check when sign-in, a launch, a running job, outputs or billing do not behave as you expect.","stage":"GA","skill":"nodus-troubleshooting","headings":[{"depth":2,"slug":"signing-in","text":"Signing in"},{"depth":2,"slug":"a-launch-is-refused","text":"A launch is refused"},{"depth":2,"slug":"a-job-does-not-start","text":"A job does not start"},{"depth":2,"slug":"a-job-fails-or-stops","text":"A job fails or stops"},{"depth":2,"slug":"outputs-and-checkpoints","text":"Outputs and checkpoints"},{"depth":2,"slug":"billing","text":"Billing"}],"text":"Troubleshooting\n\n What to check when sign-in, a launch, a running job, outputs or billing do not behave as you expect.\n\nSource: https://nodus-platform-site.pages.dev/docs/reference/troubleshooting/\nBuild revision: 211ad9f836655b1c3a2668c4693e442471f28614\n\nStart with the object itself. nodus describe / shows its phase, conditions, recent events, attempts and cost so far, and every error names a code with a fix and a docs link. When you contact support, include the requestId from the error.\n\nSigning in\n\n The browser never opens. Run nodus login --device and approve the code from any device.\n A CI machine needs a key. Create an API key in the console (or a ServiceAccount key) and pipe it in: echo \"$NODUS API KEY\" nodus login --with-token . NODUS API KEY alone also works.\n The wrong org or project. nodus whoami prints the principal, org, project and scopes in use. nodus config get-contexts lists every org you signed in to and nodus config use-context switches between them.\n\nA launch is refused\n\n Code Meaning What to do \n - - - \n InsufficientCredits (402) The hold for this launch is larger than your available credit Top up with nodus billing top-up , or lower --max-cost \n BudgetExceeded (402) A budget that covers this object would be exceeded Raise the budget or launch in a project it does not cover \n QuotaExceeded (429) An org or project quota is reached nodus get quotas ; starter limits lift at your first purchase \n CapacityUnavailable (503) No offering matches the requirements right now Allow more accelerators or regions, or wait for the ETA in the message \n Invalid (422) The spec fails validation The message names each field; nodus explain job.spec documents them \n\nEvery code has its own page under Error codes.\n\nA job does not start\n\n Pending for a long time. nodus describe job/ shows why under Conditions and Events, including the offerings considered and why others were rejected.\n The image fails to pull. Check the image reference and, for a private registry, the pull Secret it names.\n\nA job fails or stops\n\n The command exits non-zero. nodus logs job/ shows your program’s output; the phase is Failed with the exit code. nodus run exits with the same code.\n It stopped when credit ran out. A running job stops gracefully inside its reserved amount and reports it in its conditions. Add credit and it resumes from its last checkpoint.\n The CLI itself failed. Exit code 125 is a Nodus or API error and 124 a timeout, never your program’s.\n\nOutputs and checkpoints\n\n An output is missing. Only paths declared with --output NAME=/path (or spec.outputs ) are collected. List them with nodus get job/ -o jsonpath='{.status.outputs}' .\n A resumed job started from scratch. Nodus restores the files your program saved in its checkpoint directory ( NODUS CHECKPOINT DIR ); your program must load them on start. Restoring files does not restore process memory.\n\nBilling\n\n A charge looks higher than the run time. Rented machines bill from creation to confirmed deletion, including start-up and shutdown. nodus get usage --field-selector object.name= --group-by segment itemizes it.\n Where did my credit go? nodus billing shows available credit, open holds and budgets; nodus get usage --group-by project breaks spending down."}]}