Agents
View MarkdownBeta: this feature may change before general availability
An Agent is a definition: a system prompt, the Claude access it runs on and a cost cap. An AgentRun is one
conversation with it. Each run works in its own sandbox, where the model runs shell commands to check its work, and
it survives restarts of Nodus by replaying recorded results. An interrupted command or model call using your own
key parks with NeedsResolution when its outcome is uncertain; inspect its effects before starting replacement work.
Agents, parallel AgentGroups and Environment evaluation batches are in Beta. Each run is one agent conversation.
Sign in and create an agent
Section titled “Sign in and create an agent”apiVersion: nodus.dev/v1kind: Agentmetadata: name: hello-agentspec: system: | You are a careful assistant. Run shell commands in your sandbox to check your work, and answer in one short paragraph. # Nodus's Claude: each prompt is routed to Haiku, Sonnet, Opus or Fable, and billed from your credits. model: access: Nodus perRunMaxCostUSD: "0.50"nodus loginnodus apply -f agent.yamlThe agent uses Nodus’s Claude. Every prompt goes to a small router that picks the right Claude model for it (Haiku
for simple requests up to Fable for demanding ones). That model answers the whole prompt, including its commands,
so the model never changes in the middle of a turn. If the router is unavailable, Sonnet answers. Restrict the
choice with spec.model.families, for example [haiku, sonnet].
Start a run
Section titled “Start a run”apiVersion: nodus.dev/v1kind: AgentRunmetadata: name: hello-agent-runspec: agent: hello-agent input: "Use the shell to print the Python version, then tell me which one it is."nodus apply -f run.yamlnodus wait agentrun/hello-agent-run --for=jsonpath='{.status.phase}'=Succeeded --timeout=8mnodus get agentrun/hello-agent-run -o yamlOr without a file: nodus create agentrun triage --agent hello-agent --prompt "Summarize the logs". To start
from Nodus’s ready-made assistant, name the template claude-assistant.
status.turns lists each prompt with the model that served it, how it was chosen (jev or fallback) and its
cost. status.answer.preview holds the first 4 KiB of the final answer; nodus agentrun answer prints all of it.
nodus agentrun answer hello-agent-runnodus agentrun steps hello-agent-runsteps lists what the run did in order: creating its sandbox, each routing and model call, each command. A step
that shows Completed is never repeated, even after a restart.
The console’s Agents page shows the same: each run with the model of its latest turn and its cost, and on a run every turn with the model that served it, how that model was chosen and what the turn cost, then the answer and the recorded steps.
From Python
Section titled “From Python”import nodus
agent = nodus.ClaudeAgent("helper", system="You are careful.", families=["haiku", "sonnet"])print(agent.remote("Use the shell to print the Python version")) # runs to the end, returns the answer
run = agent.submit("Summarize the logs", keep_alive=True) # a handle: run.answer(), run.steps(), run.cancel()run.send("message", "Now the staging logs")ClaudeAgent creates the agent on first use. Pass api_key_secret="anthropic-key" to use your own Anthropic key,
or ClaudeAgent.from_name("claude-assistant", project="nodus") to run the ready-made template. To run many at once
from Python, see Run agents in parallel.
Keep a run open for follow-ups
Section titled “Keep a run open for follow-ups”Set spec.keepAlive: true and the run waits for messages after each turn instead of finishing. While it waits it
uses no model and keeps its sandbox, which can still accrue compute charges. Set spec.deadline to bound the run’s
lifetime. The deadline also applies while waiting for a message: the run fails with DeadlineExceeded and its
sandbox is deleted without another model call.
nodus apply -f run.yaml # with keepAlive: truenodus agentrun send hello-agent-run "Now do the same for the staging logs" --key staging-1Sending the same --key again delivers the message once. A run that is still working reads the message after its
current turn; a finished run refuses it with AgentRunFinished. Cancel a run with nodus cancel agentrun/NAME;
its sandbox is deleted.
The same calls are REST: POST …/agentruns/NAME/messages with {"payload": "…", "messageKey": "…"}, and GET
on …/steps and …/answer.
Run agents in parallel
Section titled “Run agents in parallel”An AgentGroup runs many runs of one agent at once. You choose how many run at the same time and how much the whole group may spend. The other runs wait their turn, and one command cancels or deletes all of them. Create the agent and the group, then one run for each task:
apiVersion: nodus.dev/v1kind: Agentmetadata: name: parallel-agents-workerspec: system: | Follow the task's requested answer format; otherwise answer in one short sentence. # Nodus's Claude, limited to Haiku so each run stays far below its cap. model: access: Nodus families: [haiku] maxTokens: 1024 perRunMaxCostUSD: "0.10"apiVersion: nodus.dev/v1kind: AgentGroupmetadata: name: parallel-agentsspec: agent: parallel-agents-worker limits: maxActive: 2 # at most two runs go at once maxCostUSD: "0.60" # a run starts only while its own cap still fits under what is leftapiVersion: nodus.dev/v1kind: AgentRunmetadata: name: parallel-agents-aspec: agent: parallel-agents-worker group: parallel-agents taskKey: a input: "What is the capital of France?"---apiVersion: nodus.dev/v1kind: AgentRunmetadata: name: parallel-agents-bspec: agent: parallel-agents-worker group: parallel-agents taskKey: b input: "What is the capital of Japan?"---apiVersion: nodus.dev/v1kind: AgentRunmetadata: name: parallel-agents-cspec: agent: parallel-agents-worker group: parallel-agents taskKey: c dependsOn: [a, b] # c waits until a and b have succeeded input: "Say in one sentence that both questions have been answered."nodus apply -f agent.yamlnodus apply -f group.yamlnodus apply -f runs.yamlnodus patch ag/parallel-agents --patch '{"spec":{"sealed":true}}'Or without a file: nodus create agentgroup parallel-agents --agent parallel-agents-worker --max-active 2 --max-cost 0.60.
Each run is a task of the group: spec.group names the group and spec.taskKey names the task (a DNS label that is
unique in the group). The run is named <group>-<taskKey>; leave metadata.name out or set exactly that. The run
uses the group’s agent, so spec.agent must match it. The last command seals the group, which says that no more runs
are coming (see below).
Watch progress and read the answers
Section titled “Watch progress and read the answers”nodus wait ag/parallel-agents --for=jsonpath='{.status.phase}'=Succeeded --timeout=12mnodus get agnodus get ar --field-selector spec.group=parallel-agentsnodus agentrun answer parallel-agents-cnodus get ag lists each group with its phase, RUNS (finished successfully out of all) and ACTIVE runs, and
its cost. status.counts splits the runs into queued, active, waiting (for a message or for funds),
succeeded, failed and cancelled. status.blockedReasons says why queued runs have not started and how many
wait for each reason: DependencyWait, MaxActiveReached or MaxCostReached. A queued run’s own status.reason
is DependencyWait, or GroupWait while it waits for a free slot or for room under the cost cap. The runs you list
are ordinary runs, so status.answer, nodus agentrun answer and nodus agentrun steps work on each of them.
A group is Running until it is sealed and every run has finished. It then becomes Succeeded, or Failed with
reason RunsFailed when a run failed or was cancelled because a dependency failed. A sealed group with no runs
succeeds.
How many run at once
Section titled “How many run at once”spec.limits.maxActive is how many runs the group lets run at the same time (default 10, 1 to 1000). Change it on a
live group:
nodus patch ag/parallel-agents --patch '{"spec":{"limits":{"maxActive":4}}}'spec.limits.maxPending is how many unfinished runs the group takes in (default 10000); a run beyond it is refused.
Dependencies
Section titled “Dependencies”spec.dependsOn lists the task keys a run waits for. It starts only after every one of them has succeeded, and
until then it is queued with reason DependencyWait. A dependency must already be in the group, so create the runs
it depends on first (earlier in the same file works). When a dependency fails, the run is cancelled with reason
DependencyFailed; when a dependency was cancelled, the run is cancelled with reason DependencyCancelled.
The group cost cap
Section titled “The group cost cap”spec.maxCostUSD is one cap for the whole group. A run starts only when its own cap (the agent’s
perRunMaxCostUSD) still fits under what is left: the cap minus what the group has spent and minus the caps of the
runs that have started and not finished. The members together can therefore never spend more than the cap, and runs
that do not fit yet wait with MaxCostReached. In the example, the $0.60 cap has room for the caps of six runs at
$0.10 each before any run has spent anything. Raise the cap on a live group with nodus patch; it cannot be lowered.
If the cap is smaller than one run’s cap, no run starts: raise it or lower the agent’s perRunMaxCostUSD.
status.cost shows totalUSD and limitUSD. The cap covers the runs’ Claude usage (model and routing calls), the
spend each run’s perRunMaxCostUSD limits; the sandbox each run works in is billed on its own and is not counted
against either cap.
Seal, cancel and delete
Section titled “Seal, cancel and delete”A group takes new runs until you seal it with spec.sealed: true. Sealing cannot be undone, and a sealed group
refuses new runs with a conflict. A group finishes only once it is sealed, so seal it after you create the last run.
nodus cancel ag/parallel-agentsnodus delete ag/parallel-agentsnodus cancel cancels every run that has not finished and leaves the finished runs, with their answers, as they are.
status.phase goes through Cancelling to Cancelled. nodus delete removes the group together with all of its
runs.
Evaluate an agent
Section titled “Evaluate an agent”Create an evaluation group with an Environment at an explicit version. Nodus generates the fixed task batch, runs it through ordinary AgentRuns, and grades completed answers using the Environment’s isolated grader. There is no manual submit or seal step. Set a project Budget to limit total spend: the group cap covers model usage, while agent and grading Sandbox compute is billed separately.
After deploying the example’s parallel-agents-worker Agent above, run this Python example. It creates two
held-out tasks, waits for a terminal result, prints per-case outcomes and deletes the evaluation group afterward.
"""Evaluate the deployed example agent on a fixed, versioned task batch."""
import nodus
def main() -> None: # Model caps exclude Sandbox compute; set a project Budget before running this example. evaluation = nodus.AgentGroup.create( "parallel-agents-eval", agent="parallel-agents-worker", max_active=2, max_cost="0.20", evaluation={ "environment": "nodus/arithmetic-v2@2.0.0", "split": "test", "tasks": 2, "repetitions": 1, "seed": 42, "timeout": "10m", }, ) try: status = evaluation.wait(timeout=660) results = evaluation.results() print( f"{status.phase}: cases={len(results)}, pending={status.evaluation.pending}" ) print("pass rate:", status.evaluation.get("passRate")) print("mean reward:", status.evaluation.get("meanReward")) for result in results: print(result.task_id, result.state, result.get("verdict")) finally: evaluation.delete()
if __name__ == "__main__": main()The batch is at most 100 cases including repetitions. Every member pins the same Agent revision; results record the Environment version, resolved image digest, split and seed. Only task prompts reach the agent, never hidden answers. The timeout defaults to 30 minutes, may be 1 minute to 24 hours, and includes manifest preparation, queueing and funding waits from group creation. Each member receives that absolute deadline. Grading also stops at the deadline; completed answers that could not be scored are shown as evaluation timeouts, excluded from pass rate and mean reward.
Open the group in the console to see scored cases, pass rate, mean reward, pending cases and execution/grading failures separately. The Results view links each case to its AgentRun and shows grader evidence. A failed run, missing answer or grader failure is unscored and does not become a zero reward. A wrong or invalid answer is a scored result. Per-case records share the member journal’s 30-day payload retention; deleting the group removes its runs and records. Canceling or finishing the group cleans up its grading pool.
Use your own Anthropic key
Section titled “Use your own Anthropic key”Store the key in a Secret under the key ANTHROPIC_API_KEY and point the agent at it. Nodus then charges nothing for the model, and your Anthropic
account pays for it.
spec: model: access: BYOK name: claude-sonnet-5-5 apiKeySecret: anthropic-keyWhat a run costs
Section titled “What a run costs”On Nodus’s Claude each model call, and the routing call before each turn, is billed from your credits at cost
divided by 0.875. spec.perRunMaxCostUSD (default $1.00) caps what a run spends on the model. Before each call
the run checks that the call, at its largest (the whole conversation as input and spec.maxTokens of output), still
fits under the cap. When it does not, the run waits (reason: MaxCostReached) and makes no more calls. The cap is
copied from the agent when the run is created, so changing the agent does not change a run that has started: to
continue the work, cancel the run and start a new one. If your credits run out, the run waits instead and
continues when you add credits. A run on your own Anthropic key is not charged by Nodus and has no cap to reach.
Sandbox compute is billed as for any sandbox. status.cost shows the model spend so far. Runs in a group also
share the group’s cap (see Run agents in parallel).
Limits
Section titled “Limits”- One run works one prompt at a time, with at most
spec.maxToolCallscommands per turn (default 50). - The input and each message are at most 256 KiB. A run’s recorded payloads are kept 30 days after it ends.
- A run created for a stopped agent (
spec.state: Stopped) is refused withAgentStopped. - A group takes at most 10000 unfinished runs. Create runs in batches of up to 100 by posting an
AgentRunListto…/agentrunswith anIdempotency-Keyheader; a batch is all-or-nothing, and its runs are created in list order, so list each run after the runs it depends on.
See Durable execution for what survives a restart.