Skip to content

Beta: this feature may change before general availability

An Agent is a definition: a system prompt, the Claude access it runs on and a cost cap. An AgentRun is one conversation with it. Each run works in its own sandbox, where the model runs shell commands to check its work, and it survives restarts of Nodus by replaying recorded results. An interrupted command or model call using your own key parks with NeedsResolution when its outcome is uncertain; inspect its effects before starting replacement work.

Agents, parallel AgentGroups and Environment evaluation batches are in Beta. Each run is one agent conversation.

agent.yaml
apiVersion: nodus.dev/v1
kind: Agent
metadata:
name: hello-agent
spec:
system: |
You are a careful assistant. Run shell commands in your sandbox to check your work, and answer in one short
paragraph.
# Nodus's Claude: each prompt is routed to Haiku, Sonnet, Opus or Fable, and billed from your credits.
model:
access: Nodus
perRunMaxCostUSD: "0.50"
Terminal window
nodus login
nodus apply -f agent.yaml

The agent uses Nodus’s Claude. Every prompt goes to a small router that picks the right Claude model for it (Haiku for simple requests up to Fable for demanding ones). That model answers the whole prompt, including its commands, so the model never changes in the middle of a turn. If the router is unavailable, Sonnet answers. Restrict the choice with spec.model.families, for example [haiku, sonnet].

run.yaml
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
name: hello-agent-run
spec:
agent: hello-agent
input: "Use the shell to print the Python version, then tell me which one it is."
Terminal window
nodus apply -f run.yaml
nodus wait agentrun/hello-agent-run --for=jsonpath='{.status.phase}'=Succeeded --timeout=8m
nodus get agentrun/hello-agent-run -o yaml

Or without a file: nodus create agentrun triage --agent hello-agent --prompt "Summarize the logs". To start from Nodus’s ready-made assistant, name the template claude-assistant.

status.turns lists each prompt with the model that served it, how it was chosen (jev or fallback) and its cost. status.answer.preview holds the first 4 KiB of the final answer; nodus agentrun answer prints all of it.

Terminal window
nodus agentrun answer hello-agent-run
nodus agentrun steps hello-agent-run

steps lists what the run did in order: creating its sandbox, each routing and model call, each command. A step that shows Completed is never repeated, even after a restart.

The console’s Agents page shows the same: each run with the model of its latest turn and its cost, and on a run every turn with the model that served it, how that model was chosen and what the turn cost, then the answer and the recorded steps.

import nodus
agent = nodus.ClaudeAgent("helper", system="You are careful.", families=["haiku", "sonnet"])
print(agent.remote("Use the shell to print the Python version")) # runs to the end, returns the answer
run = agent.submit("Summarize the logs", keep_alive=True) # a handle: run.answer(), run.steps(), run.cancel()
run.send("message", "Now the staging logs")

ClaudeAgent creates the agent on first use. Pass api_key_secret="anthropic-key" to use your own Anthropic key, or ClaudeAgent.from_name("claude-assistant", project="nodus") to run the ready-made template. To run many at once from Python, see Run agents in parallel.

Set spec.keepAlive: true and the run waits for messages after each turn instead of finishing. While it waits it uses no model and keeps its sandbox, which can still accrue compute charges. Set spec.deadline to bound the run’s lifetime. The deadline also applies while waiting for a message: the run fails with DeadlineExceeded and its sandbox is deleted without another model call.

Terminal window
nodus apply -f run.yaml # with keepAlive: true
nodus agentrun send hello-agent-run "Now do the same for the staging logs" --key staging-1

Sending the same --key again delivers the message once. A run that is still working reads the message after its current turn; a finished run refuses it with AgentRunFinished. Cancel a run with nodus cancel agentrun/NAME; its sandbox is deleted.

The same calls are REST: POST …/agentruns/NAME/messages with {"payload": "…", "messageKey": "…"}, and GET on …/steps and …/answer.

An AgentGroup runs many runs of one agent at once. You choose how many run at the same time and how much the whole group may spend. The other runs wait their turn, and one command cancels or deletes all of them. Create the agent and the group, then one run for each task:

agent.yaml
apiVersion: nodus.dev/v1
kind: Agent
metadata:
name: parallel-agents-worker
spec:
system: |
Follow the task's requested answer format; otherwise answer in one short sentence.
# Nodus's Claude, limited to Haiku so each run stays far below its cap.
model:
access: Nodus
families: [haiku]
maxTokens: 1024
perRunMaxCostUSD: "0.10"
group.yaml
apiVersion: nodus.dev/v1
kind: AgentGroup
metadata:
name: parallel-agents
spec:
agent: parallel-agents-worker
limits:
maxActive: 2 # at most two runs go at once
maxCostUSD: "0.60" # a run starts only while its own cap still fits under what is left
runs.yaml
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
name: parallel-agents-a
spec:
agent: parallel-agents-worker
group: parallel-agents
taskKey: a
input: "What is the capital of France?"
---
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
name: parallel-agents-b
spec:
agent: parallel-agents-worker
group: parallel-agents
taskKey: b
input: "What is the capital of Japan?"
---
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
name: parallel-agents-c
spec:
agent: parallel-agents-worker
group: parallel-agents
taskKey: c
dependsOn: [a, b] # c waits until a and b have succeeded
input: "Say in one sentence that both questions have been answered."
Terminal window
nodus apply -f agent.yaml
nodus apply -f group.yaml
nodus apply -f runs.yaml
nodus patch ag/parallel-agents --patch '{"spec":{"sealed":true}}'

Or without a file: nodus create agentgroup parallel-agents --agent parallel-agents-worker --max-active 2 --max-cost 0.60.

Each run is a task of the group: spec.group names the group and spec.taskKey names the task (a DNS label that is unique in the group). The run is named <group>-<taskKey>; leave metadata.name out or set exactly that. The run uses the group’s agent, so spec.agent must match it. The last command seals the group, which says that no more runs are coming (see below).

Terminal window
nodus wait ag/parallel-agents --for=jsonpath='{.status.phase}'=Succeeded --timeout=12m
nodus get ag
nodus get ar --field-selector spec.group=parallel-agents
nodus agentrun answer parallel-agents-c

nodus get ag lists each group with its phase, RUNS (finished successfully out of all) and ACTIVE runs, and its cost. status.counts splits the runs into queued, active, waiting (for a message or for funds), succeeded, failed and cancelled. status.blockedReasons says why queued runs have not started and how many wait for each reason: DependencyWait, MaxActiveReached or MaxCostReached. A queued run’s own status.reason is DependencyWait, or GroupWait while it waits for a free slot or for room under the cost cap. The runs you list are ordinary runs, so status.answer, nodus agentrun answer and nodus agentrun steps work on each of them.

A group is Running until it is sealed and every run has finished. It then becomes Succeeded, or Failed with reason RunsFailed when a run failed or was cancelled because a dependency failed. A sealed group with no runs succeeds.

spec.limits.maxActive is how many runs the group lets run at the same time (default 10, 1 to 1000). Change it on a live group:

Terminal window
nodus patch ag/parallel-agents --patch '{"spec":{"limits":{"maxActive":4}}}'

spec.limits.maxPending is how many unfinished runs the group takes in (default 10000); a run beyond it is refused.

spec.dependsOn lists the task keys a run waits for. It starts only after every one of them has succeeded, and until then it is queued with reason DependencyWait. A dependency must already be in the group, so create the runs it depends on first (earlier in the same file works). When a dependency fails, the run is cancelled with reason DependencyFailed; when a dependency was cancelled, the run is cancelled with reason DependencyCancelled.

spec.maxCostUSD is one cap for the whole group. A run starts only when its own cap (the agent’s perRunMaxCostUSD) still fits under what is left: the cap minus what the group has spent and minus the caps of the runs that have started and not finished. The members together can therefore never spend more than the cap, and runs that do not fit yet wait with MaxCostReached. In the example, the $0.60 cap has room for the caps of six runs at $0.10 each before any run has spent anything. Raise the cap on a live group with nodus patch; it cannot be lowered. If the cap is smaller than one run’s cap, no run starts: raise it or lower the agent’s perRunMaxCostUSD. status.cost shows totalUSD and limitUSD. The cap covers the runs’ Claude usage (model and routing calls), the spend each run’s perRunMaxCostUSD limits; the sandbox each run works in is billed on its own and is not counted against either cap.

A group takes new runs until you seal it with spec.sealed: true. Sealing cannot be undone, and a sealed group refuses new runs with a conflict. A group finishes only once it is sealed, so seal it after you create the last run.

Terminal window
nodus cancel ag/parallel-agents
nodus delete ag/parallel-agents

nodus cancel cancels every run that has not finished and leaves the finished runs, with their answers, as they are. status.phase goes through Cancelling to Cancelled. nodus delete removes the group together with all of its runs.

Create an evaluation group with an Environment at an explicit version. Nodus generates the fixed task batch, runs it through ordinary AgentRuns, and grades completed answers using the Environment’s isolated grader. There is no manual submit or seal step. Set a project Budget to limit total spend: the group cap covers model usage, while agent and grading Sandbox compute is billed separately.

After deploying the example’s parallel-agents-worker Agent above, run this Python example. It creates two held-out tasks, waits for a terminal result, prints per-case outcomes and deletes the evaluation group afterward.

evaluation.py
"""Evaluate the deployed example agent on a fixed, versioned task batch."""
import nodus
def main() -> None:
# Model caps exclude Sandbox compute; set a project Budget before running this example.
evaluation = nodus.AgentGroup.create(
"parallel-agents-eval",
agent="parallel-agents-worker",
max_active=2,
max_cost="0.20",
evaluation={
"environment": "nodus/arithmetic-v2@2.0.0",
"split": "test",
"tasks": 2,
"repetitions": 1,
"seed": 42,
"timeout": "10m",
},
)
try:
status = evaluation.wait(timeout=660)
results = evaluation.results()
print(
f"{status.phase}: cases={len(results)}, pending={status.evaluation.pending}"
)
print("pass rate:", status.evaluation.get("passRate"))
print("mean reward:", status.evaluation.get("meanReward"))
for result in results:
print(result.task_id, result.state, result.get("verdict"))
finally:
evaluation.delete()
if __name__ == "__main__":
main()

The batch is at most 100 cases including repetitions. Every member pins the same Agent revision; results record the Environment version, resolved image digest, split and seed. Only task prompts reach the agent, never hidden answers. The timeout defaults to 30 minutes, may be 1 minute to 24 hours, and includes manifest preparation, queueing and funding waits from group creation. Each member receives that absolute deadline. Grading also stops at the deadline; completed answers that could not be scored are shown as evaluation timeouts, excluded from pass rate and mean reward.

Open the group in the console to see scored cases, pass rate, mean reward, pending cases and execution/grading failures separately. The Results view links each case to its AgentRun and shows grader evidence. A failed run, missing answer or grader failure is unscored and does not become a zero reward. A wrong or invalid answer is a scored result. Per-case records share the member journal’s 30-day payload retention; deleting the group removes its runs and records. Canceling or finishing the group cleans up its grading pool.

Store the key in a Secret under the key ANTHROPIC_API_KEY and point the agent at it. Nodus then charges nothing for the model, and your Anthropic account pays for it.

spec:
model:
access: BYOK
name: claude-sonnet-5-5
apiKeySecret: anthropic-key

On Nodus’s Claude each model call, and the routing call before each turn, is billed from your credits at cost divided by 0.875. spec.perRunMaxCostUSD (default $1.00) caps what a run spends on the model. Before each call the run checks that the call, at its largest (the whole conversation as input and spec.maxTokens of output), still fits under the cap. When it does not, the run waits (reason: MaxCostReached) and makes no more calls. The cap is copied from the agent when the run is created, so changing the agent does not change a run that has started: to continue the work, cancel the run and start a new one. If your credits run out, the run waits instead and continues when you add credits. A run on your own Anthropic key is not charged by Nodus and has no cap to reach. Sandbox compute is billed as for any sandbox. status.cost shows the model spend so far. Runs in a group also share the group’s cap (see Run agents in parallel).

  • One run works one prompt at a time, with at most spec.maxToolCalls commands per turn (default 50).
  • The input and each message are at most 256 KiB. A run’s recorded payloads are kept 30 days after it ends.
  • A run created for a stopped agent (spec.state: Stopped) is refused with AgentStopped.
  • A group takes at most 10000 unfinished runs. Create runs in batches of up to 100 by posting an AgentRunList to …/agentruns with an Idempotency-Key header; a batch is all-or-nothing, and its runs are created in list order, so list each run after the runs it depends on.

See Durable execution for what survives a restart.