# Agents

> Run an agent on Claude in its own sandbox, send it follow-up prompts, and pay for the model from your credits.

Source: https://nodus-platform-site.pages.dev/docs/guides/agents/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

**Beta:** this feature may change.

An **Agent** is a definition: a system prompt, the Claude access it runs on and a cost cap. An **AgentRun** is one conversation with it. Each run works in its own sandbox, where the model runs shell commands to check its work, and it survives restarts of Nodus by replaying recorded results. An interrupted command or model call using your own key parks with `NeedsResolution` when its outcome is uncertain; inspect its effects before starting replacement work.

Agents, parallel AgentGroups and Environment evaluation batches are in Beta. Each run is one agent conversation.

## Sign in and create an agent

agent.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Agent
metadata:
  name: hello-agent
spec:
  system: |
    You are a careful assistant. Run shell commands in your sandbox to check your work, and answer in one short
    paragraph.
  # Nodus's Claude: each prompt is routed to Haiku, Sonnet, Opus or Fable, and billed from your credits.
  model:
    access: Nodus
  perRunMaxCostUSD: "0.50"
```

Terminal window

```bash
nodus login
nodus apply -f agent.yaml
```

The agent uses Nodus’s Claude. Every prompt goes to a small router that picks the right Claude model for it (Haiku for simple requests up to Fable for demanding ones). That model answers the whole prompt, including its commands, so the model never changes in the middle of a turn. If the router is unavailable, Sonnet answers. Restrict the choice with `spec.model.families`, for example `[haiku, sonnet]`.

## Start a run

run.yaml

```yaml
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
  name: hello-agent-run
spec:
  agent: hello-agent
  input: "Use the shell to print the Python version, then tell me which one it is."
```

Terminal window

```bash
nodus apply -f run.yaml
nodus wait agentrun/hello-agent-run --for=jsonpath='{.status.phase}'=Succeeded --timeout=8m
nodus get agentrun/hello-agent-run -o yaml
```

Or without a file: `nodus create agentrun triage --agent hello-agent --prompt "Summarize the logs"`. To start from Nodus’s ready-made assistant, name the template `claude-assistant`.

`status.turns` lists each prompt with the model that served it, how it was chosen (`jev` or `fallback`) and its cost. `status.answer.preview` holds the first 4 KiB of the final answer; `nodus agentrun answer` prints all of it.

Terminal window

```bash
nodus agentrun answer hello-agent-run
nodus agentrun steps hello-agent-run
```

`steps` lists what the run did in order: creating its sandbox, each routing and model call, each command. A step that shows `Completed` is never repeated, even after a restart.

The console’s Agents page shows the same: each run with the model of its latest turn and its cost, and on a run every turn with the model that served it, how that model was chosen and what the turn cost, then the answer and the recorded steps.

## From Python

```python
import nodus

agent = nodus.ClaudeAgent("helper", system="You are careful.", families=["haiku", "sonnet"])
print(agent.remote("Use the shell to print the Python version"))   # runs to the end, returns the answer

run = agent.submit("Summarize the logs", keep_alive=True)          # a handle: run.answer(), run.steps(), run.cancel()
run.send("message", "Now the staging logs")
```

`ClaudeAgent` creates the agent on first use. Pass `api_key_secret="anthropic-key"` to use your own Anthropic key, or `ClaudeAgent.from_name("claude-assistant", project="nodus")` to run the ready-made template. To run many at once from Python, see [Run agents in parallel](https://nodus-platform-site.pages.dev/docs/guides/python/agents/#run-agents-in-parallel).

## Keep a run open for follow-ups

Set `spec.keepAlive: true` and the run waits for messages after each turn instead of finishing. While it waits it uses no model and keeps its sandbox, which can still accrue compute charges. Set `spec.deadline` to bound the run’s lifetime. The deadline also applies while waiting for a message: the run fails with `DeadlineExceeded` and its sandbox is deleted without another model call.

Terminal window

```bash
nodus apply -f run.yaml                                   # with keepAlive: true
nodus agentrun send hello-agent-run "Now do the same for the staging logs" --key staging-1
```

Sending the same `--key` again delivers the message once. A run that is still working reads the message after its current turn; a finished run refuses it with `AgentRunFinished`. Cancel a run with `nodus cancel agentrun/NAME`; its sandbox is deleted.

The same calls are REST: `POST …/agentruns/NAME/messages` with `{"payload": "…", "messageKey": "…"}`, and `GET` on `…/steps` and `…/answer`.

## Run agents in parallel

An **AgentGroup** runs many runs of one agent at once. You choose how many run at the same time and how much the whole group may spend. The other runs wait their turn, and one command cancels or deletes all of them. Create the agent and the group, then one run for each task:

agent.yaml

```yaml
apiVersion: nodus.dev/v1
kind: Agent
metadata:
  name: parallel-agents-worker
spec:
  system: |
    Follow the task's requested answer format; otherwise answer in one short sentence.
  # Nodus's Claude, limited to Haiku so each run stays far below its cap.
  model:
    access: Nodus
    families: [haiku]
  maxTokens: 1024
  perRunMaxCostUSD: "0.10"
```

group.yaml

```yaml
apiVersion: nodus.dev/v1
kind: AgentGroup
metadata:
  name: parallel-agents
spec:
  agent: parallel-agents-worker
  limits:
    maxActive: 2          # at most two runs go at once
  maxCostUSD: "0.60"      # a run starts only while its own cap still fits under what is left
```

runs.yaml

```yaml
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
  name: parallel-agents-a
spec:
  agent: parallel-agents-worker
  group: parallel-agents
  taskKey: a
  input: "What is the capital of France?"
---
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
  name: parallel-agents-b
spec:
  agent: parallel-agents-worker
  group: parallel-agents
  taskKey: b
  input: "What is the capital of Japan?"
---
apiVersion: nodus.dev/v1
kind: AgentRun
metadata:
  name: parallel-agents-c
spec:
  agent: parallel-agents-worker
  group: parallel-agents
  taskKey: c
  dependsOn: [a, b]       # c waits until a and b have succeeded
  input: "Say in one sentence that both questions have been answered."
```

Terminal window

```bash
nodus apply -f agent.yaml
nodus apply -f group.yaml
nodus apply -f runs.yaml
nodus patch ag/parallel-agents --patch '{"spec":{"sealed":true}}'
```

Or without a file: `nodus create agentgroup parallel-agents --agent parallel-agents-worker --max-active 2 --max-cost 0.60`.

Each run is a task of the group: `spec.group` names the group and `spec.taskKey` names the task (a DNS label that is unique in the group). The run is named `<group>-<taskKey>`; leave `metadata.name` out or set exactly that. The run uses the group’s agent, so `spec.agent` must match it. The last command seals the group, which says that no more runs are coming (see below).

### Watch progress and read the answers

Terminal window

```bash
nodus wait ag/parallel-agents --for=jsonpath='{.status.phase}'=Succeeded --timeout=12m
nodus get ag
nodus get ar --field-selector spec.group=parallel-agents
nodus agentrun answer parallel-agents-c
```

`nodus get ag` lists each group with its phase, `RUNS` (finished successfully out of all) and `ACTIVE` runs, and its cost. `status.counts` splits the runs into `queued`, `active`, `waiting` (for a message or for funds), `succeeded`, `failed` and `cancelled`. `status.blockedReasons` says why queued runs have not started and how many wait for each reason: `DependencyWait`, `MaxActiveReached` or `MaxCostReached`. A queued run’s own `status.reason` is `DependencyWait`, or `GroupWait` while it waits for a free slot or for room under the cost cap. The runs you list are ordinary runs, so `status.answer`, `nodus agentrun answer` and `nodus agentrun steps` work on each of them.

A group is `Running` until it is sealed and every run has finished. It then becomes `Succeeded`, or `Failed` with reason `RunsFailed` when a run failed or was cancelled because a dependency failed. A sealed group with no runs succeeds.

### How many run at once

`spec.limits.maxActive` is how many runs the group lets run at the same time (default 10, 1 to 1000). Change it on a live group:

Terminal window

```bash
nodus patch ag/parallel-agents --patch '{"spec":{"limits":{"maxActive":4}}}'
```

`spec.limits.maxPending` is how many unfinished runs the group takes in (default 10000); a run beyond it is refused.

### Dependencies

`spec.dependsOn` lists the task keys a run waits for. It starts only after every one of them has succeeded, and until then it is queued with reason `DependencyWait`. A dependency must already be in the group, so create the runs it depends on first (earlier in the same file works). When a dependency fails, the run is cancelled with reason `DependencyFailed`; when a dependency was cancelled, the run is cancelled with reason `DependencyCancelled`.

### The group cost cap

`spec.maxCostUSD` is one cap for the whole group. A run starts only when its own cap (the agent’s `perRunMaxCostUSD`) still fits under what is left: the cap minus what the group has spent and minus the caps of the runs that have started and not finished. The members together can therefore never spend more than the cap, and runs that do not fit yet wait with `MaxCostReached`. In the example, the $0.60 cap has room for the caps of six runs at $0.10 each before any run has spent anything. Raise the cap on a live group with `nodus patch`; it cannot be lowered. If the cap is smaller than one run’s cap, no run starts: raise it or lower the agent’s `perRunMaxCostUSD`. `status.cost` shows `totalUSD` and `limitUSD`. The cap covers the runs’ Claude usage (model and routing calls), the spend each run’s `perRunMaxCostUSD` limits; the sandbox each run works in is billed on its own and is not counted against either cap.

### Seal, cancel and delete

A group takes new runs until you seal it with `spec.sealed: true`. Sealing cannot be undone, and a sealed group refuses new runs with a conflict. A group finishes only once it is sealed, so seal it after you create the last run.

Terminal window

```bash
nodus cancel ag/parallel-agents
nodus delete ag/parallel-agents
```

`nodus cancel` cancels every run that has not finished and leaves the finished runs, with their answers, as they are. `status.phase` goes through `Cancelling` to `Cancelled`. `nodus delete` removes the group together with all of its runs.

## Evaluate an agent

Create an evaluation group with an Environment at an explicit version. Nodus generates the fixed task batch, runs it through ordinary AgentRuns, and grades completed answers using the Environment’s isolated grader. There is no manual submit or seal step. Set a project Budget to limit total spend: the group cap covers model usage, while agent and grading Sandbox compute is billed separately.

After deploying the example’s `parallel-agents-worker` Agent above, run this Python example. It creates two held-out tasks, waits for a terminal result, prints per-case outcomes and deletes the evaluation group afterward.

evaluation.py

```python
"""Evaluate the deployed example agent on a fixed, versioned task batch."""

import nodus


def main() -> None:
    # Model caps exclude Sandbox compute; set a project Budget before running this example.
    evaluation = nodus.AgentGroup.create(
        "parallel-agents-eval",
        agent="parallel-agents-worker",
        max_active=2,
        max_cost="0.20",
        evaluation={
            "environment": "nodus/arithmetic-v2@2.0.0",
            "split": "test",
            "tasks": 2,
            "repetitions": 1,
            "seed": 42,
            "timeout": "10m",
        },
    )
    try:
        status = evaluation.wait(timeout=660)
        results = evaluation.results()
        print(
            f"{status.phase}: cases={len(results)}, pending={status.evaluation.pending}"
        )
        print("pass rate:", status.evaluation.get("passRate"))
        print("mean reward:", status.evaluation.get("meanReward"))
        for result in results:
            print(result.task_id, result.state, result.get("verdict"))
    finally:
        evaluation.delete()


if __name__ == "__main__":
    main()
```

The batch is at most 100 cases including repetitions. Every member pins the same Agent revision; results record the Environment version, resolved image digest, split and seed. Only task prompts reach the agent, never hidden answers. The timeout defaults to 30 minutes, may be 1 minute to 24 hours, and includes manifest preparation, queueing and funding waits from group creation. Each member receives that absolute deadline. Grading also stops at the deadline; completed answers that could not be scored are shown as evaluation timeouts, excluded from pass rate and mean reward.

Open the group in the console to see scored cases, pass rate, mean reward, pending cases and execution/grading failures separately. The Results view links each case to its AgentRun and shows grader evidence. A failed run, missing answer or grader failure is unscored and does not become a zero reward. A wrong or invalid answer is a scored result. Per-case records share the member journal’s 30-day payload retention; deleting the group removes its runs and records. Canceling or finishing the group cleans up its grading pool.

## Use your own Anthropic key

Store the key in a Secret under the key `ANTHROPIC_API_KEY` and point the agent at it. Nodus then charges nothing for the model, and your Anthropic account pays for it.

```yaml
spec:
  model:
    access: BYOK
    name: claude-sonnet-5-5
    apiKeySecret: anthropic-key
```

## What a run costs

On Nodus’s Claude each model call, and the routing call before each turn, is billed from your credits at cost divided by 0.875. `spec.perRunMaxCostUSD` (default $1.00) caps what a run spends on the model. Before each call the run checks that the call, at its largest (the whole conversation as input and `spec.maxTokens` of output), still fits under the cap. When it does not, the run waits (`reason: MaxCostReached`) and makes no more calls. The cap is copied from the agent when the run is created, so changing the agent does not change a run that has started: to continue the work, cancel the run and start a new one. If your credits run out, the run waits instead and continues when you add credits. A run on your own Anthropic key is not charged by Nodus and has no cap to reach. Sandbox compute is billed as for any sandbox. `status.cost` shows the model spend so far. Runs in a group also share the group’s cap (see [Run agents in parallel](https://nodus-platform-site.pages.dev/docs/guides/agents/#run-agents-in-parallel)).

## Limits

* One run works one prompt at a time, with at most `spec.maxToolCalls` commands per turn (default 50).
* The input and each message are at most 256 KiB. A run’s recorded payloads are kept 30 days after it ends.
* A run created for a stopped agent (`spec.state: Stopped`) is refused with `AgentStopped`.
* A group takes at most 10000 unfinished runs. Create runs in batches of up to 100 by posting an `AgentRunList` to `…/agentruns` with an `Idempotency-Key` header; a batch is all-or-nothing, and its runs are created in list order, so list each run after the runs it depends on.

See [Durable execution](https://nodus-platform-site.pages.dev/docs/concepts/durable-execution/) for what survives a restart.
