nodus run
View Markdownnodus run
Section titled “nodus run”Run a command as a Job: estimate, phases, logs and the final cost
Synopsis
Section titled “Synopsis”Uploads the current directory (unless –no-source), creates a Job, prints its estimate and phases, streams its logs and ends with the cost line. The command’s exit code becomes nodus’s; 125 means a Nodus or API error, 124 a timeout and 130 an interrupt. On a terminal Ctrl+C asks before canceling; -d detaches.
nodus run [flags] -- COMMAND [ARGS...] | run APP.py[::FUNCTION] [ARGS...]Examples
Section titled “Examples” nodus run --gpu L4 --image nodus/pytorch -- python train.py nodus run --gpu H100:8 --max-cost 40 -d -- torchrun train.py nodus run app.py::mainOptions
Section titled “Options” --allow-large-source Upload source above 500 MiB compressed --checkpoint string Checkpoint path (default /nodus/state) --complete-by string Finish-by time (RFC 3339) for placement --completions int Indexed completions --continuity string Recovery: Checkpointed, Restartable or Ephemeral --cpu string vCPUs -d, --detach Return once the Job is created --disk string Disk, such as 200Gi --dry-run Print the estimate only --env stringArray Environment variable K=V --expected-duration duration Expected runtime, for the estimate --gpu string Accelerator TYPE[:COUNT], such as L4 or H100:8 --gpus-per-node int Beta: GPUs per gang node -h, --help help for run --idempotency-key string Override the generated idempotency key --image string Image, such as nodus/pytorch --interruptible string[="prefer"] Interruptible capacity: allow, prefer or never --keep Keep the Job after it finishes (no 30-day TTL) -l, --label strings Labels k=v --launcher string Beta: plain, torchrun, ray or verl --max-cost string Spend cap in USD --memory string Memory, such as 64Gi --name string Job name (default: generated run-xxxxx) --network string Beta: colocated, regional or global --no-source Do not upload the current directory --no-warm Release capacity at once instead of keeping it warm for 60 s --nodes int Beta: gang nodes --output stringArray Output NAME=/path --parallelism int Indexes running at once --profile string Placement profile: Balanced, Cost or Speed --region string Region class, such as us or eu --secret stringArray Secret to expose as environment variables --startup-timeout duration Beta: gang assembly budget --timeout duration Wall-clock limit, such as 2h --total-gpus int Beta: total GPUs across the gang --transport string Beta: direct or auto --volume stringArray Volume NAME:/path[:ro]Options inherited from parent commands
Section titled “Options inherited from parent commands” --context string Context from the config file to use --org string Organization (selects the context for that org) -p, --project string Project to work in -v, --verbose Log each API request (never credentials)SEE ALSO
Section titled “SEE ALSO”- nodus - Run compute Jobs on Nodus
Tested examples
Section titled “Tested examples”These invocations run in CI against the CLI’s fake API.
nodus run app.py::main --epochs 3nodus run fail.pynodus run --gpu L4 --image nodus/pytorch -- python train.pynodus run --gpu L4 --no-source -- sh -c 'exit 3'nodus run --gpu L4 --no-source --name late -- python deadline.pynodus run -d --no-source --gpu L4 --keep --name detached -- python train.pynodus run --dry-run --no-source --gpu L4 -- python train.pynodus run -d --no-source --name gang --gpu H100 --gpus-per-node 8 --nodes 2 --launcher torchrun --network regional --max-cost 40 --env A=1 --interruptible -- torchrun train.pynodus run --dry-run --no-source --gpu H100 --gpus-per-node 8 --nodes 2 -- torchrun train.pynodus run --no-source --gpu none -- python train.py