List or watch trainingruntimes
const url = 'https://example.com/apis/nodus.dev/v1beta1/namespaces/example/trainingruntimes';const options = {method: 'GET'};
try { const response = await fetch(url, options); const data = await response.json(); console.log(data);} catch (error) { console.error(error);}curl --request GET \ --url https://example.com/apis/nodus.dev/v1beta1/namespaces/example/trainingruntimesParameters
Section titled “ Parameters ”Path Parameters
Section titled “ Path Parameters ”The project.
Query Parameters
Section titled “ Query Parameters ”Responses
Section titled “ Responses ”OK
object
TrainingRuntime is a signed trainer image plus the contract TrainingJobs compile against: the launcher, the parameter JSON Schema, the defaults, the resource presets and the measured step times (Beta).
object
object
object
object
object
object
object
TrainingRuntimeSpec is the contract a TrainingJob compiles against.
object
RuntimeCheckpointing configures recovery state.
object
Path is the state directory (default /nodus/state).
Strategy is HFTrainer (default) or Custom.
Command is the trainer entrypoint.
RuntimeDefaults apply when a TrainingJob leaves the field unset.
object
InitCommand is a preflight run before the command.
object
Command is the argv.
Timeout is at most 30m.
Parameters are merged under the TrainingJob’s parameters.
Resources are capacity floors. For GPU work the offering’s host shape applies if larger.
object
CPU is the vCPU floor.
Disk is the ephemeral disk floor.
GPURequest asks for accelerators. A family (H100) matches any of its variants and never another family.
object
Count is the GPUs per node: 1, 2, 4 or 8.
Exact pins the exact variant instead of matching the family.
Interconnect is Any or NVLink.
MinMemory is the per-GPU memory floor.
Type lists 1–8 accelerator ids or families from the catalog.
Memory is the host memory floor.
Nodes above 1 is shorthand for spec.distributed.nodes and is stored in that form.
Sidecars start before the trainer (for example a vLLM rollout server) and stop after it.
Sidecar is a helper container with native-sidecar semantics.
object
Command is the sidecar’s argv.
Env adds environment variables.
EnvVar is one environment variable: a literal value or a Secret key.
object
Name matches ^[A-Za-z_][A-Za-z0-9_]*$.
Value is the literal value.
EnvVarSource names where an env value comes from.
object
SecretKeySelector selects one key of a Secret.
object
Key is the key within it.
Name is the Secret.
Image defaults to the Job’s image.
Name is unique within the Job.
StartupTimeout bounds how long the sidecar may take to become ready.
EnvironmentTasks is AttemptAPI when this exact image fetches public manifests through the live attempt. Empty means the runtime is incompatible with GPU-host environment execution.
RuntimeEstimate is measured by the nightly runs. Keys are accelerator classes such as h100-80g or rtx-4090.
object
SecondsPerStep is used by training runtimes: expected duration is steps × seconds per step, scaled by how much more batchSize, gradientAccumulation and maxCompletionLength a TrainingJob sets than the defaults.
object
SecondsPerTask is used by evaluation runtimes: expected duration is tasks × seconds per task.
object
Image is the digest-pinned trainer image; managed runtimes are cosign-verified at release and seed.
Launcher is Plain (default), Torchrun, Accelerate or Deepspeed. Multi-node runs use it as the gang launcher.
Modes lists Train (default) and Evaluate.
Outputs are the downloadable results the runtime writes.
RuntimeOutput is one downloadable result.
object
Name is the output name.
Path is where the trainer writes it.
Parameters is the JSON Schema of spec.parameters on a TrainingJob (at most 64 KiB).
Presets choose resources for a model, method and quantization.
RuntimePreset is the resources measured for one model, method and quantization.
object
Method is Full, LoRA or QLoRA.
Model is a Hugging Face repository, for example Qwen/Qwen3-0.6B.
Quantization is empty or a scheme such as nf4.
Resources are capacity floors. For GPU work the offering’s host shape applies if larger.
object
CPU is the vCPU floor.
Disk is the ephemeral disk floor.
GPURequest asks for accelerators. A family (H100) matches any of its variants and never another family.
object
Count is the GPUs per node: 1, 2, 4 or 8.
Exact pins the exact variant instead of matching the family.
Interconnect is Any or NVLink.
MinMemory is the per-GPU memory floor.
Type lists 1–8 accelerator ids or families from the catalog.
Memory is the host memory floor.
Nodes above 1 is shorthand for spec.distributed.nodes and is stored in that form.
RuntimeRequirements say which inputs a TrainingJob must set.
object
Data requires spec.data.
Environment requires spec.environment.
Model requires spec.model (default true).
Teacher requires spec.teacher.
Task is what the runtime trains: SFT, Pretrain, DPO, ORPO, KTO, Reward, GRPO, Distill, Evaluate or Custom.
TrainingRuntimeStatus is written by the TrainingRuntime controller.
object
Conditions include Verified (SignatureVerified for managed runtimes).
object
ImageDigest is the verified image digest.
ImageUser is the user the verified image’s config runs its process as (root when it names none).
Managed is true for catalog runtimes in project nodus; it is derived, never set by a client.
ObservedGeneration is the generation the status describes.
Phase is Pending until the image digest is verified, then Ready, or Failed.
object
object
Example generated
{ "apiVersion": "example", "items": [ { "apiVersion": "example", "kind": "example", "metadata": { "annotations": { "additionalProperty": "example" }, "creationTimestamp": "example", "deletionGracePeriodSeconds": 1, "deletionTimestamp": "example", "finalizers": [ "example" ], "generateName": "example", "generation": 1, "labels": { "additionalProperty": "example" }, "managedFields": [ { "apiVersion": "example", "fieldsType": "example", "fieldsV1": { "additionalProperty": "example" }, "manager": "example", "operation": "example", "subresource": "example", "time": "example" } ], "name": "example", "namespace": "example", "ownerReferences": [ { "apiVersion": "example", "blockOwnerDeletion": true, "controller": true, "kind": "example", "name": "example", "uid": "example" } ], "resourceVersion": "example", "selfLink": "example", "uid": "example" }, "spec": { "checkpointing": { "path": "example", "strategy": "example" }, "command": [ "example" ], "defaults": { "initCommand": { "command": [ "example" ], "timeout": "example" }, "parameters": "example", "resources": { "cpu": "example", "disk": "example", "gpu": { "count": 1, "exact": true, "interconnect": "example", "minMemory": "example", "type": [ "example" ] }, "memory": "example", "nodes": 1 }, "sidecars": [ { "command": [ "example" ], "env": [ { "name": "example", "value": "example", "valueFrom": { "secretKeyRef": { "key": "example", "name": "example" } } } ], "image": "example", "name": "example", "startupTimeout": "example" } ] }, "environmentTasks": "example", "estimate": { "secondsPerStep": { "additionalProperty": 1 }, "secondsPerTask": { "additionalProperty": 1 } }, "image": "example", "launcher": "example", "modes": [ "example" ], "outputs": [ { "name": "example", "path": "example" } ], "parameters": "example", "presets": [ { "method": "example", "model": "example", "quantization": "example", "resources": { "cpu": "example", "disk": "example", "gpu": { "count": 1, "exact": true, "interconnect": "example", "minMemory": "example", "type": [ "example" ] }, "memory": "example", "nodes": 1 } } ], "requires": { "data": true, "environment": true, "model": true, "teacher": true }, "task": "example" }, "status": { "conditions": [ { "lastTransitionTime": "example", "message": "example", "observedGeneration": 1, "reason": "example", "status": "example", "type": "example" } ], "imageDigest": "example", "imageUser": "example", "managed": true, "observedGeneration": 1, "phase": "example" } } ], "kind": "example", "metadata": { "continue": "example", "remainingItemCount": 1, "resourceVersion": "example", "selfLink": "example", "shardInfo": { "selector": "example" } }}default
Section titled “ default ”An error: a metav1.Status whose reason is a registered code.
Status is the error body of every API response: a Kubernetes metav1.Status (so kubectl and client-go understand it) plus three top-level extensions that those clients ignore (ADR-028).
object
object
object
Docs is the URL of the code’s docs page.
Fix says what to do next, for example a CLI command or the field to change.
object
object
RequestID identifies the request in logs and support tickets.
Example generated
{ "apiVersion": "example", "code": 1, "details": { "causes": [ { "field": "example", "message": "example", "reason": "example" } ], "group": "example", "kind": "example", "name": "example", "retryAfterSeconds": 1, "uid": "example" }, "docs": "example", "fix": "example", "kind": "example", "message": "example", "metadata": { "continue": "example", "remainingItemCount": 1, "resourceVersion": "example", "selfLink": "example", "shardInfo": { "selector": "example" } }, "reason": "example", "requestId": "example", "status": "example"}