Training previews, tracking and recovery
- A TrainingJob dry-run (
?dryRun=All, the SDK’splan = job.preview()) returnsstatus.compiledJob: the image, command, inputs, resources and recovery of the Job the TrainingJob would create, soplan.jobshows it before you run. A created TrainingJob names its Job instatus.job. grpo-loraand environment evaluations render prompts with thinking turned off (enable_thinking=False), so Qwen3 models answer within the completion budget instead of spending it on reasoning. The runtime’s defaults now follow the GRPO recipe: learning rate1e-6, batch 1 with 4 accumulation steps for 4 sampled completions, and 64-token completions. Setparametersto change any of them.- The catalog TrainingRuntimes report metrics to your MLflow server (
spec.tracking.mlflow) or Weights & Biases project (spec.tracking.wandb) when you set one; before, the trainer images ignored both. - Recovery state records the runtime image digest and merged parameters. A run refuses state from a different configuration, and GRPO reuses a saved baseline after recovery instead of grading it again. Output provenance includes the runtime digest compiled by the controller.
- Dataset previews now run the verified runtime image with your dataset mounted read-only in an isolated CPU Job. A preview waits at most one minute for capacity, has a 45-second execution limit and a $0.25 spend cap. When capacity is unavailable, the request fails so you can retry; it does not return invented sample data.
Recovery now verifies the complete run configuration, including model and dataset references, before reusing checkpoints. A corrupt saved GRPO baseline fails with recovery instructions instead of silently grading the same tasks again. Environments requiring a package build fail explicitly when no package builder is installed; prebuilt environment images remain supported.
Training admission pins dataset, model, and teacher Volumes to committed revisions. A later upload cannot silently change a run on recovery. Missing revisions fail before the run starts.
exportTo copies a managed runtime’s complete outputs into a Ready ReadWriteOnce Volume through a bounded CPU Job under the training cap. Completion identifies the revision committed by that export attempt. Export jobs follow parent suspension, cancellation, and funding stops. A failed final Volume commit now fails the node attempt instead of reporting success.