GPU-host reinforcement learning and grading

Environment-backed TrainingJobs prepare tasks and run isolated graders on their assigned GPU host without a separate CPU machine. Graders cannot access the model, trainer credentials, API socket or GPU devices. Managed runtimes fetch public task manifests after the GPU attempt starts; custom runtimes use the TrainingJob tasks endpoint. Baseline and final model evaluation select the local CUDA device before generating answers.

Each GPU attempt admits one grader at a time, including requests arriving through different API replicas. Failed container or network removal keeps that slot reserved and retries cleanup; neither the helper nor its parent reports completion while helper resources remain allocated.

Network cleanup failures remain retryable without forgetting their firewall state. Restart cleanup retains the helper container until its network has been removed.

Managed trainer images use protocol version 2.0.0. Updating a runtime waits for verification of the current specification before accepting training, so an older image cannot inherit the new task-fetch capability.