Quantized training uses each worker's GPU
Each training worker now selects its assigned local GPU before loading a model. This prevents QLoRA workers from all loading quantized weights onto GPU 0 before the training framework configures distributed execution. The same behavior applies to fine-tuning, preference training, reward models and distillation. Explicit model device maps remain unchanged.
The fix takes effect when the updated managed runtime image is published. Local tiny-model and CPU distributed checks do not establish CUDA or cross-provider qualification.