Explicit legacy letter-counting protocols

Two versioned letter-counting Environments preserve the legacy evaluation and RL task slices separately. Their catalog examples specify the pinned model, system instruction and generation settings. GRPO now accepts explicit generation batch size and top-k sampling, and both runtimes accept an Environment evaluation batch size. Historical scores were removed from the differently generated multi-family benchmark. The new profiles have no claimed baseline or learning improvement until a matched GPU run supplies those measurements.

Choosing a catalog example in the console preserves its training/held-out task counts and catalog seed, matching the SDK and generated manifest. Drafts from another example no longer replace the selected configuration.