Agent evaluation groups
AgentGroups can now generate and score a bounded Environment benchmark through ordinary AgentRuns. Each batch pins one Agent revision and Environment image, keeps hidden answers out of agent inputs, and records per-case verdicts in encrypted run journals. The console separates pass rate and reward from execution failures, missing answers and grading failures, with case links and grader evidence.
Configure evaluation in YAML or the Python SDK, then read group.results(). Batches contain at most 100 cases
including repetitions and have a queue-inclusive timeout, defaulting to 30 minutes. Model caps keep their
existing meaning; use a project Budget to limit model plus agent/grader compute spend.