Wire
Google opens 20+ production agent metrics
Google made Agent and Model Evaluations generally available with more than 20 pre-built metrics spanning quality, safety, grounding, tool use, trajectories, summarization, and translation. The Gemini Enterprise Agent Platform announcement says the same versioned metrics can score local experiments and sampled production traces, with score-over-time charts and drift alerts; code-based metrics add no charge, while model judges and stored artifacts use standard rates. For teams assessing Google’s integrated agent stack, this closes a real observability gap, but it also makes the evaluation registry another source of cloud lock-in—keep cases, rubrics, traces, and scores exportable.