Wire
Supabase Evals puts Codex at 100%
Supabase open-sourced Evals as its live leaderboard gives Codex with GPT-5.6 Sol a 100% total score, versus 95% for Claude Code with Opus 5 or Sonnet 5 and OpenCode with Kimi K3. The official benchmark methodology runs agents against containerized Supabase projects and grades database, deployment, investigation, and repair work with deterministic checks plus an LLM judge; current results show GPT-5.4 mini at 79%. For teams deciding between goal-driven and outcome-verified coding agents, the reusable lesson is stronger than the ranking: test the running system and user access, not whether the diff looks plausible.