skip to content
The Weighted Average

Wire

TCS-BENCH gives proof verifiers 90% accuracy

TCS-BENCH evaluates LLMs on research-level theoretical-computer-science proof generation and reports a reference verifier with more than 90% accuracy against expert-labeled target-statement and proof pairs. The arXiv paper builds tasks from STOC, FOCS, and SODA papers, supplies the context needed for self-contained proofs, and checks generated proofs with a verification agent. Teams using LLMs for mathematical research should file this as evidence that proof generation and proof checking need separate benchmarks, complementing Claude’s reproducibility gate without proving that frontier models can independently produce correct research.