skip to content
The Weighted Average

Wire

ExplorationBench gives agents 140 alien-world tasks

ExplorationBench gives AI systems 55 discovery targets across two executable “alien worlds” and 140 held-out tasks, with every answer checked by code, according to the paper’s benchmark description. Across ten systems, four exploration rounds lifted the best AlienCode trajectory to 87.6%, while identical turns without environmental feedback stayed at or below 11.0%; the authors also report that continued exploration can undo earlier gains. Builders testing research or planning agents should add unfamiliar-rule tasks—not only known benchmarks—to their eval suite; Octobench’s harness-cost analysis shows why evaluation design changes what an agent score means.