Wire
CaReBench releases 1,000-video reasoning test
CaReBench released a 1,000-video test split for compositional reasoning, with each record carrying an overall caption plus separate spatial and temporal descriptions. The canonical dataset API shows eight Parquet shards and fields for video, three caption views, category, and subcategory, although its dataset card publishes no model scores or methodology. Teams testing multimodal agents behind a bounded pilot should file it as a fresh probe for whether a system understands where and when events happen—not yet as an auditable leaderboard.