MEGA Hub

TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability

Authors

Do you know Vincent Cohen-Addad?You can claim authorship or link another user.Do you know Dimitris Paparas?You can claim authorship or link another user.Do you know Ernest van Wijland?You can claim authorship or link another user.Do you know Max Springer?You can claim authorship or link another user.Do you know Julien Canitrot-Paradis?You can claim authorship or link another user.Do you know Honghao Lin?You can claim authorship or link another user.Do you know David Woodruff?You can claim authorship or link another user.Do you know Adarsh Kumarappan?You can claim authorship or link another user.Do you know Rajesh Jayaram?You can claim authorship or link another user.Do you know Rudrajit Das?You can claim authorship or link another user.Do you know Lalit Jain?You can claim authorship or link another user.Do you know Ola Svensson?You can claim authorship or link another user.Do you know Silvio Lattanzi?You can claim authorship or link another user.Do you know Mislav Balunovic?You can claim authorship or link another user.Do you know Theophane Weber?You can claim authorship or link another user.Do you know Vahab Mirrokni?You can claim authorship or link another user.

Abstract

We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof generation. TCS-Bench consists of theorem-proving tasks from papers published at top theoretical computer science venues (STOC, FOCS, and SODA). Each task provides the necessary context to derive a self-contained proof for a target result. We evaluate state-of-the-art models on this benchmark. We verify the correctness of generated proofs via a verification agent, and further benchmark the verifier against human-expert proof judgements on a set of target statements and generated proofs pairs. Our reference verifier achieves over 90% accuracy on the expert labeled set.

Community

00