🏆 PutnamBench Leaderboard 🏆
Benchmarking formal mathematical reasoning on the Putnam Mathematical Competition.
★ Featured ranking
346 of the 672 problems contain an answer placeholder that the prover must infer.
📝 Notes
- A method attached with a heart emoji (💚) is fully open-sourced, while a method attached with a blue heart emoji (💙) is partially open-sourced.
- The no-answer variant removes answer clues and requires systems to infer answer placeholders where present. Results should identify the benchmark commit, any modified inputs, and the answer-faithfulness review procedure.
- Some new problems have been added since original release, benchmarked methods have not been rerun on those problems, but PutnamBench version used in the eval is mentioned in the repo's `results.json`.
- We are open to suggestions for better indicating the differing compute budgets of approaches on the leaderboard. Please reach out to us with your ideas!
- We thank the EvalPlus team for providing the leaderboard template.