🏆 PutnamBench Leaderboard 🏆

Benchmarking formal mathematical reasoning on the Putnam Mathematical Competition.

github paper
346 of the 672 problems contain an answer placeholder that the prover must infer.

📝 Notes

  1. A method attached with a heart emoji (💚) is fully open-sourced, while a method attached with a blue heart emoji (💙) is partially open-sourced.
  2. The no-answer variant removes answer clues and requires systems to infer answer placeholders where present. Results should identify the benchmark commit, any modified inputs, and the answer-faithfulness review procedure.
  3. Some new problems have been added since original release, benchmarked methods have not been rerun on those problems, but PutnamBench version used in the eval is mentioned in the repo's `results.json`.
  4. We are open to suggestions for better indicating the differing compute budgets of approaches on the leaderboard. Please reach out to us with your ideas!
  5. We thank the EvalPlus team for providing the leaderboard template.