In the past few months, we’ve seen SOTA LLMs saturating basic coding benchmarks with short and simplified coding tasks. It's time to enter the next stage of coding challenge under comprehensive and realistic scenarios!
-- Here comes BigCodeBench, benchmarking LLMs on solving practical and challenging programming tasks!
So, can LLMs solve these tasks?
- Not yet!
🏆 Pass@1: Humans ace 97%, GPT-4o only hits 50-60%, but DeepSeek-Coder-V2 is tighy at its heels!
Check out our leaderboard, data, code, and paper:
bigcode-bench.github.io/
1/🧵