SWE-bench
A benchmark for evaluating language models on resolving real-world software engineering issues from GitHub repositories.
Explore more about Benchmarks
Related terms
A benchmark that evaluates the code generation capabilities of AI models by measuring their ability to generate functionally correct code that passes predefined unit tests for a collection of programming tasks.
MBPPMBPPA benchmark for evaluating the ability of language models to generate Python code from natural language descriptions.
LiveCodeBenchA contamination-aware benchmark for evaluating large language models on recent competitive programming problems.