MBPP
MBPPAlso called: Mostly Basic Python Problems
A benchmark for evaluating the ability of language models to generate Python code from natural language descriptions.
Explore more about Benchmarks
Related terms
A benchmark that evaluates the code generation capabilities of AI models by measuring their ability to generate functionally correct code that passes predefined unit tests for a collection of programming tasks.
LiveCodeBenchA contamination-aware benchmark for evaluating large language models on recent competitive programming problems.
SWE-benchA benchmark for evaluating language models on resolving real-world software engineering issues from GitHub repositories.
Benchmark RunA single execution of a benchmark used to measure and record a model or system's performance under defined conditions.