Article
What Makes a Programming Problem Hard for a Language Model? An Empirical Study of Item Difficulty Across Code LLMs on Two Benchmarks
2026-07-13
Abstract excerpt
<title>Abstract</title> <p>Large language models (LLMs) for code are now evaluated with benchmarks such as Human Evaland MBPP, and they are increasingly used to generate, grade, and explain exercises in introductory programming courses. Yet the standard reporting practice, a single aggregate pass@kscore per model, treats every problem as interchangeable and tells us nothing about which problems are hard or why. W...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 90691a38-a21e-57d0-bb90-2f37d8e26e29
- DOI
- 10.21203/rs.3.rs-10241058/v1
