Article
Investigating Reproducibility Challenges in LLM Bugfixing on the HumanEvalFix Benchmark
2025-05-29
Abstract excerpt
Benchmark results for Large Language Models often show inconsistencies across different studies. This paper investigates the challenges of reproducing these results in automatic bugfixing using LLMs, on the HumanEvalFix benchmark. To determine the cause of the differing results in the literature, we attempted to reproduce a subset of them by evaluating 11 models in the DeepSeekCoder, CodeGemma, and CodeLlama model...
Topics
Open a Topic to create a Post that cites this publication.
Identifiers and source
- Literature Corpus work
- 05683a2f-0354-5c60-83c7-2f82a2f94d39
- DOI
- 10.20944/preprints202505.2321.v1
