Back to search

Article

Investigating Reproducibility Challenges in LLM Bugfixing on the HumanEvalFix Benchmark

2025-05-29

Abstract excerpt

Benchmark results for Large Language Models often show inconsistencies across different studies. This paper investigates the challenges of reproducing these results in automatic bugfixing using LLMs, on the HumanEvalFix benchmark. To determine the cause of the differing results in the literature, we attempted to reproduce a subset of them by evaluating 11 models in the DeepSeekCoder, CodeGemma, and CodeLlama model...

Topics

Open a Topic to create a Post that cites this publication.

Identifiers and source

Literature Corpus work
05683a2f-0354-5c60-83c7-2f82a2f94d39
DOI
10.20944/preprints202505.2321.v1
Open publication

Related research

Semantic proximity does not establish scientific evidence.

Click a neighbor to travelStep 1 · 12 closest
Interactive article relationship graphSelect a related publication card to move it into the centre and load its closest explainable connections. Solid lines are source-backed structured connections. Dashed lines are semantic discovery signals and are not scientific evidence.
Investigating Reproducibility Challenges in LLM Bugfixing on the HumanEvalFix BenchmarkDOI 10.20944/preprints202505.2321.v1
Select a neighboring publication to make it the new centre.