When Does Context Help? A Controlled Study ofLLM-Based Bug Fixing

Authors

DOI:

https://doi.org/10.4108/eetiot.15050

Keywords:

Bug Fixing, Contextual Information, Large Language Models, Test Cases, Feedback, Problem Descriptions

Abstract

Large language models (LLMs) have shown promise for automated program repair, but it remains unclear which debugging signals are most useful and when additional context becomes distracting, costly, or ineffective. We present a controlled empirical study of LLM-based bug fixing on FIXEVAL, comparing three model families–GPT, Claude, and DeepSeek–under five prompt settings: code-only repair, problem description, failed test cases, passed and failed test cases, and all available context. We further investigate whether structured execution feedback improves repair in a second attempt after an initial patch fails. Across 400 benchmark instances spanning Wrong Answer, Runtime Error, Time Limit Exceeded, and Memory Limit Exceeded bugs, we evaluate repair accuracy, response time, computation cost and second-attempt recovery. The results show that contextual information generally improves first-attempt accuracy, but its benefit depends strongly on bug type and model family. Full context is often most effective for Wrong Answer, Runtime Error, and Time Limit Exceeded bugs, whereas Memory Limit Exceeded bugs remain difficult across settings, suggesting that resource-limit failures often require deeper algorithmic redesign. Passed tests provide limited benefit unless paired with stronger semantic or failing-test signals. Structured failure feedback improves second-attempt repair most when it exposes concrete behavioral mismatches or runtime failures, but it is less effective for coarse resource-limit signals. These findings clarify when context helps LLM-based repair and provide practical guidance for designing cost-aware, feedback-driven repair pipelines.

Downloads

Download data is not yet available.

References

[1] Prenner JA, Babii H, Robbes R. Can OpenAI’s Codex fix bugs? an evaluation on QuixBugs. In: IEEE/ACM International Workshop on Automated Program Repair (APR); 2022. p. 69-75.

[2] Sobania D, Briesch M, Hanna C, Petke J. An Analysis of the Automatic Bug Fixing Performance of ChatGPT. In: IEEE/ACM International Workshop on Automated Program Repair (APR); 2023. p. 23-30.

[3] Xia CS, Zhang L. Automated program repair via conversation: Fixing 162 out of 337 bugs for $0.42 each using chatgpt. In: Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis; 2024. p. 819-31.

[4] Haque MMA, Ahmad WU, Lourentzou I, Brown C. FixEval: Execution-based Evaluation of Program Fixes for Programming Problems. In: IEEE/ACM International Workshop on Automated Program Repair (APR); 2023. p. 11-8.

[5] Jimenez CE, Yang J, Wettig A, Yao S, Pei K, Press O, et al. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:231006770. 2023.

[6] Yang J, Jimenez CE, Wettig A, Lieret K, Yao S, Narasimhan K, et al. Swe-agent: Agent-computer interfaces enable automated software engineering. Advances in Neural Information Processing Systems. 2024;37:50528-652.

[7] Zhang Y, Ruan H, Fan Z, Roychoudhury A. Autocoderover: Autonomous program improvement. In: Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis; 2024. p. 1592-604.

[8] Yin X, Ni C, Wang S, Li Z, Zeng L, Yang X. Thinkrepair: Self-directed automated program repair. In: Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis; 2024. p. 1274-86.

[9] Bouzenia I, Devanbu P, Pradel M. RepairAgent: An Autonomous, LLM-Based Agent for Program Repair. In: Proceedings of the IEEE/ACM 47th International Conference on Software Engineering. ICSE ’25. IEEE Press; 2025. p. 2188–2200. Available from: https://doi.org/10.1109/ICSE55347.2025.00157.

[10] Fu M, Tantithamthavorn C, Nguyen V, Le T. ChatGPT for Vulnerability Detection, Classification, and Repair: How Far Are We? In: Asia-Pacific Software Engineering Conference (APSEC); 2023. .

[11] Ribeiro F, de Macedo JNC, Tsushima K, Abreu R, Saraiva J. GPT-3-Powered Type Error Debugging: Investigating the Use of Large Language Models for Code Repair. In: Proceedings of the 16th ACM SIGPLAN International Conference on Software Language Engineering; 2023. p. 111-24.

[12] Guo Q, Cao J, Xie X, Liu S, Li X, Chen B, et al. Exploring the potential of ChatGPT in automated code refinement: An empirical study. In: 46th IEEE/ACM International Conference on Software Engineering; 2024. p. 1-13.

[13] Qu X, Zuo F, Li X, Rhee J. Context matters: Investigating its impact on ChatGPT’s bug fixing performance. In: 2024 IEEE/ACIS 22nd International Conference on Software Engineering Research, Management and Applications (SERA). IEEE; 2024. p. 17-23.

[14] Puri R, Kung DS, Janssen G, Zhang W, Domeniconi G, Zolotov V, et al.. CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks; 2021.

[15] Lin D, Koppel J, Chen A, Solar-Lezama A. QuixBugs: A multi-lingual program repair benchmark set based on the Quixey Challenge. In: Proceedings Companion of the 2017 ACM SIGPLAN international conference on systems, programming, languages, and applications: software for humanity; 2017. p. 55-6.

[16] Zuo F, Qian G, Qu X, Rhee J, Fu J. Revisiting the Capability of GPT in Solving Coding Problems: A Lesson from Programming with Recursion. In: Proceedings of the 26th ACM Annual Conference on Cybersecurity & Information Technology Education; 2025. p. 134-40.

[17] Crandall AS, Sprint G, Fischer B. Generative Pre-Trained Transformer (GPT) Models as a Code Review Feedback Tool in Computer Science Programs. Journal of Computing Sciences in Colleges. 2023;39(1):38-47.

[18] Jalil S, Rafi S, LaToza TD, Moran K, Lam W. ChatGPT and software testing education: Promises & perils. In: IEEE International Conference on Software Testing, Verification and Validation Workshops (ICSTW); 2023. p. 4130-7.

[19] Wuisang MC, Kurniawan M, Santosa KAW, Gunawan AAS, Saputra KE. An Evaluation of the Effectiveness of OpenAI’s ChatGPT for Automated Python Program Bug Fixing using QuixBugs. In: 2023 International Seminar on Application for Technology of Information and Communication (iSemantic); 2023. p. 295-300.

[20] Ye H, Martinez M, Durieux T, Monperrus M. A comprehensive study of automatic program repair on the QuixBugs benchmark. Journal of Systems and Software. 2021;171:110825.

[21] Do Viet T, Markov K. Using Large Language Models for Bug Localization and Fixing. In: 12th International Conference on Awareness Science and Technology (iCAST); 2023. p. 192-7.

[22] Chandramohan M, Jancic J, Zhang Y, Krishnan P. From Benchmark Data To Applicable Program Repair: An Experience Report. arXiv preprint arXiv:250816071. 2025.

[23] Rondon P, Wei R, Cambronero J, Cito J, Sun A, Sanyam S, et al. Evaluating agent-based program repair at google. In: 2025 IEEE/ACM 47th International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE; 2025. p. 365-76.

[24] Hidvégi D, Etemadi K, Bobadilla S, Monperrus M. Cigar: Cost-efficient program repair with llms. arXiv preprint arXiv:240206598. 2024.

[25] Le-Cong T, Le B, Murray T. Can LLMs reason about program semantics? a comprehensive evaluation of LLMs on formal specification inference. In: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); 2025. p. 21991-2014.

[26] Tang H, Hu K, Zhou JP, Zhong S, Zheng WL, Si X, et al. Code repair with llms gives an exploration-exploitation tradeoff. Advances in Neural Information Processing Systems. 2024;37:117954-96.

Downloads

Published

22-09-2026

How to Cite

1.
Gummadavelli DK, Qu X, Li X, Zhang X, Song Y, Yang B. When Does Context Help? A Controlled Study ofLLM-Based Bug Fixing. EAI Endorsed Trans IoT [Internet]. 2026 Sep. 22 [cited 2026 Sep. 22];11. Available from: https://publications.eai.eu/index.php/IoT/article/view/15050

Most read articles by the same author(s)