Singhal-959492.pdf
The paper investigates how Large Language Models (LLMs) can be used to automatically repair software vulnerabilities and how specialized prompt engineering can significantly improve their performance on real-world code. The authors first evaluate baseline LLM capabilities on 5,826 code samples of varying sizes, finding that LLMs succeed on simple snippets (<40 lines) but fail on longer, more complex code. They then introduce Control Flow Graph (CFG) prompts, which modestly raise the repair success rate to about 14.4% on previously unsolvable cases. An analysis of the remaining failures reveals three challenge categories: (1) misidentification of vulnerable code, (2) correct identification but incorrect fixes, and (3) missing dependencies or incomplete context. For each category the authors design a dedicated prompt pattern that supplies targeted information—such as simplified code, explicit vulnerability explanations, dependency specifications, and precise hints. The patterns are evaluated on synthetic datasets (Juliet) and on real-world C/C++ projects (Exiv2) using the BugsCpp framework. Results show that the prompt patterns raise repair success to over 85% across all categories, far surpassing the CFG baseline. The study also discusses limitations, notably the heavy reliance on human expertise to craft prompts and the limited benefit of iterative feedback without detailed explanations. Future work is suggested to automate prompt generation and test the approach on other LLMs. The paper contributes a systematic error taxonomy for LLM‑based code repair and demonstrates that carefully engineered prompts can bridge the gap between LLM baseline abilities and the complexities of real-world vulnerability fixing.
Topics
Research Motivation and Questions: Limits of LLMs in Real-World Vulnerability Repair
The authors articulate the growing interest in using LLMs for automated code repair, but highlight a gap: existing evaluations lack graduated complexity and systematic prompt engineering.