Where in the code is this bug?
A benchmark on 918 real GitHub issues that measures how well different search methods find the file that needs fixing.
The problem
Before anyone, human or AI tool, can fix an issue, they have to find the right place in the code. Lots of methods claim to do that well. Few are compared fairly on real data.
What I built
A university team project with two engineers and a research group. My part was the ground-truth pipeline and its validation, the framework for comparing query modes, the statistics, the analysis of code written before and after ChatGPT, and the tooling that lets a local LLM expand a search query.
What came out
Letting an LLM add likely function and file names to the query was the real lever, not the fancier two-stage search, and the difference is statistically significant. Validating the ground truth also showed that the first dataset contained only 149 unique issues out of 206 entries, which changed several conclusions.
Built with
Python, Elasticsearch, embedding models, a local LLM, bootstrap statistics