- The paper introduces BLAgent, which combines AST-aware path-augmented chunking, dual-perspective queries, and two-phase agentic reranking to localize bugs efficiently.
- BLAgent achieves up to 86.7% Top-1 accuracy and 0.900 MRR with Claude-4.6, outperforming LocAgent while costing over 18 times less per instance.
- The paper shows that improved localization raises Agentless repair rates from 27.6% to 34.0%–38.3%, while retrieval recall and function- or line-level localization remain key bottlenecks.
Motivation and problem statement
File-level bug localization is the first stage of hierarchical maintenance pipelines that feed root cause analysis, triage, and automated program repair (APR). Prior empirical evidence cited by the authors shows that removing file-level localization from a multi-granularity framework causes a 94% drop in Top-5 accuracy and a 96% reduction in MAP at the statement level, making it the most consequential component of such pipelines. The paper's motivating example (django-10924) demonstrates that Agentless mislocalizes a fault to serializer.py, producing an incorrect patch, whereas identifying __init__.py enables the same APR system to resolve the issue.
Existing approaches fall into two camps with complementary weaknesses. Conventional RAG pipelines use holistic code-text embeddings and naive text chunking, which fail to capture code structure. Agentic methods such as LocAgent perform graph-guided multi-hop traversal over heterogeneous repository graphs, achieving strong accuracy but at prohibitive cost, or require LLM fine-tuning. BLAgent positions itself between these extremes: bounded agentic reasoning over a compact, retrieval-filtered candidate set.
The BLAgent pipeline
BLAgent comprises four stages:
Repository encoding. Source files are segmented using AST-aware chunking (via CodeSplitter) so that no chunk cuts across an AST construct; oversized functions are recursively subdivided at nested syntactic boundaries rather than character offsets. Each chunk is prepended with its relative file path. Path augmentation closes a systematic vocabulary mismatch: bug reports frequently contain package-qualified identifiers (django.db.models.fields), traceback entries, and module references that structurally mirror file paths but are absent from content-only chunk embeddings. Chunks are embedded with nomic-embed-text-v1 and stored in a vector database.
Dual-perspective query transformation. Each bug report is decomposed via one-shot prompts into a structural query (T0​) emphasizing identifiers, modules, and traceback tokens, and a behavioral query (T1​) capturing expected-versus-observed runtime behavior. Both queries retrieve independently, and their top-m lists are merged into a deduplicated candidate pool capped at 15 files.
Two-phase agentic reranking. Phase 1 (Skeleton-Based Agent Scoring, SAS) instantiates a ReAct agent that iteratively inspects structural skeletons—class/function signatures and docstrings—of candidate files through a ReadFileSkeleton tool and assigns 0–10 relevance scores. Phase 2 (Evidence-Anchored Reranking, EAR) is a single inference pass over the top-5 Phase-1-scored files, where each file's context is pruned to preserve global structure while expanding only the retriever-highlighted top-5 chunks to full implementations. The authors justify the two-phase split on cost and attention-degradation grounds: reranking all 15 candidates with full contexts would inflate context size and risk mid-context attention loss.
Localization results
On SWE-bench-Lite (300 Python instances), BLAgent with GPT-OSS-120B achieves MRR 0.851 and Top-1 accuracy of 78.6%, surpassing all baselines at their best published configurations—including LocAgent (Claude-3.5) at Top-1 77.7%—without fine-tuning, dependency graphs, or static analysis infrastructure. Under a controlled same-model comparison with Claude-4.6-Sonnet, BLAgent reaches MRR 0.900 and Top-1 86.7%, versus LocAgent's 82.4% Top-1 on the 182 instances LocAgent completed within a $300 budget. Notably, LocAgent under Claude-4.6 exhibits degraded Top-3/Top-5 accuracy relative to its Claude-3.5 configuration due to overconfident single-file outputs, while BLAgent improves across all thresholds—an indication that the architecture, not the backbone model, drives the gains.
Ablations decompose the contribution of each design element:
| Configuration |
MRR |
Top-1 |
| Dense retrieval only (path-augmented chunks) |
0.553 |
0.417 |
| + LLM reranking (basic RAG) |
0.734 |
0.673 |
| + SAS (Phase 1) |
0.769 |
0.685 |
| + EAR (Phase 2, full BLAgent) |
0.819 |
0.762 |
Path augmentation alone yields a 16.9% MRR improvement over code-aware chunking without paths, and up to 20.4% over text-based chunking. Query transformation lifts dense retrieval MRR by 22.9% for $T_0.Thetwotransformationsaredemonstrablycomplementary:theirunionretrievesthecorrectfileinTop−10for94.3T_0)and88.7T_1)individually,with17casesuniquelyrecoveredbyT_0and7byT_1.Candidateordering(T_0−firstvs.T_1$-first) has negligible effect, confirming no positional dependency.</p>
<p>An important finding concerns model scale: <a href="https://www.emergentmind.com/topics/qwen3-32b" title="" rel="nofollow" data-turbo="false" class="assistant-link" x-data x-tooltip.raw="">Qwen3-32B</a> and <a href="https://www.emergentmind.com/topics/gpt-oss-120b-d61913a9-5002-4c7e-9230-cedfc142f64a" title="" rel="nofollow" data-turbo="false" class="assistant-link" x-data x-tooltip.raw="">GPT-OSS-120B</a> perform near-identically in Phase 1 despite a 3.8× size difference, indicating that reasoning protocol adherence—not parametric capacity—is the dominant factor at the skeleton-inspection level. <a href="https://www.emergentmind.com/topics/frontier-models" title="" rel="nofollow" data-turbo="false" class="assistant-link" x-data x-tooltip.raw="">Frontier models</a> benefit less from Phase 2 because they produce better-separated Phase-1 scores.</p>
<p>The architecture also generalizes to function-level localization with a single prompt change: BLAgent (Claude-4.6) achieves Top-1 72.6%, improving over LocAgent (Claude-4.6) by 88% at Top-1. The authors attribute this primarily to file-level precision—BLAgent places the patch file in its Top-5 candidates for 90.7% of instances—rather than to the prompt modification itself.</p>
<h2 class='paper-heading' id='cost-analysis'>Cost analysis</h2>
<p>Bounded reasoning translates directly into cost efficiency. With GPT-OSS-120B, BLAgent consumes roughly 24.5K prompt tokens per instance at approximately $T_1$00.09 per instance ($T_1$1300 budget after 182 instances—an effective cost of $1.65 per instance, over 18× higher than BLAgent—with several instances timing out or exceeding 800K tokens/minute due to unbounded graph traversal. An ablation on EAR hyperparameters reinforces the "narrow context beats large context" principle: increasing candidates from 5 to 10 files raises token usage by 86% with no Top-1 improvement.
Impact on end-to-end repair
Integrating BLAgent as the localization module in Agentless (with GPT-OSS-120B throughout) improves resolution rates across all patch-selection configurations: 27.6% → 34.0% (majority voting), 28.6% → 36.0% (+regression tests), and 32.0% → 38.3% (+reproduction tests), a relative improvement exceeding 20%. A paired Wilcoxon signed-rank test over three runs confirms statistical significance (T1​2). Empty patches drop from 38 to 14, suggesting better-grounded repair context.
Failure attribution is instructive. Of 177 unresolved cases analyzed from the worst run, 159 were correctly localized at the file level; function-level localization failures account for 44.6% of unresolved cases, line-level failures 26.6%, and patch-generation failures 18.1%. Approximately 82% of unresolved issues trace to some stage of hierarchical localization. A case study (django-11133) shows that substituting the correct line number into the pipeline yields a correct patch, demonstrating that line-level mislocalization—not synthesis capability—caused the failure. Conversely, unique repairs achieved by BLAgent coincide with baseline localization failures in 35–50% of cases, while the reverse pattern is nearly absent, supporting a causal link between localization quality and repair success.
Limitations and open questions
The paper is candid about several constraints. All 14 retrieval failures under Claude-4.6 stem from the correct file being absent from the Top-15 candidate pool entirely, partitioned into feature requests (2), hidden dependencies (5), and vocabulary mismatches (7). Ten of these failures concentrate in sympy (13.0% repository failure rate vs. 3.5% for django), reflecting a systematic divergence between user-visible mathematical symptoms and low-level implementation sites whose naming bears no lexical resemblance to bug reports. Extending retrieval to Top-50 recovers 8 of 12 non-feature-request failures but inflates the candidate pool to ~63 files, which is impractical for agentic reranking—leaving open how to improve embedding-level recall without unbounded candidate expansion.
Evaluation is restricted to SWE-bench-Lite (Python-only), and baselines are reported at their best published configurations rather than uniformly reproduced; the controlled same-model comparison covers only BLAgent and LocAgent. The line-level accuracy metric relies on a heuristic approximation (SequenceMatcher similarity ≥ 0.6 within a positional window), and 23 matplotlib instances could not execute regression/reproduction tests due to an obsolete SWE-bench interface, meaning reported repair rates may understate true potential. Finally, evaluation uses APR as a proxy downstream task; generalization to developer-centric debugging workflows would require user studies the paper does not conduct. The ReAct architecture itself remains sequential, so early incorrect hypotheses can propagate uncorrected through subsequent inspections.
Conclusion
BLAgent demonstrates that combining AST-aware, path-augmented chunking, dual-perspective query transformation, and two-phase bounded agentic reranking achieves state-of-the-art file-level and function-level localization on SWE-bench-Lite at a fraction of the cost of graph-traversal agents, and that these gains translate into statistically significant end-to-end APR improvements. The central empirical lesson—that dense retrieval recall bounds the entire pipeline, and that well-selected compact context outperforms expanded context—frames the residual bottleneck clearly: function- and line-level localization now dominate unresolved cases, motivating the authors' proposed extension toward unified multi-granularity reasoning within a single pipeline.