Papers
Topics
Authors
Recent
Search
2000 character limit reached

AutoAnalyzer: Autonomous Program Analysis

Updated 16 July 2026
  • AutoAnalyzer is a family of autonomous program-analysis systems that automate data collection, performance debugging, and root-cause diagnosis in various computing environments.
  • It employs techniques such as automated instrumentation, clustering algorithms, decision-tree synthesis, and resource-aware tuning to optimize code analysis.
  • These systems balance precision, coverage, and computational cost while integrating validation layers and autonomous repair methods for reliable analysis.

Searching arXiv for papers on “AutoAnalyzer” and closely related automated program-analysis systems. AutoAnalyzer denotes both a specific automatic performance-debugging system for SPMD-style parallel programs and a broader design pattern for autonomous program-analysis systems. In the original sense, AutoAnalyzer automates data collection, performance behavior analysis, bottleneck localization, and root-cause diagnosis for SPMD programs, without apriori knowledge and with lightweight performance data (Liu et al., 2010, Liu et al., 2011). In later usage, the term is extended to systems that infer analyzer specifications, tailor abstract interpreters to code and resource budgets, learn transfer rules from data, classify and repair static-analysis warnings, or perform repository-level auditing with LLM-based agents (Shemetova et al., 5 May 2025, Mansur et al., 2020, Bielik et al., 2016, Joos et al., 15 Sep 2025, Guo et al., 30 Jan 2025).

1. Origins in automatic performance debugging

The earliest work using the name AutoAnalyzer presents an automatic performance debugging system for SPMD-style parallel programs. Its architecture has four major components—automatic instrumentation, data collector, data management, and data analysis—and it decomposes a program into a code region tree whose nodes are functions, subroutines, loops, and nested regions (Liu et al., 2010, Liu et al., 2011). The system collects wall clock time, CPU clock time, MPI communication time and volume, disk I/O quantity, and hardware metrics such as clock cycles, instructions retired, L1 and L2 cache misses, and derived CPI. The 2010 formulation emphasizes lightweight collection: if there are nn code regions and mm processes, the primary data needed for bottleneck detection is only 125×n×m125 \times n \times m bytes (Liu et al., 2010).

The analysis stage distinguishes two bottleneck classes. For process behavior dissimilarity, AutoAnalyzer builds per-process vectors of code-region CPU times and applies a simplified OPTICS clustering algorithm; if multiple clusters appear, the system reports dissimilarity bottlenecks. For code-region disparity, it defines the code region normalized metric

CRNM=CRWTWPWT×CPICRNM = \frac{CRWT}{WPWT} \times CPI

and uses k-means to classify regions into severity categories from very low to very high (Liu et al., 2011). A top-down search over the code region tree then identifies critical code regions and their cores, while rough set analysis maps these bottlenecks to root-cause attributes such as L2 cache miss rate, disk I/O quantity, network I/O quantity, or instructions retired (Liu et al., 2010).

The reported evaluations established the term in a concrete HPC setting. On ST, AutoAnalyzer identified a dissimilarity bottleneck in code region 11 and internal bottlenecks in regions 8 and 11; eliminating dissimilarity bottlenecks yielded about 40% speedup, eliminating disparity bottlenecks yielded 90%, and combined optimizations yielded about 170% improvement over the original code (Liu et al., 2010, Liu et al., 2011). On NPAR1WAY, the system reported no dissimilarity bottleneck but identified internal bottlenecks dominated by network I/O and instructions retired, and the resulting optimization produced about 20% speedup (Liu et al., 2010). On MPIBZIP2, it localized a computation-heavy compression region and a network-heavy send region, but concluded that there was no easy optimization path (Liu et al., 2011).

2. Specification mining and learned analyzer construction

A second lineage treats AutoAnalyzer as a system that manufactures the analyzer itself or its semantic summaries. LAMeD explicitly frames itself as an intelligent “front-end” for an AutoAnalyzer: static leak analyzers for C and C++ depend on function specifications such as allocation sources and deallocation sinks, but manual annotation is laborious, error-prone, hard to maintain, and especially difficult for third-party and large libraries (Shemetova et al., 5 May 2025). LAMeD uses zero-shot prompting over function bodies, optionally enriched with callee context extracted by Joern, to infer allocated_variables and deallocated_variables, then translates the resulting JSON into analyzer-specific annotations for Cooddy and CodeQL. On a manually annotated cJSON dataset of 152 functions, Codestral + post-filtering achieved TP=28TP=28, FP=2FP=2, FN=20FN=20, with precision ≈0.933\approx 0.933 and recall ≈0.583\approx 0.583; on 43 real leaks from DiverseVul projects, both CodeQL and Cooddy improved from 5 detected target bugs to 10 when equipped with LAMeD annotations, although warnings rose substantially, especially on libxml2 (Shemetova et al., 5 May 2025).

“Learning a Static Analyzer from Data” pushes the idea further by synthesizing the analyzer rules themselves from examples rather than annotations. The method represents an analysis as a program in a loop-free DSL with if-then-else, Move operations over syntax and call traces, and Write operations that record contextual features; synthesis is performed by an ID3-like decision-tree learner whose leaves are candidate transfer programs and whose internal nodes are guards selected by information gain (Bielik et al., 2016). The learner is embedded in a counterexample-guided loop: after a candidate analyzer is synthesized, an oracle mutates programs using equivalence-modulo-abstraction transformations and global jumps, then searches for examples on which the candidate is incorrect. The implementation learned JavaScript points-to rules for this in built-in APIs and allocation-site rules over NodeJS v4.2.6 and the ECMAScript test262 suite; the paper reports that the system automatically discovered practical and useful inference rules for many cases that are tricky to manually identify and are missed by state-of-the-art, manually tuned analyzers (Bielik et al., 2016).

These systems share a common architectural move: instead of hard-coding all semantic knowledge into the analyzer core, they treat semantic summaries, library models, or transfer rules as artifacts that can themselves be generated. This suggests an AutoAnalyzer can be built as a pipeline in which inference of analyzer knowledge is a first-class stage rather than a one-time manual engineering task.

3. Configurable precision and code-specific tailoring

Another major interpretation of AutoAnalyzer concerns automatic tailoring of analyzer precision to code and resource constraints. TAILOR formalizes analyzer configuration as an optimization problem over “ingredients” and “recipes”: an ingredient is an abstract domain together with its internal settings, and a recipe is a finite sequence of ingredients, each run in turn with previously proved assertions turned into assumptions for the next ingredient (Mansur et al., 2020). Even when restricting to recipes of at most length 3, the paper reports more than 6M configurations. Its cost function minimizes the fraction of unproved assertions while enforcing a time budget and using runtime only as a secondary tie-breaker; the evaluation considers 1 second and 5 minute budgets, corresponding to editor-style and CI-style usage scenarios (Mansur et al., 2020).

The results show that the best configuration varies substantially by file and by resource limit. Across 120 LLVM files drawn from curl, darknet, ffmpeg, git, php-src, and redis, all four search strategies—random sampling, domain-aware random sampling, simulated annealing, and hill climbing with restarts—roughly doubled the number of assertions proved compared with Crab’s default recipe (Mansur et al., 2020). The “most precise” baseline was frequently impractical: on 21 of 120 files it ran out of 264GB of memory, and on the files where it did terminate within the 5 minute budget it still proved fewer assertions than the best tailored recipes (Mansur et al., 2020). The study also found that most learned configurations remain effective across several subsequent code versions.

AutoAlias represents a related but earlier precision-control mechanism for alias analysis. Implemented inside EiffelStudio, it is a context-sensitive, call-site-sensitive, flow-sensitive analyzer for object-oriented programs, based on alias diagrams and duality semantics, with precision controlled by a constant governing the number of fixpoint iterations at creation sites (Rivera et al., 2018). Higher values increase precision and computation time, while smaller values summarize heap growth earlier. The reported performance was about 25 seconds for roughly 8000 lines of intricate code in EiffelBase 2 and about 232 seconds for roughly 150000 lines in EiffelVision (Rivera et al., 2018).

Taken together, these systems recast AutoAnalyzer not as a single fixed algorithm but as a configurable analysis platform whose effective behavior depends on learned summaries, tailored domain combinations, or explicit precision dials.

4. Autonomous warning triage, repair, and repository auditing

Recent work turns AutoAnalyzer into an end-to-end agent that not only detects issues but also classifies, repairs, validates, and reports them. CodeCureAgent is presented explicitly as such a system for Java/SonarQube: for each warning it runs a classification sub-agent, a repair sub-agent, and a change approver that enforces a three-step heuristic—build the project, rerun SonarQube and ensure the target warning disappears without introducing new warnings, and run the test suite (Joos et al., 15 Sep 2025). The evaluation covers 1,000 SonarQube warnings from 106 Java projects and 291 distinct rules. The system produced plausible fixes for 968 of 1000 warnings, corresponding to a 96.8% plausible-fix rate, and manual inspection on the 291-rule sample reported an 86.3% correct-fix rate; the average processing cost was about 2.9 cents per warning and the mean time was 4.4 minutes (Joos et al., 15 Sep 2025). The ablation study is particularly diagnostic: full checks yielded 968 plausible fixes and 0 wrongly accepted, whereas disabling all checks yielded 751 plausible fixes and 249 wrongly accepted (Joos et al., 15 Sep 2025).

RepoAudit extends the same autonomy to repository-level bug auditing. It combines an Initiator that finds source values, an Explorer that performs demand-driven path-sensitive analysis with agent memory, and a Validator that checks control-flow alignment and inter-procedural path feasibility (Guo et al., 30 Jan 2025). Its memory is explicitly structured as

M(f,v@s)={(pi,Ri)},\mathcal{M}(f, v@s) = \{(p_i, R_i)\},

where each mm0 is a feasible path in function mm1 and each mm2 is a set of data-flow facts of the form mm3. The abstract states that RepoAudit detects 40 true bugs across 15 real-world benchmark projects with a precision of 78.43%, requiring on average only 0.44 hours and $m$42.54 per project (Guo et al., 30 Jan 2025). In both descriptions, the defining features are on-demand inter-procedural traversal, explicit agent memory, and validator-driven hallucination control.

These systems illustrate a shift from report-only analysis to autonomous remediation. A plausible implication is that, in modern usage, AutoAnalyzer increasingly denotes a closed loop: detect, interpret, act, and validate.

5. Verification-oriented, structural, resource, and pedagogical analyzers

AutoAnalyzer also names or inspires systems outside bug triage. In formal verification, the SPARK proof-tool literature describes an analyzer that helps users help the analyzer through layered feedback roles—“The Nurse,” “The Investigator,” “The Magician,” and “The Surgeon.” GNATprove provides counterexamples, identifies the smallest failing sub-property, points to likely root causes such as missing loop invariants or postconditions, and exposes verification conditions for expert inspection (Moy, 2021). Although this work is not branded as a standalone AutoAnalyzer product, it explicitly develops design principles for an AutoAnalyzer in verification settings.

For heap-shape analysis, forest automata provide a fully automated analyzer that learns hierarchical “boxes” representing repetitive graph patterns. The 2013 forest-automata paper removes the need for manual box design by automatically learning knots during abstraction, enabling a fully automated analysis that handles data structures as complex as skip lists, with performance comparable to state-of-the-art fully automated tools based on separation logic that specialize in linked lists (Holik et al., 2013). The implementation in Forester verifies nested lists, trees with parent pointers, and 2- and 3-level skip lists without user-provided boxes (Holik et al., 2013).

For resource analysis, the reusable machine-calculus work introduces a CBPV abstract machine with a polymorphic and linear type system enhanced with a first-order logical fragment, then implements Automated Amortized Resource Analysis from scratch in this framework (Suzanne et al., 2023). Resource usage is encoded as a state-passing effect via an ST(p,q) token type with debit, credit, and slack constructors, and the inferred bounds are closed-form multivariate polynomials over the integers (Suzanne et al., 2023). The system is positioned as the basis for an experimental toolkit for automated memory analysis of functional languages.

In education, static analysis is reconfigured as an AutoAnalyzer for novice programmers by extending PMD with a “Novice” ruleset and novice-oriented feedback messages. When evaluated on 24 student Java projects and six Apache Commons Codec components, student code exhibited 164% more novice-rule warnings per kLoC than expert code; all 24 student projects generated at least one novice warning, while only two of six professional projects did so (Blok et al., 2017). Rule V, “instance variable not being used globally within the class,” accounted for 78.4% of novice warnings (Blok et al., 2017). Here AutoAnalyzer functions less as a verifier or bug finder than as a concept probe for programming misconceptions.

6. Recurring trade-offs, misconceptions, and outlook

Across these lineages, AutoAnalyzer is not a single algorithm or product but a family of systems organized around high automation, analyzer self-configuration, and structured validation. One recurring misconception is that more automation simply means replacing static analysis with a LLM. The literature points in the opposite direction: LAMeD uses LLMs only for specification mining while retaining classical analyzers such as Cooddy and CodeQL in the loop; RepoAudit augments LLM exploration with control-flow alignment and path-feasibility validators; CodeCureAgent requires build, re-analysis, and tests before accepting a patch (Shemetova et al., 5 May 2025, Guo et al., 30 Jan 2025, Joos et al., 15 Sep 2025).

Another recurring theme is the precision–coverage–cost trade-off. LAMeD doubled the number of known leaks detected by CodeQL and Cooddy on its real-life dataset, but the warning volume rose sharply, with a strong correlation between the number of annotations and warnings (Shemetova et al., 5 May 2025). TAILOR shows that default and “most precise” analyzer configurations are both often inferior to code-specific, budget-aware recipes (Mansur et al., 2020). AutoAlias exposes this trade-off directly through a precision constant; higher values mean better precision and higher computation time (Rivera et al., 2018). RepoAudit bounds call-chain depth at 4 and is explicitly not sound, trading completeness for repository-scale practicality (Guo et al., 30 Jan 2025).

Validation and explainability therefore emerge as central design constraints. In CodeCureAgent, dropping the approval checks leads to a steep increase in wrongly accepted patches; in RepoAudit, removing validators or abstraction substantially degrades precision; in the SPARK setting, the analyzer must explain abstraction boundaries such as loop invariants, contracts, and initialization obligations if users are to trust and extend the proof (Joos et al., 15 Sep 2025, Guo et al., 30 Jan 2025, Moy, 2021). A plausible implication is that the mature form of AutoAnalyzer is hybrid: it combines automatic synthesis, program-structure-aware reasoning, and explicit validation layers rather than relying on any single inference mechanism.

Viewed historically, AutoAnalyzer began as an HPC performance-debugging system that clustered performance vectors and used rough sets for root-cause analysis (Liu et al., 2010, Liu et al., 2011). It has since become a broader research category encompassing learned analyzers, inferred annotations, configurable abstract interpretation, autonomous repair agents, repository-scale LLM auditors, shape analyzers, resource analyzers, and pedagogical feedback systems. What unifies these otherwise diverse systems is the attempt to reduce manual analyzer engineering while preserving enough semantic structure, validation, and technical depth to remain useful on real software.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to AutoAnalyzer.