- The paper introduces MLmisFinder, an extensible metamodel-based static analyzer that detects seven machine learning cloud-service misuses across AWS, Azure, and Google Cloud systems.
- MLmisFinder achieved 96.7% average precision, 97% recall, and 96.8% F1 across 107 validated repositories containing 340 misuse occurrences, while averaging about 50 seconds per analysis.
- The prevalence study found missing data-drift monitoring in 98% of 817 repositories and missing schema-mismatch safeguards in 97%, highlighting priorities for CI/CD quality checks despite static-analysis limitations.
Overview
This paper presents MLmisFinder, a fully automated static analysis approach for detecting misuses of ML cloud services in software systems. The work addresses a gap in prior research: while code smells, ML antipatterns, and deep learning API misuses have been studied extensively, misuses specific to ML cloud services—such as those offered by AWS, Azure, and Google Cloud—have received little automated detection support. MLmisFinder is grounded in a metamodel that unifies the representation of ML service-based systems across providers, and applies rule-based detection algorithms to seven misuse types. Evaluation on 107 open-source GitHub repositories yields an average precision of 96.7% and recall of 97%, substantially outperforming the only available state-of-the-art baseline. A prevalence study across 817 repositories shows that certain misuses, particularly around data drift monitoring and schema validation, occur in over 96% of systems.
Background and motivation
The paper builds on the authors' prior multivocal study, which produced a catalog of 20 ML service misuses derived from gray and academic literature, empirical examination of GitHub systems, and a survey of 50 ML practitioners [benamor2025mlmisuses]. From this catalog, seven misuses were selected for automated detection based on three criteria: coverage of multiple stages of the ML pipeline, detectability through static analysis of code and service configurations, and documented impact on maintainability and software quality.
The seven misuse types are:
- Not using batch API for data processing: invoking per-document APIs inside loops when batch endpoints exist, causing increased network traffic, memory pressure, and cost.
- Not using training checkpoints: failing to save or restore checkpoints, forcing full training restarts after failures.
- Non-specification of early stopping criteria: allowing training to run excessive epochs, increasing cost and overfitting risk.
- Ignoring testing schema mismatch: not configuring or disabling schema consistency alerts between training and evaluation data.
- Misinterpreting model output: relying on a subset of required output metrics, e.g., using only the sentiment score while ignoring magnitude in Google's Natural Language API.
- Improper handling of ML API limits: absence of retry/backoff or rate-limit monitoring around ML service calls.
- Ignoring monitoring for data drift: absence of drift-detection instrumentation despite provider tooling.
The related work positioning is careful. Wan et al.'s Output Misinterpretation Checker [wan2021machine] targets only pretrained-model APIs from Google and AWS and relies on hard-coded, provider-specific heuristics. Wei et al.'s LLMAPIDet [wei2024demystifying] targets TensorFlow/PyTorch DL APIs, and MLScent [shivashankar2025mlscent] detects ML code smells at the framework level rather than the cloud service level. Neither addresses ML cloud service misuses, which motivates the new approach.
Approach
MLmisFinder takes a GitHub repository as input, clones it, parses the source code with Python's ast module, and instantiates a metamodel capturing the data required for detection. The metamodel's constituents include the System (root), Environment (injected variables), Configuration (service configuration files), ML Cloud Provider (AWS, Azure, or Google Cloud, identified via SDK invocation patterns such as boto3, azureml, and vertexai), Git Repository, and Code, which comprises imports, HTTP requests, database artifacts (train/test data identified via calls such as train_test_split, fit, predict), and a call graph derived from the AST.
On top of the instantiated model, one detection algorithm per misuse type applies rule-based checks. For example, the batch API algorithm compares observed API calls against a known list of batch endpoints and flags invocations occurring inside loops; the checkpointing algorithm verifies that checkpoint save and restore calls are present; the output misinterpretation algorithm checks whether the code consumes only a subset of the output properties required for correct interpretation. The design is explicitly extensible: new misuse types can be captured by extending or recomposing metamodel elements, and new providers can be added by extending the invocation pattern library.
Experimental design
The evaluation uses a dataset of 817 open-source ML service-based systems from GitHub. For accuracy assessment, the authors computed a required sample of 87 projects (95% confidence, ±10% margin) and manually analyzed 107 projects with three independent evaluators, achieving a Cohen's Kappa of 84.7%. The resulting ground truth contains 340 misuse occurrences and is, to the authors' knowledge, the most comprehensive ground truth for ML service misuses; it is publicly released as part of the replication package.
The baseline selection is justified by three criteria: overlap with at least one of the seven misuses, peer review with open-source implementation, and support for ML cloud providers. MLScent was excluded because its early-stopping detection operates at the framework level rather than the cloud service level, leaving Wan et al.'s Output Misinterpretation Checker as the only comparable baseline.
Effectiveness (RQ1)
Across the 107 validated systems, MLmisFinder achieved per-misuse precision between 80% and 100% (average 96.7%) and recall between 76.2% and 100% (average 97%). "Ignoring monitoring for data drift" was detected with perfect precision and recall, while "Not using training checkpoints" (precision 80%) and "Non-specification of early stopping criteria" (precision 81%) achieved perfect recall with some false positives. The lowest recall, 76.2%, occurred for "Misinterpreting output," attributed partly to regular-expression-based patterns that miss checks spread across multiple lines.
| Misuse type |
Precision |
Recall |
F1 |
| Misinterpreting output |
100% |
76.2% |
86.5% |
| Not using batch API |
90% |
90% |
90% |
| Not using training checkpoints |
80% |
100% |
88.9% |
| Non-specification of early stopping |
81% |
100% |
89.5% |
| Ignoring testing schema mismatch |
99% |
100% |
99.5% |
| Improper handling of API limits |
100% |
92.6% |
96.2% |
| Ignoring monitoring for data drift |
100% |
100% |
100% |
| Overall |
96.7% |
97% |
96.8% |
The baseline comparison is stark. On the 74 repositories using Amazon and Google services (the subset supported by the baseline), Wan et al.'s tool processed 68 successfully—six failed on files exceeding 1,000 lines—and achieved 17.3% precision and 56.2% recall (F1 26.5) on "Misinterpreting output," versus MLmisFinder's 100% precision on the same category. The baseline's low precision stems from rigid, provider-coupled heuristics. This is a strong claim that rests on a single shared misuse category, so generalization of the superiority claim beyond "Misinterpreting output" is not directly established by the comparison.
Efficiency (RQ2)
Average end-to-end execution time across validation projects (up to 19,879 LOC) is about 50 seconds, with execution time scaling linearly with lines of code and file count. Most individual detection algorithms complete in roughly one second; "Misinterpreting output" takes about 8 seconds—roughly nine times faster than the baseline—while "Not using batch API" averages 35 seconds due to call graph and loop-structure analysis. Variability increases for larger projects (≥5,170 LOC), but without exponential growth. These results support practical integration into development workflows such as CI/CD pipelines, which the authors highlight as a primary use case.
Prevalence (RQ3)
Running MLmisFinder on all 817 repositories revealed that misuses are widespread. "Ignoring monitoring for data drift" appears in 98% of repositories (803 occurrences) and "Ignoring testing schema mismatch" in 97% (789 occurrences). The authors candidly qualify these figures: absence of drift monitoring in source code does not necessarily mean it is ignored, since it may live in external monitoring services, logging infrastructure, or IaC configurations that static analysis cannot see; similarly, schema validation may be genuinely unnecessary for projects using simple or standard datasets. These caveats are important when interpreting the near-universal prevalence figures.
Less prevalent but still notable are "Not using batch API" (48% of repositories), "Improper handling of API limits" (36%), "Non-specification of early stopping criteria" (24%), "Not using training checkpoints" (17%), and "Misinterpreting output" (1.5%).
Limitations and threats to validity
The paper acknowledges three principal limitations of the tool itself. First, several detection algorithms depend on manually curated lists of cloud ML libraries built from provider documentation; this curation is time-consuming, prone to omissions, and requires updates for deprecated or newly released services—the source of the 90% recall for batch API detection, since retired API documentation left some endpoints unsupported. Second, the static, rule-based analysis cannot capture runtime behavior; misuses such as data drift monitoring and API limit handling could be better detected with hybrid static-dynamic analysis. Third, coverage is limited to seven misuses, three providers, and Python; the authors note that none of the 107 validation projects relied on Infrastructure-as-Code for the relevant configurations, but broader validation across languages and providers remains open. On the construct side, the reliance on Python's ast module required a preprocessing step to skip syntactically invalid lines, which could in principle omit relevant code.
Conclusion
MLmisFinder contributes a metamodel-based, extensible static analysis approach for detecting seven ML cloud service misuse types, with strong empirical results: 96.7% average precision and 97% average recall on a manually validated ground truth of 340 occurrences, linear scalability to projects of nearly 20K LOC, and a large-scale prevalence study over 817 repositories showing near-universal neglect of drift monitoring and schema validation. The comparison with the existing baseline is convincing but limited to a single shared misuse category, and the high-prevalence findings should be read in light of the acknowledged blind spots of static analysis. The paper leaves open the extension to additional providers, languages, and misuse types, hybrid static-dynamic detection, and automated refactoring of detected misuses.