Papers
Topics
Authors
Recent
Search
2000 character limit reached

MLmisFinder: A Specification and Detection Approach of Machine Learning Service Misuses

Published 18 Mar 2026 in cs.SE | (2603.17330v1)

Abstract: Machine Learning (ML) cloud services, offered by leading providers such as Amazon, Google, and Microsoft, enable the integration of ML components into software systems without building models from scratch. However, the rapid adoption of ML services, coupled with the growing complexity of business requirements, has led to widespread misuses, compromising the quality, maintainability, and evolution of ML service-based systems. Though prior research has studied patterns and antipatterns in service-based and ML-based systems separately, automatic detection of ML service misuses remains a challenge. In this paper, we propose MLmisFinder, an automatic approach to detect ML service misuses in software systems, aiming to identify instances of improper use of ML services to help developers properly integrate ML components in ML service-based systems. We propose a metamodel that captures the data needed to detect misuses in ML service-based systems and apply a set of rule-based detection algorithms for seven misuse types. We evaluated MLmisFinder on 107 software systems collected from open-source GitHub repositories and compared it with a state-of-the-art baseline. Our results show that MLmisFinder effectively detects ML service misuses, achieving an average precision of 96.7\% and recall of 97\%, outperforming the state-of-the-art baseline. MLmisFinder also scaled efficiently to detect misuses across 817 ML service-based systems and revealed that such misuses are widespread, especially in areas such as data drift monitoring and schema validation.

Summary

  • The paper introduces MLmisFinder, an extensible metamodel-based static analyzer that detects seven machine learning cloud-service misuses across AWS, Azure, and Google Cloud systems.
  • MLmisFinder achieved 96.7% average precision, 97% recall, and 96.8% F1 across 107 validated repositories containing 340 misuse occurrences, while averaging about 50 seconds per analysis.
  • The prevalence study found missing data-drift monitoring in 98% of 817 repositories and missing schema-mismatch safeguards in 97%, highlighting priorities for CI/CD quality checks despite static-analysis limitations.

Overview

This paper presents MLmisFinder, a fully automated static analysis approach for detecting misuses of ML cloud services in software systems. The work addresses a gap in prior research: while code smells, ML antipatterns, and deep learning API misuses have been studied extensively, misuses specific to ML cloud services—such as those offered by AWS, Azure, and Google Cloud—have received little automated detection support. MLmisFinder is grounded in a metamodel that unifies the representation of ML service-based systems across providers, and applies rule-based detection algorithms to seven misuse types. Evaluation on 107 open-source GitHub repositories yields an average precision of 96.7% and recall of 97%, substantially outperforming the only available state-of-the-art baseline. A prevalence study across 817 repositories shows that certain misuses, particularly around data drift monitoring and schema validation, occur in over 96% of systems.

Background and motivation

The paper builds on the authors' prior multivocal study, which produced a catalog of 20 ML service misuses derived from gray and academic literature, empirical examination of GitHub systems, and a survey of 50 ML practitioners [benamor2025mlmisuses]. From this catalog, seven misuses were selected for automated detection based on three criteria: coverage of multiple stages of the ML pipeline, detectability through static analysis of code and service configurations, and documented impact on maintainability and software quality.

The seven misuse types are:

  • Not using batch API for data processing: invoking per-document APIs inside loops when batch endpoints exist, causing increased network traffic, memory pressure, and cost.
  • Not using training checkpoints: failing to save or restore checkpoints, forcing full training restarts after failures.
  • Non-specification of early stopping criteria: allowing training to run excessive epochs, increasing cost and overfitting risk.
  • Ignoring testing schema mismatch: not configuring or disabling schema consistency alerts between training and evaluation data.
  • Misinterpreting model output: relying on a subset of required output metrics, e.g., using only the sentiment score while ignoring magnitude in Google's Natural Language API.
  • Improper handling of ML API limits: absence of retry/backoff or rate-limit monitoring around ML service calls.
  • Ignoring monitoring for data drift: absence of drift-detection instrumentation despite provider tooling.

The related work positioning is careful. Wan et al.'s Output Misinterpretation Checker [wan2021machine] targets only pretrained-model APIs from Google and AWS and relies on hard-coded, provider-specific heuristics. Wei et al.'s LLMAPIDet [wei2024demystifying] targets TensorFlow/PyTorch DL APIs, and MLScent [shivashankar2025mlscent] detects ML code smells at the framework level rather than the cloud service level. Neither addresses ML cloud service misuses, which motivates the new approach.

Approach

MLmisFinder takes a GitHub repository as input, clones it, parses the source code with Python's ast module, and instantiates a metamodel capturing the data required for detection. The metamodel's constituents include the System (root), Environment (injected variables), Configuration (service configuration files), ML Cloud Provider (AWS, Azure, or Google Cloud, identified via SDK invocation patterns such as boto3, azureml, and vertexai), Git Repository, and Code, which comprises imports, HTTP requests, database artifacts (train/test data identified via calls such as train_test_split, fit, predict), and a call graph derived from the AST.

On top of the instantiated model, one detection algorithm per misuse type applies rule-based checks. For example, the batch API algorithm compares observed API calls against a known list of batch endpoints and flags invocations occurring inside loops; the checkpointing algorithm verifies that checkpoint save and restore calls are present; the output misinterpretation algorithm checks whether the code consumes only a subset of the output properties required for correct interpretation. The design is explicitly extensible: new misuse types can be captured by extending or recomposing metamodel elements, and new providers can be added by extending the invocation pattern library.

Experimental design

The evaluation uses a dataset of 817 open-source ML service-based systems from GitHub. For accuracy assessment, the authors computed a required sample of 87 projects (95% confidence, ±10% margin) and manually analyzed 107 projects with three independent evaluators, achieving a Cohen's Kappa of 84.7%. The resulting ground truth contains 340 misuse occurrences and is, to the authors' knowledge, the most comprehensive ground truth for ML service misuses; it is publicly released as part of the replication package.

The baseline selection is justified by three criteria: overlap with at least one of the seven misuses, peer review with open-source implementation, and support for ML cloud providers. MLScent was excluded because its early-stopping detection operates at the framework level rather than the cloud service level, leaving Wan et al.'s Output Misinterpretation Checker as the only comparable baseline.

Effectiveness (RQ1)

Across the 107 validated systems, MLmisFinder achieved per-misuse precision between 80% and 100% (average 96.7%) and recall between 76.2% and 100% (average 97%). "Ignoring monitoring for data drift" was detected with perfect precision and recall, while "Not using training checkpoints" (precision 80%) and "Non-specification of early stopping criteria" (precision 81%) achieved perfect recall with some false positives. The lowest recall, 76.2%, occurred for "Misinterpreting output," attributed partly to regular-expression-based patterns that miss checks spread across multiple lines.

Misuse type Precision Recall F1
Misinterpreting output 100% 76.2% 86.5%
Not using batch API 90% 90% 90%
Not using training checkpoints 80% 100% 88.9%
Non-specification of early stopping 81% 100% 89.5%
Ignoring testing schema mismatch 99% 100% 99.5%
Improper handling of API limits 100% 92.6% 96.2%
Ignoring monitoring for data drift 100% 100% 100%
Overall 96.7% 97% 96.8%

The baseline comparison is stark. On the 74 repositories using Amazon and Google services (the subset supported by the baseline), Wan et al.'s tool processed 68 successfully—six failed on files exceeding 1,000 lines—and achieved 17.3% precision and 56.2% recall (F1 26.5) on "Misinterpreting output," versus MLmisFinder's 100% precision on the same category. The baseline's low precision stems from rigid, provider-coupled heuristics. This is a strong claim that rests on a single shared misuse category, so generalization of the superiority claim beyond "Misinterpreting output" is not directly established by the comparison.

Efficiency (RQ2)

Average end-to-end execution time across validation projects (up to 19,879 LOC) is about 50 seconds, with execution time scaling linearly with lines of code and file count. Most individual detection algorithms complete in roughly one second; "Misinterpreting output" takes about 8 seconds—roughly nine times faster than the baseline—while "Not using batch API" averages 35 seconds due to call graph and loop-structure analysis. Variability increases for larger projects (≥5,170 LOC), but without exponential growth. These results support practical integration into development workflows such as CI/CD pipelines, which the authors highlight as a primary use case.

Prevalence (RQ3)

Running MLmisFinder on all 817 repositories revealed that misuses are widespread. "Ignoring monitoring for data drift" appears in 98% of repositories (803 occurrences) and "Ignoring testing schema mismatch" in 97% (789 occurrences). The authors candidly qualify these figures: absence of drift monitoring in source code does not necessarily mean it is ignored, since it may live in external monitoring services, logging infrastructure, or IaC configurations that static analysis cannot see; similarly, schema validation may be genuinely unnecessary for projects using simple or standard datasets. These caveats are important when interpreting the near-universal prevalence figures.

Less prevalent but still notable are "Not using batch API" (48% of repositories), "Improper handling of API limits" (36%), "Non-specification of early stopping criteria" (24%), "Not using training checkpoints" (17%), and "Misinterpreting output" (1.5%).

Limitations and threats to validity

The paper acknowledges three principal limitations of the tool itself. First, several detection algorithms depend on manually curated lists of cloud ML libraries built from provider documentation; this curation is time-consuming, prone to omissions, and requires updates for deprecated or newly released services—the source of the 90% recall for batch API detection, since retired API documentation left some endpoints unsupported. Second, the static, rule-based analysis cannot capture runtime behavior; misuses such as data drift monitoring and API limit handling could be better detected with hybrid static-dynamic analysis. Third, coverage is limited to seven misuses, three providers, and Python; the authors note that none of the 107 validation projects relied on Infrastructure-as-Code for the relevant configurations, but broader validation across languages and providers remains open. On the construct side, the reliance on Python's ast module required a preprocessing step to skip syntactically invalid lines, which could in principle omit relevant code.

Conclusion

MLmisFinder contributes a metamodel-based, extensible static analysis approach for detecting seven ML cloud service misuse types, with strong empirical results: 96.7% average precision and 97% average recall on a manually validated ground truth of 340 occurrences, linear scalability to projects of nearly 20K LOC, and a large-scale prevalence study over 817 repositories showing near-universal neglect of drift monitoring and schema validation. The comparison with the existing baseline is convincing but limited to a single shared misuse category, and the high-prevalence findings should be read in light of the acknowledged blind spots of static analysis. The paper leaves open the extension to additional providers, languages, and misuse types, hybrid static-dynamic detection, and automated refactoring of detected misuses.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.