Papers
Topics
Authors
Recent
Search
2000 character limit reached

StableAML: Machine Learning for Behavioral Wallet Detection in Stablecoin Anti-Money Laundering on Ethereum

Published 19 Feb 2026 in cs.CR and cs.CE | (2602.17842v1)

Abstract: Global illicit fund flows exceed an estimated $3.1 trillion annually, with stablecoins emerging as a preferred laundering medium due to their liquidity. While decentralized protocols increasingly adopt zero-knowledge proofs to obfuscate transaction graphs, centralized stablecoins remain critical "transparent choke points" for compliance. Leveraging this persistent visibility, this study analyzes an Ethereum dataset and uses behavioral features to develop a robust AML framework. Our findings demonstrate that domain-informed tree ensemble models achieve higher Macro-F1 score, significantly outperforming graph neural networks, which struggle with the increasing fragmentation of transaction networks. The model's interpretability goes beyond binary detection, successfully dissecting distinct typologies: it differentiates the complex, high-velocity dispersion of cybercrime syndicates from the constrained, static footprints left by sanctioned entities. This framework aligns with the industry shift toward deterministic verification, satisfying the auditability and compliance expectations under regulations such as the EU's MiCA and the U.S. GENIUS Act while minimizing unjustified asset freezes. By automating high-precision detection, we propose an approach that effectively raises the economic cost of financial misconduct without stifling innovation.

Summary

  • The paper introduces a labeled dataset of 16,433 Ethereum wallets and 68 stablecoin-specific behavioral features spanning USDT and USDC transfers, network exposure, interactions, and timing.
  • CatBoost delivers the strongest three-class results, reaching 0.9775 Macro-F1, 0.9857 accuracy, and 0.9757 recall while distinguishing Normal, Cybercrime, and Blocklisted wallets.
  • The study finds engineered tree-based models outperform GraphSAGE because stablecoin-only transaction graphs are sparse, while direct contract interactions, transfer thresholds, and multi-hop exposure drive detection and interpretation.

Motivation and positioning

This paper addresses a specific gap in blockchain anti-money laundering (AML) research: the absence of supervised detection frameworks designed explicitly for stablecoin ecosystems. The authors note that stablecoins have become the dominant vehicle for illicit value transfer, accounting for over 84% of verified crypto fraud volumes in 2025, while existing methodological work focuses either on native assets such as Ether and Bitcoin or on graph-learning approaches that presuppose continuous transaction chains. Their central argument is that stablecoins constitute "transparent choke points": because issuers like Tether and Circle must preserve fiat convertibility under regulatory regimes such as the EU's MiCA and the U.S. GENIUS Act, these tokens retain auditable ledger histories even as the broader Ethereum ecosystem adopts zero-knowledge privacy layers.

The regulatory context is substantive rather than decorative in this paper. Issuer-level wallet freezes—over $4 billion in USDT and more than $1 billion in USDC immobilized—have demonstrably raised evasion costs but also pushed offenders toward longer multi-hop paths and cross-chain exits. The paper positions its detection framework as the analytical counterpart to this enforcement environment.

Dataset construction

The StableAML dataset is built from Transfer event logs of the official USDT and USDC contracts on Ethereum spanning 2017-11-28 to 2025-08-08. The choice of event logs over raw transactions is deliberate: emitted Transfer events record the true economic counterparties even in meta-transaction and relayer scenarios where the raw ledger attributes initiation to an intermediary. This is a defensible design decision with a clear consequence—the resulting graph excludes all non-stablecoin legs of laundering workflows, which later explains the GNN's underperformance.

Labels are curated from three sources: Etherscan community reports, disclosures from security firms (SlowMist, PeckShield), and OFAC SDN designations plus issuer freezes. The final dataset comprises 16,433 wallets across three classes: Normal (~48.7%), Cybercrime (~36.5%, covering exploits and scams/phishing), and Blocklisted (~14.8%, covering sanctioned and frozen addresses). This three-way distinction between behavioral illicitness (Cybercrime) and regulatory status (Blocklisted) is one of the paper's more useful taxonomic contributions, as the two categories exhibit measurably different transactional signatures.

Feature engineering produces 68 attributes in four groups: Interaction features (protocol-level contacts with swaps, lending, CEXs), Derived network features (second- and third-degree exposure to flagged wallets and high-value flows), Transfer-based features (volume thresholds at $1k/$5k/$10k, repeated identical-value transfers), and Temporal/Direct features (burst activity, EOA status, contract verification, longevity).

Modeling approach and headline results

The evaluation compares multinomial logistic regression, Random Forest, XGBoost, LightGBM, CatBoost, a two-layer DNN, and a two-layer GraphSAGE model, using an 80/20 wallet-level split with Macro-F1 as the primary metric given class imbalance.

The core empirical finding is unambiguous:

Model AUROC Accuracy Macro-F1 Recall
Logistic Regression 0.9385 0.8387 0.7831 0.7499
Random Forest 0.9976 0.9824 0.9714 0.9708
LightGBM 0.9976 0.9845 0.9755 0.9742
CatBoost 0.9974 0.9857 0.9775 0.9757
XGBoost 0.9974 0.9851 0.9766 0.9760
DNN 0.9617 0.9192 0.8708 0.8504
GNN (GraphSAGE) 0.9683 0.8290 0.8048 0.7699

CatBoost achieves the best balance, with per-class F1 scores of 0.9963 (Normal), 0.9852 (Cybercrime), and 0.9509 (Blocklisted). Two secondary findings deserve emphasis. First, the paper demonstrates that AUROC is systematically misleading here: logistic regression exceeds 0.93 AUROC on every class yet recovers only 0.49 recall on Blocklisted wallets—a caution against AUROC-driven model selection in skewed AML settings that the authors make explicit. Second, the GNN not only fails to beat the tabular ensembles but falls below the plain DNN (Macro-F1 ≈ 0.805 vs 0.871), which is a notable negative result given the prominence of graph learning in the prior literature.

The authors attribute the GNN failure to structural sparsity induced by their own scoping decision. Because laundering typically involves swapping stablecoins into volatile assets via DEXs, restricting the graph to USDT/USDC transfers yields density below 0.01, breaking the connectivity that message passing requires. Ensemble models circumvent this because engineered aggregates such as 2ndWithSwap encode the multi-hop information that the graph itself cannot express. The honest reading is that this is partly a property of the dataset boundary rather than a universal verdict against GNNs; the paper concedes this implicitly by framing it as a data constraint "specific to the stablecoin ecosystem."

Interpretability and typological dissection

A consensus ranking pipeline aggregating Gini importance, permutation importance on the holdout set, and SHAP values identifies direct smart-contract interactions (receivedFromSC, sentToSC) and volume thresholds (transferOver1k, transferOver5k) as the most stable predictors, followed by second-degree exposure features such as 2ndWithMultipleSameValue and 2ndWithOver10k.

Class-specific importance analysis yields the paper's most interesting qualitative claim: Cybercrime wallets are characterized by deep, structured signals—cluster membership derivatives and indirect relational exposure—consistent with automated layering by professionalized groups, whereas Blocklisted wallets show concentrated reliance on direct interactions (sentToSC, receivedFromCex) reflecting constrained post-enforcement behavior. The authors map these signatures onto the classical placement–layering–integration stages, providing a data-driven validation of the three-class taxonomy: the classes are not merely differently labeled but behaviorally distinct.

Error analysis tempers these results appropriately. The dominant false-positive source is legitimate high-frequency traders and arbitrage bots whose CEX/DeFi flow patterns structurally mimic placement and layering; residual confusion occurs between Cybercrime and Blocklisted classes due to overlapping obfuscation behaviors prior to enforcement.

Binary benchmark comparison

Collapsing to Normal vs. Suspicious, the tree ensembles reach F1 scores up to 0.9997 (LightGBM) with AUROC of 1.000, substantially exceeding Elliptic-dataset benchmarks (e.g., RevClassify at 0.953 F1, GCN at 0.973). These figures should be interpreted cautiously: the binary task on a curated labeled dataset is considerably easier than the Elliptic benchmark setting, and near-ceiling metrics of this magnitude suggest limited headroom for discriminating among strong models. Still, the comparison supports the authors' argument that domain-specific feature engineering outperforms purely topological subgraph mining in account-based stablecoin networks.

Limitations and open questions

Several constraints bear directly on the reported results. The labeled corpus of 16,433 wallets is a curated subsample drawn from reporting firms and sanctions lists, so label quality inherits the biases of those sources; the paper does not evaluate performance under label noise or adversarial label manipulation. The graph analysis is confined to the labeled subgraph with cumulative-aggregate features over the full observation window, meaning the evaluation is static and does not test temporal generalization to post-deployment laundering behavior—an important omission given the paper's own emphasis on adversarial adaptation. Whether the ensemble advantage persists when the transaction graph includes cross-asset edges (ETH, WBTC, bridges) remains open, as does robustness against adaptive adversaries who specifically optimize against the published feature set. Finally, the false-positive mechanism involving arbitrage bots raises a deployment question the paper flags but does not resolve: how to calibrate thresholds so that precision gains translate into reduced unjustified asset freezes in practice.

Conclusion

StableAML contributes a labeled stablecoin-specific dataset, a feature engineering pipeline encoding token mechanics absent from native-asset models, and a benchmark demonstrating that domain-informed tree ensembles decisively outperform both linear and graph-based deep learning on this task, with CatBoost reaching 0.9775 Macro-F1 in the three-class setting. The typological separation between high-velocity Cybercrime layering and constrained Blocklisted footprints, mapped onto classical AML stages, adds interpretive value beyond raw predictive performance. The main caveats—curated labels, static evaluation, and a deliberately sparse graph—bound the generality of the anti-GNN conclusion, but within its stated scope of deterministic, permissioned compliance tunnels, the framework offers a credible and reproducible template for stablecoin-scale AML detection.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.