- The paper introduces MemeTrans, the first large-scale dataset of 41,470 Pump.fun tokens that migrated to Raydium, covering 218.5 million transactions and 122 pre-migration risk features.
- The paper combines price-collapse thresholds with a neural manipulation detector, labeling 84.13% of launches as high-risk and showing that bundle-adjusted ownership, market activity, and early concentration reveal insider coordination.
- The paper finds feature-based models outperform standalone LSTM and Transformer baselines, while an MLP-guided selection strategy reduces simulated losses from 60.71% to 26.64%, a 56.1% reduction.
Motivation and problem setting
Launchpad-based token issuance has replaced direct DEX liquidity provision as the dominant mechanism for memecoin creation on Solana. Under this paradigm, tokens are sold through a bonding-curve program (e.g., Pump.fun) and only migrate to a constant-product AMM pool (e.g., Raydium) after a sale threshold is reached. This shift invalidates prior rug-pull detection methodology: because the launchpad custodies liquidity during the sale, creators no longer control pool-level operations, so existing detectors that monitor LP withdrawals or burns cannot apply. The dominant threat instead manifests as early buyers accumulating cheap inventory at low bonding-curve tiers and unwinding it into DEX liquidity after migration, draining base assets from later buyers. The authors report that investors lost over $500 million to memecoin scams in 2024 alone, motivating systematic study of pre-migration signals.
The paper introduces MemeTrans, described as the first dataset for studying and detecting high-risk memecoin launches on Solana. It covers 41,470 Pump.fun launches that successfully migrated to Raydium between December 1, 2024 and March 1, 2025, comprising 218.5 million transactions in total — 30.8 million during launchpad sales and 187.7 million in the first hour post-migration — with 1.9 million unique pre-migration accounts and 6.5 million post-migration accounts.
Dataset construction
Construction proceeds in five stages. Memecoin collection identifies migration transactions initiated by the official Pump.fun creator account transferring SOL to the Raydium fee account, then resolves mint addresses and metadata via Metaplex program-derived addresses. Transaction collection uses the Google BigQuery Solana public dataset, with a balance-change-based parser that classifies transactions into swaps, transfers, wash trades, and mints rather than relying on heterogeneous program logs. Two findings from the parsing are notable: 21.4% of pre-migration transactions are wash trades, and among "mint" transactions, 98.7% combine creation and purchase in one transaction, indicating that developers almost universally pre-acquire inventory at the earliest price tiers.
Bundle data collection addresses multi-account coordination through three heuristics: multiple accounts purchasing within a single atomic transaction; shared funder relationships (excluding centralized-exchange funders); and shared Jito bundle IDs retrieved via web crawling. After clustering and merging overlapping bundles, bundled accounts hold 36.5% of total token supply across all launches — evidence that raw per-address holding statistics substantially understate true concentration.
Feature engineering yields 122 features across five groups computed strictly from pre-migration data: contextual information (SOL price, calendar time), holding concentration (developer, sniper, top-10/20 holder shares), market activity (sale duration, transaction/trader/holder counts), bundle statistics (concentration recomputed after bundle clustering), and price/volume time series.
Risk annotation
Labels combine two components. The statistical indicator is min_price_ratio: the minimum token price within y=20 minutes post-migration normalized by the migration price, with the window chosen empirically to balance insider sell-off completion against natural speculative decay. The distribution is heavily skewed toward collapse: 60.26% of tokens fall below 0.2 of their migration price within the window, and over 73% drop below 0.4.
Because manipulators can actively manage prices to keep min_price_ratio high via structured sell-buy cycles, the authors add an ML-based manipulation detector. They manually annotate 1,555 memecoins (599 manipulated, 956 non-manipulated), release this subset, and train a TCN on post-migration price-volume trajectories, achieving 87.85% accuracy, F1 of 0.8649, and AUC of 0.9320 on the held-out test split.
The annotation rule marks a token high-risk if min_price_ratio < 0.3 or predicted manipulation score ≥ 0.7; low-risk requires both min_price_ratio ≥ 0.7 and score < 0.3; the remainder are medium-risk. Under this rule, 84.13% of launches are labeled high-risk, 11.34% medium, and 4.53% low. This extreme skew is itself a substantive finding — under these criteria, most migrated memecoins impose losses on normal buyers — but it also means the labels inherit the assumptions of both the 20-minute window and the detector's threshold choices, which the authors acknowledge as configurable rather than canonical.
Empirical observations
High-risk launches differ measurably at the launchpad stage. The first 10 and 20 buyers of high-risk tokens hold 17 and 19 percentage points more supply than those of low-risk tokens. High-risk sales complete faster, involve fewer buy transactions and fewer holders at migration, and exhibit larger average per-buyer volume. After bundle clustering, median top-10 holding percentages increase by 24%, 9%, and 6% for high-, medium-, and low-risk tokens respectively, indicating that insiders of high-risk tokens disproportionately use bundled accounts to conceal concentration. These patterns establish that pre-migration behavior carries predictive signal for post-migration outcomes.
Detection benchmark
The detection task binarizes labels (high-risk vs. combined medium/low) and filters out launches shorter than one minute or with fewer than 100 holders, since 94.9% of such cases are high-risk; this mirrors practical trader screening and mitigates class imbalance, yielding 16,048 high-risk and 5,587 normal samples. Training uses class-weighted binary cross-entropy with a 7:3 train/test split and five-fold cross-validation for hyperparameter tuning.
Results favor tabular feature-based models over sequence encoders: MLP attains the best single-model AUPRC (0.5729) versus LSTM at 0.5023 and Transformer at 0.4841 among time-series models. Ensemble combinations improve further, with MLP+LSTM reaching AUPRC 0.5827 and macro F1 0.7027. The authors conjecture that short, noisy trading sequences lack the long-range dependencies Transformers exploit. Ablation shows removing any feature group degrades performance; removing market activity causes the largest drop, and removing bundle statistics hurts more than removing raw holding-concentration features, confirming the value of bundle-revealed ownership. Removing both concentration groups together drops AUPRC to 0.5268.
In a selection-strategy application, ranking test-set tokens by predicted risk and buying the top-k reduces simulated losses from 60.71% (random, top-100) to 26.64% with MLP-guided selection — a 56.1% loss reduction — while raising non-high-risk precision from 25.4% to 80%. Notably, models were not optimized for loss minimization, so this understates achievable gains; conversely, the simulation assumes execution at migration price with random exit timing within one hour, which abstracts away slippage and MEV considerations.
Limitations and open questions
Several constraints bound the results. The dataset covers a single launchpad (Pump.fun) and a four-month window on Solana, so generalization to other launchpads, chains, or later market regimes is untested. Labels derive from heuristic thresholds (min_price_ratio cutoffs, detector score cutoffs, the 20-minute window) whose sensitivity is not systematically analyzed. The manipulation detector is trained on 1,555 manually annotated examples, and its ~90% accuracy implies label noise propagates into risk annotations. Bundle identification relies on three heuristics that likely miss coordinated accounts funded through more oblique paths, meaning reported concentration figures are lower bounds. Finally, the loss-reduction evaluation is a retrospective simulation on held-out data rather than live trading, leaving open whether predictions retain edge under adversarial adaptation by insiders aware of such detectors.
Conclusion
MemeTrans provides the first large-scale resource for launchpad-era memecoin risk research, combining 41k+ launches, 200M+ transactions, bundle traces exposing hidden multi-account coordination, 122 engineered features, and hybrid statistical-plus-ML risk annotations. Benchmarks show engineered pre-migration features outperform raw time-series modeling, and model-guided token selection cuts simulated investment losses by up to 56%. The dataset, code, and pipeline are publicly released, supporting follow-on work in fraud detection, entity matching, and trading strategy design on high-frequency Solana markets.