- The paper empirically maps Starlink TLE error across 24,641 TLE pairs from 501 satellites and finds median 3D position errors rising from about 0.94 km with SGP4 at 6 hours to 38.5 km at 7 days.
- The paper shows that high-fidelity propagation initialized from public TLEs generally performs worse than SGP4, winning only 25.1%–35.3% of shell-pooled comparisons because model upgrades cannot correct mean-element, drag, and operator-orbit-determination biases.
- The paper identifies a practical exception in the v2-mini cohort, where high-fidelity propagation wins 72.6% of 7-day comparisons at 550 km, while concluding that improved initial states are essential for most operational applications.
Overview and motivation
This paper presents an empirical characterisation of public Two-Line Element (TLE) propagation accuracy on the Starlink megaconstellation, comparing the analytic SGP4 model against a high-fidelity numerical propagator (NASA's GMAT, configured with EGM2008 70×70 gravity, NRLMSISE-00 drag, Sun/Moon third bodies, and conical-shadow SRP). The study sweeps 24,641 starting-TLE/next-TLE pairs across 501 satellites stratified by altitude shell (540, 550, 560 km) and platform generation (v1.0, v1.5, v2-mini) over April 2026, using the operator's subsequently released TLE as proxy truth — the next-TLE-as-truth methodology of Vallado et al., here applied at constellation scale for the first time in this regime. The motivating observation is that Starlink TLEs publish at roughly six updates per satellite per day, so the operationally relevant staleness question is measured in hours rather than weeks; the corpus samples staleness horizons of 6 h, 1 d, 3 d, and 7 d.
Three hypotheses structure the analysis: (H1) how position error scales with staleness across cohorts; (H2) whether high-fidelity propagation improves accuracy when initialised from a public TLE; and (H3) whether short-window solar-flux modulation of thermospheric density is detectable in per-satellite error-growth coefficients.
Corpus construction and truth-floor calibration
The corpus is built from Space-Track gp_history data over a locked one-month window, with one starting TLE drawn deterministically per satellite-day to avoid within-satellite correlation inflation. Pairs are matched at four target offsets within a ±2 h tolerance, and a semi-major-axis-jump maneuver filter (100 m threshold, calibrated empirically after Lemmens' technique) removes 19.1% of candidate pairs whose intervals enclose station-keeping events. The authors quantify the selection effect of the matching tolerance directly: unmatched fractions (~50% overall) are dominated by Poisson-window statistics rather than tracking-gap selection, since the 95th-percentile worst per-satellite gap (~42 h) sits well inside the longest target offset.
A key methodological contribution is an empirical truth-floor diagnostic. Because the next-TLE proxy inherits the operator's orbit-determination residual at both endpoints, no comparison against it can drop below an O(1) km floor regardless of propagator quality. Propagating earlier TLEs forward with the same GMAT configuration yields per-cohort medians of 0.81–1.37 km at ~5 h arcs, pooled at 1.11 km (IQR 0.23–2.89 km). Notably, this floor exceeds the SGP4 pooled median at 6 h (~0.94 km), which the authors attribute to partial cancellation between adjacent operator OD residuals sharing common-mode biases. The consequence is stated plainly: the 6 h headline is a measurement constrained by public-TLE delivery precision, not by propagator skill, uniformly across the corpus.
Staleness curves and power-law fits (H1)
Per-cell fits of ∥Δr(Δt)∥≈AΔtk, estimated by per-pair OLS with satellite-level bootstrap confidence intervals (1,000 resamples preserving within-satellite correlation), yield exponents spanning both sides of unity: sub-linear (k≈0.75–$0.86$) on SGP4 v1.x cells, super-linear (k≈1.15–$1.44$) on all v2-mini cells and on the high-fidelity v1.x arm at 540 and 560 km, with the high-fidelity v1.x cell at 550 km sitting at the k=1 boundary (the k=1 null not rejected, p=0.76). The likelihood-ratio test rejects both the O(1)0 and O(1)1 nulls elsewhere at O(1)2, supporting a mixture interpretation: O(1)3 encodes the relative weight of constant mean-motion bias (linear along-track displacement from O(1)4 mis-fit) versus constant unmodelled in-track acceleration (quadratic displacement from drag-area mis-specification).
Pooled L₂ medians grow from ~0.94 km (SGP4) / 1.49 km (high-fid) at 6 h to 38.5 / 76.0 km at 7 d. On v1.x the high-fidelity arm tracks SGP4's curve shape but sits a factor of 1.6–3.6 higher throughout; on v2-mini the ordering flips with staleness, with the high-fidelity median falling below SGP4 at 3–7 d. A linear mixed-effects cross-check confirms slope agreement across all populated cells.
Pair-by-pair propagator comparison (H2)
The central negative result: the high-fidelity propagator initialised from a public TLE does not improve over SGP4 at any shell-pooled horizon. High-fidelity wins on only 33.7% [32.2, 35.3] of 6 h pairs, 25.1% at 1 d, 28.7% at 3 d, and 35.3% [31.9, 38.8] at 7 d in 3D L₂ norm — statistically distant from the 50% null everywhere, and robust to the choice of metric since the along-track-only variant returns nearly identical fractions beyond 1 d. Median high-fidelity error exceeds SGP4 by factors of roughly 1.6–2.0 across buckets.
The paper decomposes this bulk negative result into three mechanisms that partition cleanly by horizon:
- Mechanism 1 (short Δt): operator-OD residual dominance inflates absolute norms at 6 h but cancels from the paired differential, so it does not bias win counts.
- Mechanism 2 (6 h–1 d): truth-construction kernel alignment. Both the SGP4 prediction and the proxy truth pass through the same analytic kernel, so SGP4's systematic biases mirror and cancel; the high-fidelity arm integrates a pseudo-osculating initial condition living on SGP4's mean-element manifold into EGM2008/NRLMSISE-00 osculating dynamics without compensation, carrying an O(1)5 km Brouwer short-period mismatch.
- Mechanism 3 (3–7 d): preferential spacecraft-property bias amplification. A systematic O(1)6 bias feeds directly into the high-fidelity drag acceleration, whereas SGP4 absorbs the equivalent bias into the operator-fitted O(1)7. A ±20% O(1)8 sensitivity shifts v2-mini medians by at most ±17%, insufficient to invert the H2 sign even at the favourable 0.8× factor.
The RSW decomposition confirms the entire H2 gap lives along-track: the along-track median sits two decades above radial and cross-track at every bucket, and by 7 d the L₂ norm is indistinguishable from the along-track component.
Cohort-resolved exception: v2-mini at long staleness
The single regime where high-fidelity overtakes SGP4 on a majority of pairs is cohort-specific, not shell-specific. On v2-mini alone, high-fidelity wins 72.6% [69.9, 75.4] of 7 d pairs at 550 km and 55.1% [49.9, 60.4] at 560 km, while on v1.x it loses badly at both shells (16.1% and 11.9%). The apparent shell-fragility of this exception in pooled statistics is shown to be a blending artefact of differing v2-mini population ratios across shells. The physical reading is coherent: v2-mini carries super-linear SGP4 exponents (O(1)9 at 550 km), the signature of a residual in-track acceleration the operator-fitted ∥Δr(Δt)∥≈AΔtk0 has not absorbed, so where SGP4's averaged drag fit fails most visibly, an explicit-drag integrator has the most room to gain — more so at lower altitude.
The authors flag three caveats honestly: the v2-mini drag area is bus-size-scaled rather than fitted to ballistic data; correlated biases in the newer cohort's ∥Δr(Δt)∥≈AΔtk1 cannot be separated from public TLEs alone; and the 560 km × v2-mini × 7 d cell carries only ∥Δr(Δt)∥≈AΔtk2 pairs, small enough that a halved corpus could move its win fraction across the 50% null.
Solar modulation (H3)
An exploratory ANCOVA regression of the per-satellite SGP4 coefficient ∥Δr(Δt)∥≈AΔtk3 against daily observed F10.7, with generation intercepts absorbed as nuisance covariates, returns a positive slope clearing conventional significance at only one of three shells: ∥Δr(Δt)∥≈AΔtk4 sfu⁻¹ [0.001, 0.031], ∥Δr(Δt)∥≈AΔtk5 at 560 km, versus direction-consistent but non-significant results at 540 km (∥Δr(Δt)∥≈AΔtk6) and a zero-consistent result at 550 km (∥Δr(Δt)∥≈AΔtk7). The interaction test rejects the additive restriction at 560 km (∥Δr(Δt)∥≈AΔtk8), flagged as either a real cohort-specific F10.7 response or a window-end sampling artefact. The 81-day-centred robustness fits return physically wrong-signed slopes attributable to the predictor's compressed ~1.3 sfu span, and are correctly demoted to limit-of-leverage diagnostics. The verdict is appropriately conservative: direction-consistent with the LEO density-gradient expectation, but not a calibrated F10.7-modulation measurement on a 30-day window spanning only ~17 sfu.
Limitations
The analysis is explicit about its constraints: a single constellation (Starlink); a 30-day moderate-activity window limiting H3 power; reliance on NRLMSISE-00 where NRLMSIS 2.1 or JB2008 could shift modelled densities by 10–20%; truth construction bounded by operator OD residuals unavailable publicly; the unscaled v2-mini spacecraft-property prior; and the maneuver filter's blindness to purely continuous low-thrust drift modes below the SMA-jump detection threshold. None of these undermines the headline H2 result, which survives the ∥Δr(Δt)∥≈AΔtk9, maneuver-threshold, and predictor-choice sensitivity tests quantitatively.
Conclusion
This study delivers a cohort-resolved staleness atlas for public Starlink TLEs with three practitioner-facing conclusions. Public TLEs are adequate inside the operator's ~4-hour update interval, with ~1 km median error dominated by OD residual rather than propagator choice. High-fidelity propagation from public-TLE inputs should not be invested in unless improved initial states accompany it — the gain from force-model upgrade alone is negative in every shell-pooled cell except the v2-mini long-staleness regime, where mechanism-2's kernel-alignment advantage erodes against SGP4's own fit failure. The per-cell k≈0.750 table is released as a reusable benchmark target for enhanced-propagator work (SGP4-XP, differentiable SGP4, ML correctors) and as the deterministic baseline against which future covariance-realism studies must be consistent. The principal open questions left by the paper are whether the three-mechanism partition holds under alternative atmosphere models and cross-constellation replication, and whether a longer corpus straddling a solar-cycle transition would convert the H3 signal from directional evidence into a calibrated density-modulation measurement.