Determine the clean-cache full-read cost for a 140 GB lazy-mounted model

Determine the deferred full-read cost of a 140 GB model served through an eStargz lazy mount under a clean cache with ample capacity, in order to complete the deferred-cost curve beyond the measured 2 GB and 14 GB cases.

Background

The paper measures the cost of fully reading lazily mounted model images at 2 GB and 14 GB, finding that the deferred read becomes slower than the eager pull at 14 GB. A clean, adequately provisioned 140 GB full-read experiment would be needed to establish how this cost scales at the largest evaluated model size.

The authors report that four attempts failed for environment-related reasons and therefore leave the 140 GB point quantitatively unresolved, although they state that the trend is qualitatively confirmed.

References

We could not obtain a clean 140\,GB full read with an amply sized cache (four attempts failed for environment reasons documented in the artifact release), so the deferred-cost curve is quantified at 2 and 14\,GB and qualitatively confirmed at 140\,GB.

— The Lazy Pod That Lies: Deferred Cost and Failure Semantics of Lazy Container Image Pulling for Model Serving on Kubernetes  (2608.19412 - Kliukovkin, 19 Aug 2026) in Section 5.1, “Full-read cost and the FUSE tax”