Papers
Topics
Authors
Recent
Search
2000 character limit reached

Closing the Prediction Gap: Completing Machine Shape So That Predicted Time, Power, Energy, and Mapping Match What Real Hardware Does

Published 2 Oct 2026 in cs.DC | (2610.03197v1)

Abstract: A companion paper derives a weighted, communication free partition for heterogeneous, multi institution device ensembles from each device measured shape, rho_machine(d), validated on real hardware. Its experiments show rho_machine is incomplete: fp32 to fp16 speedup and power shifts are not fully indexed, and memory capacity predictions rely on an assumed, not measured, reservation overhead. This paper closes that gap, indexing five unmeasured fields in the psi selection vocabulary: memory capacity Mi<sup>capM_i<sup>{cap}, power draw PiP_i, execution unit capability XdX_d, compiler directive configuration CdC_d, and memory reservation overhead RdR_d. We introduce reached(d), a Boolean for whether compiled code dispatches to specialized hardware, explaining the NVIDIA fp16 power shift that hardware availability alone cannot predict. MoA fixes a computation independently of its target, so compiler directives get a two part admissibility criterion: select a real hardware quantity without discretionary reordering, and be deterministic, not compiler overridable. Across Intel, AMD, NVIDIA, OpenMP, OpenACC, and Open MPI, this admits only cache placement controls, deterministic vector or tile widths, and occupancy caps; heuristics like autotuning are excluded and absorbed into a residual epsilon(k)d(k)_d. We validate on five devices: A100, H100, V100, MI100, and Max 1550, measuring or establishing all five fields on each. Reachability proves real but partial, occupancy limits are device specific, and reservation overhead ranges about 0.004 to 0.030. A residual of negative 9.8 percent for the excluded Triton autotuner shows an unvalidated heuristic can act opposite a presumed penalty. These results close all five fields for every device in the companion ensemble, converting assumed behavior into measured, target specific quantities.

Authors (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.