Exact Quantile Balancing and Load-Error Injection for Mixture-of-Experts
Abstract: Mixture-of-Experts (MoE) training requires global load balance to prevent expert under-utilization and local balance for efficient expert-parallel execution. Existing distributed Quantile Balancing (QB) uses shard-dependent or approximate global quantiles, while token-independent expert biases cannot ensure microbatch-level balance. We introduce Exact Quantile Balancing (EQB), which computes exact global-batch BF16 quantiles with negligible communication, and Load-Error Injection (LEI), which injects local load errors directly into router-score gradients. On 7.5B-parameter MoEs trained for up to 500B tokens, EQB improves global balance and downstream performance over naive QB, while LEI improves local balance and outperforms the GShard loss at comparable quality.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies how to make Mixture-of-Experts (MoE) LLMs use their different “experts” more evenly.
An MoE model contains many smaller networks called experts. For each word or token, a router chooses only a few experts to process it. This saves computing power while allowing the model to have many total parameters.
However, a problem can occur: the router may send too many tokens to some experts and too few to others. The busy experts become overloaded, while the unused experts waste resources. The paper introduces two techniques to solve this:
- Exact Quantile Balancing (EQB): improves balance across the entire training batch.
- Load-Error Injection (LEI): improves balance inside smaller groups of tokens processed at the same time.
2. What questions are the researchers asking?
The researchers mainly ask:
- Can they calculate the exact global balance information needed by an MoE model, even when the training data is spread across many computers?
- Can they keep expert usage balanced not only across a whole training step, but also within each smaller local batch?
- Do these improvements make training more efficient or improve the model’s performance on tasks such as reading comprehension, mathematics, and programming?
The paper points out that global balance and local balance are different. Imagine a school with four lunch lines:
- Across the whole school day, each line might serve about the same number of students.
- But during one short lunch break, one line might be extremely crowded while another is almost empty.
The first problem is like global imbalance. The second is like local imbalance.
3. How did the researchers approach the problem?
Mixture-of-Experts routing
For every token, the router gives each expert a score. It then chooses the top few experts with the highest scores. In the experiments, each token was sent to 6 out of 256 experts.
The goal is for each expert to receive roughly the same number of tokens.
Exact Quantile Balancing
The first method, EQB, improves an earlier method called Quantile Balancing (QB).
A quantile is a value that divides a group of numbers into a certain proportion. For example, the middle value in a sorted list is the 50th percentile, or median.
QB uses quantiles to decide how much to adjust each expert’s bias. An expert that is chosen too often can be made slightly less attractive, while an expert that is rarely chosen can be made more attractive.
The difficulty is that training is distributed across many computers. Each computer sees only part of the tokens. Simply averaging each computer’s quantile is not correct, because:
The average of several local quantiles is usually not the same as the quantile of all the data together.
EQB solves this using a technique called radix selection. In everyday terms, this is like finding a particular number in a huge sorted list by sorting its digits in stages:
- First, the method looks at the higher-order part of each number.
- It finds which group contains the desired value.
- It then looks at the lower-order part inside that group.
The numbers are stored in a 16-bit format called BF16, so EQB can find the exact value in two stages using groups of 256 possible values.
Importantly, the computers communicate only small counting tables, not every token. This keeps communication cheap even when the batch contains a very large number of tokens.
Load-Error Injection
EQB handles balance across the global batch, but it cannot fully fix differences between local microbatches. A single expert bias is shared by all microbatches, even though different microbatches may contain different kinds of tokens.
LEI works differently. It checks how much each expert’s local load differs from its desired load:
- If an expert receives too many tokens, LEI pushes its router scores downward.
- If an expert receives too few tokens, LEI pushes its router scores upward.
This correction is sent directly to the router’s learning signals, or gradients. A gradient is information that tells the model how to change its settings to improve.
The paper compares LEI with the commonly used GShard auxiliary loss. The researchers argue that GShard mixes together the errors of different experts, while LEI gives each expert a more direct and separate correction.
Because very large corrections can make training unstable, the researchers limit the size of the LEI correction. This is similar to putting a speed limit on a car: the car can still move in the correct direction, but it is less likely to become uncontrollable.
4. What experiments were performed?
The researchers trained a 7.5-billion-parameter MoE LLM.
The model had:
- 256 experts
- 6 experts selected for each token
- 48 MoE layers
- Training runs of up to 500 billion tokens
They compared several systems, including:
- Rank-averaged QB
- EQB alone
- EQB with the GShard loss
- EQB with LEI
- Normalized versions of the GShard loss and LEI
They measured:
- Global MaxVio: how overloaded the busiest expert was across a whole optimizer step.
- Local MaxVio: how overloaded the busiest expert was inside local microbatches.
- Language-model quality on tasks involving general knowledge, mathematics, reading comprehension, and programming.
A smaller MaxVio value means better balance. For example, a Local MaxVio of 5.96 means the busiest expert received about 6.96 times its ideal share in the worst layer.
5. What did they find?
EQB improved global balance
After training on 100 billion tokens:
| Method | Global imbalance | Local imbalance |
|---|---|---|
| Rank-averaged QB | 0.92 | 5.96 |
| EQB | 0.74 | 5.38 |
| EQB + normalized LEI | 0.60 | 3.52 |
Compared with rank-averaged QB, EQB reduced global imbalance from 0.92 to 0.74. It also slightly improved local balance and improved results on the tested language tasks.
This shows that using the exact global quantile is better than averaging quantiles calculated separately by different computers.
LEI improved local balance
LEI was especially helpful for local balance. At 100 billion tokens:
- EQB with the GShard loss had local imbalance of 4.71.
- EQB with normalized LEI had local imbalance of 3.52.
At 500 billion tokens:
- EQB alone had local imbalance of 5.14.
- EQB with normalized LEI reduced this to 4.15.
The model’s average accuracy also increased slightly, from 48.25% to 48.56%.
The correction had to be controlled
Without limiting the size of the injected correction, both GShard and LEI sometimes caused unstable attention values during training. The normalized version avoided this problem while keeping most of the balancing benefits.
The two methods solve different problems
EQB and LEI are not competing methods that do exactly the same thing:
- EQB balances expert usage over the whole global batch.
- LEI reacts to imbalances in individual local batches.
Using them together gives the model better control at both levels.
6. Why is this important?
MoE models can be much larger than ordinary neural networks without requiring the same amount of computation for every token. But this advantage depends on experts being used efficiently.
If some experts receive too many tokens:
- They may become a computing bottleneck.
- Other experts may sit unused.
- Communication between computers may become slower.
- Training may become less stable or less effective.
The paper’s methods could help LLMs:
- Use their hardware more efficiently.
- Reduce wasted expert capacity.
- Train more smoothly.
- Possibly improve performance on language, reasoning, and programming tasks.
7. Simple conclusion
The paper presents two improvements for keeping MoE LLMs balanced.
EQB finds the exact global information needed to distribute tokens fairly, without sending huge amounts of data between computers. LEI gives direct local corrections when some experts receive too many or too few tokens.
The experiments suggest that the methods improve expert balance and can maintain or slightly improve model quality. However, the results should be interpreted carefully: the researchers tested only one model size and one training setup, and they did not measure the actual increase in hardware speed directly.
Overall, the research could make future large MoE models more efficient, more stable, and better able to use all of their experts.
Knowledge Gaps
The paper leaves the following knowledge gaps, limitations, and open questions unresolved:
- Limited model and architecture generality: EQB and LEI are evaluated only on one 7.5B-parameter decoder-only MoE with 256 experts, top-6 routing, sigmoid router scores, and a specific attention architecture; their effectiveness for different model scales, expert counts, routing values of , router parameterizations, shared-expert designs, and dense baselines remains unknown.
- No assessment across random seeds: The reported results use a single seed and do not provide confidence intervals, variance estimates, or statistical significance tests, so it is unclear whether the balance and quality improvements are robust or could reflect run-to-run noise.
- Incomplete comparison with competing balancing methods: The experiments do not comprehensively compare EQB or LEI against load-balancing controllers, auxiliary-loss-free methods, sequence-wise balancing, expert-choice routing, capacity-aware routing, or optimization-based assignment methods under matched training settings.
- Unclear contribution of exactness versus other implementation differences: EQB is compared primarily with rank-averaged QB, but the paper does not isolate whether its gains arise specifically from exact global quantiles, from differences in bias-update dynamics, from BF16 representation, or from other implementation details.
- No direct comparison with histogram resolution trade-offs: Although EQB is motivated as an exact alternative to histogram-based quantiles, the experiments do not quantify the accuracy, communication, runtime, and downstream-quality trade-offs across different histogram resolutions.
- Unresolved effect of BF16 quantization: EQB is exact only with respect to the BF16-quantized routing margins. The paper does not measure how BF16 rounding, ties, signed-zero behavior, or conversion from higher-precision router logits affects quantile accuracy and expert assignments.
- Assumption of divisible global load is restrictive: The derivation assumes is an integer. The paper does not specify how EQB handles global batches for which is not divisible by , variable-length sequences, padding, dropped tokens, or dynamically changing batch sizes.
- Tie handling is not fully resolved in practice: The theoretical discussion acknowledges that exact quantiles can produce ties and that arbitrary top- tie-breaking need not realize the balanced assignment. The paper does not evaluate how often ties occur numerically or provide a practical tie-breaking rule that preserves global balance.
- The simultaneous QB update is not theoretically characterized: The appendix states that one finite-batch quantile update is not generally a complete solution of the coupled dual problem, but convergence properties, fixed points, oscillation behavior, and conditions for repeated QB updates to stabilize remain unproved.
- Bias-update dynamics are not analyzed: The interaction between the one-step-delayed QB bias, changing router representations, optimizer updates, and LEI gradients is empirically observed but not theoretically modeled. In particular, the causes of the reported increase in global MaxVio after adding LEI at 500B tokens remain unclear.
- No convergence or stability guarantees for LEI: LEI is introduced as a straight-through gradient correction, but the paper does not establish when it decreases local imbalance, when it may oscillate, or how its behavior depends on batch size, router saturation, learning rate, optimizer momentum, or score distributions.
- The normalization rule is heuristic: The bounded residual transformation using is motivated by instability observations, but its hyperparameter , effects on optimization, and relationship to gradient clipping or adaptive scaling are not theoretically or systematically studied.
- Sensitivity to LEI hyperparameters is underexplored: Only a small sweep of is reported, and the study does not provide scaling rules for across model sizes, numbers of experts, local batch sizes, router types, or training phases.
- Separate coefficient tuning weakens the comparison with GShard: GShard and LEI use separately tuned coefficients, while the paper does not report equivalent tuning budgets, optimality criteria, or sensitivity curves for both methods. This leaves the fairness of the quality and balance comparison uncertain.
- The relation between local batch size and LEI effectiveness is unknown: Because LEI uses local hard-load estimates, its variance should depend strongly on the number of tokens per expert-parallel microbatch. The paper does not quantify how local balance and downstream quality change with microbatch size.
- Proxy metrics are not validated against actual systems performance: Local MaxVio is explicitly treated as a proxy for dispatch cost, but the paper does not report expert-parallel throughput, communication volume, padding or token-dropping rates, memory use, latency, or end-to-end training efficiency.
- EQB communication claims omit total systems overhead: The reported all-reduce payload is independent of token count, but the paper does not measure kernel execution time, synchronization latency, GPU memory traffic, radix-counting overhead, overlap with backpropagation, or scaling across different interconnects and data-parallel sizes.
- The scalability of the two-pass implementation is untested: Results use experts and 48 MoE layers. The cost and practical viability of EQB for thousands of experts, substantially more layers, larger data-parallel groups, or heterogeneous hardware remain unresolved.
- Communication cost is characterized only in payload size: The paper does not compare EQB’s two all-reduces with a single approximate histogram all-reduce or alternative distributed selection algorithms in wall-clock terms.
- No analysis of failure modes under highly skewed or nonstationary data: It is unclear how EQB and LEI behave under domain shifts, mixture-of-domains training, curriculum schedules, rare-token distributions, adversarial routing patterns, or abrupt changes in expert specialization.
- Uniform expert load may not be optimal: The methods target , but the paper does not investigate heterogeneous expert capacities, unequal expert costs, specialization-aware target loads, or routing distributions in which balanced utilization reduces model quality.
- Potential conflict between balancing and expert specialization is not studied: The reported benchmark changes do not establish whether stronger balancing suppresses useful specialization, alters expert diversity, or changes the semantic or domain structure of expert assignments.
- The causal relationship between balance and downstream quality remains unclear: EQB and LEI sometimes improve both balance and accuracy, but the paper does not determine whether quality gains result from better compute utilization, optimization regularization, reduced expert collapse, altered specialization, or confounding hyperparameter effects.
- Training loss and optimization diagnostics are incomplete: The paper reports benchmark metrics and MaxVio but does not analyze language-model loss trajectories, gradient norms, router entropy, expert parameter utilization, token-dispatch overflow, or the evolution of expert specialization.
- Attention-logit instability is only partially investigated: The paper observes runaway attention LSE with unnormalized score-side corrections, but does not identify the mechanism, establish whether instability occurs outside sliding-window attention, or compare LEI normalization with standard gradient clipping, loss scaling, or router normalization.
- Generalization beyond sigmoid routing is unresolved: The LEI derivation is presented for unnormalized sigmoid scores, and the paper does not establish whether the method remains effective or requires modification for softmax, top- softmax, normalized sigmoid, or other router score functions.
- The use of unbiased scores for mixture weights is not systematically evaluated: Selection uses biased logits while expert mixture weights use unbiased scores. The paper notes possible misalignment for GShard but does not test alternative designs in which the bias affects both selection and weighting.
- Effects on token dropping and capacity constraints are not reported: The experiments do not clarify whether improved MaxVio translates into fewer dropped or overflowed tokens under finite expert capacity, nor how EQB and LEI interact with capacity factors.
- Long-horizon evidence is limited: Only EQB and normalized LEI are trained to 500B tokens, and the paper does not provide long-horizon comparisons against all baselines, late-training behavior, or evidence that the observed improvements persist at larger token budgets.
- Evaluation coverage is narrow relative to the training claim: The downstream evaluation uses a limited set of mostly English and German benchmarks. Robustness across languages, generation quality, calibration, instruction following, safety, and production workloads remains untested.
- Reproducibility details are insufficient for independent verification: Important implementation information—such as exact optimizer settings, router initialization, bias initialization, token-dropping policy, precision choices throughout the computation, tie-breaking behavior, and hardware configuration—is not fully specified in the main paper.
- Potential implementation and notation ambiguities remain: Several displayed equations and definitions appear malformed or incomplete in the provided manuscript, making it difficult to verify the exact loss, gradient, and algorithm implementations without access to source code.
- No study of adaptive or learned balancing targets: The methods use fixed uniform targets and fixed normalization rules. Future work is needed to determine whether targets or correction strengths should adapt to expert capacity, specialization, training stage, or observed throughput.
- Interaction with optimizer state is unexplored: Since LEI modifies router-score gradients rather than adding a conventional loss, its effect on momentum, adaptive second moments, weight decay, and optimizer state differs from standard auxiliary losses; this interaction is not analyzed.
Practical Applications
Immediate Applications
- Large-scale MoE training infrastructure — EQB integration
- Replace rank-averaged quantiles or coarse global histograms with Exact Quantile Balancing (EQB) in distributed Mixture-of-Experts training.
- A training framework can implement the two-pass BF16 radix procedure as a router-balancing primitive: each data-parallel rank builds 256-bin local histograms, performs two all-reduces per MoE layer, and reconstructs the exact global quantile without gathering token-level scores.
- This is directly applicable to LLMs, multimodal transformers, code models, and other sparse neural networks.
- The reported 7.5B-parameter experiments reduced global
MaxViofrom0.92to0.74relative to rank-averaged QB and improved benchmark BPB. - Dependencies and assumptions: the router margins must be represented in an order-preserving BF16 format; the global target should be integral or require a clearly defined quantile convention; the communication fabric must support reliable low-latency all-reduce operations; results are demonstrated on one model scale and seed.
- Expert-parallel dispatch optimization — LEI integration
- Add Load-Error Injection (LEI) to the backward pass of MoE routers to reduce load skew within local expert-parallel microbatches.
- LEI can be implemented as a router-gradient hook that computes each expert’s local relative load error and adds the bounded correction
gradient += eta * normalized_load_error- to the corresponding router-score gradients.
- This can reduce padding, token dropping, idle expert capacity, and straggler effects during all-to-all dispatch.
- In the reported experiments, normalized LEI reduced local
MaxViofrom5.14to4.15at 500B tokens, although global imbalance increased from0.44to0.63. - Dependencies and assumptions: local
MaxViois used as a proxy rather than direct measured throughput; the coefficientetaand residual boundcrequire tuning; gradient clipping or normalization is necessary to avoid attention-logit instability.
- Production MoE training libraries and compiler kernels
- EQB and LEI could be packaged as optional modules in systems such as PyTorch-based MoE libraries, distributed training platforms, and accelerator-specific router kernels.
- A practical workflow would expose:
- 1. global balance mode: rank-averaged QB, histogram QB, or EQB;
- 2. local balance mode: no correction, GShard, or normalized LEI;
- 3. monitoring of global/local
MaxVio, expert capacity utilization, dispatch padding, and all-to-all latency. - This supports engineering trade-offs between communication overhead, model quality, and hardware utilization.
- Dependencies and assumptions: integration must preserve the distinction between the biased scores used for top- selection and the unbiased scores used for mixture weighting; numerical behavior must be validated across GPU and accelerator implementations.
- Training diagnostics and capacity planning
- Use separate global and local load metrics to diagnose whether an MoE system suffers from:
- persistent expert specialization or collapse, indicated by poor global balance; or
- microbatch routing skew, indicated by poor local balance despite acceptable global balance.
- This enables targeted interventions instead of increasing expert capacity or adding auxiliary losses indiscriminately.
- Dependencies and assumptions: operational teams should measure actual expert-parallel throughput, memory pressure, token padding, and dropped-token rates in addition to
MaxVio, since the paper does not establish a direct quantitative mapping fromMaxVioto throughput.
- Academic research baselines for MoE routing
- EQB provides a partition-invariant baseline for evaluating distributed quantile-balancing methods, while LEI provides a simpler alternative to the GShard auxiliary loss.
- Researchers can use these methods in experiments on routing stability, expert specialization, scaling laws, multilingual models, multimodal models, and sparse reinforcement-learning systems.
- Dependencies and assumptions: comparisons should use multiple seeds, model sizes, routing functions, and batch configurations because the paper reports one principal model scale and does not estimate variance.
- Policy and infrastructure benchmarking for efficient AI
- Organizations evaluating the energy or cost efficiency of large-model training can use EQB and LEI as candidate controls for reducing wasted expert capacity and communication inefficiency.
- A useful audit workflow would compare energy per training token, all-to-all traffic, GPU utilization, expert idle time, and final model quality under rank-averaged QB, EQB, GShard, and LEI.
- Dependencies and assumptions: energy savings are not demonstrated directly; they depend on whether improved routing balance translates into reduced padding, fewer stragglers, or higher accelerator utilization on the target hardware.
Long-Term Applications
- Scalable foundation models with lower serving and training cost
- Combining EQB for optimizer-step/global balance with LEI for microbatch/local balance could enable larger sparse models without proportionally increasing compute or communication costs.
- Potential products include multilingual LLMs, domain-specialized assistants, code-generation models, and multimodal models in which experts can specialize while remaining operationally balanced.
- The likely workflow is:
- 1. use EQB to prevent long-term concentration on a subset of experts;
- 2. use LEI to correct short-range dispatch skew;
- 3. dynamically select the balance strength based on measured throughput and quality.
- Dependencies and assumptions: the complementary behavior observed in the paper must remain stable at substantially larger expert counts, longer contexts, different top- values, and heterogeneous hardware configurations.
- Adaptive global–local routing controllers
- A future system could dynamically adjust the relative strength of EQB and LEI based on the current training state.
- For example, it could emphasize EQB when global expert collapse is detected and increase LEI when local all-to-all latency or padding rises, creating a feedback controller for distributed routing.
- Such a controller could optimize a joint objective involving model loss, global balance, local balance, communication time, and energy consumption.
- Dependencies and assumptions: the paper establishes that the two controls address different scales, but does not provide a validated joint controller or prove that improvements in the two balance metrics always improve end-to-end throughput.
- Hardware-aware and topology-aware MoE routing
- EQB and LEI could be extended so that the target load is not uniformly $1/E$ but reflects hardware topology, expert placement, memory capacity, or communication cost.
- For example, experts located on congested nodes could receive lower target loads, while experts with more memory or compute capacity could receive higher targets.
- This could support topology-aware routing on clusters with multiple GPU types, disaggregated accelerators, or geographically distributed training systems.
- Dependencies and assumptions: modifying the uniform target changes the balanced-assignment formulation and requires new theoretical and empirical analysis; network contention and expert placement must be measured accurately.
- Direct optimization of dispatch latency and energy
- LEI could be generalized from correcting expert-count error to injecting gradients derived from measured system-level costs, such as all-to-all latency, queueing delay, memory pressure, or power consumption.
- A future router might learn to balance not only tokens per expert but also the cost of sending tokens to particular devices or nodes.
- This could produce energy-aware MoE schedulers for data centers and edge deployments.
- Dependencies and assumptions: differentiable or reliable low-variance proxies for system costs are required; noisy hardware measurements may destabilize training; optimization could trade model quality against efficiency.
- Online expert specialization with controlled load
- Because EQB stabilizes global utilization while allowing experts to receive different token distributions, it could support long-term expert specialization by domain, language, modality, or task without allowing a small number of experts to monopolize traffic.
- Possible applications include continual-learning systems, multilingual routing, medical-domain models, and enterprise models with organization-specific experts.
- Dependencies and assumptions: uniform balancing may conflict with useful specialization. Future work must determine whether target loads should be uniform, task-dependent, or learned, and whether LEI suppresses or preserves meaningful specialization.
- MoE inference and serving systems
- Although the paper focuses on training, its diagnostic framework could inform inference-time routing policies, capacity allocation, and expert placement.
- Serving systems could use historical routing statistics to place frequently used experts closer to one another, preallocate capacity, or batch requests by likely expert paths.
- EQB itself is primarily a training-time update because it computes biases for subsequent optimizer steps, while its resulting biases could potentially be retained for inference.
- Dependencies and assumptions: inference has no optimizer-step statistics in the same form as training; deployment would require calibration, request-distribution monitoring, safeguards against distribution shift, and latency-aware rather than purely uniform balancing.
- Formal methods and theory for finite-batch balanced routing
- The balanced-assignment and bias-polytope analysis provides a foundation for studying uniqueness, ties, convergence, and optimality of quantile-based routing.
- Future academic tools could include convergence guarantees for iterative EQB updates, tie-aware routing implementations that explicitly realize a balanced assignment, and generalized quantile algorithms for nonuniform expert targets.
- Dependencies and assumptions: the paper notes that a single finite-batch quantile update is not always a complete solution to the coupled dual problem; additional iterations or globally coordinated tie-breaking may therefore be needed for exact balanced assignments.
- Broader applications beyond LLMs
- The methods may transfer to sparse vision transformers, recommendation systems, retrieval models, robotics policies, scientific surrogate models, and mixture-of-specialists systems where computation is conditionally activated.
- In robotics or edge AI, local balancing could reduce uneven device workloads; in recommendation systems, it could prevent a small set of specialists from becoming bottlenecks.
- Dependencies and assumptions: these applications require evidence that the routing-score distributions, batch structure, and expert-parallel execution patterns resemble those of the decoder-only MoE evaluated in the paper.
Glossary
- All-reduce: A distributed operation that combines values from all participating processes and makes the result available to each one. “Each rank forms 256-bin counts locally, and an all-reduce locates the target bin globally.”
- Auxiliary loss: An additional training objective used to encourage a desired property, such as balanced expert utilization. “The GShard auxiliary loss provides a local training signal”
- BF16: A 16-bit floating-point format, also called bfloat16, commonly used in deep-learning computation. “We exploit the 16-bit BF16 representation”
- Bipartite matching polytope: The convex geometric representation of feasible matchings between two sets of nodes. “The convex hull of is the integral bipartite -matching polytope.”
- Coercive: A mathematical property of a function that grows without bound as the norm of its argument grows. “ is coercive on .”
- Convex hull: The smallest convex set containing a given collection of points. “The convex hull of is the integral bipartite -matching polytope.”
- Convex polytope: A bounded geometric region formed by the intersection of finitely many half-spaces. “Hence is a nonempty compact polytope.”
- Data-parallel rank: One participating worker or process in distributed training that handles a portion of the data. “We require one global order statistic per expert from margins sharded across data-parallel ranks.”
- Decoder-only model: A LLM architecture that generates outputs using only Transformer decoder blocks. “All experiments use a 7.5B-parameter decoder-only model”
- Dual attainment: The property that an optimization problem’s dual optimum is achieved by an actual feasible solution. “LP strong duality and dual attainment therefore give”
- Empirical quantile: A quantile computed directly from the observed values in a finite dataset. “the exact global-batch empirical quantile”
- Expert-parallel (EP): A distributed MoE execution strategy in which experts are partitioned across devices or processes. “Local balance within an expert-parallel (EP) microbatch reduces load skew”
- Expert under-utilization: A condition in which some experts receive substantially fewer tokens or computations than others. “Global load balance to prevent expert under-utilization”
- Finite-difference: A numerical approximation to a derivative based on function values at nearby points. “a single simultaneous finite-batch update is not in general a complete solution”
- Gradient residual: The discrepancy between a desired and observed quantity that is used to modify a gradient. “LEI injects local load residuals directly into router-score gradients.”
- Hard-load fraction: The fraction of routed assignments received by a particular expert according to discrete top- routing. “define the hard-load fraction ”
- Histogram bin: An interval or category used to count values in a histogram. “the estimate remains approximate, as its resolution is limited by the histogram bin width.”
- Integral polytope: A polytope whose relevant vertices have integer coordinates, often enabling combinatorial interpretations. “the integral bipartite -matching polytope”
- Jacobian: A matrix of first-order partial derivatives describing how a vector-valued function changes with its inputs. “Its probability-dependent Jacobian couples experts”
- Load skew: Uneven distribution of computational or routing load across experts or workers. “Local balance within an expert-parallel (EP) microbatch reduces load skew”
- Log-sum-exp (LSE): A numerically useful smooth approximation to the maximum of a set of values, computed as the logarithm of summed exponentials. “We additionally track the maximum log-sum-exp (LSE) of the pre-softmax attention logits”
- Logit: An unnormalized real-valued score produced before applying a probability-generating function. “The router produces logits $z_t=W_{\mathrm r}x_t\inR^E$”
- Margin: A score difference measuring how close an expert is to entering or leaving a routing cutoff. “the distribution of routing margins”
- Microbatch: A smaller subset of a training batch processed as one local computational unit. “A single expert bias is shared across microbatches”
- Mixture-of-Experts (MoE): A neural architecture that routes each input through a sparse subset of specialized sub-networks called experts. “Sparse MoE models increase parameter capacity without proportionally increasing per-token compute”
- Order statistic: The element occupying a specified rank after values have been sorted. “radix selection, which locates an order statistic”
- Oracle: An ideal solution or assignment that achieves the optimum of a specified optimization problem. “The oracle itself is generically unique.”
- Partition invariant: Unaffected by how data are divided among distributed workers or shards. “yielding a partition-invariant estimate of the global quantile”
- Pre-softmax: Referring to values before the softmax function converts them into normalized probabilities. “the pre-softmax attention logits”
- Quantile: A value marking a specified proportion of observations below or above it. “QB chooses the corresponding quantile over the margins”
- Radix selection: An algorithm that finds an order statistic by repeatedly grouping values according to portions of their encoded keys. “which locates an order statistic by successively histogramming groups of key bits”
- Router score: A numerical value used by an MoE router to rank or weight candidate experts for a token. “LEI applies the mean unnormalized router score”
- Sharding: Dividing data or model state across multiple devices or processes. “from margins sharded across data-parallel ranks”
- Sigmoid: A function mapping real-valued inputs to the interval , often used to produce gating scores. “selected experts are weighted by the unbiased sigmoid scores”
- Sparse routing: Routing each token to only a small subset of available experts rather than all experts. “Sparse MoE models increase parameter capacity without proportionally increasing per-token compute”
- Stop-gradient: An operation that prevents a value from contributing to gradient computation during backpropagation. “where denotes stop-gradient.”
- Straight-through estimator (STE): A technique that uses one function for the forward computation and a differentiable surrogate for the backward gradient. “To obtain a useful backward signal, consider a differentiable surrogate ”
- Strong duality: The condition that the optimal values of a primal optimization problem and its dual are equal. “LP strong duality and dual attainment therefore give”
- Subgradient: A generalized gradient for a possibly nondifferentiable convex function. “its subgradient is the hard-load residual.”
- Surrogate gradient: A differentiable approximation used to provide gradients when the original operation is discrete or nondifferentiable. “The GShard loss can be viewed as using mean normalized routing probabilities as the STE surrogate.”
- Top- assignment: The discrete selection of the highest-scoring experts for each token. “ is determined by the discrete top- assignment”
- Uniform target: The desired equal allocation of tokens or load across all experts. “let be the uniform target.”
- Warmup-stable learning-rate schedule: A training schedule designed to increase the learning rate initially and then maintain stable optimization behavior. “a warmup-stable learning-rate schedule”
