Papers
Topics
Authors
Recent
Search
2000 character limit reached

Luce-Distributed Permutations

Updated 10 July 2026
  • Luce-distributed permutations are probability models that generate rankings by sequentially sampling items with probabilities proportional to positive weights.
  • These models encompass standard, extended, and inverse variants, enabling flexible formulations in ranking, Bayesian inference, and pairwise comparison frameworks.
  • Recent research focuses on efficient computational methods, asymptotic analysis, and applications in rank aggregation, causal learning, and secretary-type online decision problems.

Luce-distributed permutations are probability distributions on permutations generated by sequential Luce choices: at each step, one samples without replacement from the remaining items with probabilities proportional to prescribed positive weights. In the statistical literature, the standard Plackett–Luce model, the Extended Plackett–Luce model, and inverse-Luce variants are all instances of this general idea, while broader Gibbs-posterior formulations recover Luce-family noisy-comparison models such as Bradley–Terry–Luce. Recent work studies these distributions from several directions: parametric ranking models, Bayesian inference, message-passing approximations, asymptotic limit theory, stochastic-gradient methods for latent permutations, and secretary-type online decision problems (Johnson et al., 2020, Cantwell et al., 2021, Borga et al., 9 Sep 2025, Pinsky, 11 May 2026, Gadetsky et al., 2019).

1. Sequential-choice definition and equivalent continuous representations

In the standard Plackett–Luce formulation, a complete ranking of KK entities is a permutation x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K, where xjx_j is the entity assigned to rank jj. Each entity kk has a positive worth parameter λk>0\lambda_k>0, and the model assigns probability

Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.

This is a discrete distribution on all K!K! permutations. In the logit parameterization used for black-box optimization, for a permutation b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k and score vector θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k),

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K0

Both expressions encode the same stagewise mechanism: the first item is chosen from all items, the second from the remaining items, and each selection obeys Luce’s choice axiom (Johnson et al., 2020, Gadetsky et al., 2019).

A weight-based formulation emphasizes sampling without replacement. For positive weights x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K1 with total mass

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K2

define

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K3

The law on x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K4 is then

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K5

Equivalently, if x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K6 is the draw order, then x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K7, so x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K8 is the draw time of label x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K9 (Borga et al., 9 Sep 2025).

Two continuous representations are central. In one, if xjx_j0 are independent exponentials with mean xjx_j1, then comparisons among the xjx_j2’s encode the Luce order. In the other, if

xjx_j3

then the event xjx_j4 has probability exactly equal to the Plackett–Luce probability. These representations are the basis of both asymptotic analysis and low-variance gradient estimators (Borga et al., 9 Sep 2025, Gadetsky et al., 2019).

The mode has a simple form in the standard model. Under standard Plackett–Luce, the modal ordering is

xjx_j5

and, in the score parameterization, the mode is the permutation that sorts scores in descending order (Johnson et al., 2020, Gadetsky et al., 2019).

2. Extended, reverse, and inverse Luce models

The Extended Plackett–Luce (EPL) model generalizes standard Plackett–Luce by introducing a choice order

xjx_j6

which specifies the order in which ranks are filled during the ranking process. Its probability mass function is

xjx_j7

Equivalently, it is the standard Plackett–Luce probability applied to permuted data

xjx_j8

The model therefore induces a flexible discrete distribution on xjx_j9 in which the sequence of Luce choices is itself permutation-valued (Johnson et al., 2020).

Two special cases are explicit. If jj0, EPL reduces to the standard forward Plackett–Luce model. If jj1, it becomes the reverse Plackett–Luce model. The worth parameters jj2 still govern selection probabilities, but for general jj3 their direct interpretation as “preference for rank jj4, rank jj5, \dots” is no longer immediate because ranking stages are traversed in the order jj6, not necessarily from best to worst. The modal ranking under EPL is

jj7

so the choice order re-labels the ranking stages (Johnson et al., 2020).

The parameterization is mixed discrete/continuous:

  • discrete component: jj8,
  • continuous component: jj9.

The likelihood is invariant under scalar multiplication,

kk0

so the model is identifiable only up to scale. The cited work notes that this is not a major substantive issue, but it can affect MCMC mixing (Johnson et al., 2020).

A distinct but closely related variant is the inverse-Luce model. For positive weights kk1, the direct Luce distribution on kk2 is

kk3

Its inverse-Luce version is the pushforward under inversion kk4: kk5 The inverse-Luce model is the one that becomes tractable in the secretary problem when the smallest number is the top rank (Pinsky, 11 May 2026).

3. Statistical inference and computational methods

Bayesian inference for EPL starts from data kk6 and likelihood

kk7

The prior is factored as

kk8

For the choice order, the cited work uses a Plackett–Luce prior

kk9

which is uniform over λk>0\lambda_k>00 if λk>0\lambda_k>01 for all λk>0\lambda_k>02. Conditional on λk>0\lambda_k>03, the worth parameters have independent Gamma priors,

λk>0\lambda_k>04

A central feature is a mode-preserving prior predictive framework: the hyperparameters λk>0\lambda_k>05 are chosen so that the modal ordering of the prior predictive distribution is preserved across different choice orders (Johnson et al., 2020).

Posterior sampling is carried out with Metropolis-coupled MCMC (MCλk>0\lambda_k>06) / parallel tempering because the posterior over λk>0\lambda_k>07 can have local modes separated by large distances in permutation space. The algorithm alternates between updating λk>0\lambda_k>08, updating λk>0\lambda_k>09, rescaling Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.0, and swapping states between tempered chains. The paper recommends geometric temperature spacing and swapping adjacent chains only. Predictive inference is explicit: Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.1 The posterior predictive mode is interpreted as the aggregate ranking, while discrepancies

Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.2

are used to detect lack of fit (Johnson et al., 2020).

A different inference perspective is given by Gibbs posteriors on permutations from pairwise comparison data. With a directed comparison graph Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.3, a uniform prior over permutations, and edge factor Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.4, the posterior is

Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.5

This is interpreted as

Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.6

The framework includes a hard step-function Hamiltonian for partial orders and a Bradley–Terry–Luce noisy-comparison model with

Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.7

Belief propagation computes approximate marginal position distributions, while the Bethe free energy is used both to approximate the number of linear extensions and to perform model selection between competing ranking models such as the step-function model, Bradley–Terry–Luce, and related alternatives (Cantwell et al., 2021).

For optimization over latent permutations, low-variance score-function estimators have been developed specifically for the Plackett–Luce distribution. The task is

Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.8

where Pr(X=xλ)=j=1Kλxjm=jKλxm.\Pr(X=x\mid \lambda)=\prod_{j=1}^{K} \frac{\lambda_{x_j}}{\sum_{m=j}^{K}\lambda_{x_m}}.9 is a permutation and K!K!0 may be non-differentiable or only available as a black box. The cited work extends REBAR and RELAX to permutations by using the Gumbel representation K!K!1, deriving a conditional distribution K!K!2 as a sequence of truncated Gumbels, and constructing PL-REBAR and PL-RELAX control variates. The support size is K!K!3, but Plackett–Luce uses K!K!4 parameters and can be sampled in K!K!5, making it usable in causal structure learning and other black-box permutation problems (Gadetsky et al., 2019).

4. Global and local asymptotic theory

A recent asymptotic theory studies permutations K!K!6 through their permuton and local limits. Set

K!K!7

and assume

K!K!8

for some positive, finite, measurable function K!K!9. Define

b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k0

The limiting permuton b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k1 is the law of

b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k2

where b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k3, b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k4, independent. Under these assumptions, b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k5 converges in probability in the permuton sense to the deterministic permuton b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k6 (Borga et al., 9 Sep 2025).

The limit is explicit. The permuton is absolutely continuous with density

b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k7

For Sukhatme weights b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k8, one has b=(b1,,bk)Skb=(b_1,\dots,b_k)\in S_k9, and the density is asymmetric and singular near θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)0. Pattern densities are also identified: for θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)1,

θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)2

where θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)3 are the order statistics of θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)4 i.i.d. uniform variables. Thus the limiting pattern law is an average of exact Luce laws with random weights. For Sukhatme weights, θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)5 and θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)6, showing a marked global bias (Borga et al., 9 Sep 2025).

The same work emphasizes that exact Luce permutations and permutations sampled from the limiting permuton can differ significantly on fine-scale statistics, even though they share the same global permuton limit. A key example is the distribution of the first or top values: a naive approximation replacing random weights θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)7 by deterministic weights is false in general. This suggests a separation between global convergence in the permuton sense and finer local or top-rank behavior (Borga et al., 9 Sep 2025).

Local convergence is formulated through consecutive patterns. For fixed θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)8 and θ=(θ1,,θk)\theta=(\theta_1,\dots,\theta_k)9, if

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K00

exists, then

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K01

in probability. For Sukhatme weights x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K02, the limiting local frequencies are uniform,

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K03

for every fixed x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K04. Hence the model is locally uniform even though it is globally non-uniform. The paper further proves a Berry–Esseen-type CLT for consecutive pattern occurrences and a CLT for inversions. For any positive weights,

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K05

with an explicit variance formula, and for Sukhatme weights

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K06

where x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K07 (Borga et al., 9 Sep 2025).

5. Applications: rank aggregation, causal structure learning, and secretary problems

Empirical studies of EPL emphasize cases where the ranking process is not naturally forward-only. In the Song data example, the posterior puts almost all mass on a nonstandard choice order,

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K08

with only tiny mass on the standard or reverse Plackett–Luce choice orders. The posterior predictive distribution under EPL assigns observed rankings much more plausible probability than standard PL, and the modal predictive ranking differs between EPL and SPL. In Formula 1 data, with x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K09 and x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K10, the posterior suggests a choice order closer to reverse PL in the later stages of the ranking process; EPL gives much more reasonable predictive counts for wins and podiums than standard PL, and its posterior predictive aggregate ranking is closer to the final championship standings in meaningful ways (Johnson et al., 2020).

For causal structure learning, Luce-distributed permutations are used as latent topological orders. A toy objective minimizes

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K11

and the cited experiments report that REINFORCE has too much variance to be useful, PL-REBAR improves substantially, and PL-RELAX performs best due to the learned control variate. In DAG learning from continuous data, optimization is over permutation matrices x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K12 through a black-box score

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K13

while for discrete Bayesian networks a non-differentiable quotient NML score is used. The reported implication is that the Plackett–Luce family is practically useful for black-box permutation optimization, including non-differentiable objectives (Gadetsky et al., 2019).

The secretary problem provides a different application. Here the arrival order of the x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K14 ranked items is not uniform, but follows either a Luce distribution or a Mallows distribution on x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K15. For inverse-Luce permutations, when the smallest label is best, the success probability of the threshold rule x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K16 is

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K17

For exponential weights x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K18, the cited work proves the exact identity

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K19

and further states that, for every x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K20 and every strategy, the secretary-problem probabilities when the smallest number is treated as highest rank under inverse-Luce weights x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K21 coincide with those for the corresponding Mallows distribution (Pinsky, 11 May 2026).

The same paper derives asymptotically optimal thresholds in several regimes. For x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K22 with x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K23,

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K24

For x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K25, x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K26,

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K27

again with limiting success probability x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K28. For fixed x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K29,

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K30

and the limiting success probability is

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K31

For Sukhatme weights x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K32,

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K33

while for reverse Sukhatme weights x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K34,

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K35

and in both cases the limiting success probability is x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K36 (Pinsky, 11 May 2026).

6. Scope of the term and distinction from unrelated permutation coordinatizations

The phrase “Luce-distributed permutations” refers in these sources to probabilistic models on x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K37: sequential weighted sampling without replacement, parametric ranking distributions such as Plackett–Luce and EPL, inverse-Luce laws, and Gibbs posteriors that include Bradley–Terry–Luce as a special case (Johnson et al., 2020, Cantwell et al., 2021, Borga et al., 9 Sep 2025, Pinsky, 11 May 2026, Gadetsky et al., 2019).

It should be distinguished from deterministic combinatorial uses of permutations that do not define probability distributions. A clear example is the theory of join-distributive lattices of length x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K38 and join-width at most x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K39. In that setting, there are two descriptions by x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K40 permutations acting on an x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K41-element set: the Edelman–Jamison construction inside a powerset lattice and a lattice-theoretic coordinatization by eligible tuples. The key result is

x=(x1,,xK)SKx=(x_1,\dots,x_K)\in \mathcal S_K42

and the paper also characterizes join-distributive lattices by trajectories. This is a theory of finite join-distributive lattices, semimodularity, meet-semidistributivity, and trajectory classes of prime intervals, rather than a Luce-type stochastic law on permutations (Czédli et al., 2012).

The distinction matters because both areas use permutation data, but for different purposes. In Luce-type models, permutations are random outcomes whose probabilities are determined by worth parameters, choice orders, or pairwise interaction factors. In the join-distributive lattice setting, permutations serve as coordinates for a deterministic representation theorem. The two subjects therefore share notation but not mathematical content (Czédli et al., 2012).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Luce-distributed permutations.