Papers
Topics
Authors
Recent
Search
2000 character limit reached

Joint Constellation Shaping

Updated 10 July 2026
  • Joint constellation shaping is the simultaneous optimization of constellation geometries and symbol probabilities to enhance mutual information in communication channels.
  • It employs autoencoder-based and iterative optimization methods to co-design transmitter, channel, and receiver settings in the presence of nonlinearities and other constraints.
  • This approach overcomes the limitations of traditional Maxwell–Boltzmann shaping by addressing both geometric and probabilistic aspects across optical, VLC, and multi-user systems.

Joint constellation shaping is the simultaneous optimization of constellation point locations and symbol occurrence probabilities. It generalizes the classical separation between geometric shaping, which changes the placement of points in signal space, and probabilistic shaping, which changes the probability mass function over a fixed alphabet. In the literature, it is typically posed as an achievable-information-rate optimization problem, often over arbitrary channel models and sometimes over receiver algorithms as well. Autoencoder-based formulations have made it possible to learn both geometry and probabilities end to end (Stark et al., 2019), while nonlinear-fiber studies show that separate or Maxwell–Boltzmann-type designs can be highly sub-optimal once nonlinearities and constellation-size constraints are imposed (Yankov, 2019).

1. Formal definition and information-theoretic basis

The canonical objective is inherited from channel coding theory: for an input distribution p(x)p(x), the design target is to maximize mutual information,

C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).

In this setting, joint shaping means optimizing both the support of p(x)p(x) and the probabilities assigned to that support, rather than fixing one and optimizing the other (Stark et al., 2019).

This distinction matters because classical probabilistic shaping usually assumes a fixed geometry, often QAM, and classical geometric shaping assumes uniform signaling. A standard baseline is the Maxwell–Boltzmann probability mass function

PX(X)exp(λX2),P_X(X) \propto \exp(-\lambda|X|^2),

which is near-optimal in linear Gaussian settings but is repeatedly reported as insufficient once the channel departs from AWGN assumptions, the receiver is mismatched, or hardware and complexity constraints are introduced (Yankov, 2019).

In coded-modulation systems, the objective is not always symbol-wise mutual information. End-to-end formulations have been developed that maximize either mutual information for symbol-metric decoding or generalized mutual information for bit-metric decoding, and related work on trainable coded modulation directly maximizes bit-wise mutual information while jointly adapting shaping, labeling, and demapping (Aref et al., 2021, Aoudia et al., 2020). This establishes joint shaping as a receiver-aware design problem rather than a purely geometric one.

2. Optimization frameworks

A dominant methodological line uses end-to-end autoencoders. In the general formulation, the transmitter learns a source distribution and a constellation mapping, the channel is inserted as a differentiable layer, and the receiver learns posterior estimates. One influential information-theoretic derivation shows that a modified cross-entropy objective can be interpreted as mutual-information maximization, enabling joint learning of the probability mass function and the constellation geometry over arbitrary channels. In that formulation, differentiable sampling is implemented with the Gumbel–Softmax trick, so that discrete probabilistic shaping becomes compatible with gradient descent (Stark et al., 2019).

Subsequent work diversified the transmitter and receiver parameterizations. Model-based optical systems used trainable constellations and trainable probabilities together with the NLIN model of Dar et al.; the transmitter samples symbols with optimized probabilities, enforces only an average-power constraint, and uses Adam to maximize mutual information (Neskorniuk et al., 2022). Other coded-modulation autoencoders replaced approximate sampling by exact sampling from the learned probability vector PMP_M, thereby avoiding Gumbel–Softmax approximation issues and supporting direct maximization of either MI or GMI (Aref et al., 2021). In realistic phase-noise settings, differentiable DSP has also been incorporated: a differentiable blind phase search replaces the non-differentiable argmin\arg\min by a softmin, allowing pilot-free joint geometric and probabilistic shaping to be trained through carrier-phase estimation (Rode et al., 2022).

Not all algorithms are gradient-based. For nonlinear fiber with constellation-size constraints, a greedy iterative algorithm alternates local optimizations over groups of points with the same multi-dimensional amplitude; each step uses a 2D grid search over probability and amplitude scaling and allows zero probability, so pruning is part of the optimization itself (Yankov, 2019). In multi-user VLC, the joint design of symbol probabilities and precoding is formulated as a non-convex sum-rate maximization and is tackled either by a firefly algorithm or by alternating optimization with zero-forcing precoding (Nguyen et al., 2024). In direct-detection optics, reinforcement learning via the score-function estimator is used to optimize discrete probabilities in the presence of a non-linear channel and a complexity-constrained CNN receiver (Fischer et al., 7 Jul 2026). In precoded ISI channels, arithmetic distribution matching is extended to Markovian symbol sequences so that the optimized non-i.i.d. transition law can be realized in practice (Ramin et al., 16 May 2025).

3. Multidimensional and nonlinear optical-fiber formulations

Optical coherent transmission is the most extensively developed application area. In wavelength-division multiplexed nonlinear fiber systems, the interaction between the input probability mass function and fiber nonlinearity makes separable optimization of PMF and geometry insufficient. A 4D joint probabilistic and geometric shaping algorithm, designed for dual-polarization signaling, was proposed specifically for this regime. It alternates probability and amplitude-scaling updates over shared 4D amplitudes, can prune points by assigning zero probability, and can enforce constellation-size constraints motivated by decoder complexity. In simulation for a 5×30 Gbaud, 50 GHz-spaced WDM system over 250 km, the method increased optimal launch power by 1 dB and yielded gains of approximately $0.25$ bits/4D for 64264^2QAM and $0.13$ bits/4D for 16216^2QAM over traditional schemes; at high powers it converged to single-amplitude constellations and at moderate powers it pruned a 4096-point format down to 784 points with only three unique amplitudes (Yankov, 2019).

Model-based deep learning refined this viewpoint. A joint geometric-and-probabilistic optical autoencoder using the NLIN model reported an extra C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).0 bits/4D-symbol mutual information over Maxwell–Boltzmann-shaped 256QAM for 64 GBd transmission over a 170 km SSMF link, and attributed the nonlinear-regime advantage to lower higher-order moments C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).1 and C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).2 (Neskorniuk et al., 2022). A separate end-to-end system trained directly over the split-step Fourier model of the fiber channel proposed a 4D 10 bits/symbol constellation and reported a 13.6% reach increase at a data rate of approximately 400 Gbits/second relative to polarization-multiplexed 32-QAM at 20% forward-error-correction overhead; in that high-cardinality regime, the learned symbol probabilities converged to uniform, implying that geometry dominated the gain (Oliari et al., 2021).

Multidimensional shaping beyond 4D has been analyzed comparatively rather than only constructively. Studies of 2D, 4D, 8D, 16D, and 32D formats over AWGN and fiber channels conclude that no constellation is universal and that required SNR and effective SNR must be jointly considered for a specific optical scenario (Chen et al., 2023). A later analytical 4D nonlinear-interference model explicitly incorporated the full 4D joint distribution, including non-i.i.d. signaling, to estimate modulation-dependent NLI power and relative shaping gains for arbitrary dual-polarization formats; the reported conclusion was that linear shaping gain and modulation-dependent NLI must be jointly considered for nonlinearity mitigation (Chen et al., 2024). Taken together, these results make multidimensional joint shaping less a single algorithm than a design principle: optimize the full 4D or higher-dimensional distribution against both linear and nonlinear impairments.

4. Coded modulation, demapping, and receiver-aware learning

Joint shaping has also become tightly coupled to coded-modulation architecture. One trainable coded-modulation scheme jointly optimized probabilistic shaping, geometric shaping, bit labeling, and demapping for a specific channel model over a wide SNR range, directly maximizing bit-wise mutual information. Unlike PAS, it was not restricted to symmetric distributions, could be optimized for any channel model, and worked with any code rate C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).3 for integer C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).4. On AWGN it was competitive with an MB-QAM baseline, while on mismatched Rayleigh block fading it achieved the highest BMI and spectral efficiency among the reported baselines (Aoudia et al., 2020).

Autoencoder-based GeoPCS for coded modulation sharpened the distinction between symbol-metric and bit-metric objectives. A fully differentiable transmitter samples exactly from the learned probability vector, maps to normalized trainable constellation points, and uses either a categorical demapper to maximize MI or a bit-wise demapper to maximize GMI. This formulation was presented as a general end-to-end method for coded-modulation systems rather than an optical-specific special case (Aref et al., 2021).

Receiver-aware optimization becomes more pronounced when the front end is non-ideal. For Wiener phase-noise channels, end-to-end optimization with a differentiable blind phase search yielded pilot-free 64-ary constellations that improved performance by at least C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).5 bit/symbol over square QAM constellations with neural demappers and by C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).6 bit/symbol over geometric-only shaping. The learned constellations deliberately broke QAM-like rotational symmetries that are problematic for blind carrier-phase estimation (Rode et al., 2022). A later shaping-encoder assisted architecture evaluated learned constellations with iterative receivers and deep unfolding of the iterative detection-and-decoding loop; relative to standard APSK or QAM, it reported BER gains up to C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).7 dB and C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).8 dB under two iterative receivers, and C=maxp(x)I(X;Y).C = \underset{p(x)}{\max}\, I(X;Y).9 dB under block fading with deep unfolding (Jayarathne et al., 26 Oct 2025). These results indicate that, in practice, joint shaping is increasingly optimized jointly with the demapper, the labeling, and even the iterative receiver schedule.

5. Coupling with waveform design, precoding, and multi-user transmission

The same joint-design logic extends beyond the constellation itself. In waveform learning, pulse shaping and constellation geometry were jointly optimized together with a neural receiver to maximize an achievable information rate under out-of-band-emission and power-envelope constraints. On an AWGN channel, the learned waveforms achieved up to orders of magnitude smaller ACLRs, with a reported example of approximately 30 dB lower ACLR than traditional filtering, and a PAPR of about 6.45 dB versus 7.53 dB for an RRC baseline under a strict power-envelope constraint, without significant information-rate loss (Aoudia et al., 2021). This suggests that “joint shaping” in a broader sense includes waveform degrees of freedom when spectral masks and amplifier efficiency are part of the design space.

In multi-user VLC broadcast channels, probabilistic constellation shaping and transmit precoding were jointly optimized to maximize sum rate under LED amplitude constraints. The exact joint problem is non-convex because each user rate depends on both the probability matrix and the precoder. A firefly algorithm was proposed for general precoding, and a lower-complexity alternating-optimization method was proposed for zero-forcing precoding. At p(x)p(x)0 dB, the joint design with PCS improved the sum rate by 19.7% for 8-PAM and 23.4% for 16-PAM over uniform signaling; the AO method was reported to converge in as few as five iterations, whereas the FA method achieved higher maximum sum rates at greater computational cost (Nguyen et al., 2024).

Precoded channels create a different coupling: shaping must account for temporal filtering. A joint precoding-and-PCS formulation introduced statistical dependencies among consecutive symbols through a stationary Markov model and optimized transition probabilities to minimize average transmit power after the precoder subject to an entropy-rate constraint. A new arithmetic distribution matching method was then proposed to realize those Markovian statistics. In strong-ISI cases, the reported shaping gains approached the ultimate shaping gain of 1.53 dB, whereas memoryless PCS with linear precoding or THP did not realize those gains (Ramin et al., 16 May 2025).

A related but more geometric multi-user formulation appears in MU-MIMO broadcast channels with perfect CSIT. There, a learning-based MAX-MIN method directly optimized a joint constellation p(x)p(x)1 so that the projections into each receiver subspace maximize the minimum user mutual information, without imposing superposition coding, linear precoding, or SIC. The reported rates exceeded those of matched-filter, ZF, and MMSE linear precoders, and in a p(x)p(x)2 scenario the minimum mutual information was up to 42% higher than the best linear-precoding design at high SNR (Vaillant et al., 2024).

6. Regime dependence, implementation constraints, and recurring conclusions

Integrated sensing and communications makes the regime dependence of joint shaping explicit. In OFDM-ISAC, communication prefers Gaussian-like signaling, whereas sensing prefers low-kurtosis, nearly constant-modulus signaling. Autoencoder-based studies compared geometric, probabilistic, and joint shaping and found that constellation shaping enables a dynamic trade-off between communications and sensing; depending on whether sensing or communications is prioritized, geometric or probabilistic shaping is preferred, while joint shaping combines the advantages of both and significantly outperforms legacy modulation formats (Geiger et al., 20 Jan 2025). A more detailed OFDM-ISAC treatment imposed a kurtosis constraint, derived lower and upper bounds on the achievable information rate, and reported that joint shaping performs best especially at intermediate sensing constraints, while generalized PAS can approach joint-shaping performance with low-complexity LLR computation (Geiger et al., 4 Sep 2025).

Complexity constraints alter the ranking of methods. In a simulated 56 GBd, 2.2 km C-band direct-detection optical system with CNN-based joint equalization and demodulation, joint probabilistic and geometric shaping improved normalized GMI by 10% for 8-PAM and about 5% for 4-PAM starting from approximately 100 real-valued multiplications per bit. The same study also reported that, for 4-PAM, geometric shaping alone is nearly as effective as joint shaping, whereas for 8-PAM joint shaping is necessary for best performance (Fischer et al., 7 Jul 2026). This mirrors the high-cardinality 4D optical result in which learned probabilities became uniform and geometry dominated (Oliari et al., 2021), and it contrasts with ISAC regimes where probability shaping can be essential at loose sensing constraints (Geiger et al., 20 Jan 2025).

Several recurrent conclusions follow. First, there is no universal jointly shaped constellation: the balance between required SNR and effective SNR depends on the channel and operating point (Chen et al., 2023). Second, Maxwell–Boltzmann shaping is a strong AWGN baseline but is not generally optimal in nonlinear or constrained settings (Yankov, 2019). Third, receiver architecture, complexity budget, and side constraints such as phase recovery, precoder memory, ACLR, PAPR, sensing kurtosis, and decoder complexity are now part of the shaping problem rather than external implementation details (Rode et al., 2022, Aoudia et al., 2021). A plausible implication is that joint constellation shaping is best understood not as a single family of constellations, but as a cross-layer optimization framework in which the input alphabet, the input distribution, and often the receiver itself are co-designed for a specific channel and operational objective.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (17)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Joint Constellation Shaping.