Shifted Gaussian Encoding
- Shifted Gaussian encoding is a design motif that deliberately translates Gaussian centers to represent structured data in domains such as positional embedding, quantum error correction, and integer encoding.
- In coordinate MLPs and GKP state synthesis, tuning parameters like shift, width, and damping controls memorization capacity, interpolation smoothness, and error fidelity.
- Applications such as smooth integer recovery and explicitly correlated Gaussians require tailored post-processing like integral cancellation and symmetry projection for optimal results.
Searching arXiv for the specified paper and closely related uses of shifted Gaussian encoding. Search results will help anchor citations and confirm related terminology across domains. Searching "Rethinking Positional Encoding" (Zheng et al., 2021), shifted Gaussian encoding, GKP Gaussian breeding, surface-GKP designed bias. Shifted Gaussian encoding denotes a family of constructions in which information is represented by Gaussian functions whose centers are deliberately translated across a domain. In the cited literature, this motif appears in several technically distinct forms: as a positional embedding for coordinate-based MLPs, with features (Zheng et al., 2021); as approximate Gottesman–Kitaev–Preskill (GKP) codewords built from displaced squeezed Gaussians in propagating light (Takase et al., 2022); as a squeezing-deformed GKP lattice used to bias logical noise in a concatenated surface–GKP code (Hänggli et al., 2020); as a smooth integer representation based on localized Gaussian bumps with alternating coefficients (Semenov, 28 Apr 2025); and as shifted-center explicitly correlated Gaussians in pre-Born–Oppenheimer variational calculations (Muolo et al., 2018). Across these settings, the operative degrees of freedom are the Gaussian centers, widths, and, where relevant, amplitudes and projection operations.
1. Mathematical forms and shared structure
A basic shifted Gaussian positional embedding associates to a coordinate an -dimensional feature vector
with equally spaced shifts and . In , a separable construction samples points per axis with standard deviation and concatenates the per-axis vectors, yielding total embedding dimension (Zheng et al., 2021).
In continuous-variable quantum information, approximate GKP codewords are superpositions of equally spaced, squeezed Gaussians in the position basis. With squeezing parameter 0 and envelope width 1, the unnormalized codewords are
2
where 3 and 4 (Takase et al., 2022).
A smooth integer encoding uses a finite sum of shifted Gaussian bumps,
5
with 6, 7, and coefficients
8
Here the integer is not an explicit parameter of a symbolic representation; it is recovered from the integral balance of the smooth function (Semenov, 28 Apr 2025).
In few-body quantum chemistry, a floating explicitly correlated Gaussian (FECG) is
9
with real symmetric positive-definite block-diagonal 0 and shift vector 1. The shifted center 2 supplies the basis with localization flexibility unavailable to origin-centered ECGs (Muolo et al., 2018).
These examples share a common translation mechanism but differ in what is being encoded: Euclidean coordinates, logical qubits, effective error bias, integers, or many-particle wavefunctions. This suggests that “shifted Gaussian encoding” is best understood as a structural motif rather than a single canonical formalism.
2. Shifted Gaussian positional encoding in coordinate MLPs
For coordinate-based MLPs, the shifted Gaussian embedder is analyzed through two metrics defined on the embedding matrix 3 formed from sampled coordinates 4. The stable rank
5
measures the effective number of non-negligible singular values, and higher stable rank gives more capacity to memorize arbitrary 6's. The embedded-distance quantity
7
is desired to be a monotonic function of the original distance 8 so that nearby coordinates remain similar and far points remain dissimilar in feature space (Zheng et al., 2021).
For the Gaussian embedder, the continuous-limit inner product satisfies
9
After zero-mean centering and normalizing, one may ignore the 0 prefactor and view
1
When 2 and 3 are large enough,
4
The bandwidth parameter 5 controls the central trade-off. Decreasing 6 raises the upper bound 7 on stable rank and therefore increases memorization capacity for high-frequency content, but it also makes 8 decay rapidly, so nearby points become nearly orthogonal. Increasing 9 reduces stable rank and can underfit high-frequency structure, while making interpolation overly smooth. The embedding dimension must satisfy 0 to reach the stable-rank ceiling, and Nyquist sampling requires 1 to preserve the inner-product integral. A practical one-dimensional rule chooses 2 so that
3
stays above a small threshold 4 for nearest-neighbor spacing 5, giving
6
In 7 dimensions, the same per-axis rule is applied under separable sampling, while keeping total dimension linear in 8 through concatenation (Zheng et al., 2021).
Empirically, the paper studies 1D and 2D image-signal reconstruction with coordinate-MLPs. Baselines include no encoding, fixed sinusoid (“basic”), Random Fourier Features (RFF), impulse, square wave, and random noise. In 1D experiments with a linear one-layer network, varying 9 over 10 seeds, the Gaussian embedder with 0 set by 1 yields test PSNR 2 with small error bars, whereas RFF is volatile at low 3 and stabilizes only at large 4. In 2D experiments with a 4-layer ReLU MLP, separable Gaussian sampling along 5 and optionally additional rotated axes matches or exceeds RFF in test PSNR, for example 6 at 7, while requiring smaller 8 for similar quality. Rank-versus-distance plots place Gaussian and RFF in a sweet spot of intermediate stable rank and reasonable distance preservation, whereas impulse and random noise over-rank and sine and square-wave embeddings under-rank. Training dynamics further show faster convergence, lower hidden-layer stable-rank growth, and higher final accuracy than unencoded or basic encodings (Zheng et al., 2021).
3. Shifted Gaussian codewords and Gaussian breeding for GKP qubits
In propagating-light implementations of the GKP code, logical basis states are approximate codewords formed from equally spaced displaced squeezed vacua with a Gaussian envelope. In the square-lattice case one often takes 9, and in the limit 0 the logical states approach ideal Dirac-comb codewords (Takase et al., 2022).
The paper’s central operation is the coherent bifurcation 1, which acts linearly on superpositions of displaced squeezed vacua according to
2
Iterating it 3 times with step size 4 yields
5
This constructs a comb of 6 peaks (Takase et al., 2022).
A physical implementation starts from two single-mode squeezers and combines them either on a beam splitter of transmittance 7 or through a QND gate 8, followed by photon-number-resolving detection on the second mode. Detecting 9 photons heralds an approximate two-peak superposition
0
up to small overlap errors, and for even 1 the wavefunction is symmetric. The QND variant 2 commutes with displacements, so linearity holds exactly and the coherent bifurcation becomes fully iterable (Takase et al., 2022).
Envelope control is implemented by heralding 3 in the QND setup, which realizes an unsharp measurement in 4 equivalent to multiplication by a Gaussian in momentum space, or operatorially
5
After a 6 phase rotation one also obtains 7. These damping steps set the global envelope 8 without changing the internal peak spacing. To prepare an arbitrary superposition 9, one first constructs a seed
0
then applies 1 to obtain
2
Via Bloch–Messiah reduction, all on-line squeezers and QND gates can be replaced by a fixed interferometer, off-line squeezing, and photon-number detection; in practice, using the same 3 at each stage reduces the beam-splitter count to 4 (Takase et al., 2022).
The performance formulas quantify scalability. The per-round success probability is
5
maximized at 6, giving
7
After 8 bifurcations with 9, the comb envelope has variance 0, and in the square-lattice case additional damping yields total squeezing parameter 1. Each two-peak building block has overlap
2
so the overall infidelity remains small provided 3. Threshold analyses cited in the summary typically require 4 of squeezing, that is 5, together with logical-state fidelity 6. With 7, one achieves 8 and state fidelities 9 after 00 rounds. For generation of 01 to 02, the example resource count is 03 photons per bifurcation, 04–05 rounds, per-round success probability 06, and total success probability 07, potentially improvable by multiplexing or quantum memory. For arbitrary magic states, the seed succeeds with probability 08 and the full process reaches 09 with 10 (Takase et al., 2022).
4. Noise-biased surface–GKP encoding through squeezing deformation
A different use of shifted Gaussian structure appears in the concatenated surface–GKP code, where each bosonic GKP mode is first transformed by a single-mode squeezing unitary
11
associated with the symplectic matrix
12
Under conjugation, the quadratures become 13 and 14, and displacement operators transform as
15
The square-lattice GKP stabilizer lattice
16
is deformed into the rectangular lattice
17
with dual lattice
18
In the rescaled coordinates, the logical Pauli operators are
19
This deformation converts isotropic Gaussian displacement noise into anisotropic noise. An original channel
20
becomes an anisotropic channel with covariance
21
Equivalently, the quadrature jitters satisfy
22
and the bias parameter is
23
For 24, the 25-quadrature noise, which causes logical 26 errors, is much larger than the 27-quadrature noise, which causes logical 28 errors (Hänggli et al., 2020).
Using nearest-lattice-point decoding at the GKP level, the residual logical-qubit error probabilities satisfy approximately
29
so 30 for large 31. The encoding is then relabeled so that the small-noise quadrature is associated with the qubit’s 32 basis; in the paper the GKP “0/1” states are mapped to 33. The resulting single-qubit channel has
34
with 35, and in the large-36 limit one finds 37. The purpose is to realize a Pauli-38-biased qubit channel that the outer surface code can exploit (Hänggli et al., 2020).
For decoding, the study uses the Bravyi–Suchara–Vargo tensor-network decoder with bond dimension 39 up to 40, which approaches maximum-likelihood decoding as 41. Monte Carlo simulations show that, even without using GKP side information, an optimal choice of 42 raises the threshold from 43 at 44 to 45 for 46. Incorporating GKP analog-syndrome information into the prior further boosts the threshold to 47; even for 48 one observes 49, and a similar effect persists on asymmetric hexagonal lattices with 50 (Hänggli et al., 2020). The paper frames this as a two-step map: single-mode squeezing reshapes isotropic displacements into an anisotropic Gaussian channel, and reinterpretation of the logical axes turns the small-variance direction into predominantly 51 errors.
5. Smooth integer encoding by shifted Gaussian integral balance
The integer-encoding construction of (Semenov, 28 Apr 2025) represents 52 by a smooth bump sum
53
with 54 bumps and coefficients
55
Because 56, the total integral tends to zero in the large-57 limit (Semenov, 28 Apr 2025).
Exact Gaussian integration gives
58
hence
59
Since 60 for some 61, the tail obeys 62, so 63 exponentially fast. Because the coefficients alternate in sign, 64 and 65 oscillate about zero with exponentially decaying amplitude. Recovery is defined by the first near-cancellation: 66 or operationally
67
The summary states that no two integers share the same small-magnitude value once 68 lies below the preceding oscillation amplitude (Semenov, 28 Apr 2025).
Several inversion procedures are given. Threshold-based inversion from a measured integral 69 uses
70
and is stable if measurement noise satisfies 71. A tabulation-and-binary-search scheme precomputes 72 and locates the smallest compatible 73 in 74 time. A spline interpolation 75 may be inverted numerically via
76
and Newton’s method can be applied to 77. The local error relation is
78
A piecewise analytical inversion is also supplied: 79 so that if 80 then
81
The construction extends to tuples 82 by summing 83-dimensional Gaussian bumps centered at 84 with coefficients such as
85
producing
86
The paper further notes that the map 87 is differentiable, or 88 under a smooth blending replacement for the floor-based extension, enabling uses such as a differentiable layer “SmoothInteger,” a regularizer for integer constraints, and soft-argmax with guaranteed exact recovery (Semenov, 28 Apr 2025).
6. Shifted-center explicitly correlated Gaussians in pre-Born–Oppenheimer calculations
In pre-Born–Oppenheimer quantum calculations, shifted Gaussian encoding appears as floating explicitly correlated Gaussian basis functions. For an 89-particle system in three dimensions, a single basis element is
90
With 91, 92, and 93, one may equivalently write
94
In practice, 95 with 96, while the shifted center 97 is a 98-vector of Gaussian centers in the laboratory frame. The role of the shift is to describe localized structures, including nuclei positions, more flexibly than origin-centered ECGs (Muolo et al., 2018).
The overlap of two FECGs is analytic: 99 where 00. For 01,
02
so the normalized basis function is
03
Because a general FECG is not an eigenfunction of total angular momentum or parity, projection is required. Rotation-inversion projection uses
04
with 05. Parity projection is
06
The projected function
07
then satisfies the requisite 08, 09, and parity eigenvalue equations (Muolo et al., 2018).
The triple Euler-angle integral is evaluated numerically by a product Gauss–Legendre scheme. With 10 nodes per angle,
11
Typical values 12–13 suffice to converge 14 to a few 15, and the naive cost scaling 16 can be reduced to 17 by exploiting idempotency and Hermiticity of the projector. For on-the-fly optimization, a nested Gauss–Kronrod rule may be used (Muolo et al., 2018).
Parameter optimization proceeds through competitive selection and Powell’s derivative-free refinement, often in a two-stage procedure: optimize non-projected FECGs by 18 minimization, then solve the linear variational problem with projected functions. Fully projected optimization is possible but more expensive. The reported benefit of shifted centers is that Gaussians can be localized at arbitrary interparticle distances, so fewer functions are needed to capture nuclear motion; the shift vector also explicitly encodes global translation, rotation, and internal equilibrium geometry (Muolo et al., 2018).
The principal application reported is the five-particle 19 ion with target state 20, 21, parity 22. Basis sets range from a small 23 test through 24 to 25. For the largest projected basis,
26
and extrapolation with 27 gives
28
The previous best non-shifted ECG result cited is 29, differing by approximately 30, while a perturbative non-adiabatic model estimate is 31 (Muolo et al., 2018).
7. Comparative interpretation and recurrent trade-offs
The cited literature does not present a single universal theory covering all uses of shifted Gaussian encoding. Instead, each field uses shifted centers to control a different structural property. In positional encoding, the decisive variables are stable rank and embedded-distance preservation, both governed by 32 and sampling density (Zheng et al., 2021). In GKP-state synthesis, the central issues are iterable superposition growth, envelope control, heralding probability, and fidelity under repeated coherent bifurcation (Takase et al., 2022). In the surface–GKP setting, squeezing-induced anisotropy is used to transform isotropic Gaussian displacement noise into a biased qubit channel that is better matched to the surface code (Hänggli et al., 2020). In smooth integer encoding, the essential mechanism is oscillatory near-cancellation of the total integral with exponentially decaying tails (Semenov, 28 Apr 2025). In shifted-center ECGs, the gain is variational flexibility at the cost of numerical projection onto symmetry sectors (Muolo et al., 2018).
A common misconception would be to treat these constructions as interchangeable merely because they use translated Gaussians. The source material instead indicates domain-specific semantics for the same geometric operation. In one case the centers 33 sample a coordinate domain; in another they mark lattice displacements of squeezed vacua; in another they are laboratory-frame centers of many-particle basis functions. This suggests that the unifying concept is not a shared application, but a shared representational device: localization through Gaussian basis elements whose positions are shifted to encode structure, constraints, or discrete alternatives.
A second recurring theme is that translation alone is insufficient; usefulness depends on accompanying control variables. Positional encoding requires an appropriate balance between rank and distance preservation. GKP breeding requires damping to shape the global envelope without changing the peak spacing. Surface–GKP biasing requires reinterpretation of logical axes after squeezing. Integer encoding requires carefully designed alternating coefficients and a recovery rule based on local minima of 34. Projected FECGs require explicit symmetry projection to restore good quantum numbers. The literature therefore presents shifted Gaussian encoding not as a generic recipe, but as a design pattern whose efficacy depends on how shifts, widths, amplitudes, and post-processing are coupled to the target problem.