---
title: 'ShaLa: Multi-Domain Methods in ML and Astronomy'
url: https://www.emergentmind.com/topics/shala
type: topic
---

# ShaLa: Multi-Domain Methods in ML and Astronomy

ShaLa is a polysemous label used for several unrelated research objects across machine learning and astronomy. In the literature considered here, it denotes a reinforcement-learning framework for aligning large language models to annotator label distributions, a generative framework for multimodal shared latent space modelling, an occasional misspelling of the Spitzer/HETDEX Exploratory Large-Area survey SHELA, and, in observatory instrumentation usage, the sodium laser guide star adaptive-optics system associated with ShaneAO at the Lick Observatory Shane 3-m telescope [2606.05376] [2508.17376] [1812.04646] [1407.8207].

## 1. Nomenclature and scope

The term has no single canonical meaning across fields. In recent machine-learning work, “SHALA-LLM” expands to “Smartly Handling Ambiguous Labels in Aligning LLMs,” while “ShaLa” in multimodal generative modelling abbreviates “Shared Latent Space Modelling.” In observational astronomy, “ShaLa” appears as an occasional misspelling of SHELA, the Spitzer/HETDEX Exploratory Large-Area survey; in adaptive optics, “ShaLa/ShaneAO” designates the sodium laser guide star system used with ShARCS on the Shane telescope [2606.05376] [2508.17376] [1812.04646] [1407.8207].

| Usage | Domain | Referent |
|---|---|---|
| SHALA-LLM | LLM alignment | RL framework for learning annotator label distributions |
| ShaLa | Multimodal generative modelling | Shared latent-space VAE plus diffusion prior |
| ShaLa / SHELA | Extragalactic survey astronomy | Occasional misspelling of the SHELA survey |
| ShaLa / ShaneAO | Adaptive optics instrumentation | Sodium laser guide star AO facility at Lick |

This disambiguation matters because the four usages are methodologically independent. Two are modern ML frameworks, one is a survey label in wide-field astronomy, and one is an AO instrumentation name. The cited works do not indicate a shared lineage beyond the reused string “ShaLa.”

## 2. ShaLa as ambiguity-aware LLM alignment

SHALA-LLM is a reinforcement learning–based alignment framework for subjective or ambiguity-sensitive tasks such as natural language inference and emotion recognition. Its core premise is that annotator disagreement should be treated as information rather than noise: instead of collapsing annotations to a single majority label, the model is trained to predict the full empirical annotator distribution and to emphasize highly ambiguous samples during optimization [2606.05376].

The formal target is the empirical annotator distribution for sample $q$ with $N$ annotators and $C$ classes,
$$
p_{q,c} = \frac{n_{q,c}}{N}, \qquad \sum_{c=1}^C p_{q,c} = 1,
$$
with the model producing a parsed probability vector $\hat{\mathbf p}_{(q,i)}$ from a JSON-like textual output. Ambiguity is quantified by normalized entropy,
$$
\tilde H(\mathbf p_q) = \frac{-\sum_{c=1}^C p_{q,c}\log p_{q,c}}{\log C} \in [0,1],
$$
and distributional agreement is measured with the Jensen–Shannon Distance. The rollout reward is
$$
r_{(q,i)}^{\mathrm{SHALA}} = \tilde H(\mathbf p_q)\left[1 - D_{\mathrm{JS}}\big(\hat{\mathbf p}_{(q,i)}, \mathbf p_q\big)\right].
$$
This reward design couples soft-label alignment to ambiguity-aware prioritization: high-entropy items receive larger effective rewards when the predicted distribution matches human disagreement well.

Optimization uses Group Relative Policy Optimization. The paper defines group-normalized advantages
$$
\hat A_{(q,i)} = \frac{r_{(q,i)} - \hat \mu_{G_{(q)}}}{\hat \sigma_{G_{(q)}} + \varepsilon},
$$
and a PPO-style clipped surrogate with importance ratio $\varphi_{(q,i):k}(\theta)$. In the reported experiments, the base model is Qwen2.5-Omni-7B; GRPO is implemented via TRL with AdamW, learning rate $1\times10^{-6}$, rollouts per sample $K=4$, temperature $1.2$, maximum completion length $128$ tokens, and KL penalty $\beta=0$. Training is run on a single node with $2\times$NVIDIA H200 GPUs under DeepSpeed ZeRO-3 and effective batch size $4$.

The empirical results emphasize both distributional fidelity and conventional classification metrics. On ChaosNLI overall, JSD decreases from $0.477$ under majority-label supervision to $0.181$ under ShaLa, a $62.1\%$ reduction; BC increases from $0.751$ to $0.966$; ACC rises from $0.699$ to $0.768$; F1 rises from $0.650$ to $0.758$; and W-F1 rises from $0.684$ to $0.767$. On MSP-Podcast, JSD improves from $0.580$ to $0.544$, BC from $0.585$ to $0.694$, ACC from $0.488$ to $0.496$, F1 from $0.233$ to $0.301$, and W-F1 from $0.415$ to $0.455$. On GoEmotions, JSD improves from $0.542$ to $0.465$ and BC from $0.638$ to $0.756$; the reported ShaLa label metrics are ACC $0.60$, F1 $0.59$, and W-F1 $0.60$.

A notable ablation removes ambiguity scaling by setting $\tilde H(\mathbf p_q)=1$ (“w/o Ambi-En”). That variant still benefits from JSD-based distributional alignment, but the full entropy-weighted method further improves JSD, BC, ACC, and F1 by focusing optimization on high-disagreement samples. The paper also reports smaller degradation across ambiguity strata than zero-shot or majority-label baselines; for NLI, ShaLa shows no statistically significant degradation across ambiguity levels ($p>0.05$), whereas majority-label supervision exhibits marked declines, including a BC drop from $0.970$ to $0.693$ across levels.

The stated limitations are equally specific. Evaluations are confined to categorical tasks with structured label spaces; empirical annotator distributions can preserve societal and demographic biases; GRPO-based RL is compute-intensive; robustness beyond NLI and emotion recognition remains to be established; and the method does not disentangle the reasons underlying disagreement, such as expertise or contextual variation.

## 3. ShaLa as multimodal shared latent space modelling

In multimodal generative modelling, ShaLa denotes a two-stage framework for learning a shared latent representation across modalities while improving synthesis quality through a latent diffusion prior. The model is explicitly motivated by limitations of prior multimodal VAEs: rigid Product-of-Experts or Mixture-of-Experts aggregation, poor expressiveness of joint posteriors, the prior-hole mismatch between the aggregated posterior and a fixed Gaussian prior, and poor scaling as the number of modalities grows [2508.17376].

The generative model is
$$
p_\theta(X,z) = p_\theta(X\mid z)\,p_0(z), \qquad
p_\theta(X\mid z) = \prod_{m\in M} p_{\theta_m}(x_m\mid z),
$$
with shared latent $z\in\mathbb R^d$ and typically $p_0(z)=\mathcal N(0,I_d)$. Each modality has a deterministic encoder $f_m:x_m\mapsto e_m$ and a decoder $g_m:z\mapsto \hat x_m$. Rather than combining stochastic unimodal posteriors through PoE or MoE, ShaLa uses deterministic architectural fusion:
$$
e_m=f_m(x_m), \qquad h=\odot(e_1,\ldots,e_M),
$$
where $\odot$ is implemented as concatenation followed by a small MLP. The joint variational posterior is then
$$
q_\phi(z\mid X)=\mathcal N\big(z;\mu_\phi(h),\Sigma_\phi(h)\big).
$$
This places the semantic alignment burden on the fused deterministic bottleneck $h$.

Stage 1 optimizes the multimodal ELBO
$$
\mathcal L(X)=\sum_{m\in M}\lambda_m\,\mathbb E_{q_\phi(z\mid X)}[\log p_{\theta_m}(x_m\mid z)]
- \mathrm{KL}\big(q_\phi(z\mid X)\,\|\,p_0(z)\big).
$$
Stage 2 replaces the Gaussian prior with an expressive latent DDPM fit to samples from the aggregated posterior $q_\phi(z)$. Using the standard forward process
$$
q(z_t\mid z_{t-1})=\mathcal N(\sqrt{\alpha_t}z_{t-1},\beta_t I), \qquad
z_t=\sqrt{\bar\alpha_t}z_0+\sqrt{1-\bar\alpha_t}\,\varepsilon,
$$
the diffusion prior is trained with
$$
\mathcal L_{\mathrm{DDPM}}
=
\mathbb E\!\left[
\left\|
\varepsilon-\varepsilon_\psi\!\left(\sqrt{\bar\alpha_t}z_0+\sqrt{1-\bar\alpha_t}\varepsilon,\ t,\ c\right)
\right\|^2
\right],
$$
where the conditioning variable $c$ can be a fused representation $h$, a single-modality embedding $e_j$, a subset embedding $e_S$, or be dropped for unconditional learning. This second stage addresses the prior-hole problem and supports cross-modal inference from incomplete modality sets.

The reported benchmarks are PolyMNIST, MNIST-SVHN-Text, CUB, and ShapeNet Cars. On PolyMNIST, unconditional coherence is $0.815$ for ShaLa versus $0.781$ for CMVAE, $0.735$ for MVEBM, $0.344$ for MMVAE+, $0.232$ for MMVAE, and $0.141$ for MoPoE. Conditional PolyMNIST coherence is $0.897$, tying CMVAE. On MST, unconditional coherence is $0.44$ versus $0.31$ for MoPoE, and conditional coherence is $0.75$ versus $0.72$ for mmJSD and $0.69$ for MoPoE. For image quality, PolyMNIST unconditional FID is $47.30$, better than CMVAE at $78.52$ and MVAE at $50.65$; PolyMNIST conditional FID is $40.18$, better than CMVAE at $74.53$; and CUB conditional FID is $25.58$, versus $28.00$ for CMVAE, $136.16$ for MVEBM, and $164.94$ for MMVAE+.

The scalability claim is most explicit on ShapeNet Cars, where 16 views are treated as modalities. ShaLa reports PSNR $24.7$ and SSIM $0.89$, compared with $19.3/0.59$ for MMVAE+ and $20.5/0.64$ for CMVAE. Against task-specific baselines, it is reported as competitive with PixelNeRF ($23.2/0.90$), EG3D ($21.8/0.71$), RenderDiffusion ($25.4/0.81$), and SyncDreamer ($21.9/0.88$). The ablations are equally diagnostic: conditioning the diffusion prior on fused $h$ yields PolyMNIST unconditional/conditional coherence of $0.815/0.897$, versus $0.584/0.612$ without $h$; a larger diffusion model improves PolyMNIST unconditional FID from $56.44$ to $45.24$; and concatenation plus MLP outperforms summation and gated fusion on conditional coherence.

The paper’s practical position is that ShaLa differs from MVAE, MMVAE, MoPoE, mmJSD, MVTCAE, MMVAE+, MVEBM, and CMVAE by jointly modifying posterior construction and prior modelling. Its contribution is therefore not merely a stronger decoder or a larger latent, but a specific decomposition: deterministic multimodal fusion for the approximate posterior, followed by a conditional latent diffusion model that matches the aggregated posterior and can condition on available modalities.

## 4. “ShaLa” as an occasional misspelling of SHELA

In survey astronomy, “ShaLa” is not a distinct program in the cited literature but an occasional misspelling of SHELA, the Spitzer/HETDEX Exploratory Large-Area survey. SHELA is a deep, wide-field multiwavelength program in SDSS Stripe 82 designed to study galaxy evolution over $0.35$–$4.5\,\mu$m. The field covers approximately $24$ deg$^2$, lies within the HETDEX footprint, and combines DECam $ugriz$ imaging with Spitzer/IRAC $3.6$ and $4.5\,\mu$m data; the IRAC component was first released as a post-cryogenic Spitzer survey, and a later catalog paper presented the DECam plus forced-photometry IRAC catalogs [1812.04646] [1603.05660] [2606.25228].

The 2018 catalog paper presents $riz$-selected DECam $ugriz$ catalogs over $17.5$ deg$^2$ of the overall field, using a single inverse-variance-weighted $riz$ detection image for SExtractor double-image photometry. The exclusion of $u$ and $g$ from the detection image is motivated by not penalizing high-$z$ dropouts in source finding. Images within each tile are PSF-matched to the worst-seeing band so that a single fixed aperture encloses the same fraction of a point source’s flux across $ugriz$. The catalogs reach $5\sigma$ depths of approximately $24.5$ AB mag for point sources in apertures enclosing about $70\%$ of the total flux, with $80\%$ completeness at $m\approx24.4$ and $50\%$ completeness at $m\approx24.7$ for $riz$-selected point sources. Fluxes are placed on a uniform AB system in nJy with zero-point $31.4$, corresponding to
$$
m_{\mathrm{AB}} = -2.5\log_{10}(f[\mathrm{nJy}]) + 31.4.
$$

Astrometric recalibration to SDSS applies typical offsets of about $100$ mas with $1\sigma$ scatter about $150$ mas. Photometric zero points are derived per tile and band through both an F0-star method and linear color relations to SDSS colors; for $ugriz$, the difference $|ZPT_{F0}-ZPT_{Lin}|$ is reported as $\lesssim0.04$ mag, while comparisons to DECaLS DR5 and DES DR1 show median zero-point offsets $\lesssim0.05$ mag. A $5\%$ systematic flux error is added in quadrature to account for zeropoint uncertainties and differences between sky-aperture and simulation-based error estimates.

A central technical component is forced IRAC photometry with The Tractor. The motivation is the coarse IRAC PSF of approximately $2''$ and the resulting heavy blending, with at least $35\%$ of DECam positions having a neighbor within $4''$. The Tractor uses DECam positions and morphology as priors, fits $20''\times20''$ IRAC cutouts with point-source, exponential, or de Vaucouleurs profiles convolved with the empirical PRF, and uses a two-pass optimization in which distant neighbors are first modeled or masked and then the target plus close neighbors are fitted simultaneously. The output includes IRAC fluxes, errors, model code ($0$ for PSF, $1$ for exponential, $4$ for de Vaucouleurs), reduced $\chi^2$, log-likelihood, and flags. About $47\%$ of sources are best fit with resolved profiles; among the resolved subset, about $61\%$ prefer exponential and about $39\%$ de Vaucouleurs. For isolated sources, comparison with the original IRAC catalog gives a median offset of about $-0.06$ mag for $m<22$, whereas blended sources show larger offsets of about $+0.45$ mag, interpreted in the paper as improved deblending.

The survey’s science utility is illustrated through number counts and photometric redshifts. Photometric redshifts are computed with EAZY using $ugrizJK+$IRAC fluxes, requiring at least five valid fluxes and applying a $K$-band prior where available. Against SDSS spectroscopy, the spectroscopic sample has mean $\langle z\rangle \approx 0.33$, median $\Delta z = z_{\mathrm{phot}}-z_{\mathrm{spec}}\approx -0.005$, $\sigma_{\mathrm{NMAD}}\approx0.042$, and a $5\sigma$ outlier fraction of about $2.3\%$. The paper also states that using all seven bands yields $\sigma_{\mathrm{TOTAL}}\approx0.05$ for $z<1$, with the $u$ band improving $z<0.4$ and IRAC, plus to a lesser extent $i$, improving $z>0.4$.

The earlier IRAC release characterizes the warm Spitzer component alone. It covers roughly $24$ deg$^2$ with three epochs separated by approximately $4$–$7$ months, reaches $1\sigma$ limiting sensitivities of $1.1\,\mu$Jy at both $3.6$ and $4.5\,\mu$m for $R=2''$ circular apertures, and is $80\%$ and $50\%$ complete at weighted-sum detection-image magnitudes $22.0$ and $22.6$ AB, respectively. The synergy target is HETDEX, which within the overlap is expected to provide about $200{,}000$ Ly$\alpha$ emitters at $1.9<z<3.5$ and an additional about $200{,}000$ [O II] emitters at $z<0.5$.

A further astronomy use of the SHELA field appears in the ODIN narrowband LAE survey. There, SHELA is one of seven ODIN regions and is observed with DECam narrowbands N419 and N501 and broadbands $g$ and $r$ over an effective area of about $15$ deg$^2$. The analysis focuses on the efficiency of a hybrid weighted double-broadband continuum estimator, using
$$
BB = -0.438\,g + 1.438\,r \quad \text{for N501},
$$
and
$$
BB = 0.856\,g + 0.144\,r \quad \text{for N419},
$$
with selection thresholds $\Delta_{\min}(N419)=0.71$ mag and $\Delta_{\min}(N501)=0.83$ mag together with $EW_0>20$ \AA. At typical SHELA depths of about $25.3$ AB mag in N419/N501 and about $26.3$ AB mag in $g$ and $r$, the paper finds that broadband data roughly one magnitude deeper than the narrowband recover nearly $80\%$ of LAEs, whereas equal-depth broadband and narrowband data recover only about $20\%$. DESI validation reports confirmation rates of about $93\%$, $96\%$, and $92\%$ at $z\approx2.4$, $3.1$, and $4.5$, respectively.

## 5. ShaLa/ShaneAO as a sodium laser guide star adaptive-optics system

In the instrumentation literature, ShaLa refers to the sodium laser guide star adaptive-optics facility at the Lick Observatory Shane 3-m telescope, comprising the ShaneAO adaptive-optics relay and the Shane Adaptive Red Camera and Spectrograph (ShARCS). The system operates behind the $3.05$-m primary at Cassegrain focus and is designed for diffraction-limited IR science from approximately $0.8$–$2.2\,\mu$m, with the as-built configuration covering $1.1$–$2.2\,\mu$m and provision to extend shortward [1407.8207].

The AO architecture is a sodium LGS system tuned to the Na D2 line at $\lambda=589$ nm, with both LGS and NGS modes. It uses a woofer–tweeter deformable-mirror configuration: an ALPAO 52-element high-stroke woofer with $\pm50\,\mu$m stroke handling low-order modes including fast tip–tilt, and a Boston Micromachines KILO-DM tweeter with $1020$ actuators for high-order correction. The abstract characterizes the correction bandwidth as “full dynamic range correction from tip/tilt to 16 cycles across the pupil,” and the paper relates this to the MEMS actuator density and wavefront-sensor sampling. Spatial-frequency control is therefore partitioned between stroke-limited low-order correction on the woofer and higher-order, lower-amplitude correction on the tweeter.

Wavefront sensing is performed with a variable-sampling Shack–Hartmann WFS on a Lincoln Labs CCID66 detector with $120\times120$ pixels and $1$–$2$ e$^{-}$ read noise at up to $1.5$ kHz. Two samplings are available: 8-across mode, corresponding to approximately $40$ cm subapertures at the primary, and 15-across mode, conventionally referred to as “16$\times$,” corresponding to approximately $20$ cm subapertures. A separate Marconi CCD39 tip–tilt sensor provides $80\times80$ pixels, $6$ e$^{-}$ noise, up to $1$ kHz operation, $0.3''$/pixel scale, and a $19''$ instantaneous field scanned over a $120''$ acquisition field. The centroiding law is explicitly given as
$$
c = \frac{A-B}{A+B+r},
$$
with a practical regularization choice $r=\sigma(\text{sky background})$.

The science instrument ShARCS has a pixel scale of $0.035''$/pixel over a $20''$ diameter AO-corrected field of view. It supports imaging on a Hawaii-2RG detector, spectroscopy at $R\approx500$ with a $0.1''$ slit and a planned upgrade to $R\approx2000$, coronagraphy, and polarimetry. The paper gives diffraction-limited scales for the $3.05$-m aperture using $\theta\approx1.22\lambda/D$: about $0.08$–$0.10''$ in Y/J, about $0.14''$ in H, and about $0.18''$ in K.

Several engineering details are emphasized because they determine practical sensitivity. Enhanced broadband protected-silver coatings increase throughput in both science and WFS paths; a cold pupil stop and baffling reduce thermal background; a fixed $6''$ field stop and a blackened disk suppress Rayleigh backscatter to sky level; and the opto-mechanical bench is designed to be exceptionally stiff for multi-hour stability. The stated flexure goals are to hold imaging within the diffraction limit for at least $1$ hour and spectroscopy within half the slit width ($0.05''$) for at least $4$ hours.

Commissioning results establish the instrument’s operating regime. First-light PSFs show clear Airy rings, and the paper reports an H-band Strehl of approximately $0.8$ relative to the internal calibrator. Using the Maréchal approximation,
$$
S \approx \exp\{-(2\pi\sigma/\lambda)^2\},
$$
the paper interprets $S\approx0.8$ at H as $\sigma\approx124$ nm RMS residual wavefront error. In LGS mode, the system closed the loop in 16$\times$ mode at $50$ Hz on two good-seeing nights and operated robustly in 8$\times$ mode under average to poor seeing. The authors state that routine 16$\times$ operation is expected with the higher-return fiber laser. Sensitivity is reported as approximately $m_K\approx20$ to $m_J\approx23$, and the combination of high Strehl, lower background, and improved throughput corresponds to an approximately $12\times$ improvement in exposure time to reach a given point-source SNR relative to the prior IRCAL-based system.

The paper positions ShaneAO as a pathfinder for next-generation AO through a specific set of design choices: woofer–tweeter control without a separate steering mirror, switchable WFS sampling, robust low-flux centroiding, throughput optimization through coatings, and flexure control sufficient for faint-object IR spectroscopy.

## 6. Conceptual distinctions and recurring themes

The most common misconception is that ShaLa names a single method or collaboration. In the cited literature, it does not. SHALA-LLM is an LLM-alignment framework built around annotator distributions and GRPO; multimodal ShaLa is a two-stage latent-variable generative model with a conditional diffusion prior; SHELA is a survey name in extragalactic astronomy, with “ShaLa” only an occasional misspelling; and ShaLa/ShaneAO is an adaptive-optics facility centered on a sodium laser guide star and ShARCS [2606.05376] [2508.17376] [1812.04646] [1407.8207].

The two ML usages do share a structural feature: both replace a simpler collapsed target with a richer distributional object. SHALA-LLM aligns to empirical annotator distributions rather than majority labels, while multimodal ShaLa replaces a fixed Gaussian latent prior with a learned diffusion prior fitted to the aggregated posterior. This suggests a common methodological intuition—retaining heterogeneity rather than suppressing it—but the papers present the systems independently and for different problem classes.

The astronomy usages differ again. SHELA is an observing program whose value derives from area, depth, multiwavelength coverage, catalog construction, forced photometry, and spectroscopic synergy with HETDEX. ShaneAO is an instrumentation platform whose value derives from wavefront correction, throughput, background control, and mechanical stability. Even where the same string appears, the objects of study are fundamentally different: galaxies and large-scale structure in one case, optical turbulence and diffraction-limited imaging in the other.

For technical reading, the disambiguating cues are straightforward. References to GRPO, JSD, ChaosNLI, GoEmotions, or Qwen2.5-Omni-7B identify SHALA-LLM. References to VAEs, diffusion priors, PolyMNIST, CUB, or ShapeNet identify multimodal ShaLa. References to Stripe 82, DECam, IRAC, HETDEX, or ODIN identify SHELA. References to sodium guide stars, ShARCS, woofer–tweeter DMs, Shack–Hartmann sensing, or Strehl identify ShaneAO. In practice, correct interpretation depends entirely on domain context rather than on the string “ShaLa” itself.

Source: https://www.emergentmind.com/topics/shala