---
title: 'ENIIGMA Fitting Tool: Infrared Ice Spectroscopy'
url: https://www.emergentmind.com/topics/eniigma-fitting-tool
type: topic
---

# ENIIGMA Fitting Tool: Infrared Ice Spectroscopy

Searching arXiv for ENIIGMA fitting tool and closely related infrared ice spectroscopy work.
ENIIGMA, short for **dEcompositioN of Infrared Ice features using Genetic Modelling Algorithms**, is a publicly available and open-source Python toolbox for spectroscopy analysis of infrared spectra that decomposes protostellar ice absorption spectra into physically meaningful combinations of laboratory ice spectra. It was introduced to automate continuum determination, silicate extraction, spectral decomposition, and statistical analysis, with explicit calculation of confidence intervals and quantification of degeneracy in the inferred ice composition [2107.08555]. The tool is designed for broad-band infrared data, and in the Elias 29 application it was used over the range \(2.5\text{--}20\,\mu\mathrm{m}\), where overlapping bands, matrix effects, thermal history, and energetic processing make manual or narrow-window fitting intrinsically ambiguous.

## 1. Scientific role and scope

ENIIGMA was developed for the interpretation of interstellar and circumstellar ice absorption bands in dense clouds and protostellar envelopes, where spectra contain broad, overlapping, and environment-dependent features from species such as H\(_2\)O, CO, CO\(_2\), CH\(_3\)OH, NH\(_3\), CH\(_4\), and NH\(_4^+\) [2107.08555]. The stated motivation is that a variety of laboratory ice spectra simulating different chemical environments, ice morphology, and thermal and energetic processing are required for an accurate interpretation of protostellar spectra, and that an automated statistically based computational approach becomes necessary to determine which combination of laboratory data best fits the observations.

The method differs from phenomenological multi-Gaussian fitting in narrow windows because it fits a linear combination of laboratory spectra across a broad wavelength interval. The tool therefore searches over a large database of laboratory analogs rather than restricting the analysis to pre-assigned carriers in a small spectral region. This suggests that ENIIGMA is best understood not as a line-by-line classifier, but as a spectral-decomposition framework in which molecular identification is coupled to laboratory analog selection and to the inferred physical environment.

Its astrophysical emphasis is explicitly protostellar infrared ice spectroscopy. The paper describes applications to ISO-SWS, Spitzer-IRS, and spectra of the same general type, and states that the toolbox is well timed with the launch of the James Webb Space Telescope [2107.08555].

## 2. Observational inputs and laboratory basis

ENIIGMA accepts three classes of input spectra in ASCII format with three columns: wavelength in \(\mu\mathrm{m}\), flux or optical depth, and error [2107.08555]. The accepted forms are flux density spectra, optical-depth spectra including silicate features, and optical-depth spectra with silicates already removed. This separation reflects the internal structure of the code, since continuum fitting and optical-depth conversion may be performed internally, while silicate removal may either be performed by the tool or supplied externally.

The internal laboratory database contains approximately 100 spectra collected from the NASA Ames ice database, the Leiden Ice Database, and the UNIVAP NKABS database [2107.08555]. The database includes pure ices, mixtures, heated samples, UV-processed ices, ion-irradiated ices, and residues at elevated temperature. The examples listed in the source material include pure H\(_2\)O, CO, CO\(_2\), CH\(_3\)OH, CH\(_4\), NH\(_3\), CH\(_3\)CHO, HCOOH, CH\(_3\)CH\(_2\)OH, and CH\(_3\)OCH\(_3\), as well as mixed compositions such as H\(_2\)O:CO:NH\(_3\):CH\(_3\)OH and H\(_2\)O:NH\(_3\):CO\(_2\):CH\(_4\).

Laboratory spectra measured in transmittance or absorbance are converted consistently. If \(T\) is transmittance and \(A\) is absorbance, then
\[
A = -\log_{10} T,
\]
and the laboratory optical depth is
\[
\tau^{\rm lab}_\lambda = 2.3\,A_\lambda.
\]
Baselines of the laboratory spectra were checked or corrected to avoid misassignment. The database is manually grouped by type—pure, pure+heated, mixtures+heated, irradiated, and residues—and is described as easily extensible by the user [2107.08555].

A common misconception is that the tool directly identifies molecules independently of the database. The paper states the opposite in practical terms: if an astrophysical component is absent from the database, ENIIGMA approximates it with the closest available analog. This means that identifications are always conditioned on the coverage and quality of the laboratory library.

## 3. Pre-processing: continuum, optical depth, and silicate extraction

For flux spectra \(F^{\rm obs}_\lambda\), ENIIGMA offers two continuum models [2107.08555]. The first is a low-order polynomial,
\[
F_{\mathrm{poly}}(\lambda) = \sum_{k=0}^{n} a_k \lambda^k,
\]
used in the paper for wavelength regions such as \(5\text{--}30\,\mu\mathrm{m}\). The second is a multi-blackbody continuum,
\[
F_{\mathrm{BB}}(\lambda, T)
= \sum_{i=1}^{m} f_i
\frac{2hc^2}{\lambda^5}\left[\exp\!\left(\frac{hc}{\lambda k_{\mathrm B} T_i}\right)-1\right]^{-1},
\]
used for \(2\text{--}4\,\mu\mathrm{m}\) and for SED-like continua. Continuum windows free of strong ice features are chosen by the user.

Once the continuum \(F_\lambda^{\rm cont}\) is determined, the observed optical depth is computed as
\[
\tau_\lambda^{\rm obs} =
-\ln\left(\frac{F_\lambda^{\rm obs}}{F_\lambda^{\rm cont}}\right).
\]
This step is operationally central because the subsequent decomposition is performed in optical-depth space.

The \(8\text{--}25\,\mu\mathrm{m}\) region is complicated by the 9.7 and \(18\,\mu\mathrm{m}\) silicate bands, which overlap with ice features such as NH\(_3\) at \(9.01\,\mu\mathrm{m}\), CH\(_3\)OH at \(9.74\,\mu\mathrm{m}\), ethanol near \(9.5\,\mu\mathrm{m}\), and the H\(_2\)O libration band near \(13\,\mu\mathrm{m}\) [2107.08555]. ENIIGMA therefore implements a synthetic silicate method. The empirical GCS 3 silicate profile is scaled to the observed 9.7 and \(18\,\mu\mathrm{m}\) peaks and decomposed into six Gaussians with fixed centers
\[
\lambda_0 = 8.3,\ 9.7,\ 11.2,\ 16.2,\ 18.0,\ 20.8\ \mu\mathrm{m}.
\]
Each Gaussian is written as
\[
G(\lambda; A,\lambda_0,\sigma) =
\frac{A}{\sigma\sqrt{2\pi}}
\exp\!\left[-\frac{(\lambda-\lambda_0)^2}{2\sigma^2}\right],
\]
and the synthetic silicate optical depth is
\[
\tau_\lambda^{\rm ss} = \sum_{j=1}^{6} G(\lambda; A_j,\lambda_{0,j},\sigma_j).
\]
The ice-only optical depth is then
\[
\tau_\lambda^{\rm ice} = \tau_\lambda^{\rm obs} - \tau_\lambda^{\rm ss}.
\]

The paper states that this synthetic method handles the 9.7 and \(18\,\mu\mathrm{m}\) silicate bands simultaneously, which improves the separation of H\(_2\)O libration and CO\(_2\) bending features in the \(11\text{--}18\,\mu\mathrm{m}\) range [2107.08555].

## 4. Genetic-modelling spectral decomposition

After pre-processing, ENIIGMA represents the ice-only optical depth spectrum as a linear combination of laboratory spectra,
\[
\tau_{\lambda}^{\rm model} = \sum_{j=0}^{m-1} w_j \tau_{\lambda,j}^{\rm lab},
\]
where the coefficients \(w_j\) are the free parameters [2107.08555]. In the genetic-algorithm formulation, each candidate solution is a chromosome whose genes are the coefficients of the selected laboratory components.

The quality of a candidate solution is measured by the squared residual
\[
F = \sum_{i=0}^{n-1} \left( \tau^{\rm obs}_{\lambda,i} - \sum_{j=0}^{m-1} w_j\,\tau^{\rm lab}_{\lambda,j} \right)^2.
\]
This objective is directly proportional to a \(\chi^2\) when a constant observational standard deviation is assumed.

The implementation uses the **Pyevolve** Python library for the genetic algorithm [2107.08555]. The cycle includes initialization, parent selection, crossover, mutation, evaluation, and termination. The paper describes roulette-wheel and tournament selection, several crossover operators, and Gaussian mutation,
\[
w_{mn}' = w_{mn} + \zeta,
\]
where \(\zeta\) is drawn from a Gaussian distribution. The reported mutation rate is typically \(10\%\), and the population-to-generation ratios used in the work are typically \(90/100\) or \(150/200\). Tournament selection with high crossover rate is reported to have performed best for complex fits.

A distinctive feature of ENIIGMA is that it searches not only the coefficients \(w_j\) but also the component set. The user first supplies an initial guess containing four laboratory spectra. A genetic fit is performed for that set, producing a residual \(F_1\). Additional laboratory spectra are then added one at a time, each augmented set is refit, and components for which the new residual \(F_2\) satisfies \(F_2 < F_1\) are kept as promising. The genetic algorithm is then rerun using groups of 7–8 or more components drawn from this filtered set to search for the global minimum [2107.08555]. This design is why the paper states that the initial guess is not binding.

The tool is therefore not restricted to fitting a single predefined mixture. It performs a database-guided search over broad combinations of pure, mixed, heated, irradiated, and residue spectra, constrained by the observed spectrum and the laboratory archive.

## 5. Statistical analysis, confidence regions, and degeneracy

ENIIGMA supplements the heuristic search with a post-processing statistical analysis intended to derive confidence intervals and quantify degeneracy [2107.08555]. The starting point is
\[
\chi^2 = \frac{F}{\sigma^2},
\]
and confidence regions are defined through
\[
\Delta\chi^2(\nu,\alpha) = \chi^2 - \chi_{\rm min}^2,
\]
where \(\nu\) is the number of free parameters and \(\alpha\) is the significance level.

The tool samples parameter space around the optimum by treating each coefficient \(w\) as a random variable drawn from a normal distribution,
\[
p(x) = \frac{1}{\sqrt{2\pi\sigma^2}}
\exp\!\left[-\frac{(x-\mu)^2}{2\sigma^2}\right],
\]
with mean \(\mu\) equal to the optimal coefficient. The paper gives specific \(\Delta\chi^2\) thresholds, including \(2.41\), \(4.61\), and \(9.21\) for 1–3\(\sigma\) with two components, and \(9.52\), \(13.36\), and \(20.09\) for 1–3\(\sigma\) with eight components [2107.08555]. The resulting 2D correlation maps expose anti-correlations and component trade-offs.

Column densities are computed from the optical-depth profile and the band strength \(\mathcal{A}\) using
\[
N_{\rm ice} = \frac{1}{\mathcal{A}} \int_{\nu_1}^{\nu_2} (w \cdot \tau_\nu)\, d\nu.
\]
Lower and upper bounds are obtained by repeating the integration with the lower and upper limits of \(w\) derived from the confidence analysis.

Degeneracy is further characterized through **recurrence plots**. ENIIGMA collects all acceptable solutions within a chosen \(\Delta\chi^2\) threshold, counts how often each laboratory component appears, and defines the recurrence
\[
R_i = \frac{f_i}{S},
\]
where \(f_i\) is the frequency of component \(i\) and \(S\) is the total number of acceptable solutions [2107.08555]. A recurrence of \(100\%\) indicates that a component appears in all acceptable solutions, whereas lower recurrence means that it can be replaced by other components without significantly degrading the fit.

The tool also constructs histograms of \(N_{\rm ice}\) over all acceptable solutions, using the Freedman–Diaconis estimator for bin widths. This is important because different decompositions can have similar \(\chi^2\) but different inferred compositions. A plausible implication is that ENIIGMA’s output should be interpreted as an ensemble of statistically acceptable decompositions rather than as a single unique molecular inventory.

## 6. Validation, Elias 29, and methodological boundaries

The paper reports three categories of validation: known pure and mixed laboratory spectra, fractionation tests with added Gaussian noise, and a synthetic ice spectrum containing eight pure ice components plus silicate absorption and Gaussian noise [2107.08555]. In the tests with known samples, the tool correctly retrieved the exact spectra and coefficients in both “fully sighted” and “fully blind” cases when the correct sample was present in the database. When the correct sample was removed, ENIIGMA selected the most similar available analog, such as the same species at a different temperature. Fractionation tests showed that strong degeneracy appears when one component is minor, for example in the \(90{:}10\) case of H\(_2\)O(15 K) + H\(_2\)O(40 K), where a pure 15 K solution remained statistically allowed.

The principal astrophysical demonstration is the Class I protostar **Elias 29**, for which ENIIGMA was applied to ISO-SWS/LWS and Spitzer-IRS data over \(2.5\text{--}30\,\mu\mathrm{m}\), with the decomposition focused on \(2.5\text{--}20\,\mu\mathrm{m}\) [2107.08555]. The continuum was modeled with a single blackbody of \(717\) K over \(2.5\text{--}4.2\,\mu\mathrm{m}\) and a third-order polynomial over \(4.1\text{--}30\,\mu\mathrm{m}\). After synthetic silicate subtraction, an initial guess of pure H\(_2\)O, CO, CO\(_2\), and NH\(_3\) led to an optimal seven-component solution.

The seven-component decomposition included a heavy-ion-processed H\(_2\)O:NH\(_3\):CO\(_2\):CH\(_4\) mixture at 35 K, a UV-processed H\(_2\)O:NH\(_3\):CH\(_3\)OH:CO:CO\(_2\) mixture, pure and mixed CO\(_2\), pure CO, pure and mixed H\(_2\)O, NH\(_3\)-bearing components including residues, and an H\(_2\)O:CH\(_3\)CH\(_2\)OH mixture [2107.08555]. The paper states that the overall fit from \(2.5\) to \(20\,\mu\mathrm{m}\) is excellent, while also noting residual systematic effects such as the underfit red wing of the \(3\,\mu\mathrm{m}\) band.

For Elias 29, the recurrence analysis within \(3\sigma\) showed that three components had \(R=100\%\): the processed H\(_2\)O:NH\(_3\):CO\(_2\):CH\(_4\) mixture, the UV-processed H\(_2\)O:NH\(_3\):CH\(_3\)OH:CO:CO\(_2\) mixture, and CO\(_2\) [2107.08555]. CO and pure H\(_2\)O had high recurrence, while the only laboratory component containing ethanol, the H\(_2\)O:CH\(_3\)CH\(_2\)OH mixture, had recurrence of approximately \(78\%\). The paper therefore describes the ethanol inference as a **tentative detection** rather than a definitive identification.

The global-minimum column densities reported for Elias 29 include H\(_2\)O \(= 33.1^{+8.0}_{-12.0} \times 10^{17}\,\mathrm{cm^{-2}}\), CO\(_2\) \(= 5.22^{+1.91}_{-1.82} \times 10^{17}\,\mathrm{cm^{-2}}\), CO \(= 1.55^{+0.26}_{-0.82} \times 10^{17}\,\mathrm{cm^{-2}}\), NH\(_3\) \(= 1.01^{+0.46}_{-0.26} \times 10^{17}\,\mathrm{cm^{-2}}\), CH\(_4\) \(= 0.38^{+0.13}_{-0.21} \times 10^{17}\,\mathrm{cm^{-2}}\), H\(_2\)CO \(= 0.38^{+0.04}_{-0.25} \times 10^{17}\,\mathrm{cm^{-2}}\), CH\(_3\)OH \(= 0.86^{+0.30}_{-0.80} \times 10^{17}\,\mathrm{cm^{-2}}\), NH\(_4^+\) \(= 1.15^{+1.97}_{-0.45} \times 10^{17}\,\mathrm{cm^{-2}}\), and CH\(_3\)CH\(_2\)OH \(= 0.08^{+0.03}_{-0.05} \times 10^{17}\,\mathrm{cm^{-2}}\) [2107.08555].

The paper also states clear limitations. Continuum subtraction and silicate removal can imprint artifacts in the \(13\text{--}20\,\mu\mathrm{m}\) region; the laboratory database is finite and biased; grain shape and size effects are not currently treated through Mie or CDE corrections; and intermolecular environment effects are represented only through the discrete mixtures present in the database [2107.08555]. These limitations are methodological rather than incidental. They imply that ENIIGMA is strongest when used as a statistically explicit decomposition tool tied to a well-curated laboratory archive, and weaker when laboratory coverage is incomplete or when dust radiative-transfer effects dominate the spectral morphology.

Source: https://www.emergentmind.com/topics/eniigma-fitting-tool