---
title: 'EXCAM: Star Acquisition & Cultural Metric'
url: https://www.emergentmind.com/topics/excam
type: topic
---

# EXCAM: Star Acquisition & Cultural Metric

Searching arXiv for the provided EXCAM/ExCAM papers to ground the article in current sources.
Search query: 2507.09059
EXCAM denotes two distinct technical constructs in recent arXiv literature. In Roman Coronagraph Instrument operations, **EXCAM** refers to the **Exoplanetary System Camera** star-acquisition method used at the start of a coronagraphic observation, before focal-plane-mask insertion, to place a target star within the capture range required for subsequent low-order wavefront sensing and control [2507.09059]. In large-language-model evaluation, **ExCAM** refers to **Explainable Cultural Awareness Metric**, a reference-free, fine-grained metric that identifies, rates, and explains cultural errors in arbitrary instruction–output pairs, and is trained using the ExCAM40k dataset assembled from nine existing cultural benchmarks [2605.29897].

## 1. Nomenclature and scope

The shared label masks a strong domain divergence. One usage belongs to observational astrophysics and spacecraft control; the other belongs to LLM evaluation and culturally aware text assessment.

| Term | Domain | Core function |
|---|---|---|
| EXCAM | Roman Coronagraph Instrument | Initial star acquisition before FPM insertion |
| ExCAM | LLM evaluation | Identification, rating, and explanation of cultural errors |

In the Roman context, EXCAM is embedded in a closed-loop acquisition chain that begins with spacecraft slewing and ends with hand-off to LOWFS/C, FSM control, and deformable-mirror-based dark-hole maintenance [2507.09059]. In the LLM context, ExCAM is an MQM-style evaluator that outputs structured error reports and an aggregate scalar score for cultural awareness in generated text [2605.29897].

A common source of confusion is the acronym itself. The Roman paper uses EXCAM as an instrument-acquisition term tied to the CGI science camera, whereas the LLM paper uses ExCAM as the name of an evaluation metric. No technical overlap is indicated between the two usages.

## 2. EXCAM in the Roman Coronagraph Instrument

Within NASA’s Nancy Grace Roman Space Telescope, the Coronagraph Instrument is a technology demonstration for direct imaging and spectroscopy of exoplanets around nearby stars at ultra-high contrast, approximately \(10^{-9}\), using active wavefront control in space with two deformable mirrors [2507.09059]. The observational objective imposes three coupled requirements: precise pointing onto the coronagraph focal-plane mask, active wavefront control to dig a dark hole in the stellar PSF, and LOWFS/C stabilization of tip/tilt and focus drifts.

EXCAM is used only at the very start of each coronagraphic observation, when the focal-plane mask is not yet inserted and the full science-camera field of view is unobscured at \(19.5'' \times 19.5''\) [2507.09059]. The observatory Attitude Control System first slews to the commanded reference star and, per the CGI–ACS interface, must place that star within a \(4''\) radius \((3\sigma)\) of EXCAM’s center pixel. EXCAM then performs direct imaging and centroiding to refine the line of sight until the star lies within the fine-point capture range required for hand-off.

This stage is operationally significant because it furnishes the initial condition for the downstream coronagraphic control stack. EXCAM acquisition alone places the star to within \(\pm 0.054''\) of the nominal LoS axis, ensuring that the target lies within the linear capture range of the LOWFS Tip/Tilt \((Z2/Z3)\) sensor [2507.09059]. The paper is explicit that EXCAM itself does **not** command the deformable mirrors; its function is acquisition and alignment, not dark-hole control.

## 3. Acquisition workflow and control interfaces

The EXCAM star-acquisition sequence is implemented in CGI flight software and proceeds through command setup, initial spacecraft pointing, frame acquisition and preprocessing, star identification, centroid computation, capture-range testing, ACS offload, repointing, iteration, and hand-off to LOWFS/C [2507.09059]. Command setup includes configuring detector gain and exposure time according to stellar visual magnitude \((V \le 5)\), updating and loading the appropriate master-dark frame, and selecting both the number of exposures \((Nframes)\) and a photometric threshold for star detection.

Image processing applies cosmic-ray cleaning and dark subtraction to each frame, followed by frame co-addition or averaging to improve SNR. Candidate stars are identified as local maxima whose peak intensity satisfies \(I_{\text{peak}} > I_{\text{thresh}}\), where \(I_{\text{thresh}}\) is set such that any confusing star \((\ge V+3)\) is rejected. The centroid of a candidate star is then computed over a local region of interest by intensity-weighted first moments:
\[
x_i \;=\; \frac{\displaystyle\sum_{j} I_{ij}\,x_j}{\displaystyle\sum_{j} I_{ij}}\,, 
\quad
y_i \;=\; \frac{\displaystyle\sum_{j} I_{ij}\,y_j}{\displaystyle\sum_{j} I_{ij}}\,.
\]
The single-frame signal-to-noise ratio is approximated as
\[
\mathrm{SNR} \;=\; \frac{S}{\sqrt{\,S + B + \sigma_{\rm read}^2\,}}\,,
\]
and \(Nframes\) are chosen so that \(\mathrm{SNR} \ge X\), with \(X \sim 10\text{–}20\), for reliable centroiding.

The fine-point capture range is \(\pm 2.5\) pixels, approximately \(0.054''\) radius. If the centroid lies outside this capture circle, EXCAM computes pointing error offsets, converts pixel offsets \(\Delta x,\Delta y\) into spacecraft LoS-frame rotations \((\Delta H,\Delta V)\) in radians via the plate scale \(19.5''/2048\) pixels and the CGI LoS-frame definition, and sends a “single-delta-HV” offload packet at \(1\,\mathrm{Hz}\) to ACS [2507.09059]. ACS uses these corrections, slews the telescope, and flags “AT Offset” to FALSE until pointing settles. On ACS “hold attitude” flag TRUE, the frame-processing and offload cycle is repeated until the centroid lies within capture range.

After acquisition success, the focal-plane mask is inserted immediately and the system switches to LOCAM. At that stage the LOWFS/C FPGA closes a high-speed \(10\,\mathrm{kHz}\) FSM inner loop for tip/tilt, performs periodic piezo offloads to ACS to prevent FSM saturation, estimates Zernike modes beyond tip/tilt from the occulted pupil image, and issues deformable-mirror commands to correct quasi-static wavefront errors and maintain the dark hole [2507.09059]. This clarifies the division of labor: EXCAM establishes the initial line-of-sight geometry; LOWFS/C and the DMs sustain coronagraphic performance.

## 4. Thermal-vacuum demonstration and operational trade-offs

The full-system thermal-vacuum campaign at JPL exercised EXCAM acquisition in HLC NFOV mode using the Coronagraph Verification Stimulus to emulate star input and ACS response [2507.09059]. The reported performance verifies the relevant acquisition requirements: ACS pre-pointing error was at most \(4''\) radius \((3\sigma)\), the EXCAM capture-range requirement of at most \(0.054''\) radius was verified, StarID met the \(<60\,\mathrm{s}\) requirement, and FindCentroid met the \(<20\,\mathrm{s}\) requirement. End-to-end acquisition time was \(4\text{–}5\) minutes for bright \((V=0)\) stars placed at initial offsets of \(2.24''\) and \(4.24''\), and final pointing was at most \(0.05''\) from the EXCAM boresight.

The TVAC error budget partitions residual acquisition uncertainty across multiple subsystems: ACS boresight and catalog calibration contribute approximately \(2''\) RMS, observatory jitter contributes approximately \(0.1''\text{–}0.2''\) RMS, and centroid error contributes approximately \(0.01''\), with SNR dependence [2507.09059]. All test cases met requirements, including bright \(V=0\text{–}2\) stars, dim \(V=5\) stars, and operation under simulated ACS disturbances such as reaction-wheel zero-crossing and HGA moves.

The paper contrasts EXCAM with the Raster Scan method. EXCAM offers direct imaging of the star, simple centroiding, the full \(19.5''\) FoV, and fast initialization of LoS pointing in \(4\text{–}5\) minutes, while maintaining lower complexity because no FSM raster pattern is required [2507.09059]. Its limitations are equally explicit: it is only viable before FPM insertion and requires relatively bright stars \((V \le 5)\) to achieve adequate SNR in a few-second exposures. Raster Scan, by contrast, operates with the FPM in place and supports mid-sequence target reacquisition, including fainter stars behind the bowtie mask, but it requires more complex raster-trajectory design, careful FSM strain-gauge calibration, dark-frame handling, and longer routines, up to approximately \(160\,\mathrm{s}\) for the raster plus loop closing.

The operational lessons are primarily software and interface lessons. Early hardware-in-the-loop testing, adjustable flight-software parameters such as photometric threshold and capture-range margin, rapid master-dark and gain updates, comprehensive telemetry logging, minimal perturbation before dark-hole preparation, and robust state-machine handling for ACS disturbance flags are all identified as practical recommendations [2507.09059]. This suggests that EXCAM’s importance lies not only in centroiding accuracy but also in system-level fault tolerance and commissioning maturity.

## 5. ExCAM as an explainable cultural-awareness metric

ExCAM, the Explainable Cultural Awareness Metric, is defined as the first reference-free, fine-grained metric designed specifically to identify, rate, and explain cultural errors in arbitrary instruction–output pairs produced by LLMs [2605.29897]. It formalizes cultural awareness as the absence of culturally problematic content, including factual mistakes, misrepresentations, stereotypes, and other errors that would mislead or offend members of a target culture.

Given an instruction \(i\) and an LLM-generated output \(o\), ExCAM returns an MQM-style error report \(R\) and a scalar quality score \(s\). Each entry in \(R\) contains an error type, a span, a severity label with minor \(=-1\) and major \(=-5\), and a natural-language explanation. The scalar score is defined by
\[
s^* \;=\;\sum_{i=1}^n \mathrm{sev}_i\quad\bigl(\le 0\bigr)
\]
and
\[
s =
\begin{cases}
s^*, & s^* < 0,\\
p(R), & s^* = 0,
\end{cases}
\]
where \(p(R)\) is the length-normalized probability assigned by the LLM to the entire error report [2605.29897]. A perfect, error-free output therefore attains \(s=1\) when the model is maximally confident, while each major error deducts \(5\) points and each minor error deducts \(1\) point.

Methodologically, ExCAM is not a single monolithic model but a suite of LLM evaluators based on Gemma 3-27B and Phi 3-medium-4k adapted via LoRA fine-tuning. Input is wrapped in a prompt of the form `Instruction: {i}\nText: {o}\nReturn an error report in JSON.`, and robustness is promoted with 50 automatically generated paraphrases of these templates [2605.29897]. The JSON schema contains `error_type`, `span`, `severity`, and `explanation`, making the metric inherently interpretable.

Negative supervision is produced with synthetic error generation. “Hard errors” are drawn from existing QA benchmarks by selecting incorrect answer options and assigning severity \(=-5\) with templated explanations. “Soft errors” are generated by prompting Qwen3.5-122B in “thinking” mode to introduce either a minor or major cultural error into a correct instruction–output pair, together with an error type, explanation, severity, and modified sample, while discarding outputs no longer judged culture-related [2605.29897]. A key misconception corrected by the paper is that cultural evaluation must rely on reference answers or manual annotation at inference time; ExCAM is explicitly reference-free at evaluation time.

## 6. ExCAM40k, empirical performance, and explainability

ExCAM40k consolidates nine human-verified cultural benchmarks into an instruction–output format and augments them with synthetic errors [2605.29897]. The sources are grouped as QA with fixed options—BLEnD, CulturalBench, INCLUDE; free-form QA—CaLMQA, NativQA; free-text generation—Mango; and impersonation tasks—Normad, GlobalOpinionQA, EPIC. After capping each source at \(5{,}000\) samples, the corpus contains roughly \(18.8\,\mathrm{k}\) error-free examples, \(12.2\,\mathrm{k}\) hard-error samples, and \(8.3\,\mathrm{k}\) soft-error samples, yielding approximately \(40{,}000\) instances split into training, development, and test partitions.

The annotation pipeline assigns empty error reports to original human-verified pairs, introduces hard or soft errors as negative samples, records error spans by token-sequence diffing with NLTK and difflib, and treats the synthetic reports as ground truth. A human validation of 100 randomly sampled soft errors found that \(93.8\%\) were truly culture-related, \(78.8\%\) of generated explanations were valid, and severity labels matched in \(51.7\%\) of minor cases and \(95\%\) of major cases [2605.29897]. These figures delimit both the promise and the residual subjectivity of synthetic cultural-error generation.

Evaluation uses scaled accuracy, Kendall’s \(\tau\), and tie-calibrated accuracy, with LLM-based baselines including Qwen3-14B, DeepSeek-R1-Distill-Qwen-14B, Gemma 3-27B, Mistral-24B, Phi 3-medium-4k, Phi 4-mini, and GPT-5, and with counting prompts outperforming binary prompts [2605.29897]. On the in-domain test set, the Gemma-based ExCAM variant achieves **0.591** scaled accuracy, **0.569** Kendall’s \(\tau\), and **0.680** tie-calibrated accuracy, compared with **0.351**, **0.349**, and **0.495** for Phi 4\(_{\text{cnt}}\), the strongest unfine-tuned baseline listed. The paper further states that the Gemma-based variant attains a raw error-detection accuracy of approximately \(80\%\) and that its Kendall correlation is significantly higher \((p<0.05)\) than that of any baseline.

Generalization is evaluated through leave-one-out LoRA training, omitting each source dataset in turn and testing on its held-out split. On six of nine benchmarks—BLEnD, CulturalBench, INCLUDE, Mango, NativQA, and Normad—ExCAM maintains significantly better scaled accuracy than any unfine-tuned baseline [2605.29897]. On a balanced 900-sample subset, ExCAM\(_{\text{Gemma}}\) significantly outperforms GPT-5 on six out of nine benchmarks. The paper interprets this as evidence that a dedicated, fine-tuned metric can exceed a general-purpose LLM in nuanced cultural-error detection.

Explainability is central rather than incidental. Because ExCAM returns an error report rather than a single scalar, it specifies what the error is, where it appears, how severe it is, and why it is wrong. In a spot-check of \(87\%\) of hard-error cases, ExCAM flagged the absence of the correct answer, and in \(48\%\) of its explanations it literally reproduced the one-to-one correct answer as the remedy [2605.29897]. In soft-error cases, human raters assigned its explanations an average Likert score of \(4.01/5\). Representative examples include free-text overgeneralization such as “No Germans like schnitzel,” labeled as an overgeneralization with severity \(-5\), and impersonation mismatch such as a Malaysian sleep-response judged atypical and labeled with severity \(-1\).

The future directions identified by the paper are severity calibration with richer human labels, real-world error corpora beyond synthetic perturbations, ensembling and debiasing across LLM families to mitigate self-bias, and extension to broader modalities such as image or multimodal outputs [2605.29897]. A plausible implication is that ExCAM’s long-term significance depends not only on benchmark performance but also on whether its explanation schema becomes stable enough to support iterative model debugging, alignment audits, and culturally conditioned error analysis across domains.

Source: https://www.emergentmind.com/topics/excam