Papers
Topics
Authors
Recent
Search
2000 character limit reached

MIEL: Multifaceted Instrumentation & Robotics Framework

Updated 9 July 2026
  • MIEL is an ambiguous research term that denotes both an optical apparatus for SiPM characterization and a robotic framework for resolving demonstrative ambiguity.
  • In optical applications, it employs advanced cryogenic microscopy and calibrated photon-yield measurements essential for detector R&D and cross-talk analysis in SiPM arrays.
  • In robotics, it integrates sound localization, semantic mapping, and interactive questioning to enhance resolution of ambiguous, exophoric user instructions.

Searching arXiv for the cited MIEL-related papers to ground the article in the current literature. MIEL is an ambiguous research term rather than a single established concept. In recent arXiv literature, it denotes at least two distinct systems: the Microscope for the Injection and Emission of Light, an apparatus for wavelength-resolved measurement of secondary-photon emission from single-photon avalanche diodes in silicon photomultipliers (Raymond et al., 2024), and Multimodal Interactive Exophora resolution with user Localization, a robotic framework for resolving ambiguous demonstrative instructions such as “Bring me that” when objects or users may be outside the robot’s field of view (Oyama et al., 22 Aug 2025). In adjacent usage, the lowercase Spanish noun miel appears in work on automated palynological analysis for honey quality control (Staforelli-Vivanco et al., 17 Apr 2026), while Code-MIE is a separate multimodal information extraction framework whose acronym is related but not identical (Liu et al., 21 Mar 2026). This terminological overlap makes contextual disambiguation essential.

1. Nomenclature and scope

The designation MIEL is explicitly expanded in two different ways in the supplied literature. In detector instrumentation, it refers to the Microscope for the Injection and Emission of Light, an experimental setup constructed to measure spectra of secondary photons emitted from Hamamatsu VUV4 and Fondazione Bruno Kessler VUV-HD3 SiPMs stimulated by laser light near operational voltages (Raymond et al., 2024). In robotics, it refers to Multimodal Interactive Exophora resolution with user Localization, a multimodal exophora resolution framework leveraging sound source localization, semantic mapping, visual-LLMs, and interactive questioning with GPT-4o (Oyama et al., 22 Aug 2025).

This reuse of the same acronym across unrelated subfields is not unusual in contemporary research, but it carries practical consequences. A plausible implication is that bibliographic retrieval, citation parsing, and acronym expansion models must rely on domain cues rather than token identity alone. The supplied corpus also shows two nearby but distinct usages: miel as the Spanish term for honey in a microscopy system for melissopalynology (Staforelli-Vivanco et al., 17 Apr 2026), and Code-MIE as a code-style multimodal information extraction framework (Liu et al., 21 Mar 2026). These are terminologically adjacent but conceptually separate.

2. MIEL as the Microscope for the Injection and Emission of Light

In (Raymond et al., 2024), MIEL is an optical and cryogenic measurement apparatus designed to characterize avalanche-induced secondary-photon emission from individual SPADs within SiPM devices. The setup is built on an Olympus IX-83 inverted microscope body with a broadband AR-coated sapphire vacuum window mounted on a custom flange above the objective turret. Its cryogenic sample stage uses a Micronix sub-micron x–y cryo-positioning stage on a copper cold finger, with cooling provided by an LN₂ suction pump (Instec LN2-P) and 25 L Dewar, and temperature control by Instec MK1000 with stability <0.1 K over 86 K…293 K (Raymond et al., 2024).

The laser injection path employs a PicoQuant 405 nm pulsed laser head (PDL 800-D) fiber-coupled via an OZ Optics variable attenuator, an off-axis parabolic mirror (Thorlabs RC12FC-P01), a manual iris aperture, and a 445 nm long-pass dichroic mirror that reflects the 405 nm beam into an Olympus LCPLN20XIR 20×, NA 0.45 IR-corrected objective. This objective focuses the beam onto a single SPAD. Avalanche bias is supplied by a Keithley 2280S-60-3 precision bias supply (0…60 V, sub-mV resolution), while the SPAD waveform is acquired on a CAEN DT730 2 GS/s digitizer using 5,000-sample (2 ns/bin) waveforms triggered by the laser sync (Raymond et al., 2024).

The same objective also recouples back-emitted photons. These photons pass through the dichroic path and a 550 nm long-pass filter before entering a Princeton Instruments HRS-300 spectrograph with 150 g/mm @800 nm grating and a Princeton Instruments PyLoN 400BR eXcelon CCD camera (1,340×400 pixels, 20 µm square, cooled to 163 K). In zeroth-order imaging mode the full SPAD area maps vertically onto the CCD, whereas in spectroscopy mode the horizontal axis is wavelength-calibrated (Raymond et al., 2024).

The operational principle is to stimulate a single SPAD at a pulse rate R=flaser=250kHzR=f_\mathrm{laser}=250\,\mathrm{kHz}, chosen to ensure complete SPAD recharge between flashes, with attenuation adjusted so that Pf50%80%P_f\approx 50\%\text{–}80\% of pulses trigger at least one avalanche. Each avalanche at the p–n junction produces hot carriers and bremsstrahlung photons in the 550 nm…1,050 nm range, and a fraction escapes through the Si–SiO₂ interface into the microscope acceptance (Raymond et al., 2024).

3. Calibration, inference, and detector-physics outputs

The MIEL apparatus in (Raymond et al., 2024) is notable for converting CCD counts into an absolute per-avalanche photon-yield estimate. The raw CCD units are converted from ADU to photo-electrons via the gain,

νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.

The apparatus spectral transmission ϵ(λ)\epsilon(\lambda) is determined from the relative shape measured with an OceanInsight HL-3 calibrated lamp and from absolute normalization at selected wavelengths using a pulsed single-mode fiber + Thorlabs PM101/S150C power meter vs. CCD counts, with combined systematic uncertainty ≲3–5 % overall (Raymond et al., 2024).

The basic photon-yield relation is

nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},

where Nγc(λ)N_\gamma^c(\lambda) is the number of photo-electrons detected on the CCD for the central SPAD and NavcN_{av}^c is the number of avalanches in that SPAD during the exposure (Raymond et al., 2024). The avalanche count is modeled as

Navc  =  Δt  R  Pf(ΔV,T)  Pc  (1+Ncda(ΔV,T)),N_{av}^c \;=\; \Delta t\;R\;P_f(\Delta V,T)\;P_c\;(1 + N_{cda}(\Delta V,T)),

with Δt\Delta t the camera exposure, RR the laser rate, Pf50%80%P_f\approx 50\%\text{–}80\%0 the probability that a laser pulse triggers at least one avalanche, Pf50%80%P_f\approx 50\%\text{–}80\%1 the confinement probability, and Pf50%80%P_f\approx 50\%\text{–}80\%2 the mean number of correlated delayed avalanches per prompt avalanche (Raymond et al., 2024). The full wavelength-, voltage-, and temperature-dependent yield is then

Pf50%80%P_f\approx 50\%\text{–}80\%3

The reported spectra show that both devices exhibit rising emission toward approximately 1,000 nm, followed by rapid fall-off above that range. The FBK VUV-HD3 shows SiO₂ interference fringes, whereas the HPK VUV4 spectrum is smooth, attributed in the source text to much thinner SiO₂. No significant change in spectral shape versus temperature is observed within errors (Raymond et al., 2024).

At 4 V over-voltage and room temperature, the integrated photons per SPAD avalanche into the NA = 0.45 objective over 550…1,000 nm are reported as Pf50%80%P_f\approx 50\%\text{–}80\%4 photons/avalanche for HPK VUV4 and Pf50%80%P_f\approx 50\%\text{–}80\%5 photons/avalanche for FBK VUV-HD3. The source yields inferred at the p–n junction over the same wavelength interval are Pf50%80%P_f\approx 50\%\text{–}80\%6 photons/avalanche for HPK VUV4 and Pf50%80%P_f\approx 50\%\text{–}80\%7 photons/avalanche for FBK VUV-HD3, corresponding to Pf50%80%P_f\approx 50\%\text{–}80\%8 Pf50%80%P_f\approx 50\%\text{–}80\%9 and νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.0 νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.1, respectively (Raymond et al., 2024). The abstract of the same paper reports νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.2 and νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.3 photons produced per avalanche at 4 volts over-voltage for VUV4 and VUV-HD3, and states that no significant temperature dependence is observed within the measurement uncertainties (Raymond et al., 2024). The coexistence of these figures reflects different reported yield summaries within the paper’s supplied materials and should therefore be interpreted in the context of the specific integration and extrapolation procedure being cited.

The authors also report the total escape yields into different media: in air, HPK 0.64 ± 0.26 νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.4/avalanche and FBK 0.79 ± 0.31; in LXe (νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.5), HPK 1.55 ± 0.62 and FBK 1.92 ± 0.74 (Raymond et al., 2024). These results are directly relevant to external cross-talk estimation in large SiPM arrays.

4. Measurement protocol, capabilities, and limitations of the detector MIEL

The measurement protocol is highly structured. At each νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.6 operating point, the system first performs focus alignment on the central SPAD, checked using the CCD image of the 405 nm reflection. It then records νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.7 using the digitizer over 100 k waveforms, acquires 5 emission frames + 25 SiPM-off frames on the CCD, subtracts backgrounds, masks cosmic rays, sums the central-SPAD pixels into a spectrum, and computes νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.8 through the calibrated yield expression (Raymond et al., 2024). Typical scans cover 1 V…5 V above νADUe=0.65  eADU.\nu^{e^-}_{\mathrm{ADU}} = 0.65\;\frac{e^-}{\mathrm{ADU}}.9, with measurements at 293 K and cryogenic points including FBK 172 K and HPK 165 K, plus intermediate temperatures for trend checks (Raymond et al., 2024).

The stated capabilities are threefold: single-SPAD confinement of avalanches via focused 405 nm laser, wavelength-resolved spectroscopy of avalanche-induced bremsstrahlung (550 nm…1,050 nm), and operation from 293 K down to 86 K with sub-0.1 K stability, with simultaneous electrical and optical readout (Raymond et al., 2024). Its stated limitations are equally explicit: the NA 0.45 objective limits collection to ∼6 % of 2π hemisphere, spectral resolution is sacrificed because the slit is widened to cover the SPAD pixel, and corrections for after-pulsing and internal cross-talk rely on literature values such as ϵ(λ)\epsilon(\lambda)0 and ϵ(λ)\epsilon(\lambda)1 (Raymond et al., 2024).

The significance claimed for SiPM characterization is that the instrument provides absolute secondary-photon spectra vs. over-voltage & temperature, informs external cross-talk models for large arrays (nEXO, DarkSide-20k, DUNE, etc.), confirms bremsstrahlung yields at the ϵ(λ)\epsilon(\lambda)2 ϵ(λ)\epsilon(\lambda)3 level, and quantifies device-to-device differences such as trench fill and SiO₂ thickness (Raymond et al., 2024). This places MIEL within detector R&D rather than general optical microscopy.

5. MIEL as Multimodal Interactive Exophora resolution with user Localization

In (Oyama et al., 22 Aug 2025), MIEL denotes a robotic framework for multimodal exophora resolution under incomplete observability. The target problem is the interpretation of ambiguous verbal instructions involving demonstratives, especially when objects or users are out of the robot’s view. The framework combines sound source localization (SSL), semantic mapping, visual-LLMs (VLMs), and interactive questioning with GPT-4o (Oyama et al., 22 Aug 2025).

Its architecture is organized into four major stages. First, preprocessing converts speech to text and derives text features with SentenceBERT and CLIP, while GPT-4o extracts the demonstrative token. Second, perception modules provide a semantic map, skeleton detection, and sound-source localization. Third, three estimators are computed: ϵ(λ)\epsilon(\lambda)4 for the linguistic query, ϵ(λ)\epsilon(\lambda)5 for the demonstrative region, and ϵ(λ)\epsilon(\lambda)6 for the pointing direction. Fourth, these are combined into a joint score, and the top candidates are optionally passed to GPT-4o for clarification (Oyama et al., 22 Aug 2025).

The semantic map stores, for each object ϵ(λ)\epsilon(\lambda)7, a class label and SentenceBERT embedding ϵ(λ)\epsilon(\lambda)8, a CLIP image feature ϵ(λ)\epsilon(\lambda)9, and a 3D position nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},0. Object detections such as Detic are fused through a standard occupancy-grid or TSDF approach. The probabilistic formulation includes the class posterior

nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},1

and a voxel occupancy update

nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},2

(Oyama et al., 22 Aug 2025).

When the user is out of view, MIEL uses a ReSpeaker v2 microphone array and a TDoA-based SSL. The estimated bearing and confidence are

nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},3

and if nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},4, corresponding to half the camera field of view, SSL is declared successful and the robot rotates toward the user (Oyama et al., 22 Aug 2025). This orientation step is not ancillary; it is part of the framework’s definition of user localization.

The language-grounding component combines a SentenceBERT-based label score and a CLIP-based attribute score: nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},5

nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},6

with

nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},7

The final joint score is

nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},8

where nγ(λ)=Nγc(λ)ϵ(λ)  Navc,n_\gamma(\lambda) = \frac{N_\gamma^{c}(\lambda)}{\epsilon(\lambda)\;N_{av}^c},9 is described as a demonstrative-region estimator (3D Gaussian on “ko/so/a” series) and Nγc(λ)N_\gamma^c(\lambda)0 as a pointing estimator (von Mises on angle between eye→wrist vs eye→object) (Oyama et al., 22 Aug 2025).

6. Interactive disambiguation and empirical performance of the robotic MIEL

A defining property of the robotic MIEL is that ambiguity is not handled solely by ranking. After computing the joint distribution and extracting the top-5 candidates, the system calls GPT-4o with a prompt that includes the user query and candidate objects. GPT-4o either selects the best ID or asks a single clarifying question such as “Is it red or blue?”; the user’s reply is appended to the query, and the resolution process is rerun once more (Oyama et al., 22 Aug 2025). The paper explicitly states that this single-turn Q/A dramatically improves resolution on minimally specified queries (Oyama et al., 22 Aug 2025).

The real-world evaluation uses a domestic living-room mock-up containing 114 objects (39 classes) mapped in advance with Detic+Objects365, a Toyota HSR robot with Intel i9-12900KF, RTX3080, ReSpeaker Mic Array v2, and Xtion PRO LIVE RGB-D camera, and a dataset of 90 Japanese queries divided into three levels of informativity: L1 class+feature+demonstrative, L2 class+demonstrative, and L3 demonstrative only (Oyama et al., 22 Aug 2025).

The reported top-1 success rates against ECRAP are summarized below.

Method Vis SR Unseen
ECRAP 0.42 0.27
MIEL 0.53 0.53

From these values, the paper states that MIEL achieves approximately 1.3 times better performance when the user is visible and 2.0 times better when the user is not visible, relative to methods without SSL and interactive questioning (Oyama et al., 22 Aug 2025). Top-5 SR rises from 0.67→0.79 in the visible condition and 0.57→0.79 in the unseen condition (Oyama et al., 22 Aug 2025).

The ablation study isolates the effects of SSL and questioning. SSL only yields 0.36 / 0.36 top-1 success for visible / unseen conditions; Q/A only yields 0.53 / 0.32; and SSL + Q/A (full MIEL) yields 0.53 / 0.53 (Oyama et al., 22 Aug 2025). The paper further states that removing interactive Q/A largely hurts L3 cases, with success rate dropping by more than a factor of two, whereas removing SSL collapses the unseen condition from 0.53 to 0.36. A paired t-test on SR across 90 trials indicates Nγc(λ)N_\gamma^c(\lambda)1 improvements of full MIEL over ECRAP in both conditions (Oyama et al., 22 Aug 2025).

These results identify MIEL not merely as a perception stack but as a decision process coupling multimodal grounding with controlled interaction. A plausible implication is that its gains derive from a division of labor: SSL restores observability, while GPT-4o reduces residual semantic ambiguity.

The supplied literature includes two neighboring but distinct terms that are relevant when interpreting “MIEL” in search and citation contexts. First, (Staforelli-Vivanco et al., 17 Apr 2026) presents an automated palynological analysis system for honey quality control. Here the connection is lexical rather than acronymic: miel denotes honey. The system integrates Nγc(λ)N_\gamma^c(\lambda)2 robust mechanical control, Nγc(λ)N_\gamma^c(\lambda)3-Net for salient object detection, and a DINOv2 ViT-S/14 backbone trained via Deep Metric Learning. It reports 95.8% classification recall, a 6× processing speedup versus manual expert analysis, and applicability to monofloral honeys, fraud detection, and regulatory certification (Staforelli-Vivanco et al., 17 Apr 2026). This is not designated “MIEL” as a system acronym in the supplied material, but it is directly associated with miel as a domain term.

Second, (Liu et al., 21 Mar 2026) introduces Code-MIE, a Code-style Multimodal Information Extraction framework that formalizes multimodal information extraction as unified code understanding and generation. It uses entity attributes, scene graphs, and a Python function-style input template, and reports state-of-the-art results including 61.03% and 60.49% on the English and Chinese datasets of MNγc(λ)N_\gamma^c(\lambda)4D, and 76.04%, 88.07%, and 73.94% on Twitter-15, Twitter-17, and MNRE, respectively (Liu et al., 21 Mar 2026). Although the acronym differs by one character, it is close enough that bibliographic confusion is plausible.

Taken together, these works show that “MIEL” does not denote a stable, field-independent object. In current arXiv usage it is best understood as a context-sensitive label that may refer to a cryogenic SiPM spectroscopy microscope (Raymond et al., 2024), a multimodal robotic exophora-resolution framework (Oyama et al., 22 Aug 2025), honey-related microscopy applications (Staforelli-Vivanco et al., 17 Apr 2026), or be confused with the neighboring acronym Code-MIE (Liu et al., 21 Mar 2026).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MIEL.