---
title: 'BRIS: Advanced Robotic Intubation Systems'
url: https://www.emergentmind.com/topics/robotic-intubation-system-bris
type: topic
---

# BRIS: Advanced Robotic Intubation Systems

Searching arXiv for the cited BRIS-related papers and closely related robotic intubation work.
The **Robotic Intubation System (BRIS)** denotes a family of robotic platforms for airway management that couple endoscopic perception, anatomical guidance, and robot control to support or automate tracheal or nasotracheal intubation. Across the cited literature, BRIS is associated with at least three tightly related technical trajectories: a **Bot-assisted Robotic Intubation System** emphasizing segmentation-driven surgical navigation [2305.11686], a platform for **autonomous nasotracheal intubation** using force-instrumented demonstrations and a Transformer-based recurrent policy [2508.01808], and the **Bab_Sak Robotic Intubation System**, a compact human-in-the-loop system for fiberoptic-guided endotracheal intubation with learning-enabled teleoperation and monocular depth-based placement guidance [2512.21983]. A related perception module for landmark localization using Deformable DeTR with Semantic-Aligned-Matching was also developed for robotic nasal airway intubation and integrated into the BRIS pipeline [2308.02845]. Taken together, these works define BRIS as a research program in robotic airway intervention centered on safe navigation, anatomy-aware perception, and closed-loop control.

## 1. Terminology, scope, and procedural targets

In the cited work, BRIS is used in the context of **robot-assisted tracheal intubation**, **autonomous nasotracheal intubation (NTI)**, and **fiberoptic-guided endotracheal intubation**. The procedural target is the establishment of an artificial airway while reducing failure modes associated with manual airway instrumentation, including difficult anatomical access, high contact forces, and inaccurate depth placement [2508.01808; 2512.21983].

The 2023 segmentation paper describes the BRIS vision module as part of a **Bot-assisted Robotic Intubation System** in which a segmentation network is integrated as a ROS node and used for path planning, collision avoidance, and tool guidance [2305.11686]. The 2025 autonomous NTI paper describes a BRIS platform built around a **KUKA iiwa 7-DOF lightweight manipulator**, an endoscopic view, and a prosthesis embedded with force sensors, with autonomy restricted to clinically relevant insertion DoFs [2508.01808]. The later 2025 paper presents the **Bab_Sak Robotic Intubation System (BRIS)** as a compact **human-in-the-loop platform** integrating a four-way steerable fiberoptic bronchoscope, an independent endotracheal-tube advancement mechanism, and a camera-augmented mouthpiece compatible with standard clinical workflows [2512.21983].

These uses are not identical. A plausible implication is that “BRIS” functions less as a single frozen hardware configuration than as a label for a sequence of related robotic intubation systems that share a common emphasis on anatomy-aware visual perception and closed-loop assistance or autonomy.

## 2. Hardware architectures and actuation strategies

The BRIS-related systems differ substantially in hardware embodiment. In the autonomous NTI configuration, the robot arm is a **KUKA iiwa 7-DOF lightweight manipulator**, with each joint instrumented with position and joint-torque sensing; the end-effector flange carries a **6-DoF pose estimate** and an integrated **3D force/torque sensor** [2508.01808]. Vision consists of **Camera 1**, a clinical endoscope providing the “eye-in-hand” view of the extraluminal portion of the tube at **30 fps**, and **Camera 2**, an external viewpoint used for recording only [2508.01808]. Motion is restricted during both tele-operation and autonomy to **three DoFs: planar $(x,z)$ translation and rotation about the $y$-axis (tube pitch)**, matching the clinically relevant insertion DoFs [2508.01808].

The same NTI system incorporates a **multi-material 3D-printed nasal/oropharyngeal model** with embedded force sensing: **one 3D sensor at the nostril entrance $(F_x,F_y,F_z)$ plus two 1D sensors in the sphenoid $(F_1)$ and pharyngeal $(F_2)$ regions** [2508.01808]. End-effector force/torque is read at **500 Hz**, prosthesis force sensors are logged at **200 Hz**, and segmentation outputs are delivered to the RACCT model at **5 Hz** [2508.01808].

The Bab_Sak BRIS adopts a distinct architecture centered on a **four-way steerable fiber-optic bronchoscope (FOB)** [2512.21983]. The bronchoscope has **total length 1500 mm**, **outer diameter 6 mm $(\pm 0.1\ \mathrm{mm})$**, and **inner lumen diameter 4 mm $(\pm 0.1\ \mathrm{mm})$**, with **coaxial compatibility with standard adult ET tubes (7.5–8.5 mm)** [2512.21983]. Sections 1 and 3–4 use **Pebax 7233**, while the articulation zone uses **Pebax 3533**; steering is implemented with **4 × 0.3 mm stainless steel** tendons in a differential-tendon configuration driven by **4 × Dynamixel XM430-W210-T servos**, enabling continuous orientation control of the distal tip over the unit sphere $S^2$ [2512.21983]. Shape sensing is provided by a **Fiber-Optic Shape Sensing (FOSS) module** reconstructing **$N=48$ backbone points** in real time [2512.21983].

The Bab_Sak end-effector includes a **dual lead-screw architecture** with separate platforms for FOB insertion/retraction and ET-tube linear translation, each driven by **NEMA 17 stepper** motors with microstepping for sub-millimeter resolution [2512.21983]. A **camera-augmented mouthpiece** based on **Macintosh blade curvature** integrates a **2.8 mm CMOS camera** with **120° FOV** [2512.21983]. The hardware is mounted to a **UR10-series robotic arm**, and the software stack includes **Jetson AGX** and **ROS Master (Linux)** linked to robotic arm, Dynamixel servos, and stepper drivers [2512.21983].

The 2023 BRIS vision work focuses less on manipulators and more on the imaging subsystem. Its vision module assumes a **1080p endoscopic sensor at 30 fps** with **latency <10 ms**, and inference on **NVIDIA Jetson AGX Xavier (GPU)** with average segmentation time **~30 ms/frame** [2305.11686]. This yields an end-to-end vision-to-actuation latency target of **$\leq 50\ \mathrm{ms}$** [2305.11686].

## 3. Perception modules: segmentation, landmarks, and depth

A defining feature of BRIS research is the use of specialized perception modules tuned to airway anatomy.

For organ-level understanding, the 2023 work develops **domain adaptive Sim-to-Real segmentation of oropharyngeal organs**—specifically the **uvula, epiglottis, and glottis**—for robot-assisted intubation [2305.11686]. To address limited real endoscopic data, it constructs a photo-realistic virtual oropharyngeal phantom in the **Simulation Open Framework Architecture (SOFA)** framework. Organ meshes are generated in Blender from **CT-derived surfaces**, imported into SOFA as **TetrahedronFEMObject**, and modeled using **linear tetrahedral elements (P1 FEM)** [2305.11686]. The material model is **Neo-Hookean hyperelasticity** with **Young’s modulus $E=\{5\ \mathrm{kPa},10\ \mathrm{kPa},15\ \mathrm{kPa}\}$** for uvula, epiglottis, and glottis, **Poisson ratio $\nu=0.45$**, and **Rayleigh damping $\alpha=0.01,\ \beta=0.05$** [2305.11686]. Virtual endoscopy uses **focal length 18 mm**, **FOV $=70^\circ$**, and **resolution $=640\times 480$ px**, with **Phong shading** and four point lights [2305.11686].

The segmentation network is **DeepLab v3+ with ResNet-101 encoder**, **output stride = 16**, **ASPP** rates $\{6,12,18\}$, and a **4-channel softmax** for background plus three organs [2305.11686]. Domain adaptation combines an **IoU-Ranking Blend (IRB)** strategy with **CycleGAN-style** style transfer using generators $G_{S\to R},G_{R\to S}$ and discriminators $D_S,D_R$ [2305.11686]. The per-class IoU is defined as
\[
\mathrm{IoU}_c = \frac{\lvert P^c \cap G^c\rvert}{\lvert P^c \cup G^c\rvert},
\]
and the IRB allocation rule is
\[
B_c = B \times \frac{r_c}{\sum_{j=1}^C r_j},
\]
with $r\in\{5,3,2\}$ for $\{\text{uvula, epiglottis, glottis}\}$ [2305.11686]. A ranking loss
\[
\mathcal{L}_\mathrm{rank} =\sum_{i<j}\max\bigl(0,\;\mathrm{IoU}_j-\mathrm{IoU}_i+\delta\bigr),
\quad \delta=0.02,
\]
is used to encourage the network to close the performance gap between classes [2305.11686].

For landmark-level detection, the 2023 nasal intubation work proposes a **Deformable Detection Transformer** augmented with **Semantic-Aligned-Matching (SAM)** to detect **nostrils** and **glottis** [2308.02845]. The backbone is **ResNet-50**, feature maps are extracted at strides **$\{8,16,32,64\}$**, the decoder uses **$N=100$ queries** of dimension **$d=256$**, and cross-attention is implemented with **multi-scale deformable attention** using **$K=4$ sampling points per level** [2308.02845]. The SAM module predicts a reference box
\[
b_i = \sigma(W_b\cdot q_i^{pos} + b_b)
\]
and fuses original and salient ROI-derived query embeddings using a learnable gate
\[
\alpha_i = \sigma(W_\alpha\cdot q_i + b_\alpha),
\]
yielding
\[
q_i' = \alpha_i \odot q_i + (1-\alpha_i)\odot \mathrm{mean}(Q_i^{new}) .
\]
The loss follows standard DETR-style Hungarian matching with classification and box terms [2308.02845].

For depth-aware airway placement, the Bab_Sak BRIS integrates **zero-shot monocular depth** using the **“Depth Anything” foundation model (Yang et al. 2024)** to produce a dense relative-scale depth map $D(u,v)$ from live **30 fps** endoscopic video [2512.21983]. Estimated tip-to-carina distance $\hat d$ is converted into **Zone I**, **Zone II**, and **Zone III** by a threshold-based classifier:
\[
\mathrm{Zone}=
\begin{cases}
\mathrm{I}, & \hat d > d_1,\\
\mathrm{II}, & d_2 < \hat d \le d_1,\\
\mathrm{III}, & \hat d \le d_2,
\end{cases}
\]
with example thresholds **$d_1\approx 4\ \mathrm{cm}$** and **$d_2\approx 1\ \mathrm{cm}$** [2512.21983]. Guidance is further supported by a passive visual servoing overlay derived from the misalignment vector
\[
\vec v = C_{lumen} - C_{cam},
\]
where $C_{lumen}$ is the centroid of maximal depth pixels and $C_{cam}$ is the optical center [2512.21983].

In the autonomous NTI paper, the perception branch is specialized to the tube rather than directly to anatomical organs. The input branch uses **endoscope image $\to$ DLO-segmentation $\to$ morphological post-processing $\to$ skeleton & curvature features**, combined with **6-DoF end-effector pose** and **3D robot-flange force** [2508.01808]. This suggests that BRIS perception has evolved toward task-specific representations: organ masks for navigation, landmarks for staged insertion, and tube skeleton/curvature for contact-aware policy learning.

## 4. Learning-enabled control and autonomy

The BRIS literature spans several control paradigms, from segmentation-assisted motion planning to learning-enabled teleoperation and full imitation-learned autonomy.

In the 2023 segmentation-integrated BRIS pipeline, segmentation masks are converted to **3D point clouds via hand-eye calibration and depth estimation** [2305.11686]. The **3D centroids of glottis pixels** define a spatial goal for tip articulation; **voxelized organ volumes** from segmentation plus stereo depth feed into **MoveIt!** for motion planning under anatomical constraints; and **uvula** and **epiglottis** masks inform active bending of the **EndoBot stylet** via a **closed-loop PID** controlling a curvature motor [2305.11686]. The software interface is organized as **ROS topic `/camera/image_raw` → segmentation node → `/segmentation/mask`** [2305.11686].

The autonomous NTI paper introduces **Recurrent Action-Confidence Chunking with Transformer (RACCT)** as the principal policy model [2508.01808]. Inputs are embedded as
\[
z_0 = \mathrm{Linear}_{img}(x_{img}) + \mathrm{Linear}_{pose}(x_{pose}) + \mathrm{Linear}_{force}(x_{force}) + \text{style-token } Z,
\]
then processed by a standard Transformer encoder using self-attention
\[
\mathrm{Attention}(Q,K,V) = \mathrm{softmax}\bigl((QK^\top)/\sqrt{d_k}\bigr)V .
\]
The decoder jointly regresses a chunk of future actions
\[
A_t[i] \in \mathbb{R}^3 \quad (\Delta x,\Delta z,\Delta \theta_y)
\]
and associated confidences
\[
C_t[i] \in (0,1),
\]
which are combined into an executed command
\[
p_t = \frac{\sum_{i=1}^k e^{-m\cdot i} C_t[i]\cdot A_t[i]}{\sum_{i=1}^k e^{-m\cdot i} C_t[i]},
\quad m=0.95 .
\]
The loss is
\[
\mathcal{L} = \sum_{i=1}^k \frac{c_i\cdot|\hat A_i-\hat y_i|}{k\cdot(\epsilon + 1 - C_t[i])}
\;-\;
\lambda\cdot \log\Bigl[\frac{1}{k}\sum_{i=1}^k C_t[i]\Bigr],
\]
with **$\epsilon=0.2$** and **$\lambda=0.1$** [2508.01808]. The model is expressly intended to handle **complex tube-tissue interactions and partial visual observations** [2508.01808].

The Bab_Sak BRIS uses a different learning-enabled controller: a learned forward dynamics model embedded in **MPC** for stable teleoperation under tendon nonlinearities and airway contact [2512.21983]. State, shape, and input are defined as
\[
x_t \in \mathbb{R}^6,\qquad
S_t=\{p_t^{(1)},\ldots,p_t^{(N)}\},\qquad
u_t=[\tau_t^{(1)},\tau_t^{(2)},\tau_t^{(3)},\tau_t^{(4)},\tau_t^{(5)}]^\top .
\]
A **Temporal Convolutional Network (TCN)** encoder maps a short window **$S_{t-L:t}\to z_t\in\mathbb{R}^{16}$** with **$L=6$** (approximately **120 ms**) [2512.21983]. A residual neural network predicts the next tip state:
\[
\hat x_{t+1} = f_\theta(x_t,z_t,u_t).
\]
Training uses the loss
\[
\mathcal{L}(\theta,\phi)=\|\hat x_{t+1}-x_{t+1}\|^2 + \beta \|z_{t+1}-z_t\|^2,
\quad \beta=0.1,
\]
optimized with **Adam**, **lr $=1\mathrm{e}{-3}$**, **batch 256**, over **120 epochs** on approximately **$6.2\times 10^4$** state-transition samples [2512.21983].

The MPC solves a finite-horizon OCP at each time step with **horizon $M=10$**, **update rate = 50 Hz**, and **solve time $\lesssim 12$ ms/step** [2512.21983]. The joystick mapping translates axes to Cartesian velocity and buttons to **ET tube retract**, **go-to-waypoint**, and **emergency-stop** [2512.21983]. This is a markedly different control philosophy from RACCT: MPC is used for **stable and intuitive teleoperation**, whereas RACCT is used for **autonomous** insertion from demonstrations.

## 5. Safety mechanisms and quantitative performance

Safety is explicit in all BRIS variants, but implemented through different mechanisms.

In the vision-guided BRIS pipeline, **end-to-end vision-to-actuation latency $\leq 50\ \mathrm{ms}$** is required to ensure real-time response [2305.11686]. A **watchdog** monitors segmentation confidence; if **mean softmax $<0.6$**, the robot halts and requests clinician override [2305.11686]. All control commands pass through a **2-layer safety filter** checking joint limits and maximum curvature rates [2305.11686].

In the autonomous NTI system, safety begins at the data-collection stage. Demonstrations are filtered using three criteria: **intubation time $t<20\ \mathrm{s}$**, **peak prosthesis force $F_{peak}<5\ \mathrm{N}$**, and impulse
\[
I=\int \max(F-F_{thr},0)\,dt,\qquad F_{thr}=1.5\ \mathrm{N},
\]
with **$\log_{10}(I)<1$**, i.e. **$I<10\ \mathrm{N\cdot s}$** [2508.01808]. Only episodes in which all three metrics fall below **70%** of these thresholds are retained for imitation learning [2508.01808]. Real-time control uses only the robot’s flange **3D force/torque**; the prosthesis sensors are used **a posteriori** for filtering, annotation, and evaluation [2508.01808].

The Bab_Sak BRIS incorporates safety into the user interface and guidance logic. Features include a **hardware E-stop** accessible to Operator B, **software velocity limits**, and **depth-triggered automatic withdrawal if Zone III is entered accidentally** [2512.21983]. The success criterion is explicit: the **ET tube tip stably in Zone II (mid-trachea)** [2512.21983].

The principal quantitative outcomes reported across the cited BRIS works are summarized below.

| System/paper | Reported task | Key outcomes |
|---|---|---|
| Vision BRIS [2305.11686] | Real-phantom segmentation | Overall Dice: **0.64** baseline vs **0.73** domain-adaptive; overall Pixel Acc: **0.88** vs **0.92** |
| Autonomous NTI BRIS [2508.01808] | Autonomous nasotracheal intubation | **100%** success rate; time **9.13 s $(\pm 0.27\ \mathrm{s})$**; peak **$F_x$** **1.56 N $(\pm 0.09)$** vs doctor **2.75 N $(\pm 0.13)$** |
| Bab_Sak BRIS [2512.21983] | Fiberoptic-guided intubation on mannequins | **48** trials; **100 %** success in standard and constrained scenarios; depth MAE **$2.4 \pm 1.1\ \mathrm{mm}$** overall |

The segmentation study reports per-class gains on the real phantom: **uvula Dice 0.65 to 0.72**, **epiglottis 0.58 to 0.68**, **glottis 0.70 to 0.78**, and **overall 0.64 to 0.73** under the domain-adaptive pipeline [2305.11686]. The conclusion states **$\Delta\mathrm{Dice}\approx +9\%$** [2305.11686].

The autonomous NTI study reports that **RACCT outperforms the ACT model in all aspects** and achieves a **66% reduction in average peak insertion force compared to manual operations while maintaining equivalent success rates** [2508.01808]. More specifically, mean outcomes over successful trials include **RACCT 100% success**, **time 9.13 s $(\pm 0.27\ \mathrm{s})$**, **doctor 5.03 s $(\pm 0.12\ \mathrm{s})$**, and peak force in **$F_2$** of **1.81 N** for RACCT versus **2.66 N** for doctor [2508.01808]. The paper additionally notes that impulse reductions show **35–60% lower continuous loading in RACCT vs. human** and that the reductions across **$N=20$ trials per condition** exceed inter-trial variability by **$>3\times$**, though **no formal p-values were reported** [2508.01808].

The Bab_Sak BRIS reports **48 total** mannequin trials, split into **24 standard** and **24 constrained** scenarios [2512.21983]. It achieved **100 % success rate in both scenarios**, with **no endobronchial or esophageal misplacements** [2512.21983]. Depth-estimation MAE was **$2.4 \pm 1.1\ \mathrm{mm}$ overall**, **$2.0 \pm 0.9\ \mathrm{mm}$ standard**, and **$2.8 \pm 1.2\ \mathrm{mm}$ constrained**, and **98 % of trials ended within $\pm 20\ \mathrm{mm}$ of target** [2512.21983]. Ablations showed **46 % reduction in joystick command variance during distal navigation** with learned MPC versus naïve teleoperation without shape sensing; turning visual guidance off in a **16-trial subset** increased distal wall contacts by **52 %** and corrective withdrawals by **35 %** [2512.21983]. Tracking RMSE was **< 1.5 mm** in free-space trajectories and **< 3 mm** under soft contact [2512.21983].

## 6. Datasets, training regimes, and experimental methodology

The BRIS literature relies on heterogeneous data sources, reflecting the scarcity of real labeled airway datasets and the safety constraints of clinical experimentation.

For segmentation, the 2023 paper explicitly motivates virtual data generation because real datasets of oropharyngeal organs are limited due to patient privacy issues [2305.11686]. Its synthetic source domain is generated in SOFA from **medical-grade CAD scans** and **CT-derived surfaces** [2305.11686]. The segmentation model is trained with **SGD**, **momentum = 0.9**, **weight decay = $1\times 10^{-4}$**, **initial learning rate 0.01**, **“poly” policy** $\mathrm{lr}_t=\mathrm{lr}_0(1-t/T)^{0.9}$, **batch size 8 (4 synthetic + 4 blended)**, and **100 epochs** [2305.11686]. Data augmentation includes **random horizontal/vertical flips**, **scaling [0.8–1.2]**, and **color jitter $\pm 20\%$** [2305.11686].

For landmark detection, two datasets are used [2308.02845]. The **nostril dataset** is derived from the **BioID face keypoints dataset** with **1 521 images**, resized to **640×480**, split into **1 021 train / 185 val / 315 test**, and automatically annotated by expanding the nostril keypoint to a **$40\times 40$ px** box [2308.02845]. The **glottis dataset** comes from the **BAGLS segmentation dataset**, with **881 nasal endoscopy frames**, split into **377 train / 88 val / 416 test** after converting masks to axis-aligned bounding boxes [2308.02845]. All models are trained for **24 epochs** with **Adam**, **$\mathrm{lr}_{backbone}=1\mathrm{e}{-5}$**, **$\mathrm{lr}_{head}=1\mathrm{e}{-4}$**, **batch 8**, in **MMDetection** [2308.02845].

The landmark detector achieves the following benchmarked performance. On **glottis detection**, **Ours (SAM Deformable)** reports **mAP [0.5:0.95] = 0.282**, **mAP@0.5 = 0.661**, **mAP@0.75 = 0.270**, **Precision@0.5 = 0.68**, and **Recall@0.5 = 0.59** [2308.02845]. On **nostril detection**, it reports **mAP [0.5:0.95] = 0.325**, **mAP@0.5 = 0.865**, **mAP@0.75 = 0.142**, **Precision@0.5 = 0.88**, and **Recall@0.5 = 0.79** [2308.02845]. GPU inference runs at **~12 fps (80 ms/image)** on an **NVIDIA RTX 2080Ti**, with a projected path to **~40 ms/image** via reduced image resolution, TensorRT fusion, and fewer queries/layers [2308.02845].

For autonomous NTI, the dataset consists of **50 human-teleoperation episodes** from **two experienced operators** on the prosthesis, filtered using the safety criteria above [2508.01808]. Per-timestep recordings include image, segmented skeleton and curvature, 6-DoF pose, 3-D flange force, and prosthesis forces for annotation [2508.01808]. Images are cropped/resized to **256×256** and normalized, with a random **80/10/10** train/val/test split [2508.01808]. Training uses **Adam**, **learning rate = $1\mathrm{e}{-5}$**, **batch size = 8**, **chunk length $k=80$**, and **20 k gradient steps**, requiring approximately **1 hour on RTX A6000** [2508.01808].

For the Bab_Sak dynamics model, training data comprise approximately **$6.2\times 10^4$** state-transition samples from teleoperated maneuvers [2512.21983]. The optimizer is **Adam**, with **lr = $1\mathrm{e}{-3}$**, **batch 256**, and **120 epochs** [2512.21983]. The training and validation loss trajectories are described as exhibiting **smooth decline to plateau** [2512.21983].

## 7. Limitations, misconceptions, and research directions

A common misconception would be to treat BRIS as a single monolithic system with one settled architecture. The cited literature does not support that simplification. Instead, BRIS appears as a set of related robotic intubation systems spanning segmentation-guided robotic navigation [2305.11686], autonomous nasotracheal insertion [2508.01808], and human-in-the-loop fiberoptic-guided endotracheal intubation with objective depth awareness [2512.21983]. This suggests an evolving platform concept rather than a single immutable device.

Another potential misconception is that robotic intubation research is exclusively a navigation problem. The later BRIS papers explicitly extend beyond navigation to **contact-force reduction** [2508.01808] and to **objective verification of tube depth relative to the carina** [2512.21983]. The 2025 Bab_Sak paper further states that existing robotic and teleoperated systems primarily focus on airway navigation and do not provide integrated control of ET-tube advancement or objective verification of tube depth relative to the carina [2512.21983].

Observed failure modes are also explicitly documented. In the segmentation-driven BRIS pipeline, **low-lighting frames** cause false negatives on the transparent epiglottis, **anatomical occlusions** such as tongue protrusion can degrade Dice by **up to 15%**, and **excessive blood or secretions** introduce specular highlights that confuse the network [2305.11686]. Proposed enhancements include a **photometric pre-processing module with specular removal and adaptive histogram equalization**, **multi-view 3D reconstruction**, **tactile/force feedback integration**, and **online continual learning** driven by uncertainty-based annotation in the OR [2305.11686].

The autonomous NTI system is limited by validation on a **rigid phantom**; the paper states that **animal and cadaver studies are needed** to assess tissue compliance variations [2508.01808]. It also notes that no real-time prosthesis-sensor feedback is used in closed loop and suggests future incorporation of **active force-control for on-the-fly safety stopping** [2508.01808]. Generalization to patient-specific anatomies and other tubular insertion tasks is also identified as requiring **domain-adaptive vision modules and expanded demo datasets** [2508.01808].

The Bab_Sak BRIS, while demonstrating reliable navigation and controlled tube placement on **high-fidelity airway mannequins**, remains a mannequin-validated platform [2512.21983]. Its core contribution is to make fiberoptic-guided intubation more compatible with standard clinical workflow while adding real-time anatomy-aware guidance and depth awareness, but the cited text does not report human or cadaver deployment [2512.21983].

Across these works, the central research direction is consistent: combine geometry-aware or anatomy-aware perception, explicit safety criteria, and learning-enabled control to reduce operator burden and tissue loading while improving consistency of airway access and final tube placement.

Source: https://www.emergentmind.com/topics/robotic-intubation-system-bris