- The paper introduces SAGO, a framework that leverages virtual drones to reframe 3D Gaussian segmentation as an online Next-Best-View planning problem, eliminating lengthy pre-processing.
- It employs mask-shaped frustum filtering and Markov state transitions to achieve sub-second interactive segmentation while maintaining or surpassing offline accuracy.
- Experimental evaluations confirm that SAGO achieves over 50× speedup and high mIoU/mAcc on multiple datasets, enabling real-time 3D asset extraction and editing.
SAGO: Online Segment 3D Gaussians via Launching Virtual Drones
Introduction and Motivation
The acceleration of 3D Gaussian Splatting (3DGS) has revolutionized real-time scene representation and interactive manipulation. Despite remarkable rendering performance, a primary bottleneck in 3DGS-based scene interaction remains the need for time-intensive scene-specific setup phases prior to segmentation—typically requiring tens of seconds to minutes per scene for mask generation, mask lifting, and feature distillation. These constraints preclude immediate online applications and degrade the user experience for real-time editing, selection, and manipulation tasks.
SAGO (Segment Any Gaussians Online) proposes a paradigm shift by eliminating the preparatory setup stage. The framework demonstrates sub-second (≤1s) segmentation latency while maintaining or exceeding the segmentation quality of both offline and optimization-free online methods. The central innovation is the adoption of "virtual drones" that recast the segmentation problem as an online Next-Best-View (NBV) planning task, enabling interactive and efficient object segmentation from raw 3DGS scenes.

Figure 1: The standard pipeline for interactive 3DGS segmentation, highlighting the computational bottleneck introduced by heavy offline setup requirements in prior approaches.
Method: Virtual Drones for Online NBV Planning
Representation and State Initialization
Raw 3DGS scenes are represented as sets of explicit 3D Gaussians, each parameterized by spatial position, covariance (rotation and scale), spherical-harmonic color, and opacity. Segmentation is formalized as partitioning Gaussians into foreground A and background B, conditioned on a user-provided prompt (e.g., click, box, or text).
SAGO introduces virtual drones into a Markovian NBV framework. Starting from a prompt-provided initial viewpoint, a 2D segmentation mask is produced (e.g., via SAM2), and Mask-shaped Frustum Filtering (MFF) projects Gaussians into the mask, pruning likely background elements. The virtual drone’s pose in object-centric spherical coordinates defines its state, consisting of center ct, radius rt, yaw ϕt, pitch θt, and a segmentation memory bank.

Figure 2: SAGO pipeline—virtual drones explore the scene in both directions, using NBV planning and Mask-shaped Frustum Filtering (MFF) at each step to incrementally refine segmentation.

Figure 3: Mask-shaped Frustum Filtering (MFF) schematic; projected centers are compared to the 2D mask for efficient foreground/background assignment.
Online NBV Planning: Markov Exploration
At each segmentation state, two virtual drones are launched around the foreground object—one clockwise and one counterclockwise—systematically exploring candidate yaw and pitch angles to maximize information gain using an EE (Exploration-Evaluation) strategy. Each NBV is chosen to maximize the number of Gaussians filtered as background while enforcing a mask-mask consistency constraint (mIoU threshold) to stabilize Markov transitions against tracking drift.
Each new drone viewpoint produces a novel 2D mask, which is used for further MFF to refine the segmentation. This iterative refinement continues until full object coverage and segmentation convergence are achieved. The key innovation is that virtual drone states and their transitions are planned in the segmentation Markov chain without reliance on the camera distribution of the input 3DGS.

Figure 4: Illustration of NBV-based segmentation via Markov state transitions (clockwise direction shown), with state updates guided by maximum expected background exclusion and mask consistency.
Efficiency: Bidirectional Coverage and Minimal State Transitions
Theoretical analysis and empirical benchmarks show that bidirectional NBV planning typically requires only 2–3 Markov states per object, yielding sub-second processing per user interaction. This stands in sharp contrast to previous online and offline pipelines, which incur substantial per-scene preprocessing and/or per-interaction latencies.
Experimental Evaluation
SAGO demonstrates strong qualitative performance across diverse datasets (SPIn-NeRF, LERF-Mask, NVOS, 3D-OVS), extracting fine-grained segmentation with high granularity control via user prompts.

Figure 5: SAGO's iterative online segmentation on 'Figurines'—granularity modulated by initial mask, with successive NBV-driven refinements yielding clean object boundaries.

Figure 6: SAGO's generalization across 10 scenes and 5 datasets, all segmented online in 500ms per interaction.
Quantitatively, SAGO achieves:
- NVOS: mIoU=92.7%, mAcc=98.7% (best among all compared methods, including those requiring dedicated scene setup)
- SPIn-NeRF: mAcc=99.3% (highest), mIoU=92.5% (second highest)
- LERF-Mask: mIoU=91.0% (leading method; up to 5.4% improvement on difficult scenes over prior SOTA)
- 3D-OVS: mean mIoU=96.2% (highest, best in most scenes)
SAGO archives dramatic computational improvements: over 50× speedup versus prior setup-free methods with equivalent or superior accuracy.
Online 3D Asset Extraction and Scene Editing
SAGO enables direct 3D asset extraction and interactive editing (e.g., object relocation, duplication) with no overhead. Users can, for example, isolate a figurine in <1s and manipulate it in the scene (move, delete, combine) at run-time.

Figure 7: SAGO’s online 3D scene editing—user-specified objects extracted and manipulated instantly in the native 3DGS representation.
Fine-Grained Control and User Experience

Figure 8: GUI-based SAGO interface for 3D scene interaction, segmentation, real-time visualization, and user-controllable MFF parameters.
Granularity and segmentation precision are fully dictated by the 2D prompt granularity, with immediate visual feedback and live adjustment.

Figure 9: Demonstration of user-driven granularity control; the 2D prompt granularity directly determines the level of detail in the 3DGS segmentation.
Ablations: NBV Strategy and MFF Variants
Ablation studies show that omitting NBV planning (using only real views) yields catastrophic mIoU drops (e.g., >50% on complex scenes), underscoring the essential role of virtual drone-generated viewpoints for robust segmentation under occlusion and sparse input view distribution. Splat-based versus center-based MFF yields modest mIoU differences, but center-based MFF offers computational advantages.

Figure 10: Ablation on mask expansion coefficient in MFF; users can tune this to balance segmentation completeness and edge sharpness.
Robustness to Occlusions
The Markovian NBV process exploits virtual viewpoints to eliminate occlusions and ensure comprehensive object coverage, even in complex clustered-object or heavily occluded scenes.

Figure 11: Virtual drone NBV planning enables sequential occlusion elimination, ensuring all object regions are progressively exposed and correctly segmented.
Theoretical and Practical Implications
SAGO demonstrates that 3D segmentation in Gaussian Splatting can be reframed as a sequential NBV planning problem, mapped to a Markov process with explicit state transitions and memory. This method decouples the segmentation pipeline from camera/view distribution constraints and offline feature preparation, making interactive 3D manipulation viable for time-critical and user-driven applications.
Practical implications extend to real-time asset editing, robotic manipulation, AR/VR scene editing, and autonomous scene understanding. The framework naturally generalizes to new segmentation backbones, multi-object settings, and complex composite scenes. The NBV abstraction lays a foundation for future extensions incorporating active learning, multi-agent exploration, and more sophisticated segmentation-tracking pipelines.
Theoretically, SAGO motivates further study into scene representation-induced view planning, online scene understanding in neural and explicit 3D representations, and unified frameworks for segmentation, tracking, and asset extraction.
Limitations and Future Directions
- Boundary fidelity: Center-based MFF may introduce boundary artifacts (jagged borders), especially for voluminous or edge-overlapping Gaussians. Integrating online boundary refinement modules (e.g., GaussianTrimmer) is a clear path for improvement.
- Dependence on initial viewpoint: Subpar user-driven prompt views can impact downstream segmentation accuracy. Dynamic strategies for virtual drone initialization or auto-selection of optimal entry views warrant exploration.
- Multi-instance and semantic segmentation: While SAGO is readily extensible, explicit support for semantic and multi-instance segmentation, especially under open-vocabulary and occluded scenarios, remains an open challenge.
Conclusion
SAGO redefines the landscape of interactive 3DGS segmentation, introducing a fully online, setup-free approach based on virtual drone NBV planning. The method matches or surpasses segmentation accuracy of leading offline and online techniques, while achieving orders-of-magnitude efficiency gains. By coupling Markovian view planning with mask-driven background pruning, SAGO enables practical, fine-grained, real-time 3D asset segmentation and manipulation. The framework's modularity, efficiency, and empirical effectiveness position it as a foundation for advanced interactive 3DGS and generalized 3D scene understanding pipelines (2607.01628).