RT-DETR-World: Real-Time Detection Meets Rich Language Semantics

This lightning talk explores RT-DETR-World, a compact open-vocabulary object detector that learns rich visual semantics from detailed language descriptions during training but uses only lightweight category matching at inference. By combining deployment-ready text encoders with training-only language model teachers, the system achieves strong zero-shot detection across diverse domains while maintaining real-time speed—demonstrating that rich semantic knowledge can be distilled into efficient detectors without requiring heavy models during deployment.
Script
Real-time object detectors are fast but semantically shallow, while large grounding models understand rich language but are too slow for robotics or edge devices. RT-DETR-World solves this by learning from detailed descriptions during training, then discarding the heavy language models entirely at inference.
The key insight is that object descriptions contain reusable semantic cues: materials like wood or metal, parts like handles or wheels, actions like running or sitting, and spatial relations within a scene. Standard detectors ignore these cues and match only category names.
RT-DETR-World uses Dual-Path Description Alignment. A lightweight MiniLM encoder handles category matching and stays active during deployment. A frozen language model teacher provides richer semantic targets during training, then vanishes—transferring its knowledge into the detector's learned representations without adding inference cost.
Not all negative pairs should be pushed equally far apart. Relation-Aware Negative Relaxation uses the language model's semantic similarity to relax repulsion between related examples—like two chairs made of different materials—while still separating unrelated objects like chairs and bicycles.
On the 1,203-category LVIS benchmark, RT-DETR-World-B improves over the previous best real-time detector by 4.1 AP points while running at 32 frames per second. Large grounding models remain more accurate, but they sacrifice the speed needed for real-time robotics and autonomous systems.
RT-DETR-World shows that rich language supervision can be distilled into compact detectors, transferring semantic depth without runtime cost. To explore how language models shape real-time vision and create your own video summaries of cutting-edge research, visit emergentmind.com.