RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection
Abstract: Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, including attributes, actions, states, and contextual relations. Existing real-time OVD methods primarily emphasize vocabulary coverage and efficient region/query--text matching; under strict efficiency constraints, compact detectors may struggle to absorb rich instance semantics and scene context. We propose RT-DETR-World, a compact DETR-style detector that transfers the rich semantics conveyed by descriptions during training while retaining lightweight query--text matching at inference. We construct GroundingCapv2 with three levels of supervision: category names for standard OVD, object descriptions conveying instance-level semantics, and image descriptions conveying object relations and scene context. These descriptions serve only as training-time semantic supervision. To help the compact detector absorb these semantics, we propose Dual-Path Description Alignment (DDA), combining a deployment-consistent MiniLM pathway with a training-only LLM teacher. MiniLM provides query--category supervision and object-description alignment, while offline teacher features supervise matched queries and global visual representations at the object and image levels, respectively. All teacher features are precomputed, and the teacher-side modules are removed after training. We further propose Relation-Aware Negative Relaxation (RNR), which uses teacher-derived semantic similarities to relax related negatives while preserving exact positives. Experiments demonstrate competitive zero-shot accuracy and a favorable accuracy--efficiency trade-off. The code will be released.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces RT-DETR-World, a computer-vision system that can find objects in images even when it was not specifically trained on those object categories.
For example, a normal object detector might be trained to recognize chairs and spoons. An open-vocabulary detector could also recognize a wooden spoon or a wheelchair by using a text description, even if those exact categories were not part of its training labels.
The main challenge is to make the detector:
- Accurate, especially on unfamiliar objects
- Fast enough for real-time use
- Able to understand more than simple object names, such as color, material, position, actions, and relationships
The researchers’ main idea is:
Teach a small, fast detector using rich descriptions from LLMs during training, but remove the expensive LLM when the detector is actually being used.
2. What questions are the researchers asking?
The paper focuses on several important questions:
- Can detailed descriptions help a detector recognize unfamiliar objects? For instance, can learning about “a wooden chair leaning against a wall” help the system later recognize a wooden spoon?
- Can a small detector learn useful visual information from a much larger LLM? This is similar to a teacher helping a student learn. The large model provides guidance during training, but the smaller model works alone later.
- Can this be done without making the detector slower? The researchers want the final system to remain suitable for robots, self-driving vehicles, cameras, and other devices that need quick decisions.
- Can the detector avoid treating similar things as completely unrelated? For example, “a red car” and “a blue car” are not the same object, but they are still related. The system should not push their meanings too far apart during learning.
3. How did the researchers approach the problem?
A three-level training dataset
The researchers created a dataset called GroundingCapv2. It contains more than 1.1 million images or samples and over 8 million marked regions.
The data provides three kinds of information:
| Type of information | Example | What it teaches |
|---|---|---|
| Category name | “chair” | The basic object identity |
| Object description | “a wooden chair leaning against a wall” | Details about one object |
| Image description | “two chairs beside a table in a room” | Relationships and the overall scene |
Some descriptions were created or checked by vision-LLMs. These models were asked to describe objects and then verify whether their descriptions matched the image. Uncertain examples were checked by a larger model or replaced with a simpler category label.
A fast detector
RT-DETR-World is based on a detector called DETR. A detector is a computer program that looks at an image and predicts:
- What objects are present
- Where each object is, using a box around it
- Which text label best describes each object
The system uses:
- DINOv3 to extract useful visual features from an image
- A small DETR decoder to locate objects
- MiniLM, a compact LLM, to compare objects with text labels
At testing time, the detector only needs category names such as “dog,” “helmet,” or “wooden spoon.” The LLM is not used then.
Learning from a large teacher model
During training, the system uses a larger LLM as a teacher. The teacher reads detailed descriptions and converts them into numerical representations called embeddings.
An embedding is like a list of numbers that represents meaning. Similar ideas receive similar numerical patterns. For example, “car” and “vehicle” would probably have more similar embeddings than “car” and “banana.”
The student detector learns from the teacher in two ways:
- Object-level learning: It learns details about individual objects, such as their material, shape, action, or condition.
- Image-level learning: It learns about the whole scene, including relationships between objects and the surrounding context.
The teacher’s results are calculated beforehand, so the large model does not need to run during normal use.
Dual-Path Description Alignment
The researchers call this learning system Dual-Path Description Alignment, or DDA.
It has two paths:
- A MiniLM path, which matches the detector’s output with descriptions using the same small LLM that will be available during testing.
- A large-teacher path, which gives the detector richer and more detailed guidance during training.
The extra training parts are removed after learning. Therefore, they improve the detector without adding extra work when it is running.
Relation-Aware Negative Relaxation
The researchers also introduce Relation-Aware Negative Relaxation, or RNR.
Normally, contrastive learning teaches a system:
- “These two matching examples belong together.”
- “All other examples should be pushed apart.”
But this can be too strict. For example, “a bicycle” and “a motorcycle” are different, but they are both vehicles with wheels. RNR uses the teacher’s understanding to recognize that some “wrong” pairs are still related. It pushes them apart less strongly while still keeping the correct pair as the best match.
An everyday analogy is sorting books. A book about dogs and a book about cats should not be placed in exactly the same position, but they probably belong closer together than a book about astronomy.
4. What did the researchers find?
Stronger recognition of unfamiliar categories
The researchers tested the system on several benchmarks, including:
- LVIS, which contains many object categories, including rare ones
- ODinW, which tests whether a detector can work on very different types of images
- COCO-O, which tests how well the system handles changes in appearance, such as cartoons, sketches, paintings, and unusual weather
On LVIS, the largest RT-DETR-World model achieved:
- 36.6 standard AP
- 40.0 Fixed AP
- 32 frames per second
Here, AP is a score for detection accuracy. A higher score means the system more often identifies the correct object and places its box accurately. Frames per second, or FPS, measures speed.
The smaller versions were faster:
- RT-DETR-World-S: 63 FPS
- RT-DETR-World-S+: 58 FPS
- RT-DETR-World-B: 32 FPS
This shows a useful balance between speed and accuracy.
Better transfer to new environments
On the ODinW benchmark, the largest model reached:
- 43.8 AP on the 13-dataset version
- 20.2 AP on the 35-dataset version
These results were better than several competing real-time detectors. This suggests that the system can transfer what it learned to new environments, object types, and visual styles.
Better robustness to changes in appearance
On COCO-O, the largest model achieved:
- 47.2 AP on ordinary COCO images
- 46.2 AP on shifted images
The shifted images included sketches, cartoons, paintings, tattoos, and handmade images. The detector therefore remained effective even when objects looked very different from the training examples.
Detailed descriptions were especially helpful
The experiments also tested which parts of the method mattered most. The results showed that:
- Category names alone were useful but limited.
- Object descriptions improved understanding of rare categories.
- Image descriptions added useful scene and relationship information.
- Combining object-level and image-level learning worked best.
- RNR performed better than several other ways of handling similar examples.
For one version of the model, adding the full combination of methods raised the score from 27.9 AP to 35.2 AP. This is a large improvement.
5. Why are these results important?
Many real-world systems need to recognize objects they have not seen exactly during training. A robot might encounter a new tool, a self-driving car might see an unusual vehicle, or a security camera might need to identify objects in unfamiliar surroundings.
The paper shows that a detector does not have to rely only on short labels such as “chair” or “car.” It can learn from richer descriptions that explain:
- What an object looks like
- What it is made of
- How many objects there are
- What an object is doing
- Where objects are located
- How objects relate to one another
The most important practical point is that the expensive LLM is used only while teaching the detector. Once training is complete, it is removed. This means the final detector keeps much of the learning benefit while remaining relatively fast.
Conclusion and possible impact
RT-DETR-World is a method for building fast object detectors that understand more detailed visual meaning. It combines ordinary category labels with descriptions of objects and entire scenes. A LLM acts as a teacher during training, while a smaller detector handles the actual image processing later.
The results suggest that this approach can improve:
- Recognition of rare and unfamiliar objects
- Performance across different datasets
- Reliability when images look unusual
- The balance between accuracy and speed
In the future, systems based on this idea could help robots, smart cameras, augmented-reality devices, and autonomous vehicles understand new situations more effectively. However, the method still has limitations: larger models are slower, and automatically generated descriptions can sometimes contain mistakes. Careful checking of the training data remains important.
Glossary
- Adjudication: Formal review and resolution of ambiguous or conflicting annotations. “routing ambiguous cases to Qwen3-VL-32B for multimodal adjudication”
- Alignment space: A shared vector space in which representations from different modalities are made comparable. “project the text and query representations into a shared alignment space”
- AP (Average Precision): A detection metric summarizing precision across recall levels. “We report standard AP and Fixed AP”
- Backbone: The primary feature-extraction network in a deep-learning model. “RT-DETR-World-S, -S+, and -B use frozen DINOv3”
- Bidirectional contrastive learning: Contrastive learning that evaluates both visual-to-text and text-to-visual matching directions. “α = 0 recovers standard bidirectional contrastive learning”
- Bounding-box regression: Prediction of the coordinates and dimensions of an object’s bounding box. “The DETR decoder refines the selected queries and regresses bounding boxes.”
- Caption-contrastive tuning: Training an encoder to associate image captions with corresponding visual representations while separating mismatched captions. “the caption-contrastively tuned LLM text encoder from LLM2CLIP”
- Category fallback: Use of a category label when a more detailed annotation is unreliable. “unreliable descriptions fall back to category names”
- Cross-modal alignment: Matching representations from different modalities, such as images and text. “SoftCLIP-style soft alignment”
- Cosine similarity: A similarity measure based on the angle between two vectors. “where sim(·, ·) denotes cosine similarity”
- DETR decoder: The transformer decoder component of Detection Transformer architectures that produces object predictions from queries. “a DETR decoder that produces object queries”
- Distribution shift: A change in the data distribution between training and evaluation. “distribution-shift robustness on COCO-O”
- Embedding: A numerical vector representation of an item such as text, an image, or an object query. “the corresponding embeddings in the dL-dimensional teacher space”
- Feature level: One resolution or scale of the hierarchical feature representations used by a detector. “valid spatial features from each projected feature level are pooled”
- Fixed AP: An average-precision metric using a fixed per-category detection budget instead of a fixed number of detections per image. “Fixed AP replaces this image-level limit with a fixed per-category detection budget.”
- Global visual representation: A vector summarizing visual information from an entire image. “global visual representations at the object and image levels”
- Grounding: Linking textual phrases or descriptions to specific regions in an image. “GLIP and Grounding DINO unify object detection and phrase grounding”
- Hungarian matching: An optimal assignment algorithm used to match predicted object queries with annotated objects. “the Hungarian-matched query assigned by Hungarian matching to region i”
- Inference: The process of using a trained model to generate predictions. “The LLM itself is absent during detector training and inference”
- In-distribution: Belonging to the same data distribution as the training data. “COCO (Lin et al., 2014) as the in-distribution reference”
- Language-guided query selection: Selection of detector queries using similarities between visual features and text representations. “Vision–text similarities guide Language-Guided Query Selection”
- Long-tailed distribution: A data distribution in which a small number of classes are frequent while many classes are rare. “LVIS contains 1,203 categories with a long-tailed distribution”
- Mask-aware average pooling: Averaging feature values only over spatial positions designated as valid by a mask. “MAP performs mask-aware average pooling over valid spatial positions”
- Multi-scale projector: A module that transforms visual features from multiple resolutions into representations suitable for detection. “a lightweight multi-scale projector”
- Negative relaxation: Reduction of the penalty assigned to semantically related but unmatched examples during contrastive learning. “Relation-Aware Negative Relaxation (RNR)”
- Offline teacher: A pretrained model whose outputs are computed in advance and used as training targets for another model. “a training-only LLM teacher”
- Open-vocabulary detection: Object detection in which textual queries can specify categories not observed during detector training. “Open-vocabulary detection (OVD) recognizes categories unseen during training”
- Open-set object detection: Detection of objects beyond the closed set of categories defined during training. “Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection”
- Parameter freezing: Preventing selected model parameters from being updated during training. “a frozen DINOv3 visual encoder”
- Phrase grounding: Localizing a natural-language phrase to its corresponding image region. “unify object detection and phrase grounding”
- Query–text matching: Comparing detector object-query representations with textual category or description representations. “real-time OVD methods favor lightweight region/query–text matching”
- Region–phrase annotation: An annotation linking a textual phrase to a specific image region. “Region–Phrase Annotations (GQA / Flickr30K Entities / LLaVA-Cap)”
- Representation capacity: The ability of a model to encode and retain information in its learned representations. “the limited representation capacity of compact detectors”
- Semantic distillation: Transfer of semantic information from a larger or more expressive model to a smaller model. “A hierarchical semantic distillation framework for open-vocabulary object detection”
- Semantic overlap: The degree to which two descriptions share meaning despite not being identical. “partial semantic overlap”
- Semantic supervision: Training signals that convey conceptual or linguistic information about visual content. “These descriptions serve only as training-time semantic supervision.”
- Soft alignment: Alignment that uses graded similarity targets rather than only exact-match targets. “SoftCLIP-style soft alignment”
- Temperature: A scalar that controls the sharpness of probability distributions derived from similarity logits. “τ > 0 is the temperature”
- Text encoder: A neural network that converts text into numerical representations. “MiniLM encodes category names and category-oriented textual prompts.”
- Token features: Vector representations produced for individual tokens by a LLM. “we pool the valid non-special-token features of dbi”
- Training-time supervision: Information used to train a model but omitted from the deployed inference system. “object and image descriptions convey instance- and scene-level semantics during training”
- Transfer learning: Reusing knowledge learned from one set of data or tasks for another. “zero-shot cross-domain transfer on ODinW”
- Zero-shot generalization: Recognition of categories or tasks for which no direct training examples were provided. “zero-shot generalization may also benefit from reusable visual–semantic factors”
- Zero-shot object detection: Detecting object categories that were not present in the detector’s training annotations. “Zero-shot object detection”