Papers
Topics
Authors
Recent
Search
2000 character limit reached

RT-DETR-World: Transferring Rich LLM Semantics to Real-Time Open-Vocabulary Detection

Published 7 Oct 2026 in cs.CV | (2610.09502v1)

Abstract: Open-vocabulary detection (OVD) recognizes categories unseen during training through textual category queries, yet achieving strong generalization with real-time efficiency remains challenging. Beyond vocabulary scaling, zero-shot generalization may benefit from reusable visual--semantic cues learned from seen data, including attributes, actions, states, and contextual relations. Existing real-time OVD methods primarily emphasize vocabulary coverage and efficient region/query--text matching; under strict efficiency constraints, compact detectors may struggle to absorb rich instance semantics and scene context. We propose RT-DETR-World, a compact DETR-style detector that transfers the rich semantics conveyed by descriptions during training while retaining lightweight query--text matching at inference. We construct GroundingCapv2 with three levels of supervision: category names for standard OVD, object descriptions conveying instance-level semantics, and image descriptions conveying object relations and scene context. These descriptions serve only as training-time semantic supervision. To help the compact detector absorb these semantics, we propose Dual-Path Description Alignment (DDA), combining a deployment-consistent MiniLM pathway with a training-only LLM teacher. MiniLM provides query--category supervision and object-description alignment, while offline teacher features supervise matched queries and global visual representations at the object and image levels, respectively. All teacher features are precomputed, and the teacher-side modules are removed after training. We further propose Relation-Aware Negative Relaxation (RNR), which uses teacher-derived semantic similarities to relax related negatives while preserving exact positives. Experiments demonstrate competitive zero-shot accuracy and a favorable accuracy--efficiency trade-off. The code will be released.

Summary

  • The paper introduces RT-DETR-World, a detector that leverages rich language semantics from descriptions, attributes, and quantities, improving the transfer learning capabilities and real-time performance of open-vocabulary detection (OVD) systems
  • RT-DETR-World integrates three core components: GroundingCapv2, a dataset with multi-level annotated semantics, Dual-Path Description Alignment (DDA) for aligning descriptions, and Relation-Aware Negative Relaxation (RNR) to appropriately reduce penalties for related but non-matching descriptions.
  • The methodology surpasses previous models in accuracy and efficiency under varying parameter configurations, achieving 40.0 Fixed AP at 32 FPS on LVIS

Problem formulation and central contribution

RT-DETR-World addresses a specific tension in open-vocabulary detection (OVD): detectors must transfer to categories absent from training while maintaining the latency and parameter constraints associated with real-time deployment. Existing efficient OVD systems primarily optimize the breadth of the training vocabulary and the cost of query–text matching. RT-DETR-World instead argues that zero-shot transfer can be improved by learning reusable visual–semantic factors from rich descriptions, including attributes, parts, quantities, actions, states, object relations, and scene context. The central design claim is that such semantics can be transferred during training without appearing in the deployed model.

The proposed detector retains a lightweight inference interface based on category names and a compact MiniLM text encoder, while using substantially richer supervision during training. The paper introduces three coupled components:

  1. GroundingCapv2, a 1.12-million-sample dataset with category-, object-, and image-level supervision;
  2. Dual-Path Description Alignment (DDA), which transfers instance and scene semantics through both a deployable MiniLM pathway and a training-only LLM teacher;
  3. Relation-Aware Negative Relaxation (RNR), which reduces the contrastive penalty assigned to semantically related but nonmatching descriptions while preserving exact-positive supervision.

The resulting detector uses a frozen DINOv3 visual encoder, a lightweight multi-scale projector, a three-layer DETR decoder, and MiniLM-based query–category matching. The LLM teacher, description-alignment projections, and auxiliary training branches are removed after optimization. Thus, the paper’s principal systems claim is not that a larger multimodal architecture should be deployed, but that rich language supervision can shape a compact detector whose inference computation remains largely conventional.

GroundingCapv2 and confidence-routed annotation

GroundingCapv2 reorganizes and augments GroundingCap-1M into an explicit three-level supervision structure. For each image, the dataset may provide concise category names for regions, object descriptions associated with individual instances, and an image description representing scene-level context. These annotation types are intentionally separated because they encode different learning targets. Category names support the standard OVD interface; object descriptions expose instance-specific properties; and image descriptions provide global relational and contextual information.

The resulting corpus contains 1,115,690 samples and 8,050,813 annotated regions. Category names are available for 95.8% of regions and object descriptions for 98.5%; 1,115,656 samples also contain image descriptions. The source mixture combines COCO, V3Det, GQA, Flickr30K Entities, and LLaVA-Cap. For COCO and V3Det, source categories are preserved and instance descriptions are generated. For the grounding and region–phrase sources, original phrases are retained and normalized categories are added only when they are supported by the visual evidence.

The annotation pipeline is a confidence-routed multimodal curation process called GroundingAgent-Opt. Qwen3-VL-8B agents propose and verify descriptions or normalized categories, while uncertain cases are routed to Qwen3-VL-32B for adjudication. For COCO and V3Det, generated descriptions are accepted for 85.5% and 87.6% of regions, respectively; the remainder use category-name fallback. This mechanism acknowledges a central data-quality problem: description supervision can introduce unsupported attributes, incorrect region identities, or category errors. The fallback policy reduces this risk, but it also means that the effective semantic density of the dataset is heterogeneous rather than uniform.

The data construction is therefore a substantive component of the method, not merely preprocessing. The reported improvements depend on model-generated annotations, verification decisions, source-specific coverage, and the assumption that the retained descriptions are sufficiently accurate to provide useful semantic targets. The paper describes auditing metadata and quality-control procedures, but the supplied results do not include an independent human accuracy estimate for the generated descriptions.

Detector architecture and dual-path alignment

The deployable architecture follows a compact DETR-style design. Frozen DINOv3 variants provide visual features, which are transformed into multi-scale detection features by a two-layer projector. A three-layer DETR decoder predicts object queries and bounding boxes. Language-Guided Query Selection uses MiniLM embeddings of category-oriented prompts to select or condition object queries, and category embeddings are matched against query representations for open-vocabulary classification.

The three detector scales differ primarily in the DINOv3 backbone:

  • RT-DETR-World-S uses DINOv3-S/16;
  • RT-DETR-World-S+ uses DINOv3-S+/16;
  • RT-DETR-World-B uses DINOv3-B/16.

All variants use 256-dimensional projected features, 300 queries at inference, and MiniLMv2-L6-H768 as the deployed text encoder. Training uses 13 groups of 300 queries, while inference retains one group. This discrepancy is relevant to the efficiency claim: the reported runtime corresponds to the reduced inference configuration rather than the full training query budget.

DDA consists of two complementary alignment paths. The first is deployment-consistent: MiniLM encodes object descriptions, and the resulting text features are aligned with Hungarian-matched detector queries through learned projection heads. Although object descriptions are discarded at inference, this pathway trains the query representations and MiniLM-compatible semantic space that remain relevant to deployed category matching.

The second path uses an LLM2CLIP-derived LLM text encoder to generate 1,152-dimensional embeddings for object and image descriptions. Object-level teacher embeddings supervise the matched query representation, while image-level teacher embeddings supervise a global visual representation obtained by spatially pooling and concatenating multi-scale features. All teacher embeddings are precomputed. Consequently, the LLM is not loaded during detector optimization or inference, although its computational and representational contribution is still present in the offline training pipeline.

The total objective combines the base detection loss with MiniLM object-description alignment, LLM object-level alignment, and LLM image-level alignment. The selected default weights are λM=0.25\lambda_M = 0.25, λo=0.1\lambda_o = 0.1, and λg=0.1\lambda_g = 0.1. This decomposition separates three forms of supervision: category identity, instance semantics, and global scene semantics. The paper’s ablations indicate that this separation is not redundant.

Relation-Aware Negative Relaxation

RNR modifies bidirectional contrastive alignment. Standard contrastive objectives treat every unmatched visual–text pair as a negative, even when two descriptions share attributes, actions, states, or contextual relations. RNR uses the frozen LLM teacher embeddings to estimate semantic similarity between descriptions. Exact positive pairs retain unit weight, while semantically related negatives receive reduced denominator weights. An unmatched pair is therefore relaxed but never converted into a positive.

This distinction is methodologically important. Soft-target approaches can blur instance correspondence by treating semantically similar examples as partial positives. RNR preserves the one-to-one matching structure required by detection while weakening only the repulsive force between related descriptions. The relaxation coefficient is set to α=0.25\alpha = 0.25 by default. At α=0\alpha = 0, the loss reduces to standard contrastive learning; excessively large values over-relax the negative structure.

The reported sensitivity study supports a nonmonotonic effect. On LVIS minival, increasing α\alpha from 0 to 0.25 improves Fixed AP from 33.6 to 35.2 and rare-category AP from 27.6 to 32.0. Increasing it further to 0.5 reduces overall AP to 33.4 and rare-category AP to 29.4. The result implies that semantic relatedness is useful as a graded correction to contrastive learning, but not as a wholesale replacement for instance-level discrimination.

Zero-shot detection and efficiency

The principal LVIS results show a consistent accuracy–efficiency trade-off across model scales.

Model FPS Standard AP Fixed AP Rare Fixed AP Parameters total
RT-DETR-World-S 63 31.4 35.2 32.0 103.1M
RT-DETR-World-S+ 58 32.6 36.6 31.1 110.2M
RT-DETR-World-B 32 36.6 40.0 37.1 187.2M
OV-DEIM-L 91 33.7 35.9 36.8 99.4M
YOLOEv11-L 17 32.4 35.2 29.1 89.4M

RT-DETR-World-S+ reaches 36.6 Fixed AP at 58 FPS, exceeding the strongest previously reported real-time result in the comparison by 0.7 points. RT-DETR-World-B reaches 40.0 Fixed AP at 32 FPS, exceeding the previous real-time best by 4.1 points. Its rare-, common-, and frequent-category Fixed AP values are 37.1, 39.6, and 40.9, respectively.

These results support the claim that rich training-time semantics can improve large-vocabulary detection without requiring an LLM during inference. They do not establish dominance across the entire efficiency frontier. OV-DEIM-L remains substantially faster at 91 FPS and reaches 35.9 Fixed AP, so RT-DETR-World primarily improves the higher-accuracy portion of the trade-off rather than uniformly outperforming all compact baselines. Moreover, the total parameter count includes the deployed MiniLM text encoder but excludes training-only language modules, which is appropriate for inference accounting but makes direct comparisons with methods using different vocabulary-caching conventions less straightforward.

The paper also reports a remaining accuracy gap relative to heavier general-purpose grounding models. RT-DETR-World-B is 1.4 Fixed AP points below MM-Grounding-DINO-T and 4.7 points below LLMDet-T. This gap is consistent with the architectural premise: removing multi-layer cross-modal interaction and retaining lightweight query–text matching produces efficiency gains at the cost of some representational capacity.

Cross-domain transfer and distribution-shift robustness

The cross-domain results on ODinW provide stronger evidence for the paper’s semantic-transfer hypothesis than LVIS alone. RT-DETR-World-S achieves 40.1 AP on ODinW13 and 18.7 AP on ODinW35. Relative to YOLO-Worldv2.1-L, this represents gains of 4.5 and 2.1 points; relative to the strongest YOLOE variant, the gains are 8.2 and 3.9 points. RT-DETR-World-B reaches 43.8 and 20.2 AP, exceeding OV-DEIM-L by 2.2 and 1.0 points despite the latter’s strong performance.

Model ODinW13 AP ODinW35 AP
YOLO-Worldv2.1-L 35.6 16.6
YOLOEv8-L 31.9 14.8
OV-DEIM-L 41.6 19.2
RT-DETR-World-S 40.1 18.7
RT-DETR-World-S+ 41.5 18.8
RT-DETR-World-B 43.8 20.2

The ODinW results imply that the model’s advantage is not limited to recognition of a large fixed vocabulary. Its performance persists across heterogeneous domains and category distributions. However, ODinW is evaluated using standard AP averaged across datasets, and the paper does not provide per-domain variance or a detailed failure analysis. The aggregate improvement therefore establishes broad transfer but does not identify which semantic supervision components are most responsible for particular domain shifts.

COCO-O evaluates robustness under sketch, weather, cartoon, painting, tattoo, and handmake shifts while retaining the COCO category space. Because GroundingCapv2 contains COCO training images, the authors correctly treat COCO-O as a distribution-shift benchmark rather than an unseen-category benchmark. RT-DETR-World-B obtains 47.2 AP on COCO and 46.2 AP on COCO-O, with effective robustness of +24.9 under the paper’s metric. It exceeds OV-DEIM-L by 2.9 COCO-O AP points and 2.2 effective-robustness points. RT-DETR-World-S also reaches 34.9 COCO-O AP, 7.6 points above YOLO-Worldv2.1-L despite lower in-distribution COCO AP.

The fact that COCO-O AP approaches or nearly matches COCO AP for the larger variants is notable, but the effective robustness metric depends on a fixed coefficient applied to COCO AP. It should therefore be interpreted as a benchmark-specific summary rather than a distribution-free robustness measure.

Ablation evidence

The supervision ablations directly support the multi-level design. Using GroundingCap-1M, adding object- and image-level LLM teacher supervision increases Fixed AP from 27.9 to 32.9, a 5.0-point gain without changing the deployed detector. Under GroundingCapv2, adding MiniLM object-description alignment to category-only supervision raises AP from 33.1 to 35.2 and rare-category AP from 25.7 to 32.0.

The two LLM alignment levels are complementary. Starting from MiniLM-only supervision at 32.2 AP, image-level teacher alignment increases AP to 33.1, object-level alignment increases it to 33.6, and combining both reaches 35.2. The object-level branch contributes more strongly to rare categories, whereas the image-level branch supplies scene-level context that cannot be recovered from isolated category labels alone. The implication is that the reported improvement is not attributable solely to a stronger text encoder; it depends on aligning distinct visual representations with descriptions at distinct spatial scales.

RNR is also responsible for a substantial portion of the final performance. Under the same auxiliary-alignment setting, hard contrastive alignment reaches 33.6 AP, SoftCLIP-style alignment 33.8, and SRCL-style similarity regulation 33.9. RNR reaches 35.2, outperforming hard alignment by 1.6 points and SRCL-style regulation by 1.3 points. The rare-category gain over hard alignment is 4.4 points. This is consistent with the hypothesis that related-negative over-separation is particularly harmful when rare categories have limited direct visual support.

The hyperparameter results indicate that the method requires moderate weighting. Increasing the LLM teacher weight from 0.1 to 0.2 raises rare-category AP from 32.0 to 32.5 but lowers overall AP from 35.2 to 34.3. Similarly, increasing the MiniLM description weight beyond 0.25 reduces overall AP. The model therefore benefits from semantic supervision as an auxiliary signal rather than as a replacement for category-level detection supervision.

Limitations and open questions

The evaluation supports the method’s empirical claims, but several limitations constrain their interpretation. First, GroundingCapv2 relies heavily on automatically generated and verified descriptions. The paper reports acceptance rates and routing procedures, but not a human-annotated estimate of description precision, coverage conditional on object type, or systematic bias across sources. It remains open whether the gains arise from rich semantics per se, from increased annotation volume, from the stronger DINOv3 backbone, or from interactions among these factors. The ablations isolate some supervision components but do not provide a fully controlled comparison using equalized annotation counts and identical visual backbones across all competing methods.

Second, the inference claim is lightweight rather than language-free. MiniLM remains part of the deployed detector, and category embeddings are cached during FPS measurement. Runtime therefore depends on the query regime and on the amortization of text encoding. The reported FPS protocol is appropriate for standard OVD evaluation, but applications requiring dynamically changing vocabularies or many independent query batches may exhibit different costs.

Third, the comparison aggregates results from papers and repositories with potentially different implementation and timing conventions. The authors carefully distinguish cached vocabulary embeddings and NMS inclusion, but the baselines do not necessarily share identical input preprocessing, hardware kernels, or post-processing. The accuracy comparisons are more robust than the absolute latency comparisons.

Fourth, the method’s semantic transfer mechanism remains partly entangled with the pretrained components. DINOv3 supplies strong visual representations, and LLM2CLIP supplies the relation geometry used by RNR. The experiments establish that the complete system works, but they do not determine whether a smaller teacher, a domain-specific teacher, or non-LLM semantic relations would preserve the gains. Nor do they test whether the teacher’s similarities encode genuinely visual relations rather than linguistic correlation.

Finally, the results do not fully characterize calibration, prompt sensitivity, duplicate predictions, or compositional generalization under controlled attribute–object splits. The qualitative discussion acknowledges low-confidence and overlapping predictions in crowded scenes. Whether RNR improves semantic discrimination at the instance level while preserving localization calibration remains an open empirical question.

Conclusion

RT-DETR-World presents a coherent strategy for transferring description-level semantics into a real-time OVD detector. GroundingCapv2 supplies category, instance, and scene supervision; DDA aligns compact detector representations with both MiniLM and offline LLM targets; and RNR moderates contrastive repulsion between related descriptions without weakening exact correspondences. The strongest model reaches 40.0 Fixed AP on LVIS at 32 FPS, 43.8 AP on ODinW13, 20.2 AP on ODinW35, and 46.2 AP on COCO-O, while removing the LLM and auxiliary branches at inference. The ablations show that object-level, image-level, deployment-consistent, and relation-aware supervision contribute complementary gains. The principal unresolved issue is how much of this performance depends on the specific automatically curated annotations and pretrained teacher representations, rather than on the general principle of rich training-time semantic supervision.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper introduces RT-DETR-World, a computer-vision system that can find objects in images even when it was not specifically trained on those object categories.

For example, a normal object detector might be trained to recognize chairs and spoons. An open-vocabulary detector could also recognize a wooden spoon or a wheelchair by using a text description, even if those exact categories were not part of its training labels.

The main challenge is to make the detector:

  • Accurate, especially on unfamiliar objects
  • Fast enough for real-time use
  • Able to understand more than simple object names, such as color, material, position, actions, and relationships

The researchers’ main idea is:

Teach a small, fast detector using rich descriptions from LLMs during training, but remove the expensive LLM when the detector is actually being used.

2. What questions are the researchers asking?

The paper focuses on several important questions:

  1. Can detailed descriptions help a detector recognize unfamiliar objects? For instance, can learning about “a wooden chair leaning against a wall” help the system later recognize a wooden spoon?
  2. Can a small detector learn useful visual information from a much larger LLM? This is similar to a teacher helping a student learn. The large model provides guidance during training, but the smaller model works alone later.
  3. Can this be done without making the detector slower? The researchers want the final system to remain suitable for robots, self-driving vehicles, cameras, and other devices that need quick decisions.
  4. Can the detector avoid treating similar things as completely unrelated? For example, “a red car” and “a blue car” are not the same object, but they are still related. The system should not push their meanings too far apart during learning.

3. How did the researchers approach the problem?

A three-level training dataset

The researchers created a dataset called GroundingCapv2. It contains more than 1.1 million images or samples and over 8 million marked regions.

The data provides three kinds of information:

Type of information Example What it teaches
Category name “chair” The basic object identity
Object description “a wooden chair leaning against a wall” Details about one object
Image description “two chairs beside a table in a room” Relationships and the overall scene

Some descriptions were created or checked by vision-LLMs. These models were asked to describe objects and then verify whether their descriptions matched the image. Uncertain examples were checked by a larger model or replaced with a simpler category label.

A fast detector

RT-DETR-World is based on a detector called DETR. A detector is a computer program that looks at an image and predicts:

  • What objects are present
  • Where each object is, using a box around it
  • Which text label best describes each object

The system uses:

  • DINOv3 to extract useful visual features from an image
  • A small DETR decoder to locate objects
  • MiniLM, a compact LLM, to compare objects with text labels

At testing time, the detector only needs category names such as “dog,” “helmet,” or “wooden spoon.” The LLM is not used then.

Learning from a large teacher model

During training, the system uses a larger LLM as a teacher. The teacher reads detailed descriptions and converts them into numerical representations called embeddings.

An embedding is like a list of numbers that represents meaning. Similar ideas receive similar numerical patterns. For example, “car” and “vehicle” would probably have more similar embeddings than “car” and “banana.”

The student detector learns from the teacher in two ways:

  • Object-level learning: It learns details about individual objects, such as their material, shape, action, or condition.
  • Image-level learning: It learns about the whole scene, including relationships between objects and the surrounding context.

The teacher’s results are calculated beforehand, so the large model does not need to run during normal use.

Dual-Path Description Alignment

The researchers call this learning system Dual-Path Description Alignment, or DDA.

It has two paths:

  1. A MiniLM path, which matches the detector’s output with descriptions using the same small LLM that will be available during testing.
  2. A large-teacher path, which gives the detector richer and more detailed guidance during training.

The extra training parts are removed after learning. Therefore, they improve the detector without adding extra work when it is running.

Relation-Aware Negative Relaxation

The researchers also introduce Relation-Aware Negative Relaxation, or RNR.

Normally, contrastive learning teaches a system:

  • “These two matching examples belong together.”
  • “All other examples should be pushed apart.”

But this can be too strict. For example, “a bicycle” and “a motorcycle” are different, but they are both vehicles with wheels. RNR uses the teacher’s understanding to recognize that some “wrong” pairs are still related. It pushes them apart less strongly while still keeping the correct pair as the best match.

An everyday analogy is sorting books. A book about dogs and a book about cats should not be placed in exactly the same position, but they probably belong closer together than a book about astronomy.

4. What did the researchers find?

Stronger recognition of unfamiliar categories

The researchers tested the system on several benchmarks, including:

  • LVIS, which contains many object categories, including rare ones
  • ODinW, which tests whether a detector can work on very different types of images
  • COCO-O, which tests how well the system handles changes in appearance, such as cartoons, sketches, paintings, and unusual weather

On LVIS, the largest RT-DETR-World model achieved:

  • 36.6 standard AP
  • 40.0 Fixed AP
  • 32 frames per second

Here, AP is a score for detection accuracy. A higher score means the system more often identifies the correct object and places its box accurately. Frames per second, or FPS, measures speed.

The smaller versions were faster:

  • RT-DETR-World-S: 63 FPS
  • RT-DETR-World-S+: 58 FPS
  • RT-DETR-World-B: 32 FPS

This shows a useful balance between speed and accuracy.

Better transfer to new environments

On the ODinW benchmark, the largest model reached:

  • 43.8 AP on the 13-dataset version
  • 20.2 AP on the 35-dataset version

These results were better than several competing real-time detectors. This suggests that the system can transfer what it learned to new environments, object types, and visual styles.

Better robustness to changes in appearance

On COCO-O, the largest model achieved:

  • 47.2 AP on ordinary COCO images
  • 46.2 AP on shifted images

The shifted images included sketches, cartoons, paintings, tattoos, and handmade images. The detector therefore remained effective even when objects looked very different from the training examples.

Detailed descriptions were especially helpful

The experiments also tested which parts of the method mattered most. The results showed that:

  • Category names alone were useful but limited.
  • Object descriptions improved understanding of rare categories.
  • Image descriptions added useful scene and relationship information.
  • Combining object-level and image-level learning worked best.
  • RNR performed better than several other ways of handling similar examples.

For one version of the model, adding the full combination of methods raised the score from 27.9 AP to 35.2 AP. This is a large improvement.

5. Why are these results important?

Many real-world systems need to recognize objects they have not seen exactly during training. A robot might encounter a new tool, a self-driving car might see an unusual vehicle, or a security camera might need to identify objects in unfamiliar surroundings.

The paper shows that a detector does not have to rely only on short labels such as “chair” or “car.” It can learn from richer descriptions that explain:

  • What an object looks like
  • What it is made of
  • How many objects there are
  • What an object is doing
  • Where objects are located
  • How objects relate to one another

The most important practical point is that the expensive LLM is used only while teaching the detector. Once training is complete, it is removed. This means the final detector keeps much of the learning benefit while remaining relatively fast.

Conclusion and possible impact

RT-DETR-World is a method for building fast object detectors that understand more detailed visual meaning. It combines ordinary category labels with descriptions of objects and entire scenes. A LLM acts as a teacher during training, while a smaller detector handles the actual image processing later.

The results suggest that this approach can improve:

  • Recognition of rare and unfamiliar objects
  • Performance across different datasets
  • Reliability when images look unusual
  • The balance between accuracy and speed

In the future, systems based on this idea could help robots, smart cameras, augmented-reality devices, and autonomous vehicles understand new situations more effectively. However, the method still has limitations: larger models are slower, and automatically generated descriptions can sometimes contain mistakes. Careful checking of the training data remains important.

Glossary

  • Adjudication: Formal review and resolution of ambiguous or conflicting annotations. “routing ambiguous cases to Qwen3-VL-32B for multimodal adjudication”
  • Alignment space: A shared vector space in which representations from different modalities are made comparable. “project the text and query representations into a shared alignment space”
  • AP (Average Precision): A detection metric summarizing precision across recall levels. “We report standard AP and Fixed AP”
  • Backbone: The primary feature-extraction network in a deep-learning model. “RT-DETR-World-S, -S+, and -B use frozen DINOv3”
  • Bidirectional contrastive learning: Contrastive learning that evaluates both visual-to-text and text-to-visual matching directions. “α = 0 recovers standard bidirectional contrastive learning”
  • Bounding-box regression: Prediction of the coordinates and dimensions of an object’s bounding box. “The DETR decoder refines the selected queries and regresses bounding boxes.”
  • Caption-contrastive tuning: Training an encoder to associate image captions with corresponding visual representations while separating mismatched captions. “the caption-contrastively tuned LLM text encoder from LLM2CLIP”
  • Category fallback: Use of a category label when a more detailed annotation is unreliable. “unreliable descriptions fall back to category names”
  • Cross-modal alignment: Matching representations from different modalities, such as images and text. “SoftCLIP-style soft alignment”
  • Cosine similarity: A similarity measure based on the angle between two vectors. “where sim(·, ·) denotes cosine similarity”
  • DETR decoder: The transformer decoder component of Detection Transformer architectures that produces object predictions from queries. “a DETR decoder that produces object queries”
  • Distribution shift: A change in the data distribution between training and evaluation. “distribution-shift robustness on COCO-O”
  • Embedding: A numerical vector representation of an item such as text, an image, or an object query. “the corresponding embeddings in the dL-dimensional teacher space”
  • Feature level: One resolution or scale of the hierarchical feature representations used by a detector. “valid spatial features from each projected feature level are pooled”
  • Fixed AP: An average-precision metric using a fixed per-category detection budget instead of a fixed number of detections per image. “Fixed AP replaces this image-level limit with a fixed per-category detection budget.”
  • Global visual representation: A vector summarizing visual information from an entire image. “global visual representations at the object and image levels”
  • Grounding: Linking textual phrases or descriptions to specific regions in an image. “GLIP and Grounding DINO unify object detection and phrase grounding”
  • Hungarian matching: An optimal assignment algorithm used to match predicted object queries with annotated objects. “the Hungarian-matched query assigned by Hungarian matching to region i”
  • Inference: The process of using a trained model to generate predictions. “The LLM itself is absent during detector training and inference”
  • In-distribution: Belonging to the same data distribution as the training data. “COCO (Lin et al., 2014) as the in-distribution reference”
  • Language-guided query selection: Selection of detector queries using similarities between visual features and text representations. “Vision–text similarities guide Language-Guided Query Selection”
  • Long-tailed distribution: A data distribution in which a small number of classes are frequent while many classes are rare. “LVIS contains 1,203 categories with a long-tailed distribution”
  • Mask-aware average pooling: Averaging feature values only over spatial positions designated as valid by a mask. “MAP performs mask-aware average pooling over valid spatial positions”
  • Multi-scale projector: A module that transforms visual features from multiple resolutions into representations suitable for detection. “a lightweight multi-scale projector”
  • Negative relaxation: Reduction of the penalty assigned to semantically related but unmatched examples during contrastive learning. “Relation-Aware Negative Relaxation (RNR)”
  • Offline teacher: A pretrained model whose outputs are computed in advance and used as training targets for another model. “a training-only LLM teacher”
  • Open-vocabulary detection: Object detection in which textual queries can specify categories not observed during detector training. “Open-vocabulary detection (OVD) recognizes categories unseen during training”
  • Open-set object detection: Detection of objects beyond the closed set of categories defined during training. “Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection”
  • Parameter freezing: Preventing selected model parameters from being updated during training. “a frozen DINOv3 visual encoder”
  • Phrase grounding: Localizing a natural-language phrase to its corresponding image region. “unify object detection and phrase grounding”
  • Query–text matching: Comparing detector object-query representations with textual category or description representations. “real-time OVD methods favor lightweight region/query–text matching”
  • Region–phrase annotation: An annotation linking a textual phrase to a specific image region. “Region–Phrase Annotations (GQA / Flickr30K Entities / LLaVA-Cap)”
  • Representation capacity: The ability of a model to encode and retain information in its learned representations. “the limited representation capacity of compact detectors”
  • Semantic distillation: Transfer of semantic information from a larger or more expressive model to a smaller model. “A hierarchical semantic distillation framework for open-vocabulary object detection”
  • Semantic overlap: The degree to which two descriptions share meaning despite not being identical. “partial semantic overlap”
  • Semantic supervision: Training signals that convey conceptual or linguistic information about visual content. “These descriptions serve only as training-time semantic supervision.”
  • Soft alignment: Alignment that uses graded similarity targets rather than only exact-match targets. “SoftCLIP-style soft alignment”
  • Temperature: A scalar that controls the sharpness of probability distributions derived from similarity logits. “τ > 0 is the temperature”
  • Text encoder: A neural network that converts text into numerical representations. “MiniLM encodes category names and category-oriented textual prompts.”
  • Token features: Vector representations produced for individual tokens by a LLM. “we pool the valid non-special-token features of dbi”
  • Training-time supervision: Information used to train a model but omitted from the deployed inference system. “object and image descriptions convey instance- and scene-level semantics during training”
  • Transfer learning: Reusing knowledge learned from one set of data or tasks for another. “zero-shot cross-domain transfer on ODinW”
  • Zero-shot generalization: Recognition of categories or tasks for which no direct training examples were provided. “zero-shot generalization may also benefit from reusable visual–semantic factors”
  • Zero-shot object detection: Detecting object categories that were not present in the detector’s training annotations. “Zero-shot object detection”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.