Papers
Topics
Authors
Recent
Search
2000 character limit reached

MCQS: Modality Competitive Query Selection

Updated 14 July 2026
  • MCQS is a competitive mechanism where modality-specific representations vie to become initial queries for downstream processing.
  • It underpins applications like infrared-visible object detection, long-video QA, and federated learning by delaying indiscriminate fusion.
  • Empirical studies show that MCQS enhances accuracy and efficiency by selecting dominant features and mitigating cross-modal interference.

Searching arXiv for papers on MCQS and closely related terminology. Modality Competitive Query Selection (MCQS) denotes a class of selection mechanisms in which candidate modality-specific representations compete before downstream reasoning, decoding, or retrieval. In the most literal usage, MCQS is the query-initialization module of DAMSDet for infrared-visible object detection, where encoded infrared and visible features compete and the Top-KK selected modality-specific features become initial object queries (Guo et al., 2024). Closely related work extends the same logic to long-video understanding, where Q-Gate treats keyframe selection as a query-conditioned routing problem over three modality experts and dynamically suppresses irrelevant modalities (Wang et al., 19 Apr 2026). Other papers do not always use the term explicitly, but they instantiate analogous selection patterns: cross-modal option synthesis for visual multiple-choice questions, joint modality ranking under communication constraints in federated learning, and competitive choice differentiation in multiple-choice reasoning (Wang et al., 26 Aug 2025, Yuan et al., 2024, Deng et al., 2024, Wong et al., 7 Jan 2026).

1. Conceptual scope

Across recent work, MCQS is not a single canonical algorithm but a recurring computational pattern: score heterogeneous candidates, impose competition, and pass forward only the most useful modality-conditioned or choice-conditioned representation. In DAMSDet, the competition is between infrared and visible encoded features for each candidate object (Guo et al., 2024). In Q-Gate, the competition is between three relevance streams—Visual Grounding, Global Matching, and Contextual Alignment—whose weights are modulated by the query (Wang et al., 19 Apr 2026). In mmFedMC, the competition is between modality models under communication constraints, with ranking driven by Shapley value, model size, and recency (Yuan et al., 2024). In CmOS, the competition is between candidate question-reason pairs and between visually plausible answer options and distractors (Wang et al., 26 Aug 2025).

Work Domain Selection target
DAMSDet (Guo et al., 2024) Infrared-visible detection Top-KK modality-specific query features
Q-Gate (Wang et al., 19 Apr 2026) Long-video QA Query-weighted expert streams for keyframe scoring
mmFedMC (Yuan et al., 2024) Multimodal federated learning Uploaded modality models and selected clients

Taken together, these formulations suggest that MCQS is best viewed as a modality-aware competitive prior-selection mechanism rather than as a fixed fusion rule. A plausible implication is that its unifying role is to delay indiscriminate fusion until the system has identified which modality, expert, or candidate is currently most reliable.

2. MCQS in infrared-visible object detection

DAMSDet introduces MCQS explicitly as a sample-adaptive query initialization mechanism for infrared-visible object detection (Guo et al., 2024). The motivation is twofold. First, modality complementarity is highly dynamic: infrared and visible streams do not contribute equally across objects or scenes. Second, modality misalignment is common: direct early fusion can inject spatially inconsistent information into object hypotheses. MCQS addresses these issues by avoiding early indiscriminate fusion and by selecting a dominant modality feature for each object instance before decoder refinement begins.

The MCQS inputs are the flattened encoded feature sequences from the infrared and visible branches, denoted II and VV. These are concatenated and scored by a linear layer, after which the highest-scoring feature points are retained as initial queries: z=Top-K⁡(Linear⁡(concat⁡(I,V))).z=\operatorname{Top-\mathit{K}}(\operatorname{Linear}(\operatorname{concat}(I, V))). The selected Top-KK features are “sourced either from the infrared or visible features, respectively,” so the competition is implicit in the Top-KK ranking over the combined feature pool rather than in a separate soft gating equation. This is a central property of the DAMSDet formulation: MCQS is a selection procedure, not a weighted average.

These selected features serve as initial object queries for the Multispectral Transformer Decoder. In DAMSDet, the query carries both content and position information, and it is iteratively refined in a cascade DETR-style decoder. For decoder layer dd, the query state zqdz_q^d is mapped into a refined reference box: $b_{q\{x, y, w, h\}^{d}=\sigma\left(M L P^{d}\left(z_{q}^{d}\right)+\sigma^{-1}\left(b_{q}^{d-1}\right)\right).$ This couples MCQS directly to the cascade architecture: the quality of the initial modality-specific query prior determines the starting point for iterative localization and refinement.

MCQS is complemented by the Multispectral Deformable Cross-attention (MDCA) module, which adaptively samples from infrared and visible feature maps at multiple semantic levels. The paper distinguishes their roles sharply: MCQS decides where to start and which modality to trust first, whereas MDCA decides how to fuse complementary information next (Guo et al., 2024). This division of labor is especially important under misalignment, because the initial query is prevented from being polluted by a poor cross-modal mixture before the decoder has a reliable object-centered hypothesis.

The ablation study on MKK0FD isolates the effect of MCQS. Relative to standard query selection, enabling MCQS yields +1.1% KK1AP50 and +0.7% KK2AP. The best overall configuration combines MCQS, MDCA, and content query selection, reaching 80.2 KK3AP50, 56.0 KK4AP75, and 52.9 KK5AP, compared with 77.8 / 56.0 / 51.6 for the multimodal baseline with standard query strategy (Guo et al., 2024). The paper’s interpretation is that a dynamic, object-specific, modality-competitive prior is superior to both abstract learnable DETR queries and naïve early multimodal fusion.

A common misconception is to equate MCQS here with generic multimodal fusion. DAMSDet states the opposite design principle: MCQS is valuable precisely because it prevents the introduction of interference from another modality in the early stages (Guo et al., 2024).

3. Query-modulated routing in long-video understanding

Q-Gate realizes an MCQS-like mechanism in long-video QA by treating keyframe selection as a dynamic modality routing problem (Wang et al., 19 Apr 2026). Rather than using a single visual-centric metric or a static fusion of heuristic scores, it decouples retrieval into three lightweight expert streams:

  • Visual Grounding for local details
  • Global Matching for scene semantics
  • Contextual Alignment for subtitle-driven narratives

Each stream produces a time-aligned score vector KK6. The LLM-based gate then assigns query-specific weights

KK7

and the final per-frame score is a weighted mixture: KK8 This is the paper’s formalization of query-conditioned competition among modalities (Wang et al., 19 Apr 2026).

The three streams are deliberately heterogeneous. Visual Grounding extracts visual entities KK9 from the query and uses YOLO-World to verify entity presence and relations, with raw score

II0

Global Matching uses a pre-trained VLM such as BLIP-2 to compute cosine similarity between a query embedding II1 and a frame embedding II2: II3 Contextual Alignment compares the query with timestamped subtitles using Sentence-BERT: II4

Because these streams have different scales and sparsity patterns, Q-Gate applies min-max scaling and then a masked temperature softmax with II5. The masking is crucial: the paper emphasizes that a standard softmax would assign nonzero probability to absent-subtitle frames because II6, thereby injecting noise (Wang et al., 19 Apr 2026). This normalization pipeline is therefore part of the selection mechanism itself, not just a preprocessing detail.

Q-Gate formalizes the gate as a Zero-Shot MoE: the streams are pretrained experts, the LLM is the gating network, and no training is required. The selected top-II7 frames are used to construct a timestamped multimodal prompt for a downstream QA VLM. This makes Q-Gate a pre-selection module that improves evidence quality and token efficiency rather than a replacement for the answering model.

The empirical findings are substantial. On LongVideoBench with Qwen3-VL, II8, Q-Gate achieves 59.40 on Long and the paper reports +1.60 over AKS. On Video-MME with **Qwen3-VL, II9, it reports **61.19* on Long, 66.13 on Medium, and 79.41 on Short, with +6.40 over AKS* on the Long split (Wang et al., 19 Apr 2026). The ablations are equally important conceptually: removing Contextual Alignment drops LongVideoBench Long from 59.40 to 54.08, and replacing dynamic gating with equal weights VV0 performs markedly worse, including 55.67 versus 70.59 on LongVideoBench Short with Qwen3-VL-32B (Wang et al., 19 Apr 2026).

The paper’s interpretability analysis shows that the gate assigns sensible weight patterns: Counting / Attribute questions emphasize Visual Grounding, Action / scene-level questions emphasize Global Matching, and Reasoning / Subtitle-specific questions emphasize Contextual Alignment, sometimes nearly half the weight. This directly supports the claim that MCQS-like routing is suppressing cross-modal negative transfer rather than merely averaging scores.

4. Cross-modal competitive options and multiple-choice reasoning

In educational MCQ generation, the competitive-selection logic appears in a different form. CmOS addresses cross-modal option synthesis, where the output includes a generated question VV1, a visual answer option VV2, and multiple visual distractors VV3 tied to textual option descriptions (Wang et al., 26 Aug 2025). The central problem is not standard text-only distractor generation, but the construction of visually plausible, semantically competitive answer sets.

CmOS proceeds in four stages: evaluating content convertibility, generating alternative questions and reasons, selecting the optimal question-reason pair, and generating option descriptions and visual options. The first stage acts as a gatekeeper for modality competition: some content, such as formula-heavy math problems, is not naturally visualizable. To support this discrimination, the system retrieves a similar exemplar from a pool VV4 of 482 ScienceQA examples using cosine similarity over text, answer, and image embeddings, then uses the exemplar in the discriminator prompt (Wang et al., 26 Aug 2025).

Selection becomes explicit again in the Optimal Question-Reason Match (OQRM) module. For candidate pairs VV5, CmOS computes a Total Match Score combining internal consistency and external consistency, with

VV6

and reports VV7 as best in experiments. The visual option generator then retrieves template images from a database and iteratively refines generated outputs whenever the similarity score falls below threshold VV8, up to three rounds (Wang et al., 26 Aug 2025). The paper characterizes competitiveness through semantic plausibility, visual similarity among options, and misleading but valid distractors.

The reported results indicate that this competitive cross-modal synthesis is operationally effective. CmOS achieves 88.2% average accuracy on content discrimination, 75.5 average BLEU-4 and 77.2 average ROUGE-L for question generation, and 59.5 SSIM with 40.2 CLIP-T for visual option generation using Wanx2.1-turbo, compared with 41.5 / 30.8 for Wanx2.1-plus (Wang et al., 26 Aug 2025).

A related but distinct line of work studies competition among answer options rather than among modalities. DCQA proposes Differentiating Choices via Commonality for MCQA, where the model identifies semantic commonalities across answer choices, removes them from the question representation, and focuses on choice-specific nuances (Deng et al., 2024). The core subtraction

VV9

formalizes the elimination of shared, non-discriminative context. This is not MCQS in the multispectral sense, but it shares the same computational motif: suppress broad common evidence and privilege the differentiating signal.

The two-stage computation study of multiple-choice answering makes a further distinction between content selection and symbol binding (Wong et al., 7 Jan 2026). It finds that the winning content position becomes decodable immediately after the final option is processed, while the output symbol is represented closer to answer emission. The paper interprets this as a content-first, symbol-second process, which is closely aligned with the idea that competitive selection can precede later routing or binding. A plausible implication is that MCQS-like mechanisms need not terminate at selection; they may also require a subsequent dereferencing or emission stage.

5. Communication-aware modality selection in federated learning

mmFedMC extends the selection principle into multimodal federated learning by jointly optimizing modality selection on clients and client selection on the server (Yuan et al., 2024). Each client z=Top-K⁡(Linear⁡(concat⁡(I,V))).z=\operatorname{Top-\mathit{K}}(\operatorname{Linear}(\operatorname{concat}(I, V))).0 has a local multimodal dataset

z=Top-K⁡(Linear⁡(concat⁡(I,V))).z=\operatorname{Top-\mathit{K}}(\operatorname{Linear}(\operatorname{concat}(I, V))).1

and trains modality-specific models

z=Top-K⁡(Linear⁡(concat⁡(I,V))).z=\operatorname{Top-\mathit{K}}(\operatorname{Linear}(\operatorname{concat}(I, V))).2

Because clients may have different modality availability and communication budgets, uploading all modality models every round is infeasible.

The client-side modality ranking combines three factors: modality importance via Shapley value, communication overhead via modality model size, and recency. After normalization, the priority score is

z=Top-K⁡(Linear⁡(concat⁡(I,V))).z=\operatorname{Top-\mathit{K}}(\operatorname{Linear}(\operatorname{concat}(I, V))).3

with z=Top-K⁡(Linear⁡(concat⁡(I,V))).z=\operatorname{Top-\mathit{K}}(\operatorname{Linear}(\operatorname{concat}(I, V))).4, while recency is maintained as part of the modality-selection design and varied in ablation through z=Top-K⁡(Linear⁡(concat⁡(I,V))).z=\operatorname{Top-\mathit{K}}(\operatorname{Linear}(\operatorname{concat}(I, V))).5 (Yuan et al., 2024). Clients upload only the top-z=Top-K⁡(Linear⁡(concat⁡(I,V))).z=\operatorname{Top-\mathit{K}}(\operatorname{Linear}(\operatorname{concat}(I, V))).6 modality models. The server then performs client selection per modality using local loss.

This is broader than MCQS in the strict DAMSDet sense, but it is strongly MCQS-like: modalities compete for upload priority under an explicit utility function that balances informativeness, communication cost, and update diversity. The main empirical claim is that mmFedMC achieves comparable or better accuracy while reducing communication overhead by over 20×. At 5 MB cumulative communication per client, reported results include 92.28% with 0.10 MB communication on ActionSense IID, 78.10% with 0.05 MB on UCI-HAR IID, 87.04% with 0.05 MB on PTB-XL IID, 55.82% with 0.16 MB on MELD IID, and 62.79% with 0.28 MB on DFC23 IID (Yuan et al., 2024).

An important corrective point from this paper is that competitive selection does not always favor hard cases. Although the formula is written using a “top loss” set, the empirical conclusion is that selecting clients with lower local loss works better in this multimodal decision-level setting, because higher local loss may reflect outliers, noisy data, or modality-specific mismatch (Yuan et al., 2024).

6. Limitations, misconceptions, and emerging directions

Several recurrent misconceptions are clarified by the literature. First, MCQS is not equivalent to static fusion. Q-Gate argues explicitly that equal-weight fusion causes cross-modal negative transfer, because irrelevant modalities still contribute to the score (Wang et al., 19 Apr 2026). Second, MCQS is not necessarily a learned soft gate. DAMSDet implements modality competitiveness through Top-z=Top-K⁡(Linear⁡(concat⁡(I,V))).z=\operatorname{Top-\mathit{K}}(\operatorname{Linear}(\operatorname{concat}(I, V))).7 selection over scored infrared-visible features rather than through a separate gating equation (Guo et al., 2024). Third, decodability does not by itself establish causal use. The two-stage MCQA study notes that linearly decodable winner information can be present earlier than the point at which it becomes causally effective for answer emission (Wong et al., 7 Jan 2026).

The current evidence base is also bounded by task and model scope. The two-stage MCQA analysis studies only two relatively small instruction-tuned models, and the authors note dependence on prompt format, dataset, and response constraints (Wong et al., 7 Jan 2026). Q-Gate is training-free and plug-and-play, but its routing quality depends on the LLM gate and on prompt rules that encode query-intent heuristics (Wang et al., 19 Apr 2026). mmFedMC warns about a single modality optimization trap and shows that recency can help, although with only two modalities it may also induce excessive cycling (Yuan et al., 2024). DAMSDet reports strong gains overall, but the VEDAI results suggest limitations of the overall transformer-based framework on very small objects, where localization sensitivity remains high (Guo et al., 2024).

A broader interpretation suggested by these papers is that MCQS-like systems separate three computational questions that are often conflated: what candidate should win, which modality should be trusted first, and when that selection should be allowed to influence the final output. In detection, this appears as modality-specific query seeding before cross-modal aggregation (Guo et al., 2024). In long-video QA, it appears as query-dependent routing before keyframe selection (Wang et al., 19 Apr 2026). In MCQA, it appears as winner selection in content space before later symbol binding (Wong et al., 7 Jan 2026). This suggests that future work may increasingly treat competitive modality or candidate selection as an explicit intermediate stage rather than as an incidental byproduct of end-to-end fusion.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Modality Competitive Query Selection (MCQS).