Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sub-Question Decomposition

Updated 30 April 2026
  • Sub-question decomposition is the process of breaking down a complex query into self-contained, simpler sub-questions that collectively address the original query.
  • It employs techniques such as LLM prompting, supervised sequence-to-sequence models, and structural parsing to isolate distinct facets and reasoning steps.
  • Applications include multi-hop question answering, retrieval augmentation, and multi-table reasoning, yielding significant empirical performance gains.

Sub-question decomposition is the process of transforming a complex, multi-faceted, or compositional question into a set of simpler sub-questions whose answers collectively address the original information need. This technique has become central in modern machine learning and information retrieval pipelines, particularly for tasks—such as multi-hop question answering (QA), video-language understanding, fact verification, retrieval-augmented generation (RAG), numerical and multi-table reasoning, and knowledge-based question answering—where direct, holistic reasoning is infeasible for current models due to limitations in token capacity, data coverage, transparency, or reasoning depth.

1. Formal Definitions and Algorithmic Foundations

Sub-question decomposition formalizes the mapping from a complex query QQ to a set (or sequence, or graph) of sub-questions {q1,...,qn}\{q_1, ..., q_n\}, each designed to isolate a distinct facet, reasoning step, or piece of evidence. The process can be captured algebraically as a function DD,

D(Q)={qi}i=1nD(Q) = \{q_i\}_{i=1}^n

where each qiq_i is ideally “self-contained” and collectively, the qiq_i’s are sufficient to answer QQ (Huang et al., 9 Oct 2025, Xie et al., 2024, Ammann et al., 1 Jul 2025).

Algorithmic approaches include:

2. Integration into Downstream Inference Pipelines

Once decomposed, sub-questions are incorporated into various pipeline architectures:

The architectural interfaces vary. Some maintain a strict API- or prompt/LLM-driven separation (e.g., D-CoDe has zero architectural modification, using only prompt engineering (Huang et al., 9 Oct 2025)), while others opt for parameter sharing or explicit hard-EM/training objectives tying decomposition and answer generation (Zhu et al., 2023, Guo et al., 2022).

3. Empirical Benefits and Evaluation

Across domains, sub-question decomposition yields significant gains:

4. Representational and Structural Variants

Decomposition can yield:

The decomposition can be deterministic (e.g., AMR-QDAMR graph segmentation), heuristic (hand-written splitting and templating (Min et al., 2019, Tang et al., 2020)), or learned (via LLMs, step-wise prompting, or neural sequence models).

5. Practical Limitations and Challenges

Despite empirical robustness, sub-question decomposition introduces several costs and open problems:

  • Latency and Cost: Each sub-question entails a separate model (often LLM) invocation, leading to inference time and compute increases of up to 5–6x vs. baseline (Huang et al., 9 Oct 2025, Ammann et al., 1 Jul 2025, Xie et al., 2024).
  • Over-decomposition and Noise: Generation can yield trivial, redundant, or off-topic sub-questions; noisy decompositions can propagate errors, particularly in open-ended or static queries (Huang et al., 9 Oct 2025, Xie et al., 2024).
  • Prompt and Model Sensitivity: Performance and compositional coverage are heavily influenced by prompt design, LLM choice/hallucinations, and the absence of explicit loss functions or coverage objectives in many frameworks (Ammann et al., 1 Jul 2025, Huang et al., 9 Oct 2025).
  • Ambiguity in Facet Coverage: Determining when decomposition is necessary, how many sub-questions to generate, and weighting their importance remains challenging. Equal weighting may not reflect user or task priorities (Xie et al., 2024, Zhang et al., 2024).
  • Domain Adaptation and Supervision: Gold decompositions (for training or evaluation) are expensive; cross-domain generalization, especially for implied sub-questions or schema-sensitive decompositions, is imperfect without substantial annotated corpora (Chen et al., 2022, Guo et al., 2022, Eyal et al., 2023).
  • End-to-end Trainability: Many architectures remain non-differentiable or are not amenable to explicit backpropagation of downstream losses through decomposition—limiting the capacity to jointly optimize sub-question quality relative to final performance (Huang et al., 9 Oct 2025, Zhu et al., 2023, Guo et al., 2022).

6. Theoretical Guarantees and Open Problems

Formally, for any composite task that admits a polynomial-depth, constant-fanin decomposition, intermediate supervision—i.e., training with sub-task labels—renders otherwise unlearnable problems tractable for standard sequence models (polynomial-time SGD convergence) (Wies et al., 2022). This provides theoretical grounding for the empirical success of chain-of-thought and decomposed reasoning approaches, demonstrating why tasks that are end-to-end unlearnable (e.g., random parities, deep composition) become learnable when decomposed and supervised at finer granularity.

However, these results rely on full sub-task annotation, teacher forcing at training time, and the existence of low-degree decompositions. Robustness to noisy, incomplete, or model-generated sub-questions at test time, and extension to problems without such efficient decompositions, are active areas of research.

7. Domain-Specific Adaptations and Successes

Sub-question decomposition exhibits broad applicability:

In all domains, decomposition is both a tool for engineering more robust models and a critical axis for evaluating compositional generalization, logical faithfulness, and coverage. The future trajectory will likely feature more integrated, learnable, and domain-adaptive decomposition modules, new evaluation metrics favoring compositional transparency, and hybrid architectures balancing explicit structure with the fluency of large generative models.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (20)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Sub-Question Decomposition.