Papers
Topics
Authors
Recent
Search
2000 character limit reached

Chain-of-Descriptions (CoDes) for LLMs

Updated 3 July 2026
  • Chain-of-Descriptions (CoDes) is a prompting methodology that integrates natural-language descriptions as intermediate steps to improve task decomposition and multi-modal reasoning.
  • It significantly boosts performance, as demonstrated by measurable accuracy gains in audio, vision, and VHDL code generation benchmarks.
  • The approach transforms LLMs into agents that articulate their understanding, offering interpretable and actionable insights for complex task resolution.

Chain-of-Descriptions (CoDes) refers to a family of prompting methodologies for LLMs in which the model generates one or more explicit, natural-language descriptions as intermediate steps before producing a final output. This approach structurally decomposes the traditional direct inference process into staged, semantically meaningful transitions, typically from representation or plan formation to answer or example generation. Contemporary research demonstrates the effectiveness of Chain-of-Descriptions in multi-modal reasoning (e.g., vision, audio), code generation, and summarization tasks, where it provides measurable gains compared to standard prompting paradigms (Guo et al., 22 Feb 2025, Vijayaraghavan et al., 16 Jul 2025).

1. Formal Framework and Definitions

Chain-of-Descriptions (CoDes, Editor's term) systems introduce intermediate semantic descriptions between the input and the answer, in contrast to classical models that attempt the downstream task directly from the raw input. In its canonical form for multi-modal inputs, Chain-of-Description (CoD) prompting factors the joint output probability via an intermediate description step:

  • Standard Prompting (SP):

A∗=arg⁡max⁡A  p(A∣X,Q)A^* = \arg\max_A \; p(A \mid X, Q)

where XX is the multi-modal input, QQ the question, and AA the answer.

  • Chain-of-Description Prompting (CoD):

(D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)

Decomposition: - Step 1: D∗=arg⁡max⁡D  p(D∣X)D^* = \arg\max_D\;p(D\mid X) (description) - Step 2: A∗=arg⁡max⁡A  p(A∣D∗,X,Q)A^* = \arg\max_A\;p(A\mid D^*, X, Q) (answer)

For code synthesis, CoDes extends this by allowing a chain D=(d1,d2,...,dk)D = (d_1, d_2, ..., d_k) of descriptive steps (plans or explanations), which are inserted into the prompt before invoking the LLM for final code or summary generation (Vijayaraghavan et al., 16 Jul 2025):

di=arg⁡max⁡npLLM(d∣X,d1,...,di−1) O=arg⁡max⁡npLLM(o∣X,D′)d_i = \arg\max_n p_\text{LLM}(d \mid X, d_1, ..., d_{i-1}) \ O = \arg\max_n p_\text{LLM}(o \mid X, D')

where OO is the final output and XX0 is a refined chain of descriptions.

This staged inference transforms the LLM from a monolithic predictor into an agent that surfaces and conditions on intermediate understanding.

2. Algorithmic Workflows and Implementation Templates

Multi-Modal Reasoning:

The CoD pipeline is instantiated as follows (Guo et al., 22 Feb 2025):

  1. Prompt the LLM: "Describe the input XX1." Output XX2 as a natural-language description.
  2. Prompt: "Given the description XX3 and question XX4, produce your answer." Output XX5.

Code Generation and Summarization:

CoDes for VHDL-centric tasks (Vijayaraghavan et al., 16 Jul 2025) involves:

  • For Generation: From a problem statement XX6, generate a stepwise plan XX7, then feed this plan together with XX8 to prompt the model for code generation.
  • For Summarization: Given code lines XX9, prompt for a description QQ0 for each code line, aggregate them (QQ1), and finally prompt for a summary using both the original code and QQ2.

Pseudocode template for multi-modal CoD: D∗=arg⁡max⁡D  p(D∣X)D^* = \arg\max_D\;p(D\mid X)3 In code domains, the construction and refinement of the intermediate chain QQ3 is task-specific and can involve multi-pass model calls.

3. Empirical Evaluation and Benchmarks

Empirical results indicate that Chain-of-Descriptions yields consistent accuracy improvements across modalities and task types.

Multi-Modal (Audio, Vision)

Audio Analysis (Qwen2-Audio on AIR-Bench-Chat):

  • Evaluation metric: alignment score QQ4, where scores are assigned by LLMs (1–10 scale).
  • Notable results (speech category): Standard Prompting QQ5 vs. CoD QQ6 (QQ7 absolute).
  • Mean improvement: QQ8 across categories.
  • Gains scale with information density of the descriptions: speech (QQ9 tokens/sec) sees greater benefit than music (AA0 tokens/sec), confirming the utility of richer descriptive content (Guo et al., 22 Feb 2025).

Vision (Qwen2-VL/Qwen2.5-VL on MMMU_Pro):

  • Metric: multiple-choice accuracy (hard-level).
  • Hard-level accuracy: Standard Prompting AA1–AA2, CoD AA3–AA4, AA5 absolute improvement.
  • Description transfer ablation: using higher-capacity model (Qwen2.5-VL) descriptions with a smaller model (Qwen2-VL-7B) increases accuracy proportionally with description quality.

VHDL Code Generation and Summarization

VHDL-Eval (Generation):

  • Pass@1 (testbench): Baseline AA6, CoDes AA7 (AA8 relative).
  • Pass@1 (sequential equivalence): Baseline AA9, CoDes (D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)0 ((D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)1).
  • Self-Consistency: (D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)2 to (D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)3 ((D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)4).

VHDL-Xform (Summarization):

  • ROUGE-L: Baseline (D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)5, CoDes (D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)6 ((D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)7).
  • LLM Preference Rate: Baseline (D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)8, CoDes (D∗,A∗)=arg⁡max⁡D,Ap(D,A∣X,Q)=arg⁡max⁡D,Ap(D∣X)⋅p(A∣D,X,Q)(D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)9 (D∗=arg⁡max⁡D  p(D∣X)D^* = \arg\max_D\;p(D\mid X)0).

Execution Strategy Impact: Multi-step (separate plan and execution) outperforms single-step prompts by up to D∗=arg⁡max⁡D  p(D∣X)D^* = \arg\max_D\;p(D\mid X)1–D∗=arg⁡max⁡D  p(D∣X)D^* = \arg\max_D\;p(D\mid X)2 relative on key benchmarks (Vijayaraghavan et al., 16 Jul 2025).

4. Modeling Rationale and Mechanistic Insights

The key principle underlying Chain-of-Descriptions is articulated as “What I can understand, I can put into words.” Forcing the model to verbalize its internal latent representations serves to:

  • Align multi-modal or code embeddings with surface language,
  • Encourage comprehensive semantic attention to input details,
  • Supply an explicit, human-interpretable context for difficult reasoning or generation tasks.

Empirical ablations reveal that:

  • Description quality tightly correlates with downstream task accuracy.
  • Benefit is maximized on high information-density tasks (complex multi-modal content, functionally intensive code).
  • The approach is model-agnostic and works in a zero-shot fashion (no model retraining required) (Vijayaraghavan et al., 16 Jul 2025).

A plausible implication is that CoDes narrows the gap between the model's internal "understanding" and the human-assessable surface forms required for effective task completion.

5. Limitations, Trade-offs, and Failure Modes

Chain-of-Descriptions is associated with several trade-offs:

  • Inference Overhead: The two-stage or multi-step pipeline increases latency and token cost; in VHDL domains, CoDes incurs 2–3× inference overhead (Vijayaraghavan et al., 16 Jul 2025).
  • Description Quality Dependency: The model’s overall output is constrained by the fidelity and coverage of the intermediate descriptions or plans. Poor or misaligned steps lead to misguidance.
  • Scalability: Line-by-line or stepwise explanations become unwieldy for large codebases or high-resolution multi-modal inputs.
  • Dataset Scope: Empirical validations have primarily targeted Qwen-family multi-modal models and Granite-Code-34B for VHDL; broader generalization requires further study (Guo et al., 22 Feb 2025, Vijayaraghavan et al., 16 Jul 2025).
  • Distracting Irrelevance: Verbose or non-salient descriptions may attenuate gains on easier tasks, potentially distracting the model from essential information.

6.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Chain-of-Descriptions (CoDes).