---
title: Chain-of-Descriptions (CoDes) for LLMs
url: https://www.emergentmind.com/topics/chain-of-descriptions-codes
type: topic
---

# Chain-of-Descriptions (CoDes) for LLMs

Chain-of-Descriptions (CoDes) refers to a family of prompting methodologies for large language models (LLMs) in which the model generates one or more explicit, natural-language descriptions as intermediate steps before producing a final output. This approach structurally decomposes the traditional direct inference process into staged, semantically meaningful transitions, typically from representation or plan formation to answer or example generation. Contemporary research demonstrates the effectiveness of Chain-of-Descriptions in multi-modal reasoning (e.g., vision, audio), code generation, and summarization tasks, where it provides measurable gains compared to standard prompting paradigms [2502.16137, 2507.12308].

## 1. Formal Framework and Definitions

Chain-of-Descriptions (CoDes, *Editor's term*) systems introduce intermediate semantic descriptions between the input and the answer, in contrast to classical models that attempt the downstream task directly from the raw input. In its canonical form for multi-modal inputs, Chain-of-Description (CoD) prompting factors the joint output probability via an intermediate description step:

- **Standard Prompting (SP):**
  $$
  A^* = \arg\max_A \; p(A \mid X, Q)
  $$
  where $X$ is the multi-modal input, $Q$ the question, and $A$ the answer.

- **Chain-of-Description Prompting (CoD):**
  $$
  (D^*,A^*) = \arg\max_{D, A} p(D, A \mid X, Q) = \arg\max_{D, A} p(D \mid X) \cdot p(A \mid D, X, Q)
  $$
  Decomposition:
  - Step 1: $D^* = \arg\max_D\;p(D\mid X)$ (description)
  - Step 2: $A^* = \arg\max_A\;p(A\mid D^*, X, Q)$ (answer)

For code synthesis, CoDes extends this by allowing a *chain* $D = (d_1, d_2, ..., d_k)$ of descriptive steps (plans or explanations), which are inserted into the prompt before invoking the LLM for final code or summary generation [2507.12308]:
$$
d_i = \arg\max_n p_\text{LLM}(d \mid X, d_1, ..., d_{i-1}) \\
O = \arg\max_n p_\text{LLM}(o \mid X, D')
$$
where $O$ is the final output and $D'$ is a refined chain of descriptions.

This staged inference transforms the LLM from a monolithic predictor into an agent that surfaces and conditions on intermediate understanding.

## 2. Algorithmic Workflows and Implementation Templates

**Multi-Modal Reasoning:**
The CoD pipeline is instantiated as follows [2502.16137]:
1. Prompt the LLM: "Describe the input $X$." Output $D$ as a natural-language description.
2. Prompt: "Given the description $D$ and question $Q$, produce your answer." Output $A$.

**Code Generation and Summarization:**
CoDes for VHDL-centric tasks [2507.12308] involves:
- **For Generation:** From a problem statement $P$, generate a stepwise plan $D'$, then feed this plan together with $P$ to prompt the model for code generation.
- **For Summarization:** Given code lines $C = [c_1, ..., c_n]$, prompt for a description $d_i$ for each code line, aggregate them ($D'$), and finally prompt for a summary using both the original code and $D'$.

Pseudocode template for multi-modal CoD:
```python
def CoD_Inference(X, Q, model):
    # Step 1: Descriptive stage
    prompt_desc = "Input: <X>\nPlease provide a detailed description of the above input."
    D = model.generate(prompt_desc)
    # Step 2: Reasoning/Answer stage
    prompt_ans = "Description: " + D + "\nQuestion: " + Q + "\nAnswer the question based on the description."
    A = model.generate(prompt_ans)
    return A
```
In code domains, the construction and refinement of the intermediate chain $D'$ is task-specific and can involve multi-pass model calls.

## 3. Empirical Evaluation and Benchmarks

Empirical results indicate that Chain-of-Descriptions yields consistent accuracy improvements across modalities and task types.

### Multi-Modal (Audio, Vision)

**Audio Analysis (Qwen2-Audio on AIR-Bench-Chat):**
- Evaluation metric: alignment score $r = s_p / s_{gt}$, where scores are assigned by LLMs (1–10 scale).
- Notable results (speech category): Standard Prompting $91.24\%$ vs. CoD $95.02\%$ ($+3.78\%$ absolute).
- Mean improvement: $+1.79\%$ across categories.
- Gains scale with *information density* of the descriptions: speech ($id=3.91$ tokens/sec) sees greater benefit than music ($id=2.52$ tokens/sec), confirming the utility of richer descriptive content [2502.16137].

**Vision (Qwen2-VL/Qwen2.5-VL on MMMU_Pro):**
- Metric: multiple-choice accuracy (hard-level).
- Hard-level accuracy: Standard Prompting $16.67\%$–$18.18\%$, CoD $21.97\%$–$23.48\%$, $+5.30\%$ absolute improvement.
- Description transfer ablation: using higher-capacity model (Qwen2.5-VL) descriptions with a smaller model (Qwen2-VL-7B) increases accuracy proportionally with description quality.

### VHDL Code Generation and Summarization

**VHDL-Eval (Generation):**
- Pass@1 (testbench): Baseline $19.2\%$, CoDes $25.4\%$ ($+32\%$ relative).
- Pass@1 (sequential equivalence): Baseline $18.7\%$, CoDes $24.6\%$ ($+32\%$).
- Self-Consistency: $23.5\%$ to $26.5\%$ ($+13\%$).

**VHDL-Xform (Summarization):**
- ROUGE-L: Baseline $38.6$, CoDes $40.6$ ($+5\%$).
- LLM Preference Rate: Baseline $35.7\%$, CoDes $40.0\%$ ($+12\%$).

**Execution Strategy Impact:** Multi-step (separate plan and execution) outperforms single-step prompts by up to $20$–$30\%$ relative on key benchmarks [2507.12308].

## 4. Modeling Rationale and Mechanistic Insights

The key principle underlying Chain-of-Descriptions is articulated as *“What I can understand, I can put into words.”* Forcing the model to verbalize its internal latent representations serves to:
- Align multi-modal or code embeddings with surface language,
- Encourage comprehensive semantic attention to input details,
- Supply an explicit, human-interpretable context for difficult reasoning or generation tasks.

Empirical ablations reveal that:
- Description quality tightly correlates with downstream task accuracy.
- Benefit is maximized on high information-density tasks (complex multi-modal content, functionally intensive code).
- The approach is model-agnostic and works in a zero-shot fashion (no model retraining required) [2507.12308].

A plausible implication is that CoDes narrows the gap between the model's internal "understanding" and the human-assessable surface forms required for effective task completion.

## 5. Limitations, Trade-offs, and Failure Modes

Chain-of-Descriptions is associated with several trade-offs:
- **Inference Overhead:** The two-stage or multi-step pipeline increases latency and token cost; in VHDL domains, CoDes incurs 2–3× inference overhead [2507.12308].
- **Description Quality Dependency:** The model’s overall output is constrained by the fidelity and coverage of the intermediate descriptions or plans. Poor or misaligned steps lead to misguidance.
- **Scalability:** Line-by-line or stepwise explanations become unwieldy for large codebases or high-resolution multi-modal inputs.
- **Dataset Scope:** Empirical validations have primarily targeted Qwen-family multi-modal models and Granite-Code-34B for VHDL; broader generalization requires further study [2502.16137, 2507.12308].
- **Distracting Irrelevance:** Verbose or non-salient descriptions may attenuate gains on easier tasks, potentially distracting the model from essential information.

## 6.

Source: https://www.emergentmind.com/topics/chain-of-descriptions-codes