---
title: SwimBird-SFT-92K Multimodal CoT Dataset
url: https://www.emergentmind.com/topics/swimbird-sft-92k
type: topic
---

# SwimBird-SFT-92K Multimodal CoT Dataset

SwimBird-SFT-92K is a large-scale, curated dataset comprising 92,300 chain-of-thought (CoT) multimodal reasoning examples, explicitly designed to enable dynamic, input-adaptive switching among three distinct reasoning patterns—text-only, vision-only, and interleaved vision-text—within Multimodal Large Language Models (MLLMs). Developed to support and evaluate the SwimBird reasoning-switchable MLLM architecture, this dataset systematically integrates textual and visual reasoning trajectories and enables fine-grained supervision of query-conditioned reasoning modality [2602.06040].

## 1. Dataset Composition

SwimBird-SFT-92K (|SwimBird-SFT-92K| = 92,300) consists of samples stratified by reasoning mode as follows:
- **Text-only (T):** $N_T = 50{,}000$ (54.2%)
- **Vision-only (V):** $N_V = 8{,}800$ (9.5%)
- **Interleaved (I):** $N_I = 33{,}500$ (36.3%)

The dataset sources and their mode-wise contributions are shown in the table below:

| Source                | Total  | Text-only | Vision-only | Interleaved  |
|-----------------------|--------|-----------|-------------|--------------|
| Zebra-CoT             | 26,300 | 0         | 5,900       | 20,400       |
| ThinkMorph            | 7,100  | 0         | 1,200       | 5,900        |
| MathCanvas-Instruct   | 8,900  | 0         | 1,700       | 7,200        |
| OpenMMReasoner        | 50,000 | 50,000    | 0           | 0            |

Text-only data stems from OpenMMReasoner and targets general VQA and math question answering, while vision-only and interleaved modalities predominantly focus on tasks such as visual search, jigsaw puzzles, mazes, geometry, chart reasoning, and spatial navigation, drawing from Zebra-CoT, ThinkMorph, and MathCanvas-Instruct.

## 2. Construction and Reasoning-Mode Curation Pipeline

The curation pipeline for SwimBird-SFT-92K proceeds through three delineated stages:

**Stage 1: Candidate Collection and Filtering**
All candidate multimodal CoT samples ($S_0$) are aggregated from Zebra-CoT, ThinkMorph, and MathCanvas. For each example $x$, pass@8 accuracy is evaluated via Qwen3VL-8B on the original question-image pair ($\mathrm{pass}_\mathrm{base}(x)$). Trivial ("easy") samples—where $\mathrm{pass}_\mathrm{base}(x) = 1.0$—are filtered out, yielding $S_1$.

**Stage 2: Reasoning-Mode Assignment**
For $x \in S_1$, pass@8 is computed in the presence of intermediate "thinking" images ($\mathrm{pass}_\mathrm{hint}(x)$). Only examples with $\mathrm{pass}_\mathrm{hint}(x) \geq \mathrm{pass}_\mathrm{base}(x)$ are retained. Modes are then assigned by:
\[
m(x)=
\begin{cases}
\text{Vision-only}, & \mathrm{pass}_\mathrm{hint}(x) \geq 0.75 \\
\text{Interleaved}, & \mathrm{pass}_\mathrm{base}(x) < \mathrm{pass}_\mathrm{hint}(x) < 0.75
\end{cases}
\]
This produces $N_V + N_I = 42,300$ high-quality multimodal samples.

**Stage 3: Integration of Text-Only CoT**
$N_T = 50,000$ text-only CoT sequences from OpenMMReasoner, already filtered by pass@8, are incorporated. The final dataset thus comprises 92,300 documents, with empirical mode frequencies $p(m) = \left\{0.542, 0.095, 0.363\right\}$ for $m \in \{\text{T}, \text{V}, \text{I}\}$.

No explicit reweighting is applied—mode proportions reflect source and filtering outcomes.

## 3. Annotation, Quality Control, and Verification

- **Reasoning Paths:** Textual chains are sourced directly from human- or model-curated CoT traces in contributing datasets. Visual "hints"—intermediate images portraying the reasoning process—are also inherited from source datasets (e.g., diagram crops, annotated sketches).
- **Verification:** An instance is accepted only if $\mathrm{pass}_\mathrm{hint}(x) \geq \mathrm{pass}_\mathrm{base}(x)$, ensuring that intermediate hints do not degrade answerability. Label correctness is assessed automatically using Qwen3-235B-Instruct by comparing prediction against the gold answer.
- **Quality Gating:** The pass@8 metric is employed for exclusion of low-quality or irrelevant CoT samples. No manual re-annotation is required, as automated filters achieve high-precision instance selection [2602.06040].

## 4. Statistical and Structural Properties

- **Chain Lengths:** Mean textual chain length is $\bar{T} \approx 42$ tokens for text-only samples (σ ≈ 15) and $\bar{T} \approx 26$ tokens (σ ≈ 12) for interleaved examples.
- **Visual Token Budget:** Each instance enforces a variable latent-token span ($K$) with $N_\mathrm{min} = 2$ and $N_\mathrm{max} = 32$. Vision-only samples average $\bar K \approx 18$ (σ ≈ 8); interleaved samples $\bar K \approx 14$.
- **Resolution Scaling:** High-resolution inputs yield $K$ near $N_\mathrm{max}$; low-resolution inputs near $N_\mathrm{min}$.
- **Mode Distribution Across Benchmarks:** For DynaMath and MathVerse_MINI, text-only constitutes ~90% of instances; for V* Bench and HR-Bench (4K/8K), the split is approximately 40% vision-only, 30% interleaved, and 30% text-only, indicating adaptive mode selection by benchmark.

## 5. Integration in Supervised Fine-Tuning

- **Base Model:** Qwen3-VL-8B (encoder-decoder) with frozen vision encoder.
- **SFT Objective:** Hybrid autoregressive loss combining next-token prediction for text and next-embedding prediction for visual latents:
  \[
  \mathcal{L} = \sum_{t=1}^T -\log p_\theta(w_t \mid w_{<t}, x_{\rm img}) + \lambda_{\rm vis} \sum_{k=1}^K \|\hat z_k - z_k\|_2^2
  \]
  with $\lambda_{\rm vis} = 0.2$, $\lambda_{\rm text} = 1.0$.
- **Training Regime:** Batch size 128, cosine learning rate schedule (initial LR $=1\times 10^{-5}$), on A100–80G GPUs. During inference, the dynamic latent span is terminated by emission of the </latent> token.

## 6. Comparative Context and Performance Impact

SwimBird-SFT-92K is the first multimodal CoT dataset constructed to supervise three reasoning patterns (text-only, vision-only, and interleaved) in a balanced, query-adaptive fashion. Its scale and diversity surpass predecessors such as Zebra-CoT (26.3K), ThinkMorph (7.1K), and MathCanvas-Instruct (8.9K), as well as text-only OpenMMReasoner (50K). Prior datasets are limited to fixed reasoning patterns or static latent budgets, precluding dynamic, sample-adaptive reasoning.

Empirically, models fine-tuned on SwimBird-SFT-92K demonstrate state-of-the-art results across multiple benchmarks, with absolute improvements of +1.2 (V* Bench), +2.0 (HR-Bench 4K), +3.6 (HR-Bench 8K), +6.5 (MMStar), +10.7 (WeMath), and +1.9 (DynaMath) over prior best methods. This suggests that explicit tri-mode supervision and dynamic latent allocation substantially enhance robustness and efficacy on both textual and vision-intensive tasks [2602.06040].

In summary, SwimBird-SFT-92K constitutes a principled, high-quality, and large-scale dataset supporting the development of MLLMs capable of flexible, context-driven multimodal reasoning.

Source: https://www.emergentmind.com/topics/swimbird-sft-92k