---
title: 'TextlessRAG: Speech-Based Document QA'
url: https://www.emergentmind.com/topics/textlessrag
type: topic
---

# TextlessRAG: Speech-Based Document QA

Searching arXiv for the specified paper and closely related visual/pixel-space RAG work to ground the article.
TextlessRAG is a fully “textless” speech-based Retrieval–Augmented Generation (RAG) pipeline over a large collection of document images. It is presented as the first end-to-end framework for speech-based question answering over large-scale document images, with the central design choice of eliminating ASR, TTS and OCR from the retrieval pipeline and instead directly interpreting speech, retrieving relevant visual knowledge, and generating answers over multimodal document content [2509.07538]. The system is paired with SV-DOC, described as the first bilingual speech–document RAG dataset featuring Chinese and English voice queries aligned with document images, and it is evaluated as a document-image QA system rather than a conventional text-centric RAG stack.

## 1. Definition and research scope

TextlessRAG addresses a problem formulation that had not previously been explored in the cited work: knowledge base question answering over visual document images with queries provided directly in speech [2509.07538]. The knowledge base consists of page images rather than extracted text, and the query modality is raw audio rather than typed text. The intended effect is to preserve document layout, charts, tables, pictures, and other visually grounded structures that are typically degraded by OCR or linearization.

The framework is explicitly positioned against prior methods that depend on cascades such as ASR \(\rightarrow\) text retrieval, OCR \(\rightarrow\) text indexing, or text-conditioned multimodal answer generation. In contrast, TextlessRAG maps speech and document images into a shared embedding space, retrieves relevant pages, refines them with layout-aware reranking, and answers with a multimodal decoder [2509.07538].

A plausible implication is that TextlessRAG belongs to a broader family of text-abstraction-free RAG systems. In the same data block, PixelRAG is described as performing retrieval and reading entirely in pixel space over web screenshots rather than document images [2606.28344]. Read together, these systems suggest a common research direction in which the native visual form of the knowledge source is retained rather than converted into text.

## 2. End-to-end architecture

The pipeline is organized into three stages: a speech encoder, a visual document retrieval and layout reranking stage, and a generator module [2509.07538].

In the first stage, raw user query audio \(q\) is mapped to a fixed-length speech embedding \(e_q\). The retrieval encoder is described as reusing “ColQwen-Omni,” a multimodal encoder capable of ingesting both document screenshot images and raw audio. Each page image \(P_i\) in the document knowledge base \(\mathcal{I}=\{P_1 \ldots P_n\}\) is pre-encoded into an image embedding \(e_i\).

Retrieval is based on MaxSim scoring in a shared \(d\)-dimensional embedding space, where the reported dimensionality is \(d=768\). The page relevance score is written as

$$
s_i \;=\; \mathrm{Score}(P_i, q)
\;=\;\mathrm{MaxSim}(e_q, e_i)
\;=\;\sum_{u \in e_q}\max_{v \in e_i} \langle u,\,v\rangle .
$$

Top-\(k\) retrieval selects a candidate set \(\mathcal{T}_k\) of pages for subsequent refinement. The page embeddings are pre-computed and stored in an approximate nearest-neighbor index, with FAISS IVFFlat given as the example index structure, and the query-time probe is described as taking \(O(\log n)\) time [2509.07538].

The generator stage then consumes the refined visual evidence \(\{P_t'\}_{t=1}^k\) together with the speech representation \(e_q\). The reported model is a transformer-based decoder, “Qwen2.5-Omni,” which accepts image inputs and a speech prefix and outputs an answer.

## 3. Retrieval refinement and multimodal decoding

A distinctive component of TextlessRAG is the layout-aware reranking mechanism. After the top-\(k\) pages are retrieved, each page \(P_t\) is decomposed into fine-grained blocks of four types, \(C=\{\text{chart, table, text, image}\}\), using DocLayout-YOLO [2509.07538]. The decomposition is written as

$$
\mathrm{YOLO}(P_t)
= \bigcup_{c\in C} \mathrm{Block}^c(P_t).
$$

For each block \(b\), the system computes a block-level speech–image similarity score \(s_b=\mathrm{MaxSim}(e_q,\mathrm{Enc\_image}(b))\), then keeps only the blocks that exceed a threshold \(\theta\):

$$
\mathrm{Block}^c_\theta
= \{\,b\in\mathrm{Block}^c(P_t)\mid s_b\ge\theta\}.
$$

The surviving blocks are sorted in descending order of \(s_b\) and packed into a refined page representation \(P_t'\). The stated rationale is that block-level reranking allows the model to focus on the most relevant sub-regions, such as a specific table or chart, rather than passing an entire page to the generator.

The generator module is described as a transformer-based multimodal decoder that takes a linearized sequence of image patch embeddings from the reranked blocks and a prefix embedding derived directly from the raw speech query. Internally, the decoder cross-attends to both audio tokens and image patch tokens. The token prediction distribution is given as

$$
y_t
= \mathrm{softmax}\bigl(W_o\,[h_t;\,\mathrm{Attn}(h_t,\,E_I\cup E_Q)]\bigr),
$$

where \(h_t\) is the decoder hidden state, \(E_I\) is the set of image patch embeddings, and \(E_Q\) is the set of audio token embeddings [2509.07538].

A common point of confusion concerns the term “textless.” The abstract states that the framework eliminates ASR, TTS and OCR in the pipeline, while the generator description states that the decoder produces a text answer, which is then vocoded back into speech, and that this TTS step can be bypassed if a text response is sufficient [2509.07538]. This suggests that “textless” is used primarily to describe the internal interpretation, retrieval, and evidence-processing pipeline, rather than to prohibit textual answer output.

## 4. Training objectives

TextlessRAG is trained end-to-end with three loss components: retrieval loss, reranking loss, and generation loss [2509.07538]. The retrieval objective is a contrastive loss over a training pair \((q, P^+)\) and sampled negatives \(P^-\), encouraging the positive page to score higher than negative pages under the MaxSim retrieval score. The reranking objective applies a similar contrastive principle at the block level, promoting relevant blocks of the positive page over blocks from negatives or low-score blocks below \(\theta\).

The generation component is standard cross-entropy over the ground-truth answer sequence \(y_{1:|T|}\):

$$
\mathcal{L}_{\mathrm{gen}}
= -\sum_{t=1}^{|T|}
\log p_\theta(y_t\mid y_{<t},\,E_I,\,E_Q).
$$

The total objective is a weighted combination of retrieval, reranking, and generation losses, with the weights \(\lambda\) set by validation [2509.07538]. The reported training formulation therefore couples retrieval quality, local evidence selection, and answer generation in a single optimization scheme rather than treating them as independently tuned stages.

## 5. SV-DOC benchmark

SV-DOC is described as the first bilingual (English+Chinese) speech–document retrieval-augmented QA benchmark [2509.07538]. It combines six English subsets derived from existing text/image QA corpora, each augmented with high-quality TTS voices, with a new Chinese Document RAG dataset, CDR.

| Dataset | #QA pairs | Pool size |
|---|---:|---:|
| ChartQA | 150 | 119 |
| InfoVQA | 1,048 | 300 |
| SlideVQA | 760 | 657 |
| DUDE | 496 | 422 |
| MMLong | 1,091 | 5,134 |
| Vidoseek | 1,142 | 5,349 |
| CDR | 1,260 | 30,583 |
| Total | 5,947 | 42,564 |

The English subsets are ChartQA, DUDE, InfoVQA, SlideVQA, MMLong, and Vidoseek. CDR is described as a new Chinese Document RAG dataset of 1,260 QA pairs over a 30,583-page pool spanning charts, tables, text blocks, and images. Across the benchmark, content types are denoted as \(C/T/I/X\), corresponding to charts, tables, images, and text blocks, and the domains include academic, infographics, slides, and open-domain settings [2509.07538].

The evaluation protocol uses nDCG@5 for retrieval, GPT-4o automatic answer scoring for QA, and end-to-end latency measured on a single NVIDIA A100-80 GB. The paper states that both the dataset and the pipeline will be made available at the listed repository [2509.07538].

## 6. Empirical performance and ablation analysis

Retrieval performance is reported in terms of nDCG@5 and compared against text-based baselines such as BM25, E5, and NV-Embed, as well as vision-based systems including CLIP, DSE, VisRAG-Ret, VDocRAG, and ViDoRAG [2509.07538]. TextLessRAG achieves 99.3 on ChartQA versus ViDoRAG at 100.0, 91.5 on DUDE versus 96.5, 91.6 on InfoVQA versus 97.8, 94.2 on SlideVQA versus 96.9, 66.5 on MMLong versus 67.0, 95.4 on Vidoseek versus 94.3, and 87.4 on CDR versus 87.7. The paper summarizes these results by noting that, even though the input is raw speech, performance is within 2 points of ViDoRAG on most subsets and surpasses it on Vidoseek.

QA performance is measured under top-5 retrieved pages and gold-page settings, with and without layout reranking. Under the top-5 raw-page condition, TextLessRAG reports 87.3 on ChartQA, 78.5 on DUDE, 74.5 on InfoVQA, 79.7 on SlideVQA, 33.4 on MMLong, 90.2 on Vidoseek, and 43.5 on CDR [2509.07538]. Adding layout reranking yields gains of \(+3\)–\(5\) points on InfoVQA, SlideVQA, MMLong, CDR, and Vidoseek. The gold-page upper bound peaks at 84.0 on DUDE, 80.6 on InfoVQA, 81.8 on SlideVQA, and 61.3 on CDR. The paper highlights two cases in particular: on InfoVQA, TextlessRAG with layout reranking reaches 79.4 and surpasses all baselines, and on Vidoseek with reranking it reaches 93.4.

Latency is reported end-to-end on one A100, excluding I/O overhead. The pure text pipeline, described as OCR+ASR+RAG+TTS, is approximately 2.3 s/query; ViDoRAG, described as OCR+RAG+TTS, is approximately 1.9 s; and TextlessRAG, described as direct speech encoding + image retrieval + generation, is approximately 1.2 s/query [2509.07538]. The paper therefore attributes a roughly 40–50% latency reduction to eliminating ASR/OCR/TTS while preserving near-SOTA accuracy.

The ablation results sharpen the interpretation of the architecture. Without layout reranking, QA accuracy drops by 3–7 points across most datasets, with examples including InfoVQA from 79.4 to 74.5 and SlideVQA from 82.6 to 79.7. Replacing direct speech with ASR\(\rightarrow\)text retrieval degrades retrieval nDCG by 2–5 points and increases end-to-end latency by 30–50%. OCR-based text retrieval using BM25 or E5 yields nDCG as low as 40–60% on chart- and table-rich pages [2509.07538]. These ablations support the claim that both direct speech conditioning and layout-aware evidence selection are functional rather than merely stylistic design choices.

## 7. Relation to pixel-native RAG and broader implications

TextlessRAG is focused on spoken queries over visual document collections, but the data block places it adjacent to other non-textual RAG formulations. PixelRAG, for example, is described as a retrieval-augmented method that represents websites in their native visual form and performs retrieval and reading entirely in pixel space, operating over a datastore of 30 million screenshot images and reporting gains over text-based baselines on tasks such as NQ, SimpleQA, MMSearch, LiveVQA, and MoNaCo [2606.28344]. A plausible implication is that TextlessRAG and PixelRAG instantiate two complementary variants of the same general thesis: the knowledge source may remain visual, and the query or reader stack need not force an early conversion into text.

Within that framing, TextlessRAG contributes three specific elements: speech-query retrieval over document images, a layout-aware reranking stage built on DocLayout-YOLO, and a bilingual benchmark that includes Chinese and English voice queries [2509.07538]. Its stated conclusion is that fully text-free RAG over visually rich document collections is feasible, and that the released SV-DOC benchmark and CDR dataset will catalyze further research in speech-driven visual QA and RAG.

The results do not imply that textual methods are uniformly obsolete. The reported comparisons show near-state-of-the-art retrieval and QA accuracy rather than universal dominance, and the strongest empirical claims are dataset-specific: surpassing ViDoRAG on Vidoseek retrieval, surpassing all baselines on InfoVQA under reranking, and materially reducing latency relative to OCR- and ASR-dependent alternatives [2509.07538]. The significance of TextlessRAG lies less in a categorical rejection of text than in establishing a concrete end-to-end alternative for document-image QA in which speech, layout, and visual evidence are modeled directly.

Source: https://www.emergentmind.com/topics/textlessrag