---
title: 'Cell2Sentence: Transformer-Based Single-Cell Modeling'
url: https://www.emergentmind.com/topics/cell2sentence-c2s
type: topic
---

# Cell2Sentence: Transformer-Based Single-Cell Modeling

Cell2Sentence (C2S) is a single-cell modeling framework that treats each cell as a “sentence” whose “words” are genes, allowing transformer-based language-model machinery to operate on ordered gene lists rather than only on conventional expression matrices. In the reported literature, C2S appears both as a decoder-only single-cell foundation model for cell-type annotation and as a semantic embedding mechanism that injects NCBI-derived gene knowledge into clustering pipelines. A subsequent interpretability study applies transcoders to C2S in order to extract sparse internal “decision circuits,” with the stated aim of tracing final predictions back to interpretable features and ultimately to gene tokens [2509.14723], [2606.13007].

## 1. Sentence formulation and biological semantics

The central premise of C2S is to serialize each cell’s highly expressed genes into a token sequence. In the decoder-only formulation, each assayed cell’s top-expressed genes are turned into a 128-token “sentence,” and downstream tasks are predicted via the final token’s representation [2509.14723]. In the clustering-oriented formulation used in scLLM-DSC, C2S views each cell as a “sentence” whose “words” are gene tokens drawn from a curated vocabulary of gene symbols and summaries, and each cell is represented by its top-$K$ genes with $K=2048$ [2606.13007].

In the scLLM-DSC setting, vocabulary construction begins with the set $G=\{g_1,\ldots,g_M\}$ of all protein-coding genes in the organism. For each gene $g_j$, NCBI provides its official symbol, a one-line functional summary, known pathways, and related fields; these are concatenated into a structured prompt,
$$
T_j = \text{“Gene: [symbol}_j\text{]. Summary: [text}_j\text{]. Pathways: […].”}
$$
The gene-level vocabulary $V$ consists of all gene symbols plus a small set of common “function” tokens such as “kinase” and “transcription\_factor,” derived from the summaries [2606.13007].

Sentence formation is explicitly rank-based. For each cell $i$, one takes its raw or log-normalized expression $x_i\in\mathbb{R}^D$, selects the top-$K$ genes by descending expression magnitude, sorts them so that $x_i[\mathrm{idx}_i[1]]\ge\cdots\ge x_i[\mathrm{idx}_i[K]]$, and forms
$$
S_i = [\mathrm{sym}(g_{\mathrm{idx}_i[1]}),\mathrm{sym}(g_{\mathrm{idx}_i[2]}),\ldots,\mathrm{sym}(g_{\mathrm{idx}_i[K]})].
$$
Only gene symbols, and not numeric counts, are used as tokens so that the language model attends to functional relationships rather than to explicit abundance values [2606.13007].

This published usage indicates that C2S is not reducible to a single fixed network specification. A plausible implication is that the unifying object is the cell-to-sequence transformation itself: expression profiles are converted into ordered gene-token strings so that pretrained or adapted transformer components can exploit gene co-occurrence, ordering, and external biological knowledge.

## 2. Decoder-only C2S as a single-cell foundation model

One reported implementation adapts the decoder-only “Pythia” transformer to the single-cell domain [2509.14723]. At its core, this C2S model consists of a learned token embedding for each gene, with vocabulary size $\approx 50{,}000$, plus standard positional embeddings for up to 128 ranks. The transformer stack contains $N=24$ decoder layers. Each layer comprises a multi-head self-attention block with $H=16$ heads, hidden dimension $d_{\mathrm{model}}=1024$, key/query/value projections of size $d_k=d_v=64$, and an output projection back to $d_{\mathrm{model}}$; it also contains a two-layer feed-forward network with expansion factor $f=4$, that is, $1024\rightarrow4096\rightarrow1024$, with each linear layer followed by a GELU nonlinearity [2509.14723].

For supervised cell-type annotation, the model uses a lightweight classification head: one learned linear projection followed by a softmax applied to the final “:” token hidden state of size $1024$ to predict one of $C$ cell-type labels. Fine-tuning minimizes the standard cross-entropy loss
$$
\mathcal{L}_{\mathrm{C2S}}=-\sum_{c=1}^{C} y_c \log p_\theta(c\mid \mathbf{h}_{\mathrm{final}}),
$$
where $\mathbf{h}_{\mathrm{final}}\in\mathbb{R}^{1024}$ is the last-token hidden state and $y$ is the one-hot cell-type label [2509.14723].

Pre-training the tokenizer and decoder stack on 57 million cells and biological abstracts is described as instilling both gene-expression and literature-derived co-occurrence statistics [2509.14723]. In that formulation, C2S is explicitly positioned as a state-of-the-art single-cell foundation model, and the later interpretability work is motivated by the observation that its decision-making processes are less interpretable than traditional methods such as differential gene expression analysis [2509.14723].

A key technical motivation for the interpretability pipeline is that the MLP sublayers within each transformer block are characterized as highly polysemantic and densely connected. Directly tracing how individual gene tokens drive a final prediction is therefore presented as intractable in the raw model, which motivates replacing each MLP with a sparse, separately trained transcoder [2509.14723].

## 3. Transcoder approximation of MLP sublayers

Following Dunefsky et al. (NeurIPS ’24), the interpretability study trains a separate transcoder for each MLP block of C2S [2509.14723]. For a single MLP block $m$, with input $\mathbf{x}\in\mathbb{R}^{1024}$ and output $\mathrm{MLP}_m(\mathbf{x})\in\mathbb{R}^{1024}$, the corresponding transcoder is a wide, sparsely activating autoencoder that approximates the original two-layer feed-forward function while exposing interpretable “features.”

The encoder is
$$
\mathbf{z}=\mathrm{ReLU}(W_{\mathrm{enc}}\mathbf{x}+\mathbf{b}_{\mathrm{enc}}),
$$
with $W_{\mathrm{enc}}\in\mathbb{R}^{8192\times1024}$, corresponding to expansion factor $8$, and producing a nonnegative activation vector $\mathbf{z}\in\mathbb{R}_{\ge 0}^{8192}$. The decoder is
$$
\hat{\mathbf{y}}=W_{\mathrm{dec}}\mathbf{z}+\mathbf{b}_{\mathrm{dec}},
$$
with $W_{\mathrm{dec}}\in\mathbb{R}^{1024\times8192}$, recovering the original output dimension [2509.14723].

Each transcoder minimizes
$$
\mathcal{L}_{\mathrm{trans}}
=
\left\lVert \mathrm{MLP}_m(\mathbf{x})-\hat{\mathbf{y}}\right\rVert_2^2
+
\lambda \lVert \mathbf{z}\rVert_1,
\qquad
\lambda=1.4\times 10^{-4}.
$$
The first term enforces faithful function approximation, and the second enforces sparsity so that only a small subset of the 8192 features activate per input [2509.14723].

Training is reported at the granularity of 60 million input tokens per transcoder, with peak learning rate $10^{-4}$, batches of 2048 tokens, and the $L_1$ regularization coefficient above [2509.14723]. Validation on the Heart Cell Atlas v2 yielded the following comparison:

- **Original MLP**: validation loss 2.48  
- **Transcoder-replaced MLP**: 4.63  
- **MLPs removed entirely**: 12.67  
- **KL divergence**: $\mathrm{KL}(\mathrm{orig}\parallel\mathrm{trans})\approx 2.406$ versus $\mathrm{KL}(\mathrm{orig}\parallel\mathrm{noMLP})\approx 10.52$

These values are presented as confirming that transcoders capture the vast majority of MLP functionality in a sparse, interpretable substrate [2509.14723].

## 4. Decision-circuit extraction and attribution formalism

Once the transcoders are trained, circuit extraction proceeds by tracing contributions from final predictions back through sparse transcoder features and attention-mediated token interactions [2509.14723]. For a target prediction such as the final token’s classification logits, the method first identifies the top transcoder feature or features in the last layer that contributed most strongly, measured by the product of the feature’s activation $z^{(L,i)}(\mathbf{x})$ and its decoder-row norm.

The recursion then follows two pathways. The first is the MLP approximation pathway, in which connections from feature $i$ in layer $l$ to feature $j$ in layer $l+1$ are captured by fixed decoder–encoder dot products,
$$
A_{(l,i)\to(l+1,j)}
=
f^{(l,i)}_{\mathrm{dec}}\cdot f^{(l+1,j)}_{\mathrm{enc}},
$$
which are input-invariant weights. Combined with the input-dependent activations $z^{(l,i)}(\mathbf{x})$, these provide quantitative attribution [2509.14723].

The second pathway runs through attention heads. Information flows across tokens via each head’s OV matrices and attention weights $\alpha_{s\to t}$. By linearizing each head’s update, the method obtains attributions from source token $s$ to transcoder feature $i$ in the next layer [2509.14723]. This produces a hybrid graph in which sparse MLP features and prompt tokens are both explicit objects of analysis.

To manage graph size, the procedure prunes to retain only the top-$K$ strongest paths at each step and then merges paths into a sparse computational subgraph termed the “decision circuit.” All retained edges are collected into an adjacency matrix $\mathbf{A}$ whose nodes are transcoder features $(l,i)$ and prompt tokens $t$. A simplified logical-gate view is written as
$$
\mathrm{output}_k \approx
\sigma\!\Bigl(
\sum_{(l,i)} w_{k,(l,i)}\,\mathbb{I}[z^{(l,i)}(\mathbf{x})>\tau]
\Bigr),
$$
where $\mathbb{I}$ thresholds sparse feature activations, $\tau$ is a small activation cutoff, and $\sigma$ is the classifier nonlinearity [2509.14723].

The significance of this construction is methodological rather than merely visual. It provides a principled attribution route from a final label back to layer-specific sparse features and, through attention-mediated transport, to the gene tokens that seeded the prediction.

## 5. Biological correspondence and interpretability evidence

A case study on an artery endothelial cell illustrates the intended biological reading of a C2S decision circuit [2509.14723]. The extracted circuit targeted the most strongly activated last-layer feature, identified as feature ID 3353. A distilled subcircuit highlighted three gene tokens driving the prediction of “endothelial cell of artery”:

- **VWF**: von Willebrand factor, described as the canonical Weibel–Palade body resident marker.  
- **PTPRB**: VE-PTP, described as a junctional tyrosine phosphatase critical for vascular integrity.  
- **SPARCL1**: hevin, described as an IBD-associated matricellular protein that stabilizes quiescent endothelium.

The study states that the presence of these genes, each with well-characterized roles in endothelial biology, among the top contributors to the predicted label demonstrates that transcoder-based circuit tracing recovers biologically plausible regulatory modules [2509.14723]. This is the central empirical claim linking internal C2S features to recognizable biology.

The work also reports a broader interpretability statistic based on human evaluation of layer 12 features. Specifically, 35% of the “live” features—defined as those with $\log_{10}E(f)\ge -4$, where $E(f)$ is the per-token activation probability—correspond to gene-level tokens in a semantically coherent way [2509.14723]. This is contrasted with the original model’s opaque neuron activations.

At the same time, limitations are stated directly. Current circuits can still be large and require manual feature-to-biology mapping [2509.14723]. A plausible implication is that C2S interpretability, in this form, is mechanistic but not fully automated: sparse structure makes analysis possible, yet domain expertise remains necessary to assign biological meaning to individual features and subgraphs.

## 6. Cross-modal semantic alignment and clustering

In scLLM-DSC, C2S is used as a semantic component within a cross-modal deep structural clustering framework rather than as a decoder-only predictor [2606.13007]. The framework establishes a semantically grounded representation by synergizing two views: a Knowledge-Driven Semantic View derived from NCBI gene priors and contextualized Cell2Sentence embeddings, and a Structure-Aware Topological View extracted via a graph-guided encoder. The stated motivation is that traditional scRNA-seq clustering collapses each gene to a mere index and is therefore semantically agnostic.

The semantic branch uses a frozen transformer-based text encoder $E_{\mathrm{LLM}}(\cdot)$, for example OpenAI text-embedding-3-small, mapping token sequences or prompts into $\mathbb{R}^{d_1}$. Gene prompts are embedded as
$$
g_j = E_{\mathrm{LLM}}(T_j)\in\mathbb{R}^{d_1},
$$
and cell sentences are embedded as
$$
Z_i^{(2)} = E_{\mathrm{LLM}}(S_i)\in\mathbb{R}^{d_1}.
$$
Internally, $E_{\mathrm{LLM}}$ is described as a multi-layer Transformer with self-attention, layer normalization, and a final pooling head, with weights kept frozen [2606.13007].

In parallel, the method constructs an abundance-weighted semantic embedding
$$
Z^{(1)} = \tilde{X}\cdot\tilde{G},
$$
so that for cell $i$,
$$
Z_i^{(1)}=\sum_{k=1}^{K}\tilde{X}_{i,k}\cdot g_{\mathrm{idx}_i[k]}.
$$
The two semantic views are fused as
$$
Z_i^{\mathrm{text}}=\omega\cdot Z_i^{(1)}+(1-\omega)\cdot Z_i^{(2)},\qquad 0\le \omega\le 1.
$$
This fused semantic representation is then aligned to a structural embedding $Z^{\mathrm{feat}}=f_{\mathrm{struc}}(X)\in\mathbb{R}^{N\times d_2}$ extracted by the scCDCG autoencoder+graph-cut backbone [2606.13007].

Alignment is implemented with two projection heads $f_\phi$ and $g_\psi$, each a Linear→ReLU→Linear MLP, which map the semantic and structural views into a shared $d$-dimensional space. After $\ell_2$ normalization of each row, the cross-modal similarity matrix is
$$
S = (\hat{Z}^{\mathrm{text}})(\hat{Z}^{\mathrm{feat}})^\top/\tau.
$$
The model minimizes a symmetric InfoNCE loss over a minibatch of $N$ cells together with a variance regularization term, producing
$$
\mathcal{L}_{\mathrm{align}}=\mathcal{L}_{\mathrm{CL}}+\lambda\mathcal{L}_{\mathrm{var}}.
$$
In the reported configuration, the structural encoder uses scCDCG with layer sizes $[256\rightarrow16]$, so $d_2=16$; the projection heads output $d=128$; the contrastive temperature is $\tau=0.1$; the variance weight is $\lambda=10^{-2}$; and the fusion weight is $\omega=0.5$ [2606.13007].

The evaluation uses six public scRNA-seq benchmarks from scCluBench: Mauro Pancreas, Sonya Liver, Sapiens Liver, Muris Brain, Muris Liver, and Muris Limb Muscle. Preprocessing uses no additional filtering or batch correction, and always selects the 2048 highest-expressed genes per cell. Training follows a two-stage procedure: pre-train the autoencoder with $\mathcal{L}_{\mathrm{NCut}}+\mathcal{L}_{\mathrm{MSE}}$ for 200 epochs, then jointly fine-tune all modules with the full objective
$$
\mathcal{L}=\alpha\mathcal{L}_{\mathrm{align}}+\beta\mathcal{L}_{\mathrm{NCut}}+\gamma\mathcal{L}_{\mathrm{MSE}}+\delta\mathcal{L}_{\mathrm{KL}},
$$
with $\alpha=\beta=\gamma=\delta=1$, for another 200 epochs, using Adam with learning rate $10^{-3}$, five seeds, and an NVIDIA A800 GPU [2606.13007].

The reported outcome is that scLLM-DSC significantly outperforms eleven state-of-the-art baselines in clustering accuracy [2606.13007]. In this context, C2S functions as the mechanism that converts priority-sorted gene lists into natural-language-like sequences so that a frozen LLM can infuse external biological knowledge into the latent space. Taken together with the transcoder study, the published record presents C2S as both a representational interface between transcriptomes and language models and a substrate for mechanistic interrogation of single-cell foundation models [2509.14723], [2606.13007].

Source: https://www.emergentmind.com/topics/cell2sentence-c2s