---
title: 'Donut: A Multifaceted Scientific Signifier'
url: https://www.emergentmind.com/topics/donut
type: topic
---

# Donut: A Multifaceted Scientific Signifier

Searching arXiv for recent and foundational papers on “DONUT” across domains to ground the article in published work.
The term **DONUT** and its lowercase variant **Donut** are used across contemporary research to denote several distinct objects: annular visual encodings, beam and point-spread-function designs, physical effects, machine-learning systems, synthetic benchmarks, and bibliographic infrastructures. In the literature, the same label identifies an OCR-free document model, a variational anomaly detector, a physics-aware diffraction network, a decoder-only trajectory forecaster, a DNS fingerprinting tool, a synthetic 3D-topology benchmark, and multiple non-acronymic “donut” structures in geometry and physics [2111.15664] [1802.03903] [2507.14038] [2506.06854].

## 1. Terminological scope and acronymic reuse

In acronymic usage, DONUT is not a single framework but a recurrent naming pattern. The expansions documented in the literature span document understanding, diffraction analysis, trajectory prediction, software fingerprinting, topology benchmarking, and topology bibliometrics.

| Usage | Expansion or meaning | Domain |
|---|---|---|
| Donut | Document understanding transformer | Visual document understanding |
| DONUT | Diffraction with Optics for Nanobeam by Unsupervised Training | X-ray nanodiffraction |
| DONUT | Decoder-Only Network for Unrolling Trajectories | Motion forecasting |
| DONUT | Domain Oriented Network Unmasking Tool | DNS-based software fingerprinting |
| DONUT | Dataset Of maNifold strUcTures | 3D topology benchmark |
| DONUT | Database of Original Non-Theoretical Uses of Topology | TDA bibliography |

These acronymic forms coexist with descriptive uses in which “donut” refers to a ring-like structure or annular pattern: a concentric-ring visualization for Spatial Social Networks, the vortex excitation beam in MINFLUX, the ring-like angular distribution in proton channeling, annular supergranular flow structures on the Sun, and a rectangle-with-hole construction in discrete geometry [2101.00929] [2410.03349] [1008.2629] [2301.07988] [2406.01147].

## 2. Geometric abstractions: concentric summaries and mathematical donuts

In Spatial Social Networks, the donut is a summary glyph rather than a node-link drawing. Sarkar and Yadav define it as a concentric-ring visualization centered at the viewport centroid, partitioned into eight equal angular sectors and three concentric distance buckets—Near, Medium, and Far—so that each sector-ring cell encodes the count of edges by direction and scale [2101.00929]. The construction aggregates the visible network into an \(8\times3\) matrix \(Count[d,b]\), annotates the center with the number of nodes, and uses fixed sector angles \(\Delta \theta = 2\pi/8 = \pi/4\) with ring radii \(R_b = (b/3)\cdot R_{\max}\). The same routine is defined for both a network-level donut, using all nodes and edges in the current map extent, and a regional-level donut, recomputed on a spatial subset. Its stated advantages are simultaneous encoding of topology and spatial structure, multi-scale reuse, and avoidance of “hairball” clutter; its stated limitations are loss of individual-edge detail, dependence on sector and bucket design, and the burden of color-only magnitude encoding [2101.00929].

A different formalization appears in discrete geometry. A mathematical donut is an ordered quadruple \(D=(a,b,x,y)\) of positive integers such that the outer rectangle of side lengths \(a,b\) contains an inner rectangle of side lengths \(x,y\), with the area constraint
\[
ab=2xy.
\]
The hole is twistable precisely when the \(90^\circ\)-rotated inner rectangle also fits, and the paper gives the necessary and sufficient criterion \(2x>a\) and \(2y>a\) [2406.01147]. For square donuts \((n,n,a,b)\), existence is equivalent to \(n\) being the sum of the three positive entries of some Pythagorean triple, with the complete parametrization
\[
n = 2kp(p+q),\qquad a = 2kp^2,\qquad b = k(p+q)^2,
\]
for \(p>q>0\), \(\gcd(p,q)=1\), and \(p,q\) not both odd [2406.01147]. This places the “donut” in an arithmetic setting rather than a radial one.

## 3. Annular patterns in optics, imaging, channeling, and solar dynamics

In MINFLUX localization microscopy, the donut is the conventional vortex excitation beam. It is obtained by choosing a pupil-plane phase \(\kappa(k_x,k_y)=\operatorname{atan2}(k_y,k_x)\), yielding an intensity PSF proportional to
\[
h(r)\propto \left[\frac{J_1(k_{\max}r)}{k_{\max}r}\right]^2,
\]
with a central zero and a first Bessel-ring pattern [2410.03349]. Under equal-shape constraints on all excitation beams, numerical optimization with 15 Zernike modes and many random restarts converges to the donut, and the theoretical analysis identifies the vortex mask \(P(k)=e^{j\phi}\) as the unique global maximizer of the isotropic gradient-norm criterion [2410.03349]. Quantitatively, for \(NA=1.42\), \(\lambda=640\,\mathrm{nm}\), \(L=100\,\mathrm{nm}\), \(N=100\) photons, and background \(b\approx5\%\), donut-only MINFLUX with \(K=7\) translated copies achieves \(\sigma_{\mathrm{CRB}}\approx3.0\,\mathrm{nm}\) at the center and \(4\)–\(5\,\mathrm{nm}\) on the full field of view, whereas two pairs of half-moon beams reach \(\approx1.5\,\mathrm{nm}\) centrally and \(\approx2.5\,\mathrm{nm}\) on average, corresponding to an approximately \(2\times\) improvement in \(\sigma_{\mathrm{CRB}}\) and a reduction from \(\sim100\) to \(\sim25\) photons for \(3\,\mathrm{nm}\) precision [2410.03349].

A related annular construction appears in phase-based stimulated emission depletion magnetic particle imaging. There, a spatially varying relaxation time induces a nonlinear phase lag \(\phi_n(r)=\arctan[n\omega_0\tau(r)]\), and the imaginary part of the harmonic PSF has a central null:
\[
PSF_D(r_\perp)\equiv G_n^I(r_\perp).
\]
Subtracting this donut-shaped focal spot from the Lorentzian focal spot defines
\[
PSF_{STED}(r_\perp)=PSF_L(r_\perp)-k\,PSF_D(r_\perp),
\]
and the reported focal-spot size reduction is up to \(4\times\) beyond the Langevin magnetization resolution barrier [2505.06490].

In proton channeling through an \((11,9)\) single-wall carbon nanotube, the donut is a ring-like angular distribution that develops as the proton incident angle \(\phi\) approaches the critical channeling angle \(\psi_c\) [1008.2629]. The paper identifies the effect with rainbow scattering, governed by singularities of the mapping from entrance-plane coordinates to scattering angles, i.e. the vanishing of the Jacobian \(J_\Theta\). Numerically, \(\psi_c \simeq 10.6\,\mathrm{mrad}\) with the image interaction and \(\psi_c \simeq 11.9\,\mathrm{mrad}\) without it; at \(\phi=10\,\mathrm{mrad}\), approximately \(\psi_c\), a full ring of high yield appears at radius \(\approx10\,\mathrm{mrad}\) [1008.2629].

Solar-physics usage is again descriptive. In 6–12 h temporal averages of the photospheric horizontal-flow modulus, Roudier et al. identify nearly circular rings of enhanced speed and call each such ring a donut [2301.07988]. The azimuthally averaged profile is modeled as
\[
u_h(r)=u_0\frac{r}{R}\exp\!\left[-\tfrac12(r/R)^2\right],
\]
with \(R\approx11\,\mathrm{Mm}\) and \(u_0\approx650\,\mathrm{m\,s^{-1}}\) [2301.07988]. Reported radii are \(10.6\)–\(11\,\mathrm{Mm}\), lifetimes range from \(2\) to \(55\,\mathrm{h}\) with a peak around \(24\,\mathrm{h}\), and the structures are largely absent from magnetized plage regions [2301.07988]. The paper interprets them as the most active convective cells associated with supergranulation and shows, through cork advection, that the strongest donuts alone can account for much of quiet-Sun magnetic-flux diffusion [2301.07988].

## 4. Donut and DONUT as machine-learning models

In visual document understanding, Donut denotes the **Document understanding transformer**, an OCR-free encoder-decoder model that maps a raw document image directly to a target token sequence convertible to JSON [2111.15664]. Its visual encoder is Swin-Transformer-B with stage depths \(L=\{2,2,14,2\}\), hidden dimensions \(d=\{96,192,384,768\}\), and attention heads \(h=\{3,6,12,24\}\), while the textual decoder is a 4-layer BART-style cross-attention decoder with \(d_{\mathrm{model}}=1024\), \(h=16\), \(d_{ff}=4096\), and learned positional embeddings for up to 1,536 decoding steps [2111.15664]. Pre-training is cast as text reading with cross-entropy
\[
L(\theta)=-\sum_{t=1}^T \log p_\theta(y_t\mid y_{<t},Z),
\]
and SynthDoG contributes approximately \(0.5\) M synthetic pages per language; combined with \(11\) M IIT-CDIP English scans, the pre-training corpus comprises \(2\times10^6\) synthetics plus \(11\times10^6\) reals [2111.15664]. Reported downstream results include \(95.30\%\) on RVL-CDIP at \(752\,\mathrm{ms}\) per image, CORD \(84.1\%\) F1 and \(90.9\%\) TED-Acc at \(1.2\,\mathrm{s}\), Ticket \(94.1\%\) F1 and \(98.7\%\) TED-Acc at \(0.6\,\mathrm{s}\), and DocVQA \(67.5\%\) ANLS, with \(72.1\%\) on the handwritten subset [2111.15664]. A domain application to construction specification tables of contents fine-tunes a public Donut base model on 200 annotated pages, using 180 for fine-tuning and 20 for testing, and reports field-wise accuracies \(0.92\) for heading number, \(0.78\) for heading title, \(0.91\) for subheading number, \(0.76\) for subheading title, and \(0.85\) on average [2403.07553]. A later compression study analyzes Donut’s decoder circuits, identifies \(C_3\) as the sole locus of transcription and \(M_3\) as negligible, and derives Donut-MINT variants at \(31\%\), \(18\%\), and \(7\%\) decoder budgets; on DocVQA, the \(7\%\) variant reports \(51\%\) ANLS, \(27\%\) EM, \(4.5\) M parameters, and \(90\,\mathrm{ms}\) latency, versus \(66\%\) ANLS, \(35\%\) EM, 257 M parameters, and \(250\,\mathrm{ms}\) for the teacher [2509.26235].

In time-series anomaly detection, Donut is an unsupervised VAE for seasonal KPIs in web applications [1802.03903]. The model uses fully connected encoder and decoder networks with two hidden layers of 100 ReLU units each, a Gaussian latent prior \(p(z)=\mathcal N(0,I)\), and sliding windows of length \(W=120\) [1802.03903]. Its central methodological additions are a modified ELBO that de-weights missing and anomalous points,
\[
\widetilde{\mathcal L}(x)=\mathbb E_{q_\phi(z\mid x)}\Big[\sum_{w=1}^W \alpha_w\log p_\theta(x_w\mid z)+\beta\log p(z)-\log q_\phi(z\mid x)\Big],
\]
missing-data injection with \(\lambda=1\%\), and MCMC-based imputation with \(M=10\) steps at inference [1802.03903]. On three KPI datasets, its best F-scores range from \(0.75\) to \(0.9\), outperforming both a supervised ensemble baseline and a baseline VAE; the paper also introduces a KDE interpretation of reconstruction probability [1802.03903].

In X-ray science, DONUT means **Diffraction with Optics for Nanobeam by Unsupervised Training**, a physics-aware autoencoder for scanning X-ray nanodiffraction microscopy [2507.14038]. A 2D CNN encoder maps each \(64\times64\) diffraction image to a latent vector \(z=(\epsilon,w,\chi)\), representing relative lattice strain, in-plane tilt, and out-of-plane tilt, while two heads consume \(z\): a mirror-symmetric CNN decoder \(D(z)\) and a fully differentiable forward scattering model \(f(z;p)\) [2507.14038]. Training uses only the raw diffraction images through the composite loss
\[
L=\frac1N\sum_i \Big[ W_{dec}\,\|x_i-D(E_0(x_i))\|_1 + W_{phys}\,\|x_i-f(E_0(x_i);p)\|_1\Big],
\]
with \(W_{dec}=1\) and \(W_{phys}=5\) [2507.14038]. The simulated dataset spans a \(41\times41\times41\) grid over \((\epsilon,w,\chi)\) for 68,921 diffraction patterns, while the experimental dataset is one \(165\times165\) scan with approximately 27,000 patterns [2507.14038]. Inference with the encoder alone runs at approximately \(0.024\pm0.001\,\mathrm{ms/frame}\) on GPU and \(0.27\pm0.07\,\mathrm{ms/frame}\) on CPU, compared with \(5.6\pm0.4\,\mathrm{ms/frame}\) for conventional fitting, and experimental agreement with correlation-library fitting is summarized by Pearson \(r(\epsilon)=0.88\), \(r(w)=0.96\), and \(r(\chi)=0.89\) [2507.14038].

## 5. Forecasting, speech, telemetry, and robotic manipulation

DONUT also names a decoder-only trajectory forecaster. **Decoder-Only Network for Unrolling Trajectories** tokenizes trajectories into sub-trajectories of \(T_{sub}=10\) time steps, embeds them in dimension \(D=128\), and autoregressively predicts future motion without a separate encoder [2506.06854]. Its factorization is
\[
p(Y_{1:T_{fut}}\mid X_{1:T_{hist}})=\prod_{i=1}^{\lceil T_{fut}/T_{sub}\rceil} p\big(Y_{(i-1)T_{sub}+1:iT_{sub}}\mid X,\hat Y_{1:(i-1)T_{sub}}\big),
\]
and training adds an auxiliary overprediction branch with horizon \(T_o=2T_{sub}\) [2506.06854]. On the Argoverse 2 single-agent benchmark, the reported hidden-test performance is \(b\)-minFDE\(_6 = 1.79\,\mathrm{m}\), minFDE\(_6 = 1.16\,\mathrm{m}\), minADE\(_6 = 0.63\,\mathrm{m}\), and MR\(_6 = 0.14\) [2506.06854].

In speech technology, DONUT is a CTC-based query-by-example wakeword detector for personalized keyword spotting [1811.10736]. Enrollment uses three user-recorded examples, a pre-trained 3-layer unidirectional GRU label model, beam search to generate phoneme-sequence hypotheses, and weighted aggregation of CTC forward scores during inference [1811.10736]. The model size is 168 k parameters. On English-Fewshot, reported EERs are \(7.8\%\), \(7.3\%\), and \(3.7\%\) across the three negative conditions shown in Table 1, versus \(24.2\%\), \(32.4\%\), and \(19.3\%\) for DTW on FBANK features [1811.10736].

In network telemetry, DONUT denotes the **Domain Oriented Network Unmasking Tool**, a rule-based system for DNS-based software fingerprinting [2208.07042]. It processes passively monitored DNS traffic through a parser, a packet-level CF-Matcher, and a host-level Set-Matcher mediated by an IP dictionary [2208.07042]. CF-rules match domain patterns and query types to labels, while Set-rules infer applications from required and optional label sets. On the performance dataset, DONUT processes approximately 15,000 packets per second, corresponding to 1,000,000 packets in 67 s; on a four-VM validation set with optimized parameters, it achieved zero false positives and zero false negatives, giving precision, recall, and F1 approximately \(1.0\) [2208.07042].

In robotics, “Make a Donut” names a zero-shot deformable-manipulation system that uses a large language model for high-level stage decomposition and EMD-space planning for low-level control [2311.02787]. For the donut task, the LLM specifies stages such as flattening dough into a disc, shaping a ring, and punching the hole, along with Python code that generates subgoal point clouds [2311.02787]. The low-level planner iteratively applies gradient steps in Earth Mover’s Distance space,
\[
p_i' = p_i - \alpha \frac{\partial}{\partial p_i}\mathrm{EMD}(\{p_i\},\{\bar p_j\}),
\]
and then optimizes tool actions through differentiable physics with a point-to-point loss \(L_{P2P}=\sum_i\|p_i'-p_i\|_1\) [2311.02787]. Reported donut performance is a normalized EMD-decrease score of \(0.346\) with \(75\%\) success, and the paper states that the method surpasses multiple baselines in dough manipulation without demonstrations [2311.02787].

## 6. DONUT as benchmark and bibliographic infrastructure for topology

Two further usages shift from algorithms to research infrastructure. In geometric deep learning, DONUT is the **Dataset Of maNifold strUcTures**, introduced to test whether 3D point-cloud encoders capture global topology [2604.22334]. The benchmark contains 29,517 distinct 3D manifold meshes with \(\beta_0\in\{1,\dots,6\}\) and total genus \(g\in\{0,\dots,10\}\), sampled nearly uniformly over that grid [2604.22334]. Each mesh is built from superquadrics, \(k\)-tori, or cones; 1,024 surface points are sampled and normalized to the unit sphere, and ground-truth \(H_1\) persistence diagrams are computed with GUDHI v3.11.0 using an \(\alpha\)-filtration, հետո thresholded to retain the top 10% most persistent pairs [2604.22334]. On the DONUT test split of 5,938 shapes, FILTR predicts persistence diagrams with best reported \(W_2 \simeq 0.16\), \(d_B \simeq 0.0098\), and PIE \(\simeq 1.11\), while linear probes on frozen pretrained encoders achieve only modest performance on \(\beta_0\) and \(g\) prediction, with best values of approximately \(57\%\) and \(26\%\), respectively [2604.22334].

In topological data analysis bibliometrics, DONUT means **Database of Original Non-Theoretical Uses of Topology**, a curated database of papers on practical TDA applications [2304.12417]. The project began from a 2017 group chat associated with the HIM Spring School and later migrated from a shared Google Sheet to a Zotero library and then to a public search interface [2304.12417]. By late 2022, the catalog had grown to 431 entries, of which 58 were labeled *innovate* [2304.12417]. Each entry carries at least one tag in each of three classes—area of application, mathematical tools, and input type—and an optional flavor label such as *confirm* or *innovate* [2304.12417]. The maintenance pipeline exports BibTeX through Pyzotero, indexes it with Xapian, and serves it through a Flask web front end [2304.12417].

## 7. Cross-disciplinary regularities

Across these usages, no single technical definition of DONUT is canonical. Instead, the name repeatedly marks either an annular geometry or a compact, end-to-end system with a domain-specific acronym. The annular cases emphasize central-null or ring-shaped structure—concentric sector summaries in Spatial Social Networks, vortex and depletion PSFs in microscopy and magnetic particle imaging, rainbow rings in channeling, and annular supergranular flows on the Sun [2101.00929] [2410.03349] [2505.06490] [1008.2629] [2301.07988]. The acronymic cases emphasize integrated pipelines: direct pixels-to-JSON document understanding, unsupervised seasonal-KPI anomaly detection, physics-aware diffraction inference, decoder-only motion forecasting, rule-based DNS software inference, and curated topology resources [2111.15664] [1802.03903] [2507.14038] [2506.06854] [2208.07042] [2304.12417].

This pattern suggests that “DONUT” functions less as a stable concept than as a reusable scientific signifier. In some fields it denotes literal ring structure; in others it denotes methodological compactness, architectural closure, or a memorable acronym. A plausible implication is that the term persists because it accommodates both visually intuitive geometry and concise naming, allowing unrelated communities to adopt it without semantic conflict.

Source: https://www.emergentmind.com/topics/donut