---
title: Topological Orthogonality Overview
url: https://www.emergentmind.com/topics/topological-orthogonality
type: topic
---

# Topological Orthogonality Overview

Topological orthogonality denotes several constructions in which orthogonality is induced, constrained, or interpreted by topological data. In one line of work it is an orthogonality relation \(\perp_{(\mathcal T,\mathcal F,A_X)}\) on a real vector space equipped only with a topology and a family of continuous scalar-valued functions; in another it is a relation on subsets derived from closure, proximity, uniformity, or coarse structure; elsewhere it appears as the vanishing of cosine similarity between persistence diagrams, as a topological mechanism for bounding orthogonal representations of graphs, and as a dynamical non-correlation condition such as Möbius orthogonality. Current arXiv usage therefore suggests a family of mathematically distinct notions linked by the common role of topology in certifying disjointness, incompatibility, or decoupling [1910.11751] [1803.09154] [2504.04361] [2110.00718] [2604.21392].

## 1. Topology-induced orthogonality in vector spaces

The paper "Orthogonality in a vector space with a topology And a generalization of Bhatia-Semrl Theorem" introduces an orthogonality relation on an arbitrary real vector space \(X\) equipped with a topology \(\mathcal T\), without requiring that \(\mathcal T\) make \(X\) a topological vector space [1910.11751]. The construction uses three ingredients: the topology \(\mathcal T\), a family \(\mathcal F\) of \(\mathbb R\)-valued \(\mathcal T\)-continuous functions, and a \(p\)-admissible set \(A_X\), where \(p\) is the projective equivalence relation on \(X\setminus\{0\}\) defined by
\[
u\,p\,v \iff u=\lambda v \text{ for some }\lambda\in\mathbb R\setminus\{0\}.
\]
A subset \(A_X\subset X\) is \(p\)-admissible if it contains exactly one nonzero vector from each \(p\)-equivalence class.

For \(u,v\in A_X\), the relation
\[
u \perp_{(\mathcal T,\mathcal F,A_X)} v
\]
holds if there exists \(f\in\mathcal F\) such that \(f(u)=\sup\{|f(z)|:z\in A_X\}\) and \(f(\lambda v)=0\) for all \(\lambda\in\mathbb R\). For arbitrary nonzero \(x,y\in X\), one declares
\[
x \perp_{(\mathcal T,\mathcal F,A_X)} y \iff a_X(x)\perp_{(\mathcal T,\mathcal F,A_X)} a_X(y),
\]
where \(a_X(x)\in A_X\) is the chosen representative of the line \(\mathbb Rx\). Everything is orthogonal to \(0\), and \(0\) is orthogonal to everything. The triple \((\mathcal T,\mathcal F,A_X)\) is called an orthogonality space.

A central result is that Birkhoff–James orthogonality is recovered as a special case. If \((X,\|\cdot\|)\) is a Banach space, \(\mathcal T\) is the norm topology, \(\mathcal F=S_{X^\ast}\), and \(A_X\subset S_X\) is a \(p\)-admissible slice of the unit sphere, then
\[
x \perp_{(\mathcal T_{\|\cdot\|},S_{X^\ast},A_X)} y
\iff
x \perp^{BJ} y,
\]
where \(x\perp^{BJ}y\) means \(\|x\|\le \|x+\lambda y\|\) for all \(\lambda\in\mathbb R\). The same recovery remains valid for the weak topology on a Banach space with \(\mathcal F=S_{X^\ast}\), and for perfectly normal topologies one may take \(\mathcal F\) to be all strictly-separating \(\mathcal T\)-continuous functions. At the opposite extreme, if \(\mathcal F\) contains the zero function, or if \(\mathcal T\) is discrete and \(\mathcal F\) is arbitrary, the induced relation is the trivial full relation \(x\perp y\) for all \(x,y\).

The paper also characterizes right additivity. Under the hypotheses that \(\mathcal F\) is a family of nonzero continuous linear functionals on \(X\) and no two members of \(\mathcal F\) are positive or negative multiples of one another,
\[
x\perp y \text{ and } x\perp z \Rightarrow x\perp (y+z)
\]
holds if and only if for each \(a\in A_X\) there is at most one \(f\in\mathcal F\) with \(f(a)=\sup\{|f(z)|:z\in A_X\}\). Specializing again to Banach spaces yields the classical equivalence between right additivity of Birkhoff–James orthogonality and smoothness of the space.

In finite-dimensional operator theory, the same framework yields a topological generalization of the Bhatia–Šemrl theorem. For \(T,A\in L(X,Y)\), with \(X\) finite-dimensional and \(Y\) topologized by finitely many seminorms \(p_1,\dots,p_m\), the paper characterizes orthogonality \(T\perp_{(P,S_{L(X,Y)}^\ast,A_{L(X,Y)})}A\) by the existence of \(x,y\in M_T\) and \(p,q\in P_T\) such that \(p(Tx)=q(Ty)=P(T)\), \(Ax\in Tx_p^+\), and \(Ay\in Ty_q^-\). The proof proceeds through an analogue of James’s lemma for \(u_p^\pm\) and a compactness-and-separation argument on \(A_{L(X,Y)}\). This places classical norm-based operator orthogonality inside a broader topological extremal-functional formalism.

## 2. Orthogonality relations on subsets and morphisms

A different tradition, developed by Dydak, treats orthogonality as a primitive relation on subsets of a set \(X\), and uses it to unify small-scale and large-scale geometry [1803.09154]. The starting point is a symmetric map
\[
\bullet:2^X\times 2^X\to 2^Y
\]
that is “bi-linear” in the sense that \(\varnothing\bullet X=\varnothing\) and \(C\bullet(D\cup E)=(C\bullet D)\cup(C\bullet E)\). When \(\bullet\) is basic, meaning that it takes only the values \(\varnothing\) and \(Y\), one defines
\[
C\perp D \iff C\bullet D=\varnothing.
\]
Conversely, any symmetric relation on subsets satisfying the corresponding monotonicity axioms determines such a basic dot-product.

Within this framework, the classical topological instance is
\[
C\perp_{\mathrm{top}} D
\quad:\Longleftrightarrow\quad
\operatorname{cl}(C)\cap \operatorname{cl}(D)=\varnothing,
\]
with dot-product
\[
C\bullet_{\mathrm{top}}D=\operatorname{cl}(C)\cap \operatorname{cl}(D).
\]
Here the Kuratowski closure operator \(\operatorname{cl}:2^X\to 2^X\) is viewed as an idempotent projection satisfying
\[
A\subset \operatorname{cl}(A),\qquad
\operatorname{cl}(\operatorname{cl}(A))=\operatorname{cl}(A),\qquad
\operatorname{cl}(A\cup B)=\operatorname{cl}(A)\cup \operatorname{cl}(B).
\]

Dydak also defines normal, or Tietze, orthogonality. If \(C\perp D\), normality requires the existence of \(C',D'\) with
\[
C\subset C',\quad D\subset D',\quad C'\cup D'=X,\quad C'\perp D,\quad D'\perp C.
\]
This enables a parallel–perpendicular decomposition analogous to linear algebra:
\[
X=P_C\cup P_C^\perp,
\]
where
\[
P_C=\{x\in X\mid \{x\}\not\perp C\},\qquad
P_C^\perp=\{x\in X\mid \{x\}\perp C\}.
\]
The significance of this viewpoint is that the same formalism captures topological orthogonality, proximity spaces, uniform spaces, and large-scale constructions such as metric coarse orthogonality, Higson-corona orthogonality, Gromov-hyperbolic orthogonality, and Freudenthal orthogonality. It also supports \(\perp\)-large-scale compactifications that recover the Čech–Stone compactification, Samuel–Smirnov compactification, Freudenthal compactification, Higson corona, and Gromov boundary.

A categorical reformulation appears in "A naive diagram-chasing approach to formalisation of tame topology" [1807.06986]. There orthogonality is Quillen-style lifting orthogonality of morphisms: for arrows \(f:A\to B\) and \(g:X\to Y\),
\[
f\perp g
\]
means that every commutative square with \(f\) on the left and \(g\) on the right admits a diagonal filler. Iterated left and right orthogonals of simple generating maps recover standard properties. For example, surjections are \((\varnothing\to\{\bullet\})^r\), injections are \((\{x,y\}\to\{\bullet\})^l\), connected spaces are characterized by orthogonality to the collapse map \(\{a,b\}\to\{a=b\}\), and similar constructions describe total disconnectedness, dense image, induced topology, \(T_0\), \(T_1\), Hausdorffness, and compactness. In that setting topological and uniform spaces are represented as simplicial objects in the category of filters. This suggests that, beyond subset disjointness, orthogonality can serve as an abstract logical operator encoding separation and extension principles.

## 3. Persistence diagrams and perfect topological dissimilarity

In topological data analysis, "On the cosine similarity and orthogonality between persistence diagrams" introduces an orthogonality notion for persistence diagrams based on persistence landscapes [2504.04361]. If \(D\) is a non-empty persistence diagram, its persistence-landscape transform is
\[
\phi:D\mapsto \lambda=(\lambda_1,\lambda_2,\dots),
\]
where each \(\lambda_j:\mathbb R\to\mathbb R\) is the \(j\)-th landscape layer. On the image of \(\phi\), the paper defines
\[
\langle \phi(D_1),\phi(D_2)\rangle
=
\sum_{j=1}^\infty \int_{\mathbb R}\lambda_{1j}(t)\lambda_{2j}(t)\,dt,
\]
\[
\|\phi(D)\|
=
\left(\sum_{j=1}^\infty \int_{\mathbb R}(\lambda_j(t))^2\,dt\right)^{1/2},
\]
and the cosine similarity
\[
\displaystyle
\ϝ(D_1,D_2)
=
\frac{\langle\phi(D_1),\phi(D_2)\rangle}
{\|\phi(D_1)\|\,\|\phi(D_2)\|}.
\]
By Cauchy–Schwarz, \(\ϝ(D_1,D_2)\in[0,1]\). Orthogonality is defined by
\[
D_1\perp D_2
\iff
\ϝ(D_1,D_2)=0
\iff
\langle \phi(D_1),\phi(D_2)\rangle=0.
\]

The paper proves an equivalent interval-disjointness criterion:
\[
\langle\phi(D_1),\phi(D_2)\rangle=0
\quad\Longleftrightarrow\quad
\forall\,(b_1,d_1)\in D_1,\,(b_2,d_2)\in D_2:\;
(b_1,d_1)\cap(b_2,d_2)=\emptyset.
\]
Thus orthogonality means that every open birth–death interval from one diagram is disjoint from every open birth–death interval from the other. The relation is symmetric and invariant under re-ordering of diagram points.

This orthogonality is stronger than separation by bottleneck or Wasserstein distances. If \(D_1\perp D_2\), then the trivial matching is a perfect matching for both the bottleneck distance \(W_\infty\) and the \(p\)-Wasserstein distance \(W_p\), and one obtains
\[
W_\infty(D_1,D_2)=\max\{W_\infty(D_1,\emptyset),W_\infty(D_2,\emptyset)\},
\]
\[
W_p(D_1,D_2)^p=W_p(D_1,\emptyset)^p+W_p(D_2,\emptyset)^p.
\]
At the same time, the paper emphasizes that \(W_\infty\) and \(W_p\) can be arbitrarily small even if supports are disjoint, so orthogonality is not equivalent to large metric distance. A common misconception is therefore that orthogonal persistence diagrams must be metrically far apart; the cited examples show that this need not hold.

The paper also gives an explicit orthogonal family. For
\[
D_3=\{(\theta+\tfrac12-\tfrac1{2m},\,\theta+\tfrac12+\tfrac1{2m}))\mid \theta=0,\dots,n-1\},
\]
\[
D_4=\{(\theta+\tfrac12-\tfrac1{2m},\,\theta+\tfrac12+\tfrac1{2m}))\mid \theta=2n,\dots,3n-1\},
\]
all intervals in \(D_3\) lie below those in \(D_4\), so every interval pair is disjoint and \(\ϝ(D_3,D_4)=0\).

For computation, the paper describes the following pipeline: build a Vietoris–Rips filtration from a finite point cloud and compute a persistence diagram \(D\); transform \(D\mapsto \phi(D)\), truncating when \(\lambda_j\equiv 0\); approximate the integrals by quadrature on the piecewise-linear graph; compute norms and inner products; and decide orthogonality when \(\ϝ(D_1,D_2)\) is below a numerical threshold \(\epsilon\approx 0\). In experiments on point-clouds sampled from a disk \(Q\), an annulus \(R\), and a circle \(S\), the cosine distance \(\ϝ^\ast=1-\ϝ\) separated \(R\) vs. \(S\) with \(\ϝ^\ast\approx 0.16\) and \(Q\) vs. \(R\) with \(\ϝ^\ast\approx 0.973\), whereas \(W_2\) and \(\|\cdot\|_2\) could not reliably do so. The method inherits shortcomings of persistence landscapes, including sensitivity to outliers, and numerical integration may create small nonzero inner products for nearly orthogonal diagrams.

## 4. Graph orthogonal representations and topological lower bounds

In graph theory, orthogonality is attached to vector assignments on vertices, and topology enters through Borsuk–Ulam-type lower-bound arguments. Haviv defines a \(t\)-dimensional orthogonal representation of a graph \(G=(V,E)\) over \(\mathbb F\) as an assignment \(v\mapsto u_v\in\mathbb F^t\setminus\{0\}\) such that distinct non-adjacent vertices receive orthogonal vectors, and the orthogonality dimension \(\xi_\mathbb F(G)\) is the minimum such \(t\) [1811.11488]. The paper proves general lower bounds on \(\xi_\mathbb F(G)\) using the Borsuk–Ulam theorem, especially for complements of generalized Kneser graphs.

For a set-system \(\mathcal F\subseteq 2^{[d]}\setminus\{\emptyset\}\), the complement \(\overline{K(\mathcal F)}\) of the generalized Kneser graph satisfies
\[
\xi_{\mathbb R}(\overline{K(\mathcal F)})\ge \operatorname{cd}_2(\mathcal F),\qquad
\xi_{\mathbb C}(\overline{K(\mathcal F)})\ge \tfrac12\,\operatorname{cd}_2(\mathcal F),
\]
where \(\operatorname{cd}_2(\mathcal F)\) is the \(2\)-colorability-defect. A geometric form of the bound uses configurations \(y_1,\dots,y_d\in S^{t-2}\) such that every open hemisphere contains the points of some \(A\in\mathcal F\), yielding
\[
\xi_{\mathbb R}(\overline{K(\mathcal F)})\ge t,\qquad
\xi_{\mathbb C}(\overline{K(\mathcal F)})\ge \tfrac t2.
\]
For ordinary Kneser graphs \(K(d,s)\), one recovers
\[
\xi_{\mathbb R}(\overline{K(d,s)})\ge d-2s+2,
\]
matching Lovász’s lower bound for chromatic number. Similar statements are obtained for Schrijver graphs and Borsuk graphs.

The paper "Local Orthogonality Dimension" shifts attention from ambient dimension to locality [2110.00718]. There an orthogonal representation of a graph \(G\) over \(\mathbb F\) is an assignment \(u_v\in\mathbb F^t\) with \(\langle u_v,u_v\rangle\neq 0\) for every vertex and \(\langle u_v,u_w\rangle=0\) whenever \(vw\in E\). This reflects a complement-based change of convention. The locality of a representation is
\[
\ell=\max_{v\in V}\dim\operatorname{span}\bigl(\{u_v\}\cup\{u_w\mid w\in N(v)\}\bigr),
\]
and the local orthogonality dimension \(\operatorname{lodim}_\mathbb F(G)\) is the minimum possible locality.

Topological methods again yield lower bounds. If a topological method implies \(\chi(G)\ge t\) for a graph \(G\) with at least one edge, then
\[
\operatorname{lodim}_\mathbb F(G)\ge \lceil t/2\rceil+1
\]
over every field. The proof uses the stronger fact of Alishahi–Meunier that any independent representation of a topologically \(t\)-chromatic graph contains a copy of \(K_{\lfloor t/2\rfloor,\lceil t/2\rceil}\) whose two sides are linearly independent. In some families this lower bound is tight, notably for Schrijver graphs. In others the local orthogonality dimension over \(\mathbb R\) equals the chromatic number: for every complement of a line graph,
\[
\operatorname{lodim}_{\mathbb R}(G)=\chi(G).
\]

The parameter also has algorithmic significance. For every fixed \(k\ge 3\) and any field \(\mathbb F\), deciding whether \(\operatorname{lodim}_{\mathbb F}(G)\le k\) is \(\mathsf{NP}\)-hard. In index coding one has
\[
\operatorname{minrk}_{\mathbb F}(G)\le \operatorname{lodim}_{\mathbb F}(\overline G)+\lceil \log_q|V|\rceil
\]
over \(\mathrm{GF}(q)\), so local orthogonality dimension furnishes upper bounds on optimum linear index-coding length. This makes topological orthogonality relevant not only to extremal graph theory but also to information theory and quantum one-round communication complexity.

## 5. Dynamical orthogonality and Möbius non-correlation

In topological dynamics, orthogonality refers to the vanishing of correlations between an orbit and an arithmetic or bounded sequence. Karagulyan defines topological Möbius orthogonality for a system \((X,T)\), with \(X\) a compact metric space and \(T:X\to X\) a homeomorphism, by the condition that for every \(f\in C(X)\) and every \(x\in X\),
\[
\lim_{N\to\infty}\frac1N\sum_{n=1}^N f(T^n x)\mu(n)=0,
\]
where \(\mu(n)\) is the classical Möbius function [1701.01968]. Sarnak’s conjecture predicts that this holds whenever the topological entropy vanishes.

The main theorem of that paper shows that Möbius orthogonality fails for subshifts of finite type with positive topological entropy. More precisely, if \(\sigma:\Sigma_A\to \Sigma_A\) is a subshift of finite type with \(h_{\mathrm{top}}(\sigma)>0\), then there exist \(\varphi\in C(\Sigma_A)\) and \(z\in\Sigma_A\) such that
\[
\limsup_{N\to\infty}\left|
\frac1N\sum_{n=1}^N \varphi(\sigma^n z)\mu(n)
\right|>0.
\]
Via Katok’s horseshoe theorem, every \(C^{1+\alpha}\) surface diffeomorphism with positive entropy also fails to be orthogonal to the Möbius function. The proof uses a specification-type loop-concatenation construction and arithmetic progressions with positive density of square-free integers.

The paper "Unveiling universality, encloseness, and orthogonality in dynamics" generalizes this perspective from the Möbius function to an arbitrary bounded sequence \(u:\mathbb N\to \mathbb D\) with mean zero [2604.21392]. It defines Cesàro orthogonality \((X,T)\perp u\) by
\[
\lim_{N\to\infty}
\frac1N\sum_{n=0}^{N-1} f(T^n x)\,u(n)=0
\quad
(\forall f\in C(X),\,x\in X),
\]
and logarithmic orthogonality \((X,T)\perp_{\log}u\) by the analogous logarithmic average. A stronger notion is the strong \(u\)-MOMO property:
\[
\lim_{K\to\infty}
\frac1{b_K}\sum_{k=1}^{K-1}
\left|
\sum_{n=b_k}^{b_{k+1}-1} f(T^n x_k)\,u(n)
\right|=0
\]
for every \(f\in C(X)\), every sequence \(x_k\in X\), and every increasing sequence \(b_k\) with \(b_{k+1}-b_k\to\infty\). The paper states that strong-MOMO implies orthogonality, and that \((X,T)\perp u\) is equivalent to all uniquely ergodic factors of \((X,T)\) enjoying strong-MOMO.

A principal lifting theorem says that if \((X,T)\) has the strong \(u\)-MOMO property and \((Z,R)\) is any topological system such that for each ergodic \(\nu\in M^e(Z,R)\) there exists an ergodic \(\nu'\in M^e(X,T)\) with \((Z,\nu,R)\) isomorphic to \((X,\nu',T)\), then \((Z,R)\perp u\). This motivates universal topological models for characteristic classes of measure-preserving systems. For the class \(\mathrm{DISP}_{ec}\) of automorphisms whose ergodic components have pure discrete spectrum, the paper constructs a universal model on
\[
G=(\mathbb T^{\mathbb N})\times(\mathbb T^{\mathbb N})\times(\mathbb T^{\mathbb N}),
\qquad
A(y,v,z)=(y,v,y+z).
\]
It also proves that if the union of all measure-theoretic eigenvalues of a zero-entropy system \((X,T)\) is countable, then Sarnak’s conjecture holds along a subsequence of full logarithmic density. A common source of confusion is that orthogonality in this literature is not geometric disjointness but cancellation of orbit-sequence correlations; the relevant topology is the topology of the dynamical model.

## 6. Operator theory, topological phases, and machine-learning recontextualizations

Several recent works use topological orthogonality language in more specialized ways. In "Orthogonality of bilinear forms and application to matrices," Roy, Senapati, and Sain characterize Birkhoff–James orthogonality in the Banach space \(C(U,E)\), where \(U\) is a compact topological space and \(E\) a real normed space [2407.13713]. For \(f,g\in C(U,E)\setminus\{0\}\), with
\[
M_f=\{u\in U:\|f(u)\|=\|f\|_\infty\},
\]
and cones
\[
x^+=\{y\in E:\|x+t\,y\|\ge \|x\| \ \forall t\ge 0\},\qquad
x^-=\{y\in E:\|x+t\,y\|\ge \|x\| \ \forall t\le 0\},
\]
the characterization is
\[
f\perp_B g
\iff
\exists\,u_1,u_2\in M_f
\text{ such that }
g(u_1)\in f(u_1)^+,\quad g(u_2)\in f(u_2)^-.
\]
If \(M_f\) is connected, this reduces to a single-point condition:
\[
f\perp_B g
\iff
\exists\,u_0\in M_f \text{ such that } f(u_0)\perp_B g(u_0).
\]
Applied to real bilinear forms and matrices, this yields an elementary proof of the real Bhatia–Šemrl theorem: for real matrices \(A,B\), \(A\perp_B B\) iff there exists a unit vector \(y_0\) such that \(\|Ay_0\|=\|A\|\) and \(\langle Ay_0,By_0\rangle=0\). Here compactness of the topological domain is what guarantees norm attainment and hence a finite orthogonality test set.

In topological phases of matter, "Anderson orthogonality catastrophe in \(2+1\)-D topological systems" studies the overlap
\[
S=\langle \Psi_0 \mid \Psi_0' \rangle
\]
between many-body ground states and shows a universal topological response term in its finite-size scaling [1905.00171]. At fixed points of \(2+1\)-dimensional topological orders,
\[
\ln|S|=-\alpha L^2-\beta L+\gamma_{\mathrm{topo}}+o(1),
\qquad
\gamma_{\mathrm{topo}}=(\chi c/12)\ln N=(\chi c/6)\ln L+O(1).
\]
Here \(\chi\) is the Euler characteristic and \(c\) is the central charge of the boundary CFT. For Laughlin wave functions, the paper finds a stronger leading behavior,
\[
\ln\langle \Psi^{(1)}\mid \Psi^{(2)}\rangle
=
-a\,N\ln N-\alpha N-\beta \sqrt N+\gamma\ln N+O(1),
\]
with \(a=(m-1)/4\) and \(\gamma=(m-1)^2/(24m)\) on the disk, and a corresponding sphere formula without the \(\sqrt N\) term. The leading \(-aN\ln N\) gives decay faster than exponential. In this context, topological orthogonality refers to universal topological structure in overlap scaling rather than to an explicitly defined bilinear relation.

A further recontextualization appears in "MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality" [2605.05646]. There topological orthogonality is a design principle for decoupling structural and semantic objectives in Transformers. Let \(\mathcal L_{\mathrm{topo}}\) be a structural loss, \(\mathcal L_{\mathrm{anchor}}\) a semantic loss, \(\theta_{QK}=(W_Q,W_K)\) the attention-topology parameters, and \(\theta_V=W_V\) the feature-value parameters. The orthogonality requirement is
\[
\bigl\langle \nabla_{\theta_{QK}}\mathcal L_{\mathrm{topo}},\,
\nabla_{\theta_{QK}}\mathcal L_{\mathrm{anchor}}\bigr\rangle=0,
\qquad
\bigl\langle \nabla_{\theta_V}\mathcal L_{\mathrm{topo}},\,
\nabla_{\theta_V}\mathcal L_{\mathrm{anchor}}\bigr\rangle=0,
\]
or, in shared-parameter form,
\[
\cos\theta_g
=
\frac{\langle g_{\mathrm{str}},g_{\mathrm{sem}}\rangle}
{\|g_{\mathrm{str}}\|\|g_{\mathrm{sem}}\|}
\approx 0.
\]
The architecture separates a topology stream
\[
Q=H_lW_Q,\qquad K=H_lW_K,\qquad
A=\mathrm{softmax}(QK^\top/\sqrt{d_k})
\]
from a semantic stream
\[
V=H_lW_V,\qquad H_{\mathrm{attn}}=A\,V,
\]
with stop-gradient operators to prevent cross-contamination. In experiments, the paper reports \(gFID=3.08\), linear probing \(85.2\%\) versus \(82.5\%\) for the InternViT-300M teacher, and structural \(mIoU=46.5\%\). Ablations show that removing topology loss destroys geometry, while removing semantic anchoring yields “semantic blindness” with zero-shot \(12.4\%\). Figure 3 reports a change in gradient cosine from \(-0.15\) in naive shared training to \(+0.04\) in MUSE. This suggests a modern computational usage in which “topological orthogonality” no longer refers to classical geometric orthogonality of vectors or sets, but to orthogonal routing of learning signals through topology-sensitive and semantics-sensitive parameter subspaces.

Across these settings, the unifying pattern is not a single invariant formula but a recurrent structural role: topology identifies when two entities should be treated as independent, non-overlapping, or non-interfering. In some cases this is literal closure disjointness or interval separation; in others it is non-correlation, locality obstruction, universal finite-size response, or architectural gradient decoupling. The phrase therefore functions less as a single doctrine than as a cross-disciplinary template for imposing orthogonality through topological structure.

Source: https://www.emergentmind.com/topics/topological-orthogonality