---
title: Cartesian Reverse Differential Categories
url: https://www.emergentmind.com/topics/cartesian-reverse-differential-categories
type: topic
---

# Cartesian Reverse Differential Categories

Searching arXiv for the cited papers and related work on Cartesian reverse differential categories.
Searching arXiv for "Cartesian reverse differential categories" and related papers.
Cartesian reverse differential categories are categorical structures that axiomatize reverse differentiation, the operation underlying reverse-mode automatic differentiation and backpropagation. Formally, a Cartesian reverse differential category is a Cartesian left-additive category equipped with a reverse differential combinator
\[
R\;:\;C[A,B]\;\longrightarrow\;C[A\times B,\,A]
\]
or equivalently \(\rho f:A\times B\to A\), intended to model “transpose–Jacobian multiplication” and to pull back a cotangent on \(B\) to one on \(A\) [1910.07065]. The theory was introduced as a direct axiomatization of reverse derivatives, then connected to forward differential structure, dagger structure on linear maps, generalized optimization, fibrational semantics, differential programming languages, Jacobians and gradients, and higher-order reverse chain rules [1910.07065], [2109.10262], [2409.05763], [2211.01835], [2101.10491], [2509.20931].

## 1. Foundational definition and axioms

The ambient structure of a Cartesian reverse differential category is a Cartesian left-additive category: a category with finite products, a terminal object, pairing and projections, and a commutative monoid structure \((+,0)\) on each hom-set such that pre-composition preserves \(+\) and \(0\), while projections are additive [1910.07065]. On this base one specifies, for each morphism \(f:A\to B\), a reverse derivative \(R[f]:A\times B\to A\).

The defining axioms are the reverse analogues of the familiar linearity, projection, sum, and chain-rule laws of ordinary reverse-mode differentiation. In the seven-axiom form, they state additivity in \(f\), additivity in the second argument, laws for identities and projections, compatibility with pairing, a reverse chain rule, and two higher-order coherence conditions. In the formulation used for reverse differential categories, [RD.1–2] say \(R\) is linear in each variable, [RD.3–5] recover the usual projection, pairing and chain-rule axioms, and [RD.6–7] encode the involutive and mixed-partial symmetry properties of repeated reverse differentiation [2101.10491].

The reverse chain rule is the central structural law. For composable \(A\xrightarrow f B\xrightarrow g C\), it has the form
\[
R[g\circ f]
\;=\;
R[f]\;\circ\;(1\times R[g])\;\circ\;\langle\,\pi_0,\langle f\circ\pi_0,\pi_1\rangle\rangle,
\]
or equivalent notational variants, expressing exactly the computational principle “pull back by \(g\) then by \(f\)” [2109.10262], [2203.12478].

A first-order presentation isolates the fragment RDC 1–5. In the fibrational treatment, a Cartesian reverse differential category structure on a Cartesian left-additive category \(\mathsf B\) is given by a reverse-derivative operator
\[
\rho \;:\;\mathsf B(A,B)\;\longrightarrow\;\mathsf B\bigl(A\times B,\;A\bigr)
\]
subject to linearity, additivity in the second argument, reverse projection laws, pairing, and the reverse chain rule [2409.05763]. This first-order fragment is sufficient for the fibrational characterization, while the seven-axiom form is needed for the full higher-order theory.

## 2. Concrete models and interpretation as reverse-mode differentiation

The prototypical example is the category of Euclidean spaces and smooth maps. If \(f:\mathbb R^n\to\mathbb R^m\), then
\[
R[f](x,\delta y)=J_f(x)^T\,\delta y,
\]
so the reverse derivative is exactly transposed Jacobian multiplication [2109.10262]. In this model, the forward differential is recovered from the reverse one, which makes precise the claim that Cartesian reverse differential categories generalize reverse-mode automatic differentiation rather than merely imitating it abstractly [1910.07065].

A second standard example is the polynomial category. If \(r\) is a commutative ring and \(P=(p_1,\dots,p_m):n\to m\) is an \(m\)-tuple of polynomials in \(n\) variables, then
\[
R[P](x,x')
=
\Bigl(\sum_{i=1}^m \partial_{x_1}p_i(x)\,x'_i,\;\dots,\;\sum_{i=1}^m \partial_{x_n}p_i(x)\,x'_i\Bigr),
\]
with the \(\partial_{x_j}p_i\) interpreted as formal partial derivatives [2109.10262]. This provides a model in which reverse differentiation is available beyond the real-analytic setting.

Other examples emphasize the algebraic breadth of the notion. For \(\mathsf{FVect}_{\mathbb R}\), one defines \(\rho f(v,w)=f^T(w)\), and the reverse chain rule follows from \((g\circ f)^T=f^T\circ g^T\) [2409.05763]. More generally, any †-category with †-biproducts, such as \(\mathrm{FinHilb}\) or \(\mathrm{MAT}(R)\), is a Cartesian reverse differential category via \([f]=f^\dagger\circ\pi_1\) [1910.07065]. The monoidal extension literature adds examples coming from relations, weighted relations, exterior algebra on \(\mathbb Z_2\), and quantum models such as Selinger–Valiron’s \(\overline{\mathbf{CPMs}^\oplus}\) or Vicary’s categorical quantum harmonic oscillator [2203.12478].

These examples justify the standard interpretation of \(R[f]\) as a back-propagator. In ordinary multivariate calculus, the forward derivative pushes tangents by \(J_f(x)\), while the reverse derivative pulls back cotangents by \(J_f(x)^T\); the CRDC axioms isolate the categorical content of that distinction [2203.12478].

## 3. Relation to forward differentiation, linear maps, Jacobians, and gradients

A basic theorem of the subject is that every Cartesian reverse differential category induces a Cartesian differential category. In the direct axiomatization, one defines
\[
D[f]\;:=\;\pi_1\;\circ\;R[R[f]]\;\circ\;(\langle1,0\rangle\times1)
\]
or an equivalent formula, and then verifies the Cartesian differential axioms [1910.07065]. The converse does not hold in general: a bare Cartesian differential category need not carry any dagger-structure on its linear maps, so one cannot reconstruct a reverse derivative [1910.07065].

This failure of converse is structurally precise. In a Cartesian differential category, a map \(f\) is linear iff \(D[f]=\pi_1\otimes f\), and the wide subcategory \(\mathrm{Lin}(X)\) of linear maps has finite biproducts [1910.07065]. In a Cartesian reverse differential category one moreover defines, for linear \(f\),
\[
f^\dagger \;:=\; \iota_1\;\circ\;[f]\;:\;B\to A,
\]
and the linear maps then form an additively enriched category with dagger biproducts [1910.07065]. Equivalently, a forward differential structure extends to a reverse one exactly when the linear fibration carries a fiber-wise dagger, stationary on objects, involutive, additive, preserving appropriate Cartesian lifts, and each fiber has †-biproducts [1910.07065].

The relation to Jacobians and gradients becomes especially transparent in linearly closed settings. A linearly closed Cartesian differential category has internal linear homs \(\mathcal L(A,B)\), a bilinear evaluation map, and linear currying for maps linear in the second argument [2211.01835]. The Jacobian is then defined by
\[
J(f):=\lambda_\ell(D[f]):A\to \mathcal L(A,B).
\]
In a Cartesian reverse differential category whose underlying forward structure is linearly closed, each reverse derivative is linear in its second argument, so one may also define
\[
\nabla(f):=\lambda_\ell(R[f]):A\to \mathcal L(B,A),
\]
and this gradient equals the transpose of the Jacobian:
\[
\nabla(f)=\tau\circ J(f),
\]
where \(\tau\) is a suitable linear transpose [2211.01835].

A common misconception is that Cartesian closedness alone should suffice to define Jacobians and gradients by currying the derivative. The linearly closed analysis shows that this excludes numerous important examples of Cartesian differential categories such as the category of real smooth functions [2211.01835]. The gradient construction therefore depends not merely on closure, but on an internal hom of linear maps together with transpose structure.

## 4. Fibrational reformulation and structural economy

A later reformulation recasts first-order reverse differential structure fibrationally. For a Cartesian left-additive category \(\mathsf B\), one constructs a fibration
\[
\mathsf{lens}_{\sf CLA}(\mathsf B)\;:\;\mathcal L\;\longrightarrow\;\mathsf B
\]
whose morphisms are “lenses” \((f^\sharp,f)\), with \(f^\sharp:A\times B'\to A'\) additive in its \(B'\)-argument [2409.05763]. A section of this lens fibration assigns to each \(f:A\to B\) the lens \((\rho f,f)\).

The key theorem states that giving \(\mathsf B\) a CRDC structure with axioms RDC 1–5 is equivalent to giving a stationary CLA-section
\[
R:\mathsf B\longrightarrow \mathsf{lens}_{\sf CLA}(\mathsf B)
\]
of the lens-fibration [2409.05763]. In this presentation, the reverse differential axioms are not postulated independently; they are exactly the equations forced by functoriality, product-preservation, and additivity in the lens fibration.

The same paper gives a conceptual proof that every reverse differential structure yields a forward one. The composite
\[
\mathsf B\;\xrightarrow{R} \mathsf{lens}_{\sf CLA}(\mathsf B)
\;\xrightarrow{\mathsf{lens}_{\sf CLA}(R)}
\mathsf{lens}_{\sf CLA}\bigl(\mathsf{lens}_{\sf CLA}(\mathsf B)\bigr)
\;\xrightarrow{\Phi}\;
S_{\sf CLA}(\mathsf B)
\]
produces a section satisfying the Cartesian differential axioms CDC 1–5 [2409.05763]. The paper presents this as a tidy proof of the classical result “every RDC induces a CDC,” eliminating the lengthy pointwise verifications in the original literature.

The fibrational perspective also clarifies adjacent constructions. Reverse tangent categories are defined as sections of the dual additive bundle fibration, and differential objects arise as an isoinserter of two section-functors \(T\simeq(-)^2\) [2409.05763]. This suggests that CRDCs are not an isolated formalism, but part of a broader fibrational theory of first-order differential structures.

## 5. Generalized optimization and category-theoretic learning theory

CRDCs support generalized optimization algorithms. In the setup of generalized optimization, one fixes an optimization domain \((Base,X)\) where \(Base\) is a CRDC and \(X\) is a chosen loss object whose hom-set \(Base[*,X]\) is an ordered commutative ring; an objective is a map \(l:A\to X\) [2109.10262].

The generalized gradient is defined by
\[
\nabla_{\!g}l \;:=\; R[l]\;\circ\;\langle\,\mathrm{id}_A,\;1_{A\to X}\rangle \;:\; A\to A.
\]
In the Euclidean case this is the usual gradient, and in \(\mathbf{Poly}\) it is the vector of formal partials [2109.10262]. Generalized gradient descent is then
\[
\mathrm{GD}(l)=-\,R[l]\circ\langle\mathrm{id},1\rangle:A\to A,
\]
with associated flow \(s\) satisfying
\[
\frac{d}{dt}s(t)=\mathrm{GD}(l)(s(t)).
\]
Generalized Newton’s method is defined by
\[
\mathrm{Newt}(l)
=
-\,R\bigl[\nabla_{\!g}l^{-1}\bigr]\circ\langle\nabla_{\!g}l,\nabla_{\!g}l\rangle:A\to A,
\]
which specializes in \(\mathbf{Euc}\) to \(-(\nabla^2 l)^{-1}\nabla l\) [2109.10262].

Several classical invariance properties survive. Generalized Newton’s method is invariant under all invertible linear transformations, while generalized gradient descent is invariant only to orthogonal linear transformations [2109.10262]. The same work also proves an inner-product-like decrease lemma:
\[
\frac{d}{dt}\bigl(l\circ s\bigr)(t)
=
-\; \bigl[R[l]_{s(t)}\bigr]^{\dagger}\;\circ\;R[l]_{s(t)}\;\circ\;1_X
\;\le\;0_X.
\]
In \(\mathbf{Euc}\) this becomes \(-\|\nabla l(s(t))\|^2\le 0\), so \(l\circ s\) is non-increasing; under mild bounded-below assumptions one deduces convergence of the flow [2109.10262].

The numerical experiments are deliberately elementary but concrete. In the integer polynomial domain \((\mathbf{Poly},1)\), the paper compares \(N\) steps of integer gradient descent with sampling \(N\) random integer points for random sums-of-squares polynomials with coefficients in \([1,10]\). For \(N=5,10,50,100\), the gradient method finds a strictly better minimum than random search with probability roughly \(0.74,\,0.76,\,0.77,\,0.81\) [2109.10262]. This suggests that the reverse differential formalism can support optimization procedures outside the usual real-valued analytic setting.

## 6. Differential programming languages and operational semantics

CRDCs also provide denotational semantics for differential programming languages. One application is to Abadi–Plotkin’s simple differential programming language, whose terms include real constants, addition, primitive operations, pairing, projections, conditionals, while-loops, and the reverse-derivative form \(v.\mathsf{rd}(x\mapsto m)(a)\) [2101.10491].

Given a reverse differential category, or more precisely an RDRC when partiality is included, the interpretation of the reverse-derivative form is defined by a categorical reverse derivative:
\[
\llbracket v.\mathsf{rd}(x\mapsto m)(a)\rrbracket
:
\Gamma_X
\to
U_X
\]
through a composite involving \(R[\llbracket m\rrbracket]\) [2101.10491]. The semantics extends compositionally to variables, constants, addition, pairing, projections, let-binding, conditionals using restriction-joins, and while-loops via joins of iterates.

The resulting metatheory includes symbolic-differentiation correctness, denotational soundness and adequacy, and under mild extra hypotheses even full abstraction [2101.10491]. In particular, the syntactic symbolic backpropagation operator agrees with the categorical reverse derivative on trace terms.

The categorical model also motivates operational improvements. Because in an RDRC the restriction idempotent of a reverse derivative is the same as that of the original map, most of the branchings introduced during symbolic differentiation of nested loops are in fact zero [2101.10491]. The paper therefore proposes a simplification of the symbolic differentiation rule for let-bindings when a variable does not occur free in the continuation, and states that this yields an exponential speed-up in common looping patterns [2101.10491]. It further observes that axioms [RD.6]–[RD.7] guarantee that converting once to the forward derivative and back to the reverse derivative recovers exactly \(R[f]\), so forward- and reverse-mode source transformations can be interleaved without loss [2101.10491].

## 7. Higher-order reverse differentiation and monoidal generalizations

The first-order reverse chain rule has a higher-order counterpart. Recent work defines partial reverse derivatives
\[
\mathsf R_j[f]:=\pi_j\circ \mathsf R[f]
\]
and higher-order reverse derivatives \(\rho^{(n)}[f]\), then proves a reverse analogue of Faà di Bruno’s formula [2509.20931]. In this formulation, the \((n+1)\)st reverse derivative of a composite is expressed by summing over partitions, with one block handled by a higher-order reverse derivative of \(f\), the remaining blocks by higher-order forward derivatives of \(f\), and the outer layer by a higher-order reverse derivative of \(g\) [2509.20931].

At \(n=0\), the reverse Faà di Bruno formula recovers exactly the first-order reverse chain rule [2509.20931]. In the smooth case, the stable rule needed for the construction holds by Clairaut’s theorem on mixed partials, so the higher-order theory specializes to the usual coordinate formulas for higher-order Jacobian-transposes [2509.20931]. A plausible implication is that higher-order reverse-mode constructions can be analyzed categorically with the same degree of modularity as first-order backpropagation.

A complementary extension shifts from Cartesian to monoidal structure. Monoidal reverse differential categories are additive, self-dual compact closed differential categories equipped with a reverse deriving map satisfying duals of the standard differential axioms [2203.12478]. One of the two fundamental facts is Monoidal\(\to\)Cartesian: if \(\mathbb X\) is an MRDC then its coKleisli category \(\mathbb X_\oc\) is a CRDC [2203.12478]. The converse reconstruction result shows that suitable self-dual compact closed differential categories also induce monoidal reverse structure [2203.12478].

This monoidal viewpoint produces examples not naturally emphasized in the purely Cartesian setting, including weighted relations and models of quantum computation [2203.12478]. The stated conclusion is that reverse differentiation is most naturally a monoidal rather than Cartesian phenomenon, with Cartesian reverse differential categories appearing as the Cartesian-side shadow of a richer monoidal reverse-mode theory [2203.12478].

Source: https://www.emergentmind.com/topics/cartesian-reverse-differential-categories