Cartesian Reverse Differential Categories
- Cartesian Reverse Differential Categories are categorical frameworks that axiomatize reverse-mode differentiation using structures like transpose–Jacobian multiplication and backpropagation.
- They enforce rigorous axioms such as reverse chain rule, additivity, projection, pairing, and higher-order coherence to bridge forward and reverse differentiation.
- Applications include generalized optimization, differential programming languages, and quantum computation, with models in Euclidean spaces, polynomial rings, and †-categories.
Searching arXiv for the cited papers and related work on Cartesian reverse differential categories. Searching arXiv for "Cartesian reverse differential categories" and related papers. Cartesian reverse differential categories are categorical structures that axiomatize reverse differentiation, the operation underlying reverse-mode automatic differentiation and backpropagation. Formally, a Cartesian reverse differential category is a Cartesian left-additive category equipped with a reverse differential combinator
or equivalently , intended to model “transpose–Jacobian multiplication” and to pull back a cotangent on to one on (Cockett et al., 2019). The theory was introduced as a direct axiomatization of reverse derivatives, then connected to forward differential structure, dagger structure on linear maps, generalized optimization, fibrational semantics, differential programming languages, Jacobians and gradients, and higher-order reverse chain rules (Cockett et al., 2019, Shiebler, 2021, Capucci et al., 2024, Lemay, 2022, Cruttwell et al., 2021, Biggin et al., 25 Sep 2025).
1. Foundational definition and axioms
The ambient structure of a Cartesian reverse differential category is a Cartesian left-additive category: a category with finite products, a terminal object, pairing and projections, and a commutative monoid structure on each hom-set such that pre-composition preserves and $0$, while projections are additive (Cockett et al., 2019). On this base one specifies, for each morphism , a reverse derivative .
The defining axioms are the reverse analogues of the familiar linearity, projection, sum, and chain-rule laws of ordinary reverse-mode differentiation. In the seven-axiom form, they state additivity in , additivity in the second argument, laws for identities and projections, compatibility with pairing, a reverse chain rule, and two higher-order coherence conditions. In the formulation used for reverse differential categories, [RD.1–2] say 0 is linear in each variable, [RD.3–5] recover the usual projection, pairing and chain-rule axioms, and [RD.6–7] encode the involutive and mixed-partial symmetry properties of repeated reverse differentiation (Cruttwell et al., 2021).
The reverse chain rule is the central structural law. For composable 1, it has the form
2
or equivalent notational variants, expressing exactly the computational principle “pull back by 3 then by 4” (Shiebler, 2021, Cruttwell et al., 2022).
A first-order presentation isolates the fragment RDC 1–5. In the fibrational treatment, a Cartesian reverse differential category structure on a Cartesian left-additive category 5 is given by a reverse-derivative operator
6
subject to linearity, additivity in the second argument, reverse projection laws, pairing, and the reverse chain rule (Capucci et al., 2024). This first-order fragment is sufficient for the fibrational characterization, while the seven-axiom form is needed for the full higher-order theory.
2. Concrete models and interpretation as reverse-mode differentiation
The prototypical example is the category of Euclidean spaces and smooth maps. If 7, then
8
so the reverse derivative is exactly transposed Jacobian multiplication (Shiebler, 2021). In this model, the forward differential is recovered from the reverse one, which makes precise the claim that Cartesian reverse differential categories generalize reverse-mode automatic differentiation rather than merely imitating it abstractly (Cockett et al., 2019).
A second standard example is the polynomial category. If 9 is a commutative ring and 0 is an 1-tuple of polynomials in 2 variables, then
3
with the 4 interpreted as formal partial derivatives (Shiebler, 2021). This provides a model in which reverse differentiation is available beyond the real-analytic setting.
Other examples emphasize the algebraic breadth of the notion. For 5, one defines 6, and the reverse chain rule follows from 7 (Capucci et al., 2024). More generally, any †-category with †-biproducts, such as 8 or 9, is a Cartesian reverse differential category via 0 (Cockett et al., 2019). The monoidal extension literature adds examples coming from relations, weighted relations, exterior algebra on 1, and quantum models such as Selinger–Valiron’s 2 or Vicary’s categorical quantum harmonic oscillator (Cruttwell et al., 2022).
These examples justify the standard interpretation of 3 as a back-propagator. In ordinary multivariate calculus, the forward derivative pushes tangents by 4, while the reverse derivative pulls back cotangents by 5; the CRDC axioms isolate the categorical content of that distinction (Cruttwell et al., 2022).
3. Relation to forward differentiation, linear maps, Jacobians, and gradients
A basic theorem of the subject is that every Cartesian reverse differential category induces a Cartesian differential category. In the direct axiomatization, one defines
6
or an equivalent formula, and then verifies the Cartesian differential axioms (Cockett et al., 2019). The converse does not hold in general: a bare Cartesian differential category need not carry any dagger-structure on its linear maps, so one cannot reconstruct a reverse derivative (Cockett et al., 2019).
This failure of converse is structurally precise. In a Cartesian differential category, a map 7 is linear iff 8, and the wide subcategory 9 of linear maps has finite biproducts (Cockett et al., 2019). In a Cartesian reverse differential category one moreover defines, for linear 0,
1
and the linear maps then form an additively enriched category with dagger biproducts (Cockett et al., 2019). Equivalently, a forward differential structure extends to a reverse one exactly when the linear fibration carries a fiber-wise dagger, stationary on objects, involutive, additive, preserving appropriate Cartesian lifts, and each fiber has †-biproducts (Cockett et al., 2019).
The relation to Jacobians and gradients becomes especially transparent in linearly closed settings. A linearly closed Cartesian differential category has internal linear homs 2, a bilinear evaluation map, and linear currying for maps linear in the second argument (Lemay, 2022). The Jacobian is then defined by
3
In a Cartesian reverse differential category whose underlying forward structure is linearly closed, each reverse derivative is linear in its second argument, so one may also define
4
and this gradient equals the transpose of the Jacobian: 5 where 6 is a suitable linear transpose (Lemay, 2022).
A common misconception is that Cartesian closedness alone should suffice to define Jacobians and gradients by currying the derivative. The linearly closed analysis shows that this excludes numerous important examples of Cartesian differential categories such as the category of real smooth functions (Lemay, 2022). The gradient construction therefore depends not merely on closure, but on an internal hom of linear maps together with transpose structure.
4. Fibrational reformulation and structural economy
A later reformulation recasts first-order reverse differential structure fibrationally. For a Cartesian left-additive category 7, one constructs a fibration
8
whose morphisms are “lenses” 9, with 0 additive in its 1-argument (Capucci et al., 2024). A section of this lens fibration assigns to each 2 the lens 3.
The key theorem states that giving 4 a CRDC structure with axioms RDC 1–5 is equivalent to giving a stationary CLA-section
5
of the lens-fibration (Capucci et al., 2024). In this presentation, the reverse differential axioms are not postulated independently; they are exactly the equations forced by functoriality, product-preservation, and additivity in the lens fibration.
The same paper gives a conceptual proof that every reverse differential structure yields a forward one. The composite
6
produces a section satisfying the Cartesian differential axioms CDC 1–5 (Capucci et al., 2024). The paper presents this as a tidy proof of the classical result “every RDC induces a CDC,” eliminating the lengthy pointwise verifications in the original literature.
The fibrational perspective also clarifies adjacent constructions. Reverse tangent categories are defined as sections of the dual additive bundle fibration, and differential objects arise as an isoinserter of two section-functors 7 (Capucci et al., 2024). This suggests that CRDCs are not an isolated formalism, but part of a broader fibrational theory of first-order differential structures.
5. Generalized optimization and category-theoretic learning theory
CRDCs support generalized optimization algorithms. In the setup of generalized optimization, one fixes an optimization domain 8 where 9 is a CRDC and $0$0 is a chosen loss object whose hom-set $0$1 is an ordered commutative ring; an objective is a map $0$2 (Shiebler, 2021).
The generalized gradient is defined by
$0$3
In the Euclidean case this is the usual gradient, and in $0$4 it is the vector of formal partials (Shiebler, 2021). Generalized gradient descent is then
$0$5
with associated flow $0$6 satisfying
$0$7
Generalized Newton’s method is defined by
$0$8
which specializes in $0$9 to 0 (Shiebler, 2021).
Several classical invariance properties survive. Generalized Newton’s method is invariant under all invertible linear transformations, while generalized gradient descent is invariant only to orthogonal linear transformations (Shiebler, 2021). The same work also proves an inner-product-like decrease lemma: 1 In 2 this becomes 3, so 4 is non-increasing; under mild bounded-below assumptions one deduces convergence of the flow (Shiebler, 2021).
The numerical experiments are deliberately elementary but concrete. In the integer polynomial domain 5, the paper compares 6 steps of integer gradient descent with sampling 7 random integer points for random sums-of-squares polynomials with coefficients in 8. For 9, the gradient method finds a strictly better minimum than random search with probability roughly 0 (Shiebler, 2021). This suggests that the reverse differential formalism can support optimization procedures outside the usual real-valued analytic setting.
6. Differential programming languages and operational semantics
CRDCs also provide denotational semantics for differential programming languages. One application is to Abadi–Plotkin’s simple differential programming language, whose terms include real constants, addition, primitive operations, pairing, projections, conditionals, while-loops, and the reverse-derivative form 1 (Cruttwell et al., 2021).
Given a reverse differential category, or more precisely an RDRC when partiality is included, the interpretation of the reverse-derivative form is defined by a categorical reverse derivative: 2 through a composite involving 3 (Cruttwell et al., 2021). The semantics extends compositionally to variables, constants, addition, pairing, projections, let-binding, conditionals using restriction-joins, and while-loops via joins of iterates.
The resulting metatheory includes symbolic-differentiation correctness, denotational soundness and adequacy, and under mild extra hypotheses even full abstraction (Cruttwell et al., 2021). In particular, the syntactic symbolic backpropagation operator agrees with the categorical reverse derivative on trace terms.
The categorical model also motivates operational improvements. Because in an RDRC the restriction idempotent of a reverse derivative is the same as that of the original map, most of the branchings introduced during symbolic differentiation of nested loops are in fact zero (Cruttwell et al., 2021). The paper therefore proposes a simplification of the symbolic differentiation rule for let-bindings when a variable does not occur free in the continuation, and states that this yields an exponential speed-up in common looping patterns (Cruttwell et al., 2021). It further observes that axioms [RD.6]–[RD.7] guarantee that converting once to the forward derivative and back to the reverse derivative recovers exactly 4, so forward- and reverse-mode source transformations can be interleaved without loss (Cruttwell et al., 2021).
7. Higher-order reverse differentiation and monoidal generalizations
The first-order reverse chain rule has a higher-order counterpart. Recent work defines partial reverse derivatives
5
and higher-order reverse derivatives 6, then proves a reverse analogue of Faà di Bruno’s formula (Biggin et al., 25 Sep 2025). In this formulation, the 7st reverse derivative of a composite is expressed by summing over partitions, with one block handled by a higher-order reverse derivative of 8, the remaining blocks by higher-order forward derivatives of 9, and the outer layer by a higher-order reverse derivative of 0 (Biggin et al., 25 Sep 2025).
At 1, the reverse Faà di Bruno formula recovers exactly the first-order reverse chain rule (Biggin et al., 25 Sep 2025). In the smooth case, the stable rule needed for the construction holds by Clairaut’s theorem on mixed partials, so the higher-order theory specializes to the usual coordinate formulas for higher-order Jacobian-transposes (Biggin et al., 25 Sep 2025). A plausible implication is that higher-order reverse-mode constructions can be analyzed categorically with the same degree of modularity as first-order backpropagation.
A complementary extension shifts from Cartesian to monoidal structure. Monoidal reverse differential categories are additive, self-dual compact closed differential categories equipped with a reverse deriving map satisfying duals of the standard differential axioms (Cruttwell et al., 2022). One of the two fundamental facts is Monoidal2Cartesian: if 3 is an MRDC then its coKleisli category 4 is a CRDC (Cruttwell et al., 2022). The converse reconstruction result shows that suitable self-dual compact closed differential categories also induce monoidal reverse structure (Cruttwell et al., 2022).
This monoidal viewpoint produces examples not naturally emphasized in the purely Cartesian setting, including weighted relations and models of quantum computation (Cruttwell et al., 2022). The stated conclusion is that reverse differentiation is most naturally a monoidal rather than Cartesian phenomenon, with Cartesian reverse differential categories appearing as the Cartesian-side shadow of a richer monoidal reverse-mode theory (Cruttwell et al., 2022).