---
title: Katz's Hypergeometric Sums
url: https://www.emergentmind.com/topics/katz-s-hypergeometric-sums
type: topic
---

# Katz's Hypergeometric Sums

Katz’s hypergeometric sums are finite-field analogues of classical hypergeometric functions, introduced independently by John Greene and Nick Katz in the 1980s, and realized sheaf-theoretically as Frobenius trace functions of Katz’s \(\ell\)-adic hypergeometric sheaves. In the modern literature they appear in several equivalent or closely related normalizations—Katz’s original trace-function form, Greene’s Gaussian hypergeometric functions, the Beukers–Cohen–Mellit \(H_q\)-functions, McCarthy-type normalizations, and Otsubo’s \(F(\alpha,\beta;\lambda)\)—and they serve as a bridge between character sums, \(\ell\)-adic monodromy, point counts of algebraic varieties, and explicit formulas for Hecke traces and eigenvalues [2408.02918] [1505.02900] [1805.02863].

## 1. Hypergeometric data, differential equations, and \(\ell\)-adic sheaves

A standard starting point is a hypergeometric datum
\[
\alpha=\{a_1,\dots,a_n\},\qquad \beta=\{b_1=1,b_2,\dots,b_n\},
\]
with \(a_i,b_j\in \mathbf{Q}\), and the classical function
\[
F(\alpha,\beta;t)
=
{}_nF_{n-1}\!\left[\begin{matrix} a_1 & a_2 & \cdots & a_n \\ b_2 & \cdots & b_n \end{matrix} \,;\, t\right]
=
\sum_{k\ge 0} \frac{(a_1)_k\cdots(a_n)_k}{(b_1)_k\cdots(b_n)_k}\, t^k.
\]
It satisfies a Fuchsian differential equation with regular singular points at \(0,1,\infty\). The local exponents are
\[
\begin{aligned}
&\text{at } t=0:\quad 0,\,1-b_2,\,\dots,\,1-b_n,\\
&\text{at } t=\infty:\quad a_1,\,\dots,\,a_n,\\
&\text{at } t=1:\quad 0,\,1,\,\dots,\,n-2,\,\gamma,
\end{aligned}
\qquad
\gamma=-1+\sum_{j=1}^n b_j-\sum_{j=1}^n a_j,
\]
and the local monodromy at \(t=1\) is a pseudoreflection, equivalently \(\operatorname{rank}(M-I)=1\) [2408.02918].

For primitive hypergeometric data \(HD=\{\alpha,\beta\}\) defined over \(\mathbf{Q}\), Katz constructs a rank-\(n\) \(\ell\)-adic hypergeometric sheaf \(\mathcal H(HD)_\ell\) on \(\mathbf G_m\), defined over \(\mathbf Q(\zeta_M)\) when \(M\) is the level. Its degree-\(1\) Frobenius traces recover finite-field hypergeometric character sums:
\[
\operatorname{Tr}\,\rho_{HD,\lambda,\ell}(\operatorname{Frob}_\wp)
=
(-1)^{n-1}\,\omega_\wp^{(N(\wp)-1)a_1}(-1)\,
\mathbb P(HD;1/\lambda;\kappa_\wp;\omega_\wp).
\]
If \(\lambda\neq 1\), all eigenvalues of \(\rho_{HD,\lambda,\ell}(\operatorname{Frob}_\wp)\) have absolute value \(N(\wp)^{(n-1)/2}\); if \(\lambda=1\), the stalk dimension drops from \(n\) to \(n-1\) [2408.02918].

In Katz’s framework, the parameter sets also control local monodromy: \(A\) governs tame local monodromy at \(0\), \(B\) at \(\infty\), while the additive character enters through an exponential twist. This interpretation underlies the trace-function identity between finite hypergeometric sums and \(\ell\)-adic sheaves [1505.02900].

## 2. Finite-field definitions and competing normalizations

Over a finite field \(\mathbf F_q\) of odd characteristic, with nontrivial additive character \(\Psi\), multiplicative character group \(\widehat{\mathbf F_q^\times}\), and the convention \(\chi(0)=0\), the basic ingredients are the Gauss sum
\[
\mathfrak g(\chi):=\sum_{x\in\mathbf F_q}\Psi(x)\chi(x)
\]
and the Jacobi sum
\[
J(A,B):=\sum_{t\in\mathbf F_q}A(t)\,B(1-t).
\]
Greene’s binomial coefficient is
\[
\binom{A}{B}:=-\,B(-1)\,J(A,B)
\]
in the normalization adopted in the Hecke-trace work, while other papers use the equivalent \(q^{-1}J\)-normalized form customary in Greene’s original notation [2408.02918] [1609.01429].

One widely used normalization is the Beukers–Cohen–Mellit finite hypergeometric function
\[
H_q(A,B\mid t)
=
\frac{1}{1-q}\sum_{m=0}^{q-2}
\prod_{i=1}^d
\left(
\frac{g(m+\alpha_i)\,g(-m-\beta_i)}{g(\alpha_i)\,g(-\beta_i)}
\right)
\omega\big((-1)^d t\big)^m,
\]
for parameter multisets \(A=(\alpha_1,\dots,\alpha_d)\), \(B=(\beta_1,\dots,\beta_d)\) disjoint modulo \(1\), with \((q-1)\alpha_i,(q-1)\beta_j\in\mathbf Z\). This is the normalization that coincides with McCarthy’s [1505.02900]. Katz’s original sums are the same traces without the product of base Gauss sums in the denominator, and Greene’s finite hypergeometric functions differ from \(H_q\) by an explicit multiplicative factor involving Gauss or Jacobi sums [1505.02900] [1805.02863].

In the notation of the Hecke-trace paper, Greene’s-type finite-field hypergeometric \(\mathbb P\)-function is
\[
{}_{n}\mathbb P_{n-1}\!\left[\begin{matrix} A_1 & \cdots & A_n \\ B_2 & \cdots & B_n \end{matrix}\,;\,\lambda;q\right]
\]
with an explicit character-sum definition in terms of \(\binom{A_i\chi}{\chi}\) and \(\binom{B_j\chi}{\chi}^{-1}\), and the specialization
\[
\mathbb P(\alpha,\beta;\lambda;\mathbf F_q;\omega)
=
{}_{n}\mathbb P_{n-1}\!\left[\begin{matrix}
\omega^{(q-1)a_1} & \cdots & \omega^{(q-1)a_n}\\
\omega^{(q-1)b_2} & \cdots & \omega^{(q-1)b_n}
\end{matrix}\,;\,\lambda;\mathbf F_q\right].
\]
Its relation to the Beukers–Cohen–Mellit normalization is
\[
\mathbb P(\alpha,\beta;\lambda;\mathbf F_q;\omega)
=
H_q(\alpha,\beta;\lambda;\omega)\cdot
\prod_{i=2}^n
\omega^{(q-1)a_i}(-1)\,
J(\omega^{(q-1)a_i},\omega^{(q-1)b_i}),
\]
which is the conversion needed to compare trace formulas across normalizations [2408.02918].

Otsubo’s formulation replaces ordinary Gauss sums by a pair \(G(\chi)\), \(G^{\circ}(\chi)\) and defines
\[
F(\alpha,\beta;\lambda)
=
(1-q)^{-1}\sum_{\nu\in K^*}(\alpha)_\nu\,(\beta))_\nu^{-1}\,\nu(\lambda).
\]
In this normalization, Katz’s hypergeometric sum is related by
\[
\operatorname{Hyp}(\{;A_0,\dots,A_m;B_0,\dots,B_n\})(\mathbf F_q,t)
=
(-1)^{m+n+1}
\Bigl(\prod_{i=0}^m G(A_i)\Bigr)
\Bigl(\prod_{j=0}^n \frac{q-1}{G^{\circ}(B_j)}\Bigr)
F(\alpha,\beta;t^{-1}),
\]
which makes the equivalence between character-sum and sheaf-trace viewpoints completely explicit [2108.06754].

## 3. Structural properties, purity, and transformation theory

The sheaf-theoretic form of Katz’s sums carries strong structural consequences. For disjoint character tuples \(\boldsymbol\chi=(\chi_1,\dots,\chi_m)\) and \(\boldsymbol\eta=(\eta_1,\dots,\eta_n)\), Katz’s normalized hypergeometric sum
\[
H(t,k;\boldsymbol\chi,\boldsymbol\eta)
=
(-1)^{m+n-1} q^{-(m+n-1)/2}
\sum_{x\in(k^\times)^m}
\sum_{\substack{y\in(k^\times)^n\\ N(x)=tN(y)}}
\boldsymbol\chi(x)\boldsymbol\eta(y)\psi(T(x)-T(y))
\]
is the trace function of a geometrically irreducible \(\ell\)-adic middle-extension sheaf on \(\mathbf A^1_k\), pointwise pure of weight \(0\) and of rank \(\max\{m,n\}\). Consequently,
\[
|H(t,k;\boldsymbol\chi,\boldsymbol\eta)|\le \max\{m,n\}
\]
for all \(t\in k^\times\) [2302.14681].

Several transformation laws mimic the classical hypergeometric calculus. In the Beukers–Cohen–Mellit framework one has the parameter-shift identity
\[
\omega(t)^\mu S_q(\mu+A,\mu+B\mid t)=S_q(A,B\mid t)
\]
and the inversion symmetry
\[
S_q(A,B\mid t)=S_q(-B,-A\mid 1/t),
\]
which are finite-field analogues of Euler- and Pfaff-type parameter transformations [1505.02900]. Otsubo’s formalism develops these analogies much further, proving finite-field versions of Euler, Pfaff, Kummer, Gauss summation at \(1\), Kummer’s evaluation at \(-1\), Dixon, Watson, Whipple, Saalschütz, quadratic transformations, and product formulas of Kummer, Ramanujan, and Clausen type, all in terms of Gauss sums, Jacobi sums, Fourier transforms on characters, and Davenport–Hasse multiplication [2108.06754].

The Gauss-sum identities that drive these transformations include the reflection and multiplication laws
\[
\mathfrak g(A)\mathfrak g(\bar A)=A(-1)\,q,\qquad
\prod_{\eta^m=1}\frac{\mathfrak g(A\eta)}{\mathfrak g(\eta)}
=
-\mathfrak g(A^m)\,A(m^{-m}),
\]
as well as Hasse–Davenport-type factorizations. In the “defined over \(\mathbf Z\)” setting these permit a rewrite of \(H_q(A,B\mid t)\) in terms of integer exponents \(p_i,q_j\), a polynomial gcd multiplicity function \(s(m)\), and a rational factor \(M\); this rewrite is central to integrality arguments and geometric realizations [2408.02918] [1505.02900].

## 4. Point counts, Galois representations, and arithmetic geometry

A basic arithmetic role of Katz’s hypergeometric sums is to encode point counts of varieties over finite fields. For parameters defined over \(\mathbf Z\), Beukers–Cohen–Mellit consider the affine variety
\[
x_1+\cdots+x_r-y_1-\cdots-y_s=0,\qquad
\lambda x_1^{p_1}\cdots x_r^{p_r}=y_1^{q_1}\cdots y_s^{q_s},
\]
with a suitable nonsingular completion \(\overline V_\lambda\), and prove
\[
|\overline V_\lambda(\mathbf F_q)|
=
P_{rs}(q)+(-1)^{r+s-1}q^{\min(r-1,s-1)}H_q(A,B\mid M\lambda),
\]
where
\[
P_{rs}(q)=
\sum_{m=0}^{\min(r-1,s-1)}
{r-1\choose m}{s-1\choose m}\,
\frac{q^{r+s-m-2}-q^m}{q-1}.
\]
Thus the hypergeometric value is precisely the oscillatory term in the point count [1505.02900].

The standard examples are already arithmetic-geometric. For the Legendre family,
\[
|E_\lambda(\mathbf F_q)|
=
q+1-(-1)^{(q-1)/2}\,
H_q(1/2,1/2;1,1\mid \lambda),
\]
while for the curve \(y^2+xy+y=\lambda x^3\),
\[
|E_\lambda(\mathbf F_q)|
=
q+1-H_q(1/3,2/3;1,1\mid 27\lambda).
\]
The same framework gives
\[
N_f(t)=1+H_q(1/3,2/3;1,1/2\mid t)
\]
for \(f(x)=x^3+3x^2-4t\), and an explicit formula
\[
|\overline{S_\lambda}(\mathbf F_q)|=q^2+3q+1+q\,H_q(A,B\mid 2^{14}3^95^5\lambda)
\]
for a rational elliptic surface with an eight-parameter hypergeometric datum [1505.02900].

Fuselier, Long, Ramakrishna, Swisher, and Tu formulate a parallel theory using period functions \({}_{n+1}\mathbb P_n\) and normalized \({}_{n+1}\mathbb F_n\), together with a \(2\)-dimensional \(\ell\)-adic representation
\[
\sigma_{\lambda,\ell}:G_K\to GL_2(\mathbf Q_\ell(\zeta_N))
\]
for suitable rational parameters \(a,b,c\) and \(\lambda\in\mathbf Q\setminus\{0,1\}\). At good primes \(\mathfrak p\),
\[
\operatorname{Tr}(\sigma_{\lambda,\ell}(\operatorname{Frob}_\mathfrak p))
=
-
{}_{2}\mathbb P_1\!\left[
\begin{matrix}
\iota_\mathfrak p(a) & \iota_\mathfrak p(b)\\
\iota_\mathfrak p(c)
\end{matrix}
;\lambda;q(\mathfrak p)
\right],
\]
the Frobenius eigenvalues have absolute value \(\sqrt{q(\mathfrak p)}\), and therefore
\[
|\operatorname{Tr}(\sigma_{\lambda,\ell}(\operatorname{Frob}_\mathfrak p))|\le 2\sqrt{q(\mathfrak p)}.
\]
This realizes hypergeometric trace functions as Frobenius traces of explicit Galois representations attached to generalized Legendre-type curves [1510.02575].

In Otsubo’s framework, the same hypergeometric apparatus governs zeta functions of elliptic curves and K3 surfaces. For the K3 family
\[
X_\lambda:\quad z^2=(1-\lambda xy)\,x(1-x)\,y(1-y),
\]
the zeta function is expressed in terms of the Frobenius eigenvalues of the elliptic curve \(E_{1-\lambda}\), whose trace is itself given by a hypergeometric value [2108.06754].

## 5. Hecke traces and arithmetic triangle groups

A major recent application is the computation of traces and eigenvalues of Hecke operators on spaces of cusp forms attached to arithmetic triangle groups. The foundational identity is the Hecke–Frobenius trace relation
\[
\operatorname{Tr}(T_p\mid S_{k+2}(U))
=
\operatorname{Tr}(\operatorname{Frob}_p\mid H^1(X_U\otimes\bar{\mathbf Q},V^k(U)_\ell)),
\]
and, by Grothendieck–Lefschetz,
\[
-\operatorname{Tr}(T_p\mid S_{k+2}(U))
=
\sum_{\lambda\in X_U(\mathbf F_p)}
\operatorname{Tr}(\operatorname{Frob}_\lambda\mid (V^k(U)_\ell)_{\bar\lambda}).
\]
This reduces global Hecke traces to local Frobenius traces on stalks [2408.02918].

The geometric input is a comparison between automorphic \(\ell\)-adic sheaves \(V^k(\Gamma)_\ell\) and Katz’s hypergeometric sheaves. Katz’s rigidity theorem identifies \(V^1(\Gamma)_\mathbf C\) or \(V^2(\Gamma)_\mathbf C\) with a complex hypergeometric local system by matching local monodromy at \(0,1,\infty\), and the comparison theorem lifts this to an \(\ell\)-adic isomorphism up to a finite-order twist \(\chi_\Gamma\). Explicitly,
\[
\mathcal H(\{1/2\},\{1\})_\ell\otimes \mathcal H(HD(\Gamma))_\ell
\cong
\chi_\Gamma\otimes V^2(\Gamma)_\ell
\]
after \(\lambda\mapsto 1/\lambda\), with
\[
\chi_\Gamma=\chi_{-1}
\]
for \((2,\infty,\infty)\), \((2,3,\infty)\), and \((2,6,\infty)\),
\[
\chi_\Gamma=\chi_{-3}
\]
for \((2,4,6)\), and trivial twist for \((2,4,\infty)\) [2408.02918].

For \(\Gamma=(2,\infty,\infty),(2,3,\infty),(2,4,\infty),(2,6,\infty),(2,4,6)\), the contribution of a non-elliptic, non-cusp \(\lambda\) to the Hecke trace is
\[
\operatorname{Tr}(\operatorname{Frob}_\lambda\mid (V^k(\Gamma)_\ell)_{\bar\lambda})
=
F_{k/2}(a_\Gamma(\lambda,p),p),
\]
with recursion
\[
F_{m+1}(S,T)=(S-T)F_m(S,T)-T^2F_{m-1}(S,T),
\]
and
\[
a_\Gamma(\lambda,p)=
\begin{cases}
\left(\frac{1-1/\lambda}{p}\right)H_p(HD(\Gamma),1/\lambda),&\Gamma\neq (2,4,6),\\[4pt]
\left(\frac{-3(1-1/\lambda)}{p}\right)p\,H_p(HD(\Gamma),1/\lambda),&\Gamma=(2,4,6).
\end{cases}
\]
Each cusp contributes \(1\). Elliptic points contribute CM terms depending on whether \(p\) is inert or split in the relevant CM field \(K_z\); when \(p\) splits, the contribution is expressed using Jacobi sums such as \(J_\omega(1/3,1/3)\) or \(J_\omega(1/4,1/4)\) [2408.02918].

For \(\Gamma=\Gamma_1(4)\) and \(\Gamma_1(3)\), the generic local contribution is
\[
\sum_{j=0}^{k/2}
(-1)^j\binom{k-j}{j}\,p^j\,
H_p(HD(\Gamma);1/\lambda)^{k-2j},
\]
with explicit special contributions at \(\lambda=0,1,\infty\). In particular, if \(p\equiv -1\pmod{N_\Gamma}\) and \(k\) is odd, then
\[
\operatorname{Tr}(T_p\mid S_{k+2}(\Gamma))=0.
\]
The same method computes Hecke eigenvalues from traces of \((\operatorname{Frob}_p)^r\) over \(\mathbf F_{p^r}\); examples include \(S_8(\Gamma_0(4))\) and \(S_{24}(2,4,6)\) [2408.02918].

## 6. Direct character-sum identities, field-of-definition questions, and variants

Katz’s hypergeometric sums also appear in direct character-sum identities. A prominent example is Katz’s mixed character sum identity
\[
P(j,k)=V(j)V(k).
\]
Katz’s original proof used rigid local systems, Kloosterman sheaves, and monodromy calculations under the hypothesis \(p>3\). Evans proved the case \(q\equiv 1\pmod 4\) directly for all odd \(p>2\), and Evans–Greene proved the remaining case \(q\equiv 3\pmod 4\), again for all odd \(p>2\) [1607.05889] [1609.01429].

In those direct proofs, Greene’s finite-field \({}_2F_1\) is the hypergeometric bridge. For \(q\equiv 1\pmod 4\), an auxiliary sum
\[
h(D,j)
=
\sum_{x\in\mathbf F_q^\times}
D(x)(1-x)\cdot D\eta\!\big(x(j+1)^2+(j-1)^2\big)
\]
is expressed as
\[
h(D,j)
=
\frac{G(D)^2G(\eta)}{G(D^2\eta)}\,
{}_2F_1(D,DA_4;A_4\mid j^4),
\]
and a finite-field quadratic transformation converts the argument structure needed for Mellin-transform comparisons [1607.05889]. In the \(q\equiv 3\pmod 4\) case, norm-restricted Jacobi sums over \(\mathbf F_{q^2}\) satisfy
\[
R(D,j)
=
-\phi(j)\,q\,D^4((j-1)^2)\,
{}_2F_1\!\left(\begin{matrix} D,\; D^2\phi \\ D\phi \end{matrix};\, -\frac{(j+1)^2}{(j-1)^2}\right)
\]
for \(j\neq \pm1\), again making the hypergeometric content explicit [1609.01429].

A different analytic-number-theoretic occurrence is the double character sum
\[
g(\chi,\eta;k)=\sum_{u,v\in k^\times}\chi\!\left(\frac{u(v+1)}{v(u+1)}\right)\eta(uv-1),
\]
which is identified as
\[
g(\chi,\eta;k)
=
\chi\eta(-1)\,T(\eta)\,|k|^{1/2}\,
H(1,k;(1,1,1),(\eta,\chi,\chi^{-1})).
\]
Since the associated hypergeometric sheaf has rank \(3\), one obtains immediately
\[
|g(\chi,\eta;k)|\le 3|k|,
\]
recovering the bound used in work of Conrey–Iwaniec and Petrow–Young [2302.14681].

The finite-field definition originally requires the denominators of the rational parameters to divide \(q-1\). Beukers circumvents this by replacing \(\mathbf F_q\) with finite commutative semisimple \(\mathbf F_q\)-algebras \(\mathcal A,\mathcal B\), defining
\[
H_q(\mathcal A,\mathcal B\mid t)
=
\frac{1}{g_{\mathcal A}(X_{\mathcal A})g_{\mathcal B}(X_{\mathcal B})}
\sum_{\substack{x\in\mathcal A^\times,\;y\in\mathcal B^\times\\ tN_{\mathcal A}(x)=N_{\mathcal B}(y)}}
\psi_{\mathcal A}(x)\psi_{\mathcal B}(-y)X_{\mathcal A}(x)X_{\mathcal B}(y),
\]
and proving the Fourier expansion
\[
H_q(\mathcal A,\mathcal B\mid t)
=
\frac{1}{g_{\mathcal A}(X_{\mathcal A})g_{\mathcal B}(X_{\mathcal B})}
\sum_{m=0}^{q-2}
g_{\mathcal A}(X_{\mathcal A}\omega_N^m)\,
g_{\mathcal B}(X_{\mathcal B}\omega_N^m)\,
\omega(N_{\mathcal B}(-1)t)^m.
\]
If \(K\) is the field generated by the coefficients of the parameter polynomials \(A(x)\) and \(B(x)\), and \(p\) splits completely in \(K\), then
\[
p^\Delta G_p(\alpha,\beta\mid t)
\]
is an algebraic integer in \(K\), where \(\Delta\) is the maximum of the corresponding floor-function invariant over Galois conjugates of the parameters [1805.02863].

A further variant is Katz’s \((A,B)\)-exponential-sum framework. Fu and Wan reprove Katz’s theorems using \(\ell\)-adic cohomology and a theorem of Denef–Loeser, remove the hypothesis \(p\nmid (A+B)\), and in the nondegenerate case obtain cohomology concentrated in degree \(n+1\) with
\[
\dim H_c^{n+1}=(A+B)(d-1)^n,
\]
together with the square-root bound
\[
\left|
\sum_{t_0\in\mathbf F_q^\times}\sum_{\mathbf t\in\mathbf F_q^n}
\chi(t_0)\psi(G(t_0,\mathbf t))
\right|
\le
(A+B)(d-1)^n q^{(n+1)/2}.
\]
They also prove generic ordinarity for the universal \((A,B)\)-family when
\[
p\equiv 1\pmod{\operatorname{lcm}(A,dB)}.
\]
This places Katz-type hypergeometric sums within the Newton-polyhedron and \(p\)-adic slope theory of exponential sums [2003.08796].

In McCarthy’s \(p\)-adic framework, explicit evaluations of character sums yield \(p\)-adic analogues of Kummer’s linear transformation and Clausen-type transformations for \({}_2G_2\) and \({}_3G_3\); through the Greene/McCarthy bridge these produce additional transformation laws and special values for finite-field hypergeometric functions, hence for Katz-type sums after normalization [1802.04503].

Source: https://www.emergentmind.com/topics/katz-s-hypergeometric-sums