---
title: 'Function Vector: Diverse Mathematical and ML Roles'
url: https://www.emergentmind.com/topics/function-vector-fv
type: topic
---

# Function Vector: Diverse Mathematical and ML Roles

Function vector (FV) is a polysemous technical term rather than a single standardized object. In functional analysis it appears in Banach \(A\)-valued function algebras, where vector-valued functions, \(A\)-characters, and vector-valued spectra extend classical Banach function algebra theory; in matrix analysis it denotes spectral matrix-functions generated by vector-fields on ordered eigenvalue tuples; in mechanistic interpretability it denotes activation-space task representations extracted from attention heads or residual streams and reused to steer large language models; and in randomized representation theory it denotes high-dimensional vectors representing functions in a reproducing kernel Hilbert space [1509.09215] [1811.08358] [2502.14010] [2109.03429]. This suggests a recurring theme: a function, task, or functional action is encoded in vectorial form, but the formal object depends entirely on the surrounding theory.

## 1. Vector-valued function algebras and \(A\)-characters

In the Banach-algebraic usage, a commutative unital Banach algebra \(A\) and a compact Hausdorff space \(X\) define the ambient algebra \(C(X,A)\) of continuous \(A\)-valued functions with uniform norm \(\|f\|_X=\sup_{x\in X}\|f(x)\|_A\). An \(A\)-valued function algebra \(\mathcal A\subset C(X,A)\) is required to contain all constant functions and to separate points of \(X\). If \(\mathcal A\) is admissible, meaning \(\{\varphi\circ f:\varphi\in M(A),\,f\in\mathcal A\}\subseteq \mathcal A\), then one can define an \(A\)-character as a homomorphism \(\Psi:\mathcal A\to A\) satisfying \(\Psi(1)=1\) and
\[
\varphi(\Psi(f))=\Psi(\varphi\circ f)\qquad(\forall f\in\mathcal A,\ \forall \varphi\in M(A)).
\]
The evaluation maps \(\mathcal E_x(f)=f(x)\) are the model examples, and from the defining property together with semisimplicity of \(A\) one obtains \(\Psi(a)=a\) for all constants \(a\in A\) [1509.09215].

This formalism generalizes ordinary scalar characters. When \(A=\mathbb C\), \(A\)-valued function algebras reduce to Banach function algebras and \(A\)-characters reduce to the usual character space \(M(\mathcal A)\). A central structural result is that, for natural \(A\)-valued algebras such as \(C(X,A)\), \(\operatorname{Lip}(X,A)\), \(P(K,A)\), \(R(K,A)\), \(H(K,A)\), and tensor-product constructions of the form \(\mathfrak A\,\widehat\otimes_\varepsilon A\), the only \(A\)-characters are the point evaluations \(\mathcal E_x\). The same framework identifies the \(A\)-valued spectrum through
\[
\operatorname{sp}_A(f)=\{\Psi(f):\Psi\in M_A(\mathcal A)\},
\]
so in natural cases the vector-valued spectrum is determined by evaluation. The paper also gives a non-admissible \(\mathbb C^2\)-valued example to show that admissibility is essential for the \(A\)-character machinery and the associated spectral description [1509.09215].

## 2. Spectral matrix-functions and multivector function calculus

A different mathematical usage arises for Hermitian matrix-functions based on vector-fields. If \(A\in\mathbb H_d\) has spectral decomposition \(A=U_A\Lambda_\alpha U_A^*\), with \(\alpha\in\mathbb R^d_{\ge}\) the ordered eigenvalue vector, a block-constant vector-field \(F:\mathbb R^d_{\ge}\to\mathbb R^d\) defines
\[
\mathcal L_F(A)=U_A\,\mathrm{diag}(F(\alpha))\,U_A^*.
\]
This strictly generalizes the scalar spectral calculus \(f(A)\), since the component \(F_m(x)\) may depend on all coordinates \(x_1,\dots,x_d\), not only on \(x_m\). The well-definedness condition is block-constancy: whenever \(x_m=x_n\), one requires \(F_m(x)=F_n(x)\). Under \(C^1\) point-symmetry, the paper proves a generalized Daleskii–Krein formula,
\[
\mathcal{L}_F'(E)=U\Big([F,\alpha]\circ \hat E+\mathrm{diag}\big((F'|_{\alpha})^o \hat{e}\big)\Big)U^*,
\]
and a Frobenius-norm Lipschitz bound
\[
\|\mathcal{L}_F(A)-\mathcal{L}_F(B)\|_2\leq \|F\|_{Lip}\,\|A-B\|_2.
\]
Classical scalar spectral functions are recovered by the special choice \(F(x)=(f(x_1),\dots,f(x_d))\) [1811.08358].

A broader algebraic extension appears in Clifford analysis, where elementary functions are extended from scalars, complex numbers, and quaternions to multivectors in \(C\ell(\mathbb R^2)\) and \(C\ell(\mathbb R^3)\). The paper defines exponentials, logarithms, powers, trigonometric functions, and hyperbolic functions of multivector variables by power series and polar decomposition. It gives the general power formula \(M^P=e^{P\log M}\), the amplitude \(|M|=\sqrt{M\bar M}\), the inverse \(M^{-1}=\bar M/(M\bar M)\), and a unified square-root expression \(M^{1/2}=\sqrt{(M+|M|)/2}\). One notable result is that a complex number raised to a vector power produces a quaternion:
\[
(\cos\theta+j\sin\theta)^v=\cos(\|v\|\theta)+j\hat v\,\sin(\|v\|\theta).
\]
Comparing dimensions, the paper identifies \(C\ell(\mathbb R^3)\) as a particularly versatile algebraic framework because the pseudoscalar \(j=e_{123}\) commutes with all elements [1409.6252].

## 3. LLM function vectors: operational definitions

In large language models, the modern mechanistic-interpretability usage treats a function vector as a task representation elicited during in-context learning. One head-level formulation defines, for task \(t\) and attention head \(a\), the mean task-conditioned activation
\[
\bar a^t=\frac{1}{|P_t|}\sum_{p_i^t\in P_t} a(p_i^t),
\]
and calls \(a\) an FV head when patching \(\bar a^t\) into corrupted prompts restores task-appropriate behavior. The associated FV score on a corrupted prompt \(\tilde p_i^t\) is
\[
S_{\text{FV}}(a\mid \tilde p_i^t)=f(\tilde p_i^t\mid a:=\bar a^t)[y]-f(\tilde p_i^t)[y].
\]
At task level, some experiments aggregate over the top-2% FV heads to form
\[
FV_t=\sum_{a\in\mathcal A}\bar a^t,
\]
which is then added to the residual stream to test whether it reproduces the task without informative demonstrations [2502.14010].

A residual-stream formulation defines one FV per task, template, and layer as a mean-difference direction between positive and negative in-context learning conditions:
\[
FV_{t,k,\ell}=\mathbb E[\mathbf h_\ell^{(\mathrm{pos})}]-\mathbb E[\mathbf h_\ell^{(\mathrm{neg})}],
\]
with steering implemented by
\[
\mathbf h_\ell'=\mathbf h_\ell+\alpha\cdot FV_{t,k,\ell}.
\]
Here the FV is extracted from the final-token residual stream, usually from 15 positive and 15 negative prompts, and can be transferred across templates for the same task [2604.02608].

A prompt-specific formulation defines the FV directly from the outputs of previously localized FV heads:
\[
v_{\mathrm{FV}}(p_n^t)=\sum_{a\in\mathcal A_{\mathrm{FV}}} a(p_n^t,t_{n+1}),
\]
where \(p_n^t=((x_1,y_1),\dots,(x_n,y_n),x_{n+1})\) is an \(n\)-shot prompt and \(t_{n+1}\) is the final separator token. In that setting, FV quality is measured causally by injecting \(v_{\mathrm{FV}}(p_n^t)\) into zero-shot prompts and maximizing task accuracy over intervention layer and scale [2605.16591].

Taken together, these works do not impose a single canonical definition. Instead, the term denotes a family of causally tested activation-space task representations at different granularities: per-head averages, residual mean-difference directions, and per-prompt sums over FV heads.

## 4. FV heads and the mechanisms of in-context learning

A central empirical question is whether in-context learning is driven primarily by induction heads or by FV heads. Across 12 decoder-only models, induction heads are defined as the top 2% of heads by induction score, and FV heads as the top 2% by FV score. The overlap is limited: 7 of 12 models have 0 overlapping heads in the top-2% sets, while the remaining models have only 5–15% overlap. Nevertheless, the two sets are correlated, and FV heads tend to occur in slightly deeper layers than induction heads [2502.14010].

Ablation results sharply distinguish their causal roles. Ablating FV heads drastically hurts few-shot ICL accuracy, whereas ablating induction heads has limited impact, especially when overlap is controlled by excluding heads that rank highly on both criteria. In models above 1B parameters, ablating non-FV induction heads leaves ICL accuracy almost unchanged relative to random ablation, while ablating non-induction FV heads causes large drops; in the 6.9B Pythia model, removal of FV heads is described as catastrophic for ICL, while removal of induction heads with exclusion is near-random. The same study reports that FV heads strongly affect few-shot accuracy but not token-loss difference, while induction heads show the reverse pattern [2502.14010].

Training dynamics further separate the two mechanisms. In Pythia checkpoints, induction scores rise sharply around step 1,000 of 143,000 total steps and then plateau or slightly decline, whereas FV scores emerge later, around step 16,000, and continue increasing through training. Many final FV heads begin with high induction scores and later transition toward the FV mechanism; the reverse transition is not observed. The paper interprets this as a unidirectional developmental pattern in which induction serves as a precursor to function-vector-based in-context learning [2502.14010].

## 5. Steering, transfer, and compositional structure

Large-scale steering studies refine the claim that FVs are task codes. In a cross-template setting comprising 4,032 source-target pairs across 12 tasks, 6 models, and 8 templates per task, FV steering succeeds even when the logit lens cannot decode the correct answer at any layer. Steering exceeds logit-lens accuracy for every task on every model, with gaps as large as \(-0.91\); only 3 of 72 task-model instances show the opposite decodable-without-steerable pattern, all in Mistral. FVs that achieve over 0.90 steering accuracy still project to incoherent token distributions under vocabulary projection, optimal interventions occur at early layers \(L2\)–\(L8\), and logit-lens decodability peaks only at late layers such as \(L28\)–\(L32\). The same paper finds that the previously reported negative cosine-transfer correlation dissolves at scale: pooled \(r\) ranges from \(-0.199\) to \(+0.126\), and cosine adds less than 0.011 in \(R^2\) beyond task identity [2604.02608].

Instruction-elicited FV work studies two additional design choices: head selection and steering location. Replacing Average Indirect Effect head selection with Layer-wise Relevance Propagation improves both speed and accuracy, with measured throughput changes of 4.57 vs 1541.93 samples/min for Llama-3.2-3B, 1.63 vs 1308.41 for Llama-3.1-8B, and 2.96 vs 1167.67 for Qwen3-4B, for an average factor of about 511×. For steering, distributed reinjection—adding each selected head’s averaged representation at its original head and layer location—outperforms simple aggregation, with gains up to 0.156 in accuracy; LRP-selected heads improve steering accuracy over AIE-selected heads by as much as 0.194, and some tasks such as capitalization and translation are reported to be meaningfully reproduced only with LRP-based extraction [2606.05079].

Few-shot composition studies then ask how multiple demonstrations combine into a single FV. Across Gemma-2, Llama-3.2, and Llama-3.1 models, an \(n\)-shot FV is well approximated by a linear combination of example-level sub-FVs obtained by attention-edge restriction. The fitted reconstruction achieves mean cosine similarity at least 0.925 and mean \(R^2\) at least 0.875, and preserves most of the causal effect under injection. Contextualization does not merely add examples; it reweights them. In normal tasks it mitigates recency bias by redistributing attention more uniformly, while in ambiguous tasks it concentrates attention on unambiguous examples, and a Shapley-style decomposition finds that Query–Key contextualization is the most consistent positive contributor to FV quality, whereas Value-mediated effects are more heterogeneous [2605.16591].

These results jointly support an interpretation already stated explicitly in the literature: FVs behave less like answer vectors than like computational instructions. A plausible implication is that the relevant task representation often lies upstream of direct vocabulary-space decodability and is assembled by additive superposition together with context-dependent attention reweighting.

## 6. Continual learning and adjacent vector-space formalisms

Function vectors have also been used to analyze catastrophic forgetting in continual instruction tuning. In that setting, a task FV is defined as
\[
\theta_T=\sum_{(l,k)\in\mathcal S}\bar h_{lk}^T,
\]
where \(\mathcal S\) is a global set of causally important heads identified by activation patching and \(\bar h_{lk}^T\) is the task-conditioned mean activation at the last token. The paper argues, both theoretically and empirically, that catastrophic forgetting in LLMs primarily stems from biases in function activation rather than the overwriting of task processing functions. Adding the original FV back into a fine-tuned model recovers lost performance, subtracting newly learned task FVs mitigates interference, and FV similarity predicts forgetting more strongly than last-layer hidden similarity or parameter \(L_2\) distance. The proposed FV-guided training objective
\[
\ell=\ell_{\mathrm{LM}}+\alpha_1\ell_{\mathrm{FV}}+\alpha_2\ell_{\mathrm{KL}},
\]
with \(\alpha_1=1\) and \(\alpha_2=0.08\), is evaluated on four benchmarks and is reported to improve both zero-shot and in-context performance while preserving task-specific performance [2502.11019].

A distinct but conceptually adjacent line of work represents functions themselves as high-dimensional vectors in a Vector Function Architecture. There, an encoding \(r\mapsto \mathbf z(r)\) is chosen so that inner products approximate a kernel,
\[
\mathbf z(r_1)^\top \overline{\mathbf z(r_2)} \to K(r_1-r_2),
\]
and a function \(f(r)=\sum_k \alpha_k K(r-r_k)\) is represented by the vector
\[
\mathbf y_f=\sum_k \alpha_k\,\mathbf z(r_k).
\]
Point evaluation becomes \(\mathbf y_f^\top \overline{\mathbf z(s)}\), addition is vector superposition, binding implements shifts, and binding of function vectors implements convolution. With fractional power encoding and uniformly distributed phases, the induced kernel is the sinc kernel, and the resulting RKHS is the space of band-limited functions [2109.03429].

The two programs are mathematically unrelated in construction, but they converge on a shared idea: a vector can be treated not merely as data, but as a portable representation of a function or task. This suggests why the same phrase has proved attractive across functional analysis, matrix perturbation theory, mechanistic interpretability, continual learning, and randomized function-space computation.

Source: https://www.emergentmind.com/topics/function-vector-fv