vLLM Hook: Inference-Time Model State Programming
- vLLM Hook is an open-source plugin that enables programming of transformer model internals during inference, supporting both passive monitoring and active interventions.
- It distinguishes between passive programming for capturing internal states and active programming for modifying tensors, enhancing features like prompt injection detection and RAG.
- The configuration-driven design uses JSON to specify hooks on selected layers and heads, with demonstrated use cases including activation steering and enhanced retrieval.
Searching arXiv for the primary paper and vLLM background. vLLM Hook is an opensource plug-in for programming the internal states of models deployed on vLLM, the open-source library used for model serving and inference of transformer-based LLMs. It is introduced as a response to a specific limitation in the current implementation of vLLM: restricted programmability of internal model states during inference. That restriction excludes a range of test-time alignment and enhancement methods, including adversarial-prompt detection based on attention patterns and activation steering. vLLM Hook v0 addresses this gap through configuration-driven interception of model internals and exposes two principal operating modes—passive programming and active programming—while demonstrating prompt injection detection, enhanced retrieval-augmented retrieval (RAG), and activation steering as initial use cases (Ko et al., 2 Feb 2026).
1. Motivation and problem setting
The central problem addressed by vLLM Hook is the inability, in standard vLLM deployment, to inspect or modify intermediate tensors during generation. In the formulation given for vLLM Hook v0, this is not treated as a minor observability issue but as a barrier to entire classes of inference-time methods. The abstract identifies two representative consequences: it prevents the detection of adversarial prompts based on attention patterns, and it prevents the adjustment of model responses based on activation steering (Ko et al., 2 Feb 2026).
Within that framing, vLLM Hook is positioned as an inference-time programming layer rather than a training-time method. Its purpose is to let operators specify, by configuration, which internal states should be captured or altered. The paper’s summary characterizes the resulting workflow as a three-part cycle—Build (offline), Probe (config file), and Program (live hooking). This suggests a deliberate attempt to connect offline interpretability or steering artifacts, such as precomputed steering vectors or selected “important heads,” with a production-oriented inference runtime.
A recurring misconception is that such a system must retrain or fine-tune the model to be useful. The vLLM Hook description states the opposite for its passive mode: generation output is identical to a vanilla vLLM call when hooks only read and save intermediate states. In that sense, the plug-in is defined by intervention capability, not by obligatory intervention.
2. Runtime architecture and forward-pass interception
The architecture centers on a wrapper class, HookLLM, which encapsulates a standard vLLM LLM instance together with a JSON configuration naming the layers, heads, or activations to hook. Under the hood, HookLLM launches the vLLM runtime with a custom GPU Worker subclass such as ProbeHookQKWorker, which overrides load_model(...) in order to install PyTorch forward hooks on selected modules (Ko et al., 2 Feb 2026).
The forward-pass interception path is described step by step. User code initializes HookLLM and calls LLM.generate(prompt), which triggers vLLM’s Worker.execute_model. Inside ProbeHookQKWorker.load_model(), the worker first calls super().load_model(...), then parses configuration fields such as layers, heads, and modes, and registers hooks on target attention modules. The example given uses module.register_forward_hook(lambda m, i, o: qkv_hook(i, name)).
Once generation begins, each hooked module invokes a user “hook” callback. In passive mode, the callback captures intermediate tensors into a disk cache; in active mode, it modifies them before downstream computation continues. The summary’s concrete example is qkv_hook(input, name), which extracts query and key tensors and appends them to a per-run cache of the form
6
followed by torch.save(cache, "hook_dir/qk_{run_id}.pt").
The underlying tensor objects are given explicitly. The hidden-state tensor at layer is
where is batch size, is sequence length, and is hidden dimension. The attention weight matrix is
with the number of heads. Query and key caches are represented as
After generation, the interface exposes LLM.analyze(...), which loads saved state files and runs analyzers such as AttntrackerAnalyzer to recompute attention weights or derive task-specific statistics. This separates capture from analysis and makes the hooking mechanism a reusable substrate rather than a single-purpose detector.
3. Configuration model and exposed API
The configuration format is JSON with at least three top-level sections: model_info, params, and one or more hook-specific blocks such as hookq, hookv, or hook_act (Ko et al., 2 Feb 2026). model_info identifies the exact vLLM model, either by name or HuggingFace identifier. params enumerates the internal components to target, including layers, heads, steering vectors, and reduction methods. The hook blocks then specify operational behavior, for example whether all tokens should be captured or only the last token, and where outputs should be written.
The summary gives an example of passive capture over selected attention heads:
7
It also gives an example of an activation-steering specification:
8
The API surface mirrors the configuration logic. Initialization is shown as
9
Generation proceeds through LLM.generate(...), and post hoc analysis through LLM.analyze(...) with an analyzer name and analyzer specification. This configuration-first design implies a declarative interface over low-level PyTorch hook registration, allowing the same deployment runtime to support both observation and intervention without changing the model-serving abstraction.
4. Passive programming and active programming
The conceptual core of vLLM Hook is the distinction between passive programming and active programming. The paper treats these as the two essential features of the plug-in and uses them to organize both methodology and examples (Ko et al., 2 Feb 2026).
| Mode | Behavior | Typical use |
|---|---|---|
| Passive programming | Hooks read and save intermediate states without altering tensors passed to subsequent layers | In-model monitoring, selective retrieval, prompt injection detection |
| Active programming | Hooks overwrite or add to internal tensors during generation | Activation steering, intervention on next-token distribution |
In passive programming, the hook only probes selected internal states—attentions, queries, activations—for subsequent analysis. The description is explicit that model generation remains intact and the output is identical to a vanilla vLLM call. This makes passive programming suitable for diagnostics, monitoring, and reranking, where internal evidence matters but behavioral perturbation is undesirable.
In active programming, the hook intervenes in the forward pass. The summary describes a simple intervention at layer as
where 0 is the pre-activation or post-activation tensor and 1 is a user-supplied steering vector of matching shape. The corresponding pseudocode is a forward hook of the form outputs + Δ, attached, for example, to model.layers[ℓ].mlp. In the manual API example, a tensor loaded from steer_vecs/layer8_delta.pt is added to the output of layer 8’s feed-forward block during generation.
The significance of this distinction is methodological. Passive programming turns inference into an instrumented measurement process; active programming turns it into a controlled intervention process. The plug-in’s claim is that both can be implemented inside the same vLLM serving pipeline.
5. Demonstrated use cases
Version 0 demonstrates three use cases: prompt injection detection, enhanced retrieval, and activation steering (Ko et al., 2 Feb 2026). These are not presented as exhaustive, but as evidence that internal-state programmability can support both analysis and control.
Prompt injection detection is implemented as a passive use case through AttntrackerAnalyzer. The procedure captures query and key tensors from important_heads, reconstructs attention weights, and computes a focus score over the instruction span versus the user query span. The summary expresses the reconstruction as
2
followed by a low-attention flagging criterion for suspect cases. The key point is that the detector does not require model modification during generation; it relies on post hoc analysis of selectively captured internals.
Enhanced retrieval is also passive. The summary labels the method a Selective Attention Reranker and states that it hooks only a small subset of heads during one forward pass, then uses aggregated attention scores to rerank retrieved documents through Contrastive Retrieval Heads. The abstract refers more broadly to enhanced retrieval-augmented retrieval (RAG). In both formulations, the important feature is selective inspection rather than full attention logging over the entire model.
Activation steering is the canonical active use case. A precomputed steering vector 3 is loaded for a particular layer 4, and during each generation step the system applies
5
The summary states that this empirically improves compliance with user prompts without full model retraining. A plausible implication is that vLLM Hook is intended to make offline steering artifacts operational at serving time, rather than merely analyzable in research code.
6. Performance characteristics, limits, and terminological ambiguity
The overview includes both asymptotic and empirical overhead statements. Latency is described as 6 additional forward callbacks per token; if 7 layers are hooked, worst-case cost is approximately 8. For memory, storing 9 for an attention hook across 0 tokens is stated as 1, whereas last_token mode reduces this to 2 (Ko et al., 2 Feb 2026).
The empirical measurements reported for initial experiments are correspondingly concrete: passive Q/K capture on an 8-layer model costs approximately 3–4 ms extra per 1 K tokens; active activation steering through single-layer injection adds less than 5 ms per 1 K tokens; and peak memory increase when caching full attention for length-512 sequences is reported as +200 MB, with last_token mode identified as the tuning mechanism. These figures indicate that the plug-in is designed to preserve the efficiency rationale of vLLM while adding controlled interception points.
Two objective limits follow from the description. First, overhead is sensitive to what is captured: full-state caching is materially more expensive than targeted capture. Second, the plug-in’s capabilities depend on the specificity of the configuration and the granularity of the installed hooks; the system is not described as automatic instrumentation of all internals, but as selective instrumentation of nominated components.
A further point of clarification concerns terminology. In the separate summary of "Can VLMs be used on videos for action recognition? LLMs are Visual Reasoning Coordinators" (Lunia, 2024), the phrase “vLLM Hook” is used differently: it refers to a pattern in which an LLM coordinates multiple VLMs through natural-language communication, collecting their textual outputs and producing a final answer. That usage describes orchestration across models, not programming of internal transformer states. The two notions share the general idea of inserting a coordination or interception layer into inference, but they are technically distinct. In the stricter sense established by "vLLM Hook v0: A Plug-in for Programming Model Internals on vLLM," vLLM Hook denotes a plug-in around vLLM’s GPU workers that declaratively probes or steers internal transformer states at inference time.