- The paper presents NightVision, a black-box attack that accurately infers hidden dimension, depth, and parameter count using minimal API outputs.
- It introduces a common-set prompting technique and arctan-transformed spectral analysis to robustly estimate hidden dimensions from noisy, quantized log-probabilities.
- Timing side channels are exploited to model depth and parameter count, achieving mean relative errors of 23% for hidden dimension and 53% for depth and total parameters.
Black-Box Architectural Inference of LLMs under Restrictive APIs
Introduction and Motivation
API-based access to LLMs is increasingly common but model providers restrict probabilistic outputs, typically exposing only log-probabilities for generated tokens and disabling logit-bias and top-k logit access. This significantly limits direct spectral attacks on model weights or dimensions as proposed by Carlini et al. Consequently, the viability of inferring core architectural parameters of deployed LLMs—namely hidden dimension d, depth L, and parameter count P—from such minimal black-box access is a central question for model auditing and threat modeling.
The paper presents NightVision, a generic architectural inference attack. NightVision leverages the minimal outputs permitted by current APIs (single-token log-probabilities) and exploits timing side channels to infer d, L, and P with significant accuracy. The core insight is that the softmax bottleneck remains recoverable using carefully crafted prompt/query structures and runtime measurements, even without explicit access to full or partial logit vectors.
Methodological Framework
Common-Set Prompting for Hidden Dimension Estimation
NightVision's first phase introduces a novel common-set prompting technique. For D high-entropy prompts, N stochastic next-token samples are drawn per prompt (at high temperature), recording realized tokens and corresponding log-probabilities. The process is designed to maximize vocabulary coverage per prompt:

Figure 2: Next-token distributions for a high-entropy prompt (top) and low-entropy prompt (bottom), with and without temperature scaling. High-entropy prompts and increased temperature drastically flatten the output distribution, improving coverage.
A common-token set C (tokens observed for every prompt) is constructed, and a d0 dense log-probability matrix is greedily extracted for spectral processing. This approach, unlike logit-bias attacks, does not require reconstruction of full logit vectors. Instead, it leverages the high-probability intersection tokens over many diverse high-entropy prompts.
The paper provides a formal sample complexity analysis of this intersection operation, upper-bounding the required d1 for coverage as a function of desired set size d2, vocabulary size d3, and prompt count d4:

Figure 3: Required samples d5 per prompt to achieve a common token set of size d6 for various prompt counts, compared to theoretical bounds. Empirical costs closely track theoretical projections.
Spectral Rank Recovery under Noise
Given quantized log-probabilities (rather than raw logits) and floating-point noise, classical SVD-based elbow techniques are insufficiently robust. Instead, NightVision employs an arctan-transformed spectrum and inflection-point detection across the ordered eigenvalues of the Gram matrix:



Figure 5: Spectra used to recover d7; the logit-bias attack yields a sharp eigenvalue drop (left), whereas the restricted NightVision setting results in a smooth decay, addressed by arctan transform and inflection-based detection (right panels).
Timing-Based Depth and Parameter Count Estimation
The second phase exploits wall-clock timing of next-token generation with long input sequences, where runtime is dominated by prefill (attention) costs. By fitting a learned linear scaling model d8 (where d9 is measured TTFT, L0 is depth, L1 is prompt length, L2 is hidden dim), depth can be inferred given independently recovered L3. The approach requires calibration against reference models but is robust as long as the dominant compute regime is reached.

Figure 1: Measured runtime versus predicted runtime using scaling relation across number of layers, hidden dimension, and input size.
Experimental Evaluation
Overall Accuracy
Evaluations on 32 open-source LLMs (both dense and MoE) reveal that NightVision recovers L4 (hidden dimension) with a mean relative error of L5 across all models (and L6 for MoEs). For large models (L73B parameters), jointly inferred L8 and L9 (with raw P0 from phase 1) are recovered within P1 mean relative error.
With increased token/query budgets, especially for large models, error steadily decreases, indicating that the spectral technique does not saturate at budget-constrained performance:

Figure 4: Mean relative error for P2 and P3 estimates versus total prompt-token budget for the timing method, assuming perfect P4.

Figure 8: NightVision’s estimation error for P5, P6, and P7 as a function of token budget, API calls, and estimated cost. P8 generally improves fastest, P9 is harder to estimate, and larger budgets yield lower errors.
Importantly, for large-scale production models, TTFT-based depth/size inference converges rapidly and is not sensitive to token/query budget. Estimation accuracy for d0 via spectrum improves monotonically, saturating once the eigenspace is sufficiently sampled (see also Figure 9 for spectrum slope visualization).
Cost Comparison to Logit-Bias Attacks
NightVision achieves attacks in the same threat model as currently deployed APIs, at a cost {\bf 12-190d1} higher in token budget than previous attacks assuming logit bias or top-d2 logit access.
The error analysis decomposes residuals by model family and platform (e.g., timing calibration between H200 and A100 GPUs):

Figure 6: Relative error for estimates of layers, parameters, and hidden dimension for various models (same-platform). Errors are lower for large models and for cases with sufficient context length.

Figure 10: Cross-platform relative error for layer and total parameter count. Transfer calibration (H200 d3 A100) maintains similar error magnitude.
The Spectral estimator is sensitive to the saturation of the common set; for some models (LLAMA-3.1-8B), the elbow is not realized at feasible budgets, but theoretical sample complexity predictions remain tight (Fig. 2, 10).
Implications and Future Work
This research demonstrates that API restrictions around logit access—though effective in disabling prior extraction methods—do not prevent significant leakage of model internals when attackers can systematically probe output entropy and API timings. The bottleneck remains the expensive token/query budget, which, while substantial, is not outside possible adversarial means.
From a practical perspective, NightVision-type attacks present viable strategies for forensic and security analysis, e.g., regulatory audits or model copyright assurance (cf. recent fingerprinting and auditing literature (2607.01313)). The theoretical foundation can serve as a basis for new red-teaming and watermarking strategies, as well as motivating more robust API hardening (e.g., timing obfuscation or output perturbation beyond simple logit restriction).
Architecturally, the generalization to architectures with non-uniform or state-space-based output distributions is a pressing next step. Additionally, future work should refine common-set search procedures, sample complexity tightness, and extend timing side channels to distributed or sharded LLM deployments.
Conclusion
NightVision demonstrates that transformer LLM architectural parameters—hidden dimension, depth, total parameter count—are recoverable to substantial accuracy solely from restrictive API access and query timing side channels. The attack is agnostic to model family, robust to platform transfer, and provides a lower bound on what can be inferred under strong API constraints. These results highlight the persistent leakage of proprietary model attributes through indirect signals and call into question the sufficiency of current API-level defense mechanisms.