Papers
Topics
Authors
Recent
Search
2000 character limit reached

Black-Box Inference of LLM Architectural Properties with Restrictive API Access

Published 1 Jul 2026 in cs.LG, cs.AI, cs.CL, and cs.CR | (2607.01313v1)

Abstract: In practice, most commercial LLM providers do not publicly release details of underlying LLM architectures. However, prior work has shown that given limited API access to an LLM (namely, top-kk logits and/or a logit bias function), one can recover certain architectural details of an LLM, such as the hidden dimension of the feed-forward network. Perhaps in response to these results, most commercial LLM providers have restricted their APIs to expose only the single logit for each decoded token, and they no longer give users the ability to bias logits. We show that even under current restrictive APIs, several architectural parameters are still recoverable. We present NightVision, an attack that uses restrictive black-box API access to estimate the hidden dimension, depth, and parameter count of an LLM. Algorithmically, NightVision relies on a novel common set prompting technique in which multiple prompts expose log probabilities for the same set of output tokens; a spectral analysis of these results is used to infer hidden dimension. NightVision additionally uses end-to-end time to first token (TTFT) measurements and the estimated hidden dimension to estimate depth and parameter count. We empirically evaluate NightVision on 32 open-source LLMs, recovering hidden dimension to within 23% average relative error across all models (9% on MoE models), and depth and parameter count to within 53% for models exceeding three billion parameters. We run extensive ablations to demonstrate how these accuracies scale with token budget and model properties. Overall, our results suggest that current LLM APIs are not sufficiently restricted to fully obfuscate the architectural details of their underlying models.

Summary

  • The paper presents NightVision, a black-box attack that accurately infers hidden dimension, depth, and parameter count using minimal API outputs.
  • It introduces a common-set prompting technique and arctan-transformed spectral analysis to robustly estimate hidden dimensions from noisy, quantized log-probabilities.
  • Timing side channels are exploited to model depth and parameter count, achieving mean relative errors of 23% for hidden dimension and 53% for depth and total parameters.

Black-Box Architectural Inference of LLMs under Restrictive APIs

Introduction and Motivation

API-based access to LLMs is increasingly common but model providers restrict probabilistic outputs, typically exposing only log-probabilities for generated tokens and disabling logit-bias and top-kk logit access. This significantly limits direct spectral attacks on model weights or dimensions as proposed by Carlini et al. Consequently, the viability of inferring core architectural parameters of deployed LLMs—namely hidden dimension dd, depth LL, and parameter count PP—from such minimal black-box access is a central question for model auditing and threat modeling.

The paper presents NightVision, a generic architectural inference attack. NightVision leverages the minimal outputs permitted by current APIs (single-token log-probabilities) and exploits timing side channels to infer dd, LL, and PP with significant accuracy. The core insight is that the softmax bottleneck remains recoverable using carefully crafted prompt/query structures and runtime measurements, even without explicit access to full or partial logit vectors.

Methodological Framework

Common-Set Prompting for Hidden Dimension Estimation

NightVision's first phase introduces a novel common-set prompting technique. For DD high-entropy prompts, NN stochastic next-token samples are drawn per prompt (at high temperature), recording realized tokens and corresponding log-probabilities. The process is designed to maximize vocabulary coverage per prompt:

Figure 1

Figure 2: Next-token distributions for a high-entropy prompt (top) and low-entropy prompt (bottom), with and without temperature scaling. High-entropy prompts and increased temperature drastically flatten the output distribution, improving coverage.

A common-token set CC (tokens observed for every prompt) is constructed, and a dd0 dense log-probability matrix is greedily extracted for spectral processing. This approach, unlike logit-bias attacks, does not require reconstruction of full logit vectors. Instead, it leverages the high-probability intersection tokens over many diverse high-entropy prompts.

The paper provides a formal sample complexity analysis of this intersection operation, upper-bounding the required dd1 for coverage as a function of desired set size dd2, vocabulary size dd3, and prompt count dd4:

Figure 3

Figure 3: Required samples dd5 per prompt to achieve a common token set of size dd6 for various prompt counts, compared to theoretical bounds. Empirical costs closely track theoretical projections.

Spectral Rank Recovery under Noise

Given quantized log-probabilities (rather than raw logits) and floating-point noise, classical SVD-based elbow techniques are insufficiently robust. Instead, NightVision employs an arctan-transformed spectrum and inflection-point detection across the ordered eigenvalues of the Gram matrix:

Figure 4

Figure 4

Figure 4

Figure 5: Spectra used to recover dd7; the logit-bias attack yields a sharp eigenvalue drop (left), whereas the restricted NightVision setting results in a smooth decay, addressed by arctan transform and inflection-based detection (right panels).

Timing-Based Depth and Parameter Count Estimation

The second phase exploits wall-clock timing of next-token generation with long input sequences, where runtime is dominated by prefill (attention) costs. By fitting a learned linear scaling model dd8 (where dd9 is measured TTFT, LL0 is depth, LL1 is prompt length, LL2 is hidden dim), depth can be inferred given independently recovered LL3. The approach requires calibration against reference models but is robust as long as the dominant compute regime is reached.

Figure 6

Figure 1: Measured runtime versus predicted runtime using scaling relation across number of layers, hidden dimension, and input size.

Experimental Evaluation

Overall Accuracy

Evaluations on 32 open-source LLMs (both dense and MoE) reveal that NightVision recovers LL4 (hidden dimension) with a mean relative error of LL5 across all models (and LL6 for MoEs). For large models (LL73B parameters), jointly inferred LL8 and LL9 (with raw PP0 from phase 1) are recovered within PP1 mean relative error.

With increased token/query budgets, especially for large models, error steadily decreases, indicating that the spectral technique does not saturate at budget-constrained performance:

Figure 5

Figure 4: Mean relative error for PP2 and PP3 estimates versus total prompt-token budget for the timing method, assuming perfect PP4.

Figure 7

Figure 8: NightVision’s estimation error for PP5, PP6, and PP7 as a function of token budget, API calls, and estimated cost. PP8 generally improves fastest, PP9 is harder to estimate, and larger budgets yield lower errors.

Importantly, for large-scale production models, TTFT-based depth/size inference converges rapidly and is not sensitive to token/query budget. Estimation accuracy for dd0 via spectrum improves monotonically, saturating once the eigenspace is sufficiently sampled (see also Figure 9 for spectrum slope visualization).

Cost Comparison to Logit-Bias Attacks

NightVision achieves attacks in the same threat model as currently deployed APIs, at a cost {\bf 12-190dd1} higher in token budget than previous attacks assuming logit bias or top-dd2 logit access.

Error Distribution and Cross-Platform Transfer

The error analysis decomposes residuals by model family and platform (e.g., timing calibration between H200 and A100 GPUs):

Figure 10

Figure 6: Relative error for estimates of layers, parameters, and hidden dimension for various models (same-platform). Errors are lower for large models and for cases with sufficient context length.

Figure 11

Figure 10: Cross-platform relative error for layer and total parameter count. Transfer calibration (H200 dd3 A100) maintains similar error magnitude.

The Spectral estimator is sensitive to the saturation of the common set; for some models (LLAMA-3.1-8B), the elbow is not realized at feasible budgets, but theoretical sample complexity predictions remain tight (Fig. 2, 10).

Implications and Future Work

This research demonstrates that API restrictions around logit access—though effective in disabling prior extraction methods—do not prevent significant leakage of model internals when attackers can systematically probe output entropy and API timings. The bottleneck remains the expensive token/query budget, which, while substantial, is not outside possible adversarial means.

From a practical perspective, NightVision-type attacks present viable strategies for forensic and security analysis, e.g., regulatory audits or model copyright assurance (cf. recent fingerprinting and auditing literature (2607.01313)). The theoretical foundation can serve as a basis for new red-teaming and watermarking strategies, as well as motivating more robust API hardening (e.g., timing obfuscation or output perturbation beyond simple logit restriction).

Architecturally, the generalization to architectures with non-uniform or state-space-based output distributions is a pressing next step. Additionally, future work should refine common-set search procedures, sample complexity tightness, and extend timing side channels to distributed or sharded LLM deployments.

Conclusion

NightVision demonstrates that transformer LLM architectural parameters—hidden dimension, depth, total parameter count—are recoverable to substantial accuracy solely from restrictive API access and query timing side channels. The attack is agnostic to model family, robust to platform transfer, and provides a lower bound on what can be inferred under strong API constraints. These results highlight the persistent leakage of proprietary model attributes through indirect signals and call into question the sufficiency of current API-level defense mechanisms.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 3 tweets with 17 likes about this paper.