NeuroDeX: Decompiling DNN Executables
- NeuroDeX is a decompiler for deep neural network executables that recovers high-level models from optimized binaries produced by compilers like TVM and GLOW.
- It combines static extraction, dynamic instrumentation, and LLM-driven semantic analysis to accurately recognize operators and recover attributes.
- The tool reconstructs end-to-end models with near-identical inference accuracy for non-quantized cases and functionally similar outcomes for quantized settings.
NeuroDeX is a decompiler for deep neural network executables that targets on-device models compiled for deployment on edge devices. Introduced in "NeuroDeX: Unlocking Diverse Support in Decompiling Deep Neural Network Executables" (Li et al., 8 Sep 2025), it is designed to recover high-level models from binaries produced by deep learning compilers such as TVM and GLOW, with explicit support for compilation optimizations, different architectures, and quantized compiled models. Its central methodological claim is that semantic understanding by LLMs, when combined with dynamic analysis, can support operator type recognition, operator attribute recovery, and end-to-end model reconstruction from optimized binaries that are difficult to analyze with static or symbolic techniques alone.
1. Problem setting and motivation
On-device deep learning models have extensive real world demands, and deep learning compilers efficiently compile models into executables for deployment on edge devices. Those executables may face the threat of reverse engineering, making decompilation relevant for IP protection and security auditing (Li et al., 8 Sep 2025).
Previous studies attempted to decompile DNN executables, but the paper identifies several persistent obstacles: accurate operator type recognition under compiler optimizations and operator fusion, limited support across architectures and model families, difficulty analyzing quantized executables, heavy reliance on symbolic execution or static analysis, and lack of adaptation to compiler and runtime diversity. NeuroDeX is positioned as a response to those constraints. It aims not merely to recover isolated functions, but to reconstruct a high-level PyTorch model from an executable while remaining effective under TVM and GLOW optimizations, on both x86 and aarch64, and in both non-quantized and quantized settings.
A recurring misconception in DNN decompilation is that executables can always be mapped back to an original source model through straightforward static inspection. NeuroDeX is built on the opposite premise: compiler transformations such as operator fusion, aggressive layout transformations, and quantization-specific code generation obscure both operator semantics and operator boundaries. The system therefore treats decompilation as a combined semantic and dynamic inference problem rather than a purely syntactic one.
2. Pipeline architecture
NeuroDeX is organized as a four-stage pipeline: operator function extraction, operator type recognition, operator attribute recovery, and model reconstruction (Li et al., 8 Sep 2025).
In operator function extraction, the binary is analyzed with Ghidra to extract operator function information from both disassembled and decompiled code. This stage collects parameter dimensions, parameter types, and layout information that will later constrain candidate operators and attribute values. The pipeline explicitly handles optimized layouts such as TVM’s form, denoted as NCHWc, and maps them back to standard forms.
Operator type recognition is split into a coarse and a fine stage. The coarse stage uses statically extracted parameter shape and type information for candidate filtering. The fine stage invokes LLMs to interpret decompiled code and determine operator semantics. For fused operators, NeuroDeX supplements LLM-based code understanding with runtime execution tracing, parameter address mapping, and register tracking or taint analysis, allowing it to reason about operators whose boundaries have been altered by compiler optimization.
Operator attribute recovery targets values such as kernel size, stride, and padding, which may be hidden by code transformations or only inferable at runtime. Dynamic analysis is performed with Intel Pin on x86 and GDB on aarch64 to monitor memory, parameter flow, tensor I/O, and execution order. LLM semantic extraction is then used for ambiguous attributes that appear as constants or code patterns in the decompiled representation.
Model reconstruction uses dynamic instrumentation to recover computational-graph topology, weight locations, and dataflow dependencies. The output is a high-level PyTorch implementation intended to reproduce the executable’s behavior, either exactly for non-quantized models or functionally for quantized models.
3. Operator semantics and attribute recovery
NeuroDeX organizes operators into a taxonomy that includes layout transformation operators such as concat and flatten, element-wise operators such as relu, add, and softmax, reduction operators such as maxpool and avgpool, and complex operators such as conv and dense, including fused forms (Li et al., 8 Sep 2025).
The paper’s central claim about operator recognition is that static extraction alone is insufficient once compiler optimizations obscure operator structure. NeuroDeX therefore uses a two-step strategy. First, parameter shape and type information narrow the candidate set. Second, an LLM is prompted with decompiled code and a candidate list to identify the operator. For fused operator recognition, dynamic analysis logs parameter addresses and tracks runtime memory and register flows, after which the LLM is used for code-pattern identification. This division of labor is important: dynamic analysis constrains the search space with runtime facts, while the LLM supplies semantic interpretation.
Attribute recovery combines numerical constraints and code semantics. For convolution, the paper gives the standard spatial relation
and states that unknown stride and padding are determined by constrained enumeration. For maxpool and avgpool, kernel size may be extracted directly from code patterns through LLM interpretation; the paper gives the example that the presence of “1/49” in code implies kernel size $7$. When direct code cues are insufficient, NeuroDeX dynamically simulates outputs under candidate settings until a match is found.
These procedures indicate that attribute recovery is not treated as a uniform reverse-mapping task. Some attributes are inferred algebraically from tensor dimensions, some are read from low-level constants, and some require runtime comparison. This suggests a heterogeneous recovery regime in which different operator families demand different combinations of static, dynamic, and semantic analysis.
4. Computational-graph recovery and quantized executables
The model reconstruction stage records function call sequences, parameter addresses, and input or output connectivity to recover computational-graph topology (Li et al., 8 Sep 2025). Weight recovery proceeds by dumping weights from parameter addresses at runtime and then transforming them according to the identified layout. In the non-quantized setting, the stated goal is recovery of nearly identical high-level models.
Quantized executables require additional analysis because integer weights and quantization-specific transformations separate the recovered representation from the original floating-point model. NeuroDeX distinguishes two quantization modes. For global scale quantization, it extracts shift or multiplication values from code and transforms dumped integer weights according to
For KL-divergence quantization, where scale factors are per-layer, the system uses substitute training: recovered weights are frozen and only scale variables are retrained on outputs, so that the reconstructed model approximates the original executable with minimal data.
This distinction is critical to interpreting the reported results. For non-quantized models, NeuroDeX reports functionally identical recovery after error correction. For quantized models, the paper is more limited and more precise: NeuroDeX can recover functionally similar high-level models. The system therefore does not claim exact structural inversion of quantized binaries; instead, it aims at behavioral recovery under quantization-induced information loss and optimization.
5. Experimental evaluation
The empirical study uses 96 DNN executables across 12 common DNN models, including 88 non-quantized executables and 8 quantized executables, with experiments conducted on x86 and aarch64 and across TVM 0.7, 0.8, 0.9dev, 0.17 and GLOW 2020, 2021, 2022 (Li et al., 8 Sep 2025). The implementation comprises about 8K LOC Python and about 1K LOC C++, uses Ghidra v11.1.2 for decompilation, Intel Pin for x86 dynamic analysis, GDB for aarch64 dynamic analysis, and PyTorch for model export. GPT-4o is the main LLM, with GPT-4.1, Deepseek-v3, Gemini, and GPT-4o mini also evaluated.
| Aspect | Result | Notes |
|---|---|---|
| Operator type recognition | 99.22% TRA on TVM; 97.62% TRA on GLOW | Non-quantized evaluation |
| Attribute recovery | Nearly 100% ARA | After recovery pipeline |
| Model inference accuracy | 100% after error correction | Non-quantized models |
| Quantized recovery | 72% average top-1; 86% top-5 | Functionally similar models |
| Efficiency | 2.8×–12.8× faster than BTD | Compared with prior SOTA |
The paper states that all non-quantized recovered models are functionally identical, with no loss in inference accuracy after error correction. For quantized models, the reported aggregate result is an average top-1 accuracy of 72% and top-5 accuracy of 86%. The quantized summary is further differentiated by quantization mode: global scale quantization yields recovered models with top-1 accuracy at least 26.4% and top-5 accuracy at least 40.6%, whereas KL-divergence quantization yields recovered models with top-1 accuracy at least 73% and top-5 accuracy at least 96.4%. The paper also reports that random re-training gives only 0.6% top-1 accuracy, which is presented as evidence that the substitute-training procedure retains executable-specific information unavailable to naive retraining.
The evaluation includes architectures such as EfficientNet, Inceptionv1, MobileNetv2, ResNet18/34, VGG16, ShuffleNetv1/v2, MnasNet, SqueezeNet, Emotion, and SuperRes. This breadth is used to support the paper’s claim of broader coverage than prior DNN executable decompilers.
6. Relation to prior decompilers and significance
NeuroDeX is compared in the paper with Libsteal, Shi et al, DND, Neuroscope, and BTD. The comparison claims that earlier systems provide only partial support for compilation optimization, cross-architecture operation, and quantized executables, whereas NeuroDeX supports all three (Li et al., 8 Sep 2025).
| Work | Optimization / Cross-Arch | Quantized |
|---|---|---|
| Libsteal | × / × | × |
| Shi et al | × / × | × |
| DND | × / ✓ | × |
| Neuroscope | × / ✓ | × |
| BTD | ✓ / × | × |
| NeuroDeX | ✓ / ✓ | ✓ |
The paper attributes its efficiency gains over BTD primarily to LLM-powered code understanding and avoidance of heavy symbolic execution. It also notes that LLM choice can affect accuracy, although top-tier models such as GPT-4o, GPT-4.1, and Deepseek-v3 all produce high-quality results. This situates NeuroDeX within a broader methodological shift: semantic analysis by general-purpose LLMs is being incorporated into reverse engineering workflows that previously depended more heavily on hand-crafted static rules or symbolic reasoning.
A plausible implication is that DNN executable decompilation is becoming less dependent on compiler-specific heuristics and more dependent on hybrid semantic-runtime pipelines. Within that interpretation, NeuroDeX is significant less because it solves decompilation in an abstract sense than because it demonstrates a concrete operating point: optimized and quantized binaries can be reverse engineered with high recovery accuracy when runtime dataflow extraction is coupled to LLM-based interpretation of low-level code.