PhyIP: Neural, Phylogenetic, and IoT Protocols
- PhyIP is a multi-domain framework defining protocols for non-invasive probing in neural models using linear decoding with metrics like Pearson ρ and MAPE.
- It also enables high-throughput phylogenetic informativeness profiling by leveraging site-specific substitution rates and efficient database summarization.
- Additionally, PhyIP supports robust IoT authentication via physical-layer identity, using maximum-entropy quantization of RF features for secure device identification.
PhyIP refers to several rigorously defined protocols and computational frameworks in distinct research areas, each denoted by the acronym “PhyIP” in their respective contexts. These span: (1) non-invasive linear probing for physical law discovery in neural world models, (2) high-throughput estimation of phylogenetic informativeness in molecular systematics, and (3) physical-layer identity mechanisms for authentication in NDN-IoT environments. The following article systematically covers each of these usages and their foundational aspects.
1. Non-Invasive Physical Probing in Neural World Models
The term “PhyIP” was introduced by Internò et al. (2026) to denote a non-invasive protocol for evaluating whether self-supervised neural world models genuinely internalize physical laws, as opposed to encoding statistical shortcuts. The main principle is to test whether physically relevant quantities are linearly decodable from the frozen neural representation—motivated by the Linear Representation Hypothesis (LRH), which asserts that in robust models, physical features correspond to linear directions in latent space (Internò et al., 12 Feb 2026).
Protocol Overview
- Feature Extraction: A world model (e.g., U-Net, UNetConvNext, FNO, TFNO) is pretrained on a next-step prediction loss, producing latent activations .
- Linear Probe Training: A linear readout (with weights ) is trained to regress the target physical quantity (e.g., total internal energy, gravitational force), with the backbone strictly frozen.
- OOD Evaluation: The probe is validated under out-of-distribution (OOD) generative parameters. Key metrics are Pearson correlation and MAPE.
A successful probe showing under OOD indicates the model contains the physical law as a linear feature.
Mathematical Foundations
The LRH formally requires that there exists such that:
For dynamics, the ability to recover robustly over OOD confirms that the model encodes the law, not just ad-hoc features.
Error decomposition bounds the linear probe’s error in terms of the self-supervised loss and the curvature of the target physical functional:
0
Experimental Results
The protocol was tested on fluid dynamics (TRL-2D, RSG-3D, SN-3D) and orbital mechanics (synthetic and real Solar System). Key findings:
| Dataset / Task | PhyIP ρ | Invasive FT ρ | PhyIP MAPE | Invasive FT MAPE |
|---|---|---|---|---|
| TRL-2D (UNetConvNext) | 0.83 | 0.80–0.81 | 36.9% | 32.5–41.2% |
| RSG-3D | 0.91 | ≤0.82 | 18.2% | — |
| SN-3D | <0.21 | 0.71 | >135% | 18.3% |
| Orbital Mechanics | 0.91 | 0.05–0.88 | 24.5% | 22.5–65.0% |
Invasive probes (fine-tuned MLPs or full model adaptation) often destroy or distort underlying latent physics, yielding deceptively high downstream scores. PhyIP, by contrast, preserves the genuine latent invariants (Internò et al., 12 Feb 2026).
Broader Implications
The protocol directly addresses the “observer effect” arising from adaptation-based evaluation: adaptation itself can collapse or overwrite meaningful representations. Consequently, high performance after fine-tuning does not guarantee internalization of fundamental laws, but may reflect relearning in the probe.
2. High-Throughput Phylogenetic Informativeness Profiling
“PhyIP” also denotes “phylogenetic informativeness profiling” in molecular phylogenetics, particularly as implemented in the TAPIR pipeline (Faircloth et al., 2012). This framework quantifies the expected phylogenetic signal provided by each site or locus as a function of evolutionary time, enabling informed marker selection and comparative analysis across thousands of loci.
Theoretical Basis
Site-specific PI (phylogenetic informativeness) for inferred substitution rate 1 at site 2 is:
3
Net informativeness at time 4 is:
5
For an epoch 6:
7
Implementation in TAPIR
TAPIR pipeline steps:
- Input Preparation: Dated tree (Newick), locus-wise alignments, times/epochs.
- Model Selection and Site-Rate Inference: For each locus, HYPHY is used to select the best-fitting model and infer site-specific rates 8.
- PI Matrix Construction: For a grid 9, compute 0, and sum to obtain 1.
- Aggregation: Summaries stored in SQLite database, enabling rapid visualization and downstream analysis.
TAPIR supports high-throughput computation—hundreds of loci per hour—far surpassing prior tools. Intermediate results (e.g., JSON site-rates) are cached for efficient re-use with new time grids.
Benchmarking and Performance
Empirical benchmarks:
| Dataset Size | Time to Completion | Throughput Rate |
|---|---|---|
| 20 loci (4 kb each) | ~4 min | ~300 loci/hour |
| 183 loci (median 500bp) | ~35 min | ~313 loci/hour |
| ~900 UCE loci | ~5 hours | ~180 loci/hour |
TAPIR’s architecture—with locus model selection, vectorized computation, and database back-end—enables systematic scaling for modern phylogenomic datasets (Faircloth et al., 2012).
3. Physical-Layer Identity for NDN-IoT Authentication
In the context of NDN-IoT security, “PhyIP” refers to a PHY-ID scheme exploiting device-specific RF imperfections as unique hardware fingerprints (Hao et al., 2019). Here, the protocol augments cryptographic authentication by using maximum-entropy quantization of estimated I/Q imbalance features for device identification in mobile edge computing environments.
Protocol Mechanics
- RF Feature Extraction: From each end-device, the gateway/MECD estimates a condensed RF imperfection (e.g., amplitude/phase mismatch).
- Offline Enrollment: The observed feature is quantized (using maximum-entropy binning) and hashed as the PHY-ID.
- Two-Step Online Authentication:
- Coarse check: Does the candidate’s feature bin match the stored device?
- Fine-grained LRT/GLRT: Multiple samples are compared statistically (using, e.g., Neyman–Pearson LRT or GLRT depending on variance knowledge).
- Offloaded Signing: Upon authentication, data signing workloads are optimally partitioned between MECD and end-device using convex optimization to minimize overall latency subject to CPU constraints.
Security and Performance
- Robustness: The scheme is intrinsically resilient to replay and key-compromise attacks since PHY-ID is hardware-bound and not forgeable via cryptographic key-theft.
- Differentiation Rate and Correct Authentication Probability (CAP): MEB quantization and two-step GLRT achieve >95% CAP in dense IoT conditions (2000 devices), substantially outperforming uniform binning.
- Efficiency: Offloading achieves 60–90% signing speedup over device-only signing under reasonable MECD loads; even with very limited MECD availability, 30–40% speedup is maintained (Hao et al., 2019).
4. Cross-Domain Summary Table
| Domain | Role of PhyIP | Key Principle | Primary Metric/Output |
|---|---|---|---|
| Neural World Models | Non-invasive probe of latent physics | Linear Representation Hypothesis | OOD Pearson ρ/MAPE |
| Comparative Phylogenetics | Informativeness profiling (TAPIR) | Locus/site PI, ML rate fitting | PI(t), SQLite summary tables |
| IoT Authentication/NDN Security | Physical-layer identity quantization | Maximum-entropy binning, two-step LRT | Authentication rate, CAP |
5. Perspectives and Future Directions
In neural world modeling, PhyIP’s rigorous non-invasive probing is now considered necessary for distinguishing genuine scientific understanding from superficial adaptation and overfitting. Adopting OOD validation, fixed-capacity probes, and symbolic regression downstream from frozen backbones is suggested as foundational for upcoming general-purpose world models.
In molecular phylogenetics, the throughput and modelling flexibility of TAPIR-style PhyIP has rendered large-scale, model-fit-aware locus selection a tractable and empirically robust standard, enabling greater precision in phylogenomic study design.
In NDN-IoT security, the integration of hardware-level PHY-ID with cryptographic schemes via maximum-entropy quantization and MEC-optimized offloading underlines a growing trend toward multi-factor, cross-domain authentication in resource-constrained and adversarial environments.
A plausible implication is that PhyIP, across all these domains, represents a shift towards protocols that minimize “observer effects,” maximize interpretability, and preserve the integrity of the underlying information—whether scientific, phylogenetic, or security-related—against both algorithmic and adversarial distortion.