Signal Fidelity Index (SFI) Overview
- Signal Fidelity Index (SFI) is a family of domain-specific metrics that evaluate how well a system preserves its informative signal structure.
- In communications, SFI is reflected in measures like EVM and constellation geometry, while in EHR modalities it aggregates diagnostic consistency indicators.
- Time-series applications employ composite signal quality indices to assess waveform morphology, periodicity, and artifact resistance for reliable analysis.
Searching arXiv for papers on “Signal Fidelity Index” and closely related signal-fidelity / signal-quality formulations. Signal Fidelity Index (SFI) denotes a fidelity-oriented assessment of how faithfully a system preserves, encodes, or represents informative signal structure, but current arXiv usage does not converge on a single canonical formula. One line of work defines SFI explicitly as a patient-level measure of diagnostic data quality for dementia prediction, computed from diagnostic specificity, temporal consistency, entropy, contextual concordance, medication alignment, and trajectory stability. Related work in communications and time-series auditing treats signal fidelity operationally through complex-baseband distortion, error vector magnitude (EVM), or composite signal quality indices rather than through a universally named SFI scalar (Cheng et al., 10 Sep 2025, Blosser et al., 2024, Gao et al., 2024).
1. Conceptual Scope
Taken together, these works suggest that “Signal Fidelity Index” functions less as a single standardized metric than as a family of fidelity summaries tied to domain-specific observables. In one case, SFI is a formal arithmetic mean over six interpretable components. In another, signal fidelity is quantified mainly through EVM versus data rate and constellation behavior. In a third, the natural SFI is a composite of signal quality indices (SQIs) or a classifier score estimating whether a signal is clean (Cheng et al., 10 Sep 2025, Blosser et al., 2024, Gao et al., 2024).
| Context | Representation of fidelity | Status of the term “SFI” |
|---|---|---|
| Dementia prediction across heterogeneous EHR data | Unweighted mean of six patient-level components in | Explicitly defined |
| Parametric-amplifier receiving antennas | Complex baseband response, constellations, and EVM vs. symbol rate | No named scalar SFI |
| Time-series signal quality auditing | Single SQIs, fused SQIs, or classifier probability of being clean | Interpreted rather than formally named |
This distinction matters methodologically. In the EHR setting, fidelity refers to the reliability of coded diagnostic evidence. In the communications setting, fidelity refers to preservation of amplitude and phase in a practical QAM waveform. In time-series auditing, fidelity is closely aligned with morphology preservation, periodicity, and resistance to acquisition artifacts. A plausible implication is that any use of the label “SFI” is only meaningful relative to the signal model, the downstream task, and the observables used to quantify degradation.
2. Communications Interpretation
In parametric-amplifier receiving antennas, signal fidelity is defined operationally: a receiver has high signal fidelity if the complex baseband signal at the output matches the transmitted QAM signal, in amplitude and phase, with minimal distortion and interference. The paper models an RF plane-wave input as , defines the complex baseband response , and constructs the received baseband signal for a QAM sequence by superposition of phase-dependent step responses:
The main quantitative fidelity measure is EVM after equalization, plotted as a function of incident data rate. The study uses a 16-QAM constellation, 2048 symbols from a pseudorandom bit sequence, a 100 MHz carrier, and a 128-tap linear equalizer with 16:1 downsampling; the example constellation plots use 0.5 Msym/s (Blosser et al., 2024).
The comparison among an LTI receiver, a nondegenerate time-varying parametric receiver (NDTV), and a degenerate time-varying parametric receiver (DTV) shows why frequency-domain bandwidth alone is not a sufficient fidelity proxy. Peak received power for the LTI configuration at 100 MHz is dBW. The DTV configuration exhibits a 4.9 dB spike at 100 MHz, and the reported dB fractional bandwidths are for LTI, for NDTV, and for DTV. Yet the DTV architecture suffers from interference produced by an in-band difference harmonic; its gain becomes phase-dependent, its pre-equalization constellation is “highly compressed along one dimension,” and its post-equalization EVM remains worse than that of NDTV despite the apparent bandwidth advantage. By contrast, NDTV lacks that detrimental overlap, shows phase-independent gain, and exhibits increased signal throughput over the reference LTI receiver (Blosser et al., 2024).
In this setting, the closest analogue to an SFI is therefore not a named scalar but a fidelity profile: phase dependence of , qualitative constellation spread, and low EVM at a given symbol rate. The paper explicitly notes that it does not introduce a named “Signal Fidelity Index” as a single scalar formula; rather, it uses standard communications constructs that can be interpreted as an SFI-type measure. This makes EVM-vs-rate the primary quantitative embodiment of signal fidelity in that work.
3. Time-Series Signal Quality as an SFI Surrogate
In time-series quality auditing, the phrase “Signal Fidelity Index” does not appear, but the paper explicitly frames signal quality indices as fidelity-related descriptors. It implements a broad collection of SQIs from the ECG literature and uses them for binary signal quality classification, outlier detection, denoising evaluation, and downstream alert discrimination. The most natural SFI in this framework is either a single SQI, such as a morphology or periodicity score, or a composite index learned from many SQIs; the strongest candidate is the combined SQI feature set with a Random Forest classifier, whose predicted probability of being “clean” can be treated as an SFI in 0 (Gao et al., 2024).
The SQI set spans morphology, spectral content, distributional statistics, energy allocation, entropy, and periodicity. The paper groups or references measures such as bSQI, pSQI, kSQI, sSQI, fSQI, basSQI, bsSQI, eSQI, hfSQI, purSQI, rsdSQI, entSQI, hfMSQI, PiCASQI, pcaSQI, OrphanidouSQI, Neurokit-based averageQRS_sqi, Zhao 2018, and geometric HR features. This structure makes fidelity a multidimensional property: stable waveform morphology, expected spectral profile, reasonable temporal regularity, and absence of flatline, baseline wander, or high-frequency contamination. On the PhysioNet 2011 ECG Quality Challenge Set A, using all SQIs as features with a Random Forest yields the best reported performance, with AUC approximately 1 and accuracy approximately 2. In unsupervised quality assessment, Isolation Forest reaches AUC approximately 3 and accuracy approximately 4 (Gao et al., 2024).
The same framework also links fidelity to denoising. The paper evaluates wavelet denoising using DWT with Daubechies-4, EMD denoising, and a CNN denoising autoencoder, using MSE against a clean reference as the direct reconstruction metric. CNN AE gives the best MSE at low SNR, wavelet denoising performs better at high SNR, and EMD is generally worse than CNN and wavelet in that evaluation. This suggests a two-level interpretation of SFI in time-series work: when ground truth exists, fidelity can be measured directly by reconstruction error; when it does not, fidelity is inferred from morphology-, spectrum-, and periodicity-based SQIs. The paper’s broader claim is that this framework can be extended to arbitrary time-series measurements in complex systems, especially when waveform morphology and periodicity are central to downstream analyses (Gao et al., 2024).
4. Formal Patient-Level SFI in EHR-Based Dementia Prediction
The clearest explicit definition of Signal Fidelity Index appears in work on dementia prediction across heterogeneous real-world data. There, SFI quantifies diagnostic data quality at the patient level under diagnostic signal decay, defined as variability in diagnostic quality and consistency across institutions. For patient 5,
6
Each component is normalized to 7, so 8, with higher values indicating more specific, stable, contextually aligned, and therapeutically supported diagnostic coding (Cheng et al., 10 Sep 2025).
The six components are defined as follows.
| Component | Definition | Interpretation |
|---|---|---|
| Specificity | 9 | Coding precision |
| Temporal Consistency | 0 | Stability over time |
| Entropy | 1 | Low coding randomness is rewarded |
| Contextual Concordance | 2 | Clinical plausibility of coding context |
| Medication Alignment | 3 | Therapeutic corroboration |
| Trajectory Stability | 4 if most common inpatient code = most common outpatient code, else 5 | Cross-setting consistency |
This construction ties fidelity to structured EHR evidence rather than waveform reconstruction. Specificity distinguishes definitive dementia diagnoses from vague or nonspecific codes. Temporal Consistency and Entropy quantify whether the diagnostic trajectory remains coherent across encounters. Contextual Concordance asks whether the code appears in clinically appropriate settings, such as neurology or inpatient encounters. Medication Alignment measures co-occurrence with dementia-specific medications, including donepezil, memantine, rivastigmine, and galantamine. Trajectory Stability reduces cross-setting agreement to a binary consistency test. The resulting SFI is explicitly label-free at deployment time because it requires only encounter-level diagnoses, care setting, and medication information, not outcome labels (Cheng et al., 10 Sep 2025).
5. SFI-Aware Calibration
The dementia paper embeds SFI in a post-hoc calibration rule designed to improve transportability across heterogeneous datasets without target labels. The simulation framework generates 2,500 synthetic datasets, each with 1,000 patients and realistic demographics, encounters, and coding patterns based on dementia risk factors. Patients have 2–20 encounters between January 1, 2020 and January 1, 2025. A reference dataset of 2,000 patients is split into 1,000 for training and 1,000 for testing, and a Random Forest classifier using age and race predicts dementia. Calibration then rescales each patient’s raw probability according to that patient’s SFI and the mean SFI in the reference dataset:
6
If 7, the calibrated probability is increased; if 8, it is decreased. Candidate 9 values range from 0 to 1 in steps of 2, and the recommended value converges to 3 (Cheng et al., 10 Sep 2025).
At 4, the paper reports statistically significant improvements across all metrics with 5. One summary gives raw versus calibrated performance as follows: AUC 6, Balanced Accuracy 7, Detection Rate 8, F1-score 9, Precision 0, and Recall 1–2, with reported gains ranging from 3 for Balanced Accuracy to 4 for Recall and 5 for F1-score. The reference models themselves achieve mean AUC 6, Balanced Accuracy 7, Detection Rate 8, F1-score 9, Precision 0, and Recall 1. The paper further reports that calibrated F1-score and Recall move to within about 2 of the reference standards, while Balanced Accuracy and Detection Rate become about 3 to 4 closer to reference. In this formulation, SFI is not merely descriptive; it is a calibration covariate that operationalizes diagnostic signal decay as a tractable source of distributional shift (Cheng et al., 10 Sep 2025).
A notable methodological property is that the adjustment is label-free in the target domain. Only structured EHR data and the model’s raw probabilities are required. This distinguishes SFI-aware calibration from approaches that assume access to target outcomes or focus only on covariate shift and label shift without modeling the quality of the diagnostic coding itself.
6. Acronym Ambiguity, Limits, and Extensions
The acronym “SFI” is heavily overloaded outside signal-fidelity contexts. In systems and security, SFI denotes Software-based Fault Isolation, a sandboxing technique based on instruction-level rewriting and static validation (Emamdoost et al., 2018). In solar physics, SFI denotes Solar Flare Index, defined from flare importance and duration as 5, and used to analyze Gnevyshev gaps and geophysical lag structure across Solar Cycles 18–24 (Takalo, 2023). These meanings are entirely unrelated to signal fidelity. A plausible implication is that technical literature searches require explicit domain qualification rather than acronym search alone.
Within signal-fidelity work itself, the principal limitation is the absence of a single domain-independent formulation. In communications, the paper does not define a variable called SFI and instead relies on EVM, constellation geometry, and phase-dependent step responses. In time-series auditing, the phrase “Signal Fidelity Index” does not appear and fidelity is represented through SQIs, anomaly scores, and denoising metrics. In dementia prediction, the index is explicit but currently validated in simulation rather than real administrative datasets. The dementia study states that real EHR validation is still required, and the time-series study notes that ECG-derived SQIs must be adapted carefully when transferred to other domains (Blosser et al., 2024, Gao et al., 2024, Cheng et al., 10 Sep 2025).
Despite that heterogeneity, a coherent pattern emerges. Signal fidelity is consistently treated as the extent to which observable data retain the structure needed for a downstream inference task: symbol recovery in QAM reception, clean morphology and periodicity in physiological time series, or clinically credible diagnostic evidence in EHR prediction. This suggests that an SFI, when explicitly defined, is best understood as a task-conditioned summary statistic rather than a universal physical invariant. The most credible future extensions in the cited literature are therefore domain-specific: disease-agnostic adaptation of the six-component EHR SFI to Parkinson’s disease, multiple sclerosis, diabetes, or chronic kidney disease; extension of SQI frameworks to arbitrary time-series measurements in complex systems; and communication-system fidelity measures that continue to privilege realistic modulated-signal performance over scalar bandwidth alone (Cheng et al., 10 Sep 2025, Gao et al., 2024, Blosser et al., 2024).