SiCmiR: Single-Cell miRNA Inference
- SiCmiR is a deep-learning regression framework designed to infer mature miRNA expression from mRNA data using a compact set of 977 LINCS L1000 landmark genes.
- It employs a two-layer neural network architecture with batch normalization, dropout, and early stopping to achieve robust performance across diverse cancer types and sparse single-cell data.
- SiCmiR Atlas complements the model as a public resource that visually and interactively maps inferred single-cell miRNA landscapes for biomarker discovery and network analysis.
Searching arXiv for SiCmiR and closely related single-cell miRNA inference work to ground the article in current literature. {"query":"SiCmiR single-cell miRNA arXiv", "max_results": 10} {"query":"miRSCAPE single-cell miRNA arXiv", "max_results": 10} SiCmiR is a deep-learning regression framework for inferring mature microRNA expression profiles from mRNA expression data, designed specifically to operate effectively on single-cell RNA-seq inputs despite the sparsity and dropout that constrain direct single-cell small-RNA profiling. Introduced together with SiCmiR Atlas, it uses only 977 LINCS L1000 landmark genes to predict 1,298 mature miRNAs, and is positioned as a bridge between the statistical power of bulk paired mRNA–miRNA data and the cellular resolution required for cancer heterogeneity, biomarker discovery, miRNA-target interaction analysis, and extracellular-vesicle-mediated communication studies (Cai et al., 6 Aug 2025).
1. Conceptual basis and problem setting
SiCmiR was developed in response to three stated limitations. First, single-cell miRNA sequencing remains technically difficult because of low capture efficiency, high sparsity, protocol limitations, and difficulty distinguishing miRNAs from other small RNAs. Second, existing inference approaches are described as suboptimal for scRNA-seq: miRSCAPE uses approximately 20,000 genes and is affected by zero inflation and dropout in single-cell data, whereas miTEA-HiRes infers miRNA activity through enrichment over canonical target lists rather than directly recovering continuous expression levels. Third, cancer biology requires single-cell resolution because miRNAs are central post-transcriptional regulators and tumor heterogeneity limits bulk-only analysis (Cai et al., 6 Aug 2025).
Within this framing, SiCmiR is a supervised multi-output regression model that predicts mature miRNA abundance indirectly from the transcriptome. The central methodological choice is a compact input representation based on the 977 LINCS L1000 landmark genes. These genes are described as perturbation-responsive, reproducible across RNA-seq datasets, and able to infer a large fraction of unmeasured transcripts. The paper attributes a substantial part of SiCmiR’s robustness in sparse single-cell settings to this dimensionality reduction.
This design also defines the scope of the method. SiCmiR is intended to infer miRNA expression rather than measure it directly. The paper therefore treats it as a computational surrogate for mature miRNA profiles in settings where direct single-cell small-RNA assays remain inaccessible or insufficiently robust.
2. Architecture and mathematical formulation
SiCmiR is described as a two-layer fully connected neural network implemented in PyTorch (Cai et al., 6 Aug 2025). Its architecture comprises an input layer with 977 genes, a hidden layer with 1,024 nodes, and an output layer with 1,298 miRNAs. The network uses batch normalization, dropout, ReLU activation, the SGD optimizer, a learning rate of 0.4, a dropout rate of 0.3, and early stopping after 20 epochs without validation improvement.
The model is formulated as a supervised learning problem over a dataset
where denotes the gene-expression vector and denotes the miRNA-expression vector. The learned mapping is a neural function taking .
Training minimizes mean squared error:
with denoting the predicted miRNA vector.
Evaluation is reported per miRNA using Pearson correlation coefficient, MSE, RMSE, and coefficient of determination:
0
The paper’s interpretation of dropout resistance is operational rather than absolute. It attributes reduced sensitivity to dropout primarily to the restricted 977-gene feature set, together with batch normalization, dropout regularization, pan-cancer training, and, in some applications, pseudo-bulk or cell-type averaging.
3. Training corpus, validation design, and predictive performance
SiCmiR was trained on 6,462 paired TCGA mRNA–miRNA samples spanning 33 cancer types, with mRNA features restricted to the L1000 landmark genes and miRNAs with all-zero counts removed; the final model predicts 1,298 miRNAs (Cai et al., 6 Aug 2025). The independent validation set contains 2,768 samples. The study also used 3-fold cross-validation, with 20% of the training data held out for validation in each run; because the feature set was fixed a priori, this cross-validation was not nested.
Among the tested architectures, the neural network outperformed ResNet and Transformer. The reported overall performance for the neural network was MSE = 0.522, RMSE = 0.722, 1, and PCC = 0.673. Using only the 977 L1000 features, the model achieved a mean miRNA PCC of about 0.75 on training and about 0.67 on test. In 3-fold cross-validation, the reported values were training PCC = 0.75 ± 0.00067 and test PCC = 0.67 ± 0.00073; sample-level PCC was 0.72 ± 0.09583 in training and 0.63 ± 0.13707 in test.
The paper presents the compact feature set as a substantive modeling choice rather than a minimal representation. Relative to larger variable-gene panels, the 977 landmark genes performed as well as or better on the test set, while reducing computational burden and improving robustness in sparse single-cell settings.
Comparison with miRSCAPE is central to the benchmark narrative. On an independent test set, SiCmiR reached PCC = 0.67, compared with PCC = 0.61 for miRSCAPE. Inference time is reported as 2.23 seconds using the pretrained SiCmiR model, versus more than 2 hours for miRSCAPE because of on-the-fly training.
4. Generalization across cancers, perturbations, and single-cell use cases
A central claim is that the pan-cancer model generalizes beyond its immediate training distribution (Cai et al., 6 Aug 2025). The paper states that it outperformed cancer-specific models overall across 33 TCGA cancer types and generalized to unseen cancer types, drug perturbations, and scRNA-seq data. The unseen-cancer example is ACTH-secreting pituitary adenoma, which was not included in the TCGA training set. The perturbational example is A549 cells treated with Cinnamomi Ramulus, where the model identified drug-responsive DEmiRs consistent with qPCR-validated trends.
Proof-of-concept analyses were performed in K562, 293T, HeLa, and A549 cell lines, where predicted patterns were presented as biomarker-like miRNA signatures. In hepatocellular carcinoma, SiCmiR recovered 19 of 24 DEmiRs from the cited study, corresponding to sensitivity = 0.79, and nine were significant at 2.
Single-cell applications occupy a major part of the paper. In pancreatic ductal adenocarcinoma scRNA-seq, SiCmiR identified dysregulated miRNAs distinguishing DC2 from DC1 and acinar cells, with 0.73 sensitivity on pooled data and 0.31 sensitivity on direct single-cell data. Among 101 dysregulated miRNAs, it correctly predicted 66 of 90 covered miRNAs. The examples highlighted include hsa-miR-30b-3p, hsa-miR-21-5p, and hsa-miR-147b-3p, the last described as a potentially novel malignancy-associated miRNA.
In ACTH-secreting pituitary adenoma scRNA-seq, pooled analysis identified 55 of 75 reported dysregulated miRNAs, for sensitivity of about 0.73; examples include hsa-miR-136-3p and hsa-miR-410-3p. In glioblastoma, the model was integrated into miRTalk to infer extracellular-vesicle-mediated intercellular miRNA signaling. The resulting high-confidence miRNA-target edges increased to 114,501, approximately a 20-fold increase over the original proxy approach; the negative-correlation proportion rose from 15.6% to 36.9%, and the average interaction score improved by about 47-fold. Notable miRNAs in this setting included hsa-miR-125b-5p, hsa-miR-10b-5p, and hsa-miR-21-5p.
These applications indicate that SiCmiR is being used not merely as a regression engine but as an upstream layer for network reconstruction, biomarker prioritization, and cell-state stratification in heterogeneous tumors.
5. Hub-miRNAs, network structure, and model interpretability
The paper defines hub-miRNAs as those with 3 on the independent test set (Cai et al., 6 Aug 2025). By this criterion, 414 miRNAs were classified as hubs. These hub-miRNAs are reported to show reproducible patterns across 33 cancers and to occupy denser cancer-related network neighborhoods than lower-correlation miRNAs.
The quantitative network contrast is explicit: hub-miRNAs had mean degree 11.12, compared with 4.63 for lower-correlation miRNAs. They were enriched for functions related to oncogenesis, progression, metastasis, and prognosis. This positioning gives SiCmiR a dual role: it is both a predictive model and a ranking mechanism for candidate regulatory nodes whose reproducibility supports downstream prioritization.
Interpretability is addressed using SHAP with gradient explainer. The reported output consists of feature-to-miRNA contribution modules, including a prominent module centered on COL1A1. The paper links these modules to metastasis-related processes such as angiogenesis, extracellular matrix remodeling, and epithelial-mesenchymal transition. It also reports that several influential genes correlate with survival in kidney cancers.
A plausible implication is that the model’s utility depends not only on average predictive accuracy but also on whether the inferred miRNA programs align with known or testable regulatory structure. In the paper’s presentation, hub-miRNA analysis and SHAP-based decomposition are the mechanisms used to argue that this alignment exists.
6. SiCmiR Atlas: resource structure, scope, and interpretive boundaries
SiCmiR Atlas is the database constructed from SiCmiR outputs and is presented as the first public database dedicated specifically to single-cell mature miRNA expression (Cai et al., 6 Aug 2025). The abstract reports 632 public datasets, 9.36 million cells, and 726 cell types. The detailed resource description reports 362 public datasets together with 9.36 million cells, 726 cell types, 189 tissues, 26 major organs, 84 physiological or disease conditions, and 12 broad disease categories. The coexistence of the 632-dataset and 362-dataset figures indicates a discrepancy in the provided record, whereas the larger-scale descriptors for cells and cell types are consistent across the description.
The atlas provides interactive visualization, including dataset browsing, UMAP/t-SNE-style visualization, and expression exploration for mRNA and miRNA. It supports biomarker identification through cell-type-specific miRNA markers, disease-associated miRNAs, and differential expression analysis. It also provides cell-type-resolved miRNA-target network construction through an in-built MTI network builder integrating TargetScan, miRWalk, miRDB, and miRTarBase. Additional functions include data integration and annotation, biomarker mining, miRNA/mRNA visualization, differential analysis, and single-cell landscape search by cell type, disease state, tissue, and condition.
The atlas is intended as a community resource for biomarker discovery and cell-type-resolved analysis of mature miRNA programs. At the same time, the paper is explicit that SiCmiR infers mature miRNA abundance from mRNA rather than directly measuring miRNA molecules. It also states that some applications use pseudo-bulk or cell-type averaging to mitigate sparsity. This suggests that SiCmiR Atlas is best understood as a prediction-derived landscape of mature miRNA expression rather than a direct single-cell small-RNA assay compendium.
In this formulation, SiCmiR and SiCmiR Atlas together define an inferential framework: bulk-derived paired mRNA–miRNA statistical structure is transferred into single-cell transcriptomic space, enabling mature miRNA analyses that would otherwise be technically inaccessible, while preserving enough structure to support hub-miRNA discovery, cancer-network analysis, and biomarker-oriented exploration.