RiskInDroid: Android Risk-Scoring Frameworks
- RiskInDroid is a family of Android risk-scoring techniques that convert diverse technical evidence into interpretable scores for app and device assessments.
- It employs varying units of analysis—from APK code and pre-installed apps to vulnerability metadata and user profiles—to adapt to different security scenarios.
- Common methodologies include static analysis, machine learning classifiers, CVSS adaptations, and Bayesian models to support actionable risk decisions.
Searching arXiv for the cited RiskInDroid-related papers and adjacent Android risk-scoring work. RiskInDroid is a label used in Android security and privacy research for several risk-assessment frameworks that translate heterogeneous technical evidence into an interpretable score or ranking for trust, triage, or user decision support. In the cited literature, the name does not denote a single canonical system. Instead, it appears in multiple formulations, including APK-level static analysis with machine-learned scoring, device-level scoring of pre-installed applications, a time-aware enhancement of CVSS for mobile vulnerabilities, and profile-aware estimation of malicious-app encounter risk (Kalhor et al., 2023, Ozbay et al., 2022, 1807.10435, Dambra et al., 2023).
1. Scope, terminology, and recurring abstractions
Across these formulations, RiskInDroid consistently operates by selecting a unit of analysis, extracting a structured evidence vector, and mapping that evidence into a scalar or categorical risk output. The unit of analysis varies substantially: in one usage it is an APK; in another, a device and its factory-installed software; in another, a vulnerability over time; and in another, a user profile inferred from installed apps. This suggests a common abstraction—risk normalization for Android ecosystems—rather than a single fixed implementation (Kalhor et al., 2023, Ozbay et al., 2022, 1807.10435, Dambra et al., 2023).
| Usage in the literature | Unit of analysis | Main output |
|---|---|---|
| Static APK scanner (Kalhor et al., 2023) | App / APK | High/Medium/Low label and |
| Pre-install scoring system (Ozbay et al., 2022) | Device via pre-installed apps | Total_Device_Score on a 0–100 scale |
| Time-aware CVSS enhancement (1807.10435) | Vulnerability / CVE | |
| Profile-aware encounter model (Dambra et al., 2023) | Device / user profile | , , and per-profile classification |
A further terminological complication is that some papers invoke RiskInDroid as a proposed enhancement or design direction rather than as a named standalone artifact. The app-choice study on Android marketplaces, for example, frames its contribution as practical guidance for “Risk-in-Droid at selection time,” emphasizing selection-time indicators and decision support rather than code-level analysis (Momenzadeh et al., 2020). For encyclopedia purposes, RiskInDroid is therefore best understood as a family of Android risk-scoring approaches united by output form and decision objective, but differentiated by data source and threat model.
2. APK-centric static analysis and machine-learned risk indexing
In the virtual-assistant security study, RiskInDroid is presented as a static-analysis pipeline with four principal components: a static-analysis engine, a vulnerability database, feature extraction plus machine-learning classification, and a reporting interface (Kalhor et al., 2023). The static-analysis engine decompiles the APK using Apktool or equivalent into Smali or pseudocode, parses AndroidManifest.xml, and scans bytecode for API calls and code patterns such as raw SQL, SSL checks, and JavaScript interfaces. The vulnerability database stores risky code patterns, permission-misuse rules, and SSL-bypass signatures derived from CWE entries, CVE records, and hand-crafted heuristics.
Feature extraction maps each app to a vector whose features include the number of non-SSL URLs, presence of HostnameVerifier overrides, raw SQL calls, WebView JavaScript interfaces, and permissions used but not declared in the manifest. RiskInDroid then applies scikit-learn’s Support Vector Machine and Multinomial Naive Bayes classifiers to assign a High/Medium/Low label and produce a continuous risk score . The Multinomial Naive Bayes formulation is stated as
with
The SVM decision function
is likewise mapped linearly to the same 0–100 range (Kalhor et al., 2023).
The operational workflow is explicit: decompile the APK, extract manifest and Smali code, search code for each vulnerability rule, build the feature vector, feed it into both SVM and MNB classifiers, and compute a normalized risk probability. The reporting layer emits an HTML or CLI report containing the overall risk index, a contribution breakdown, and concrete code locations in Smali files or manifest entries. Within the study’s testbed of eight Android-based voice assistants, the highest RiskInDroid scores were reported for Cortana and Kalliope at approximately 85–90/100, medium scores for Alexa and Extreme at approximately 70/100, and lower scores for Braina and Google Assistant at approximately 40–50/100. Representative flagged patterns included non-validation of SSL certificates, raw SQL queries, and weak AES modes such as AES/ECB/PKCS5Padding. The study further reports Spearman correlation between RiskInDroid’s overall index and MobSF’s heavier-weight security score (Kalhor et al., 2023).
The same paper also states explicit limitations. RiskInDroid is static-only, so it cannot observe dynamically loaded code or native libraries; its training set was limited to the eight virtual assistants plus a few open-source apps; and it is susceptible to false positives in pattern-heavy apps, including cases in which raw SQL is used legitimately in parameterized form. Those constraints position this version of RiskInDroid as a lightweight fusion layer over static evidence, not a complete substitute for dynamic or semantic validation (Kalhor et al., 2023).
3. Device-level scoring of pre-installed application ecosystems
A distinct formulation appears in the study of pre-installed Android applications, where RiskInDroid is a scoring system for assessing the security and privacy risks of device-resident software that ships from the factory (Ozbay et al., 2022). The empirical basis is a dataset collected with the “Pre-App Collector,” published on Google Play and ethically approved by TOBB University. Volunteers from at least 14 countries produced a corpus comprising 98 distinct device models from 22 OEMs, 77 survey participants after filtering, and 143,862 firmware-partition files, including 14,178 APKs, 418 certificates, and 58,721 native libraries. The study reports that each device shipped with 294 pre-installed APKs on average, that only 9% of these were on Google Play, and that 55% had never been updated since factory installation (Ozbay et al., 2022).
This RiskInDroid examines ten criteria. The “new” findings are privileged system-UID apps, allowBackup enabled, unsigned-by-OEM apps, apps not updated for more than two years, usesCleartextTraffic enabled, debuggable enabled, embedded tracker SDKs excluding crash reporters, and cloud-service misconfigurations such as Google Maps API key exposure, AWS S3 key disclosure, Firebase Realtime Database world-readable or writeable access, and leaked Slack webhook URLs or OAuth secrets. The “legacy” findings are exported components without required permissions and dangerous Android permissions implicitly granted to pre-installs. Each criterion is measured as the number of pre-installed apps on the device exhibiting that property, denoted 0 (Ozbay et al., 2022).
The scoring model is explicitly quantitative and inspired by CVSS. For each finding, RiskInDroid uses a probability side determined by 1, a difficulty-to-exploit coefficient 2, and a user-awareness coefficient 3, together with an impact coefficient 4. The per-finding score is
5
and the final device score is
6
Difficulty coefficients range from Easy at 1.00 to Very Hard at 0.10; awareness ranges from “User almost never aware” at 1.00 to “Almost certainly aware” at 0.10; and impact ranges from Very high at 1.00 to Low at 0.10. Example coefficients are given for trackers, with 7, 8, 9; for the debuggable flag, with 0, 1, 2; and for dangerous permissions, with 3, 4, 5 (Ozbay et al., 2022).
The study’s interpretation layer is also notable. It reports that the Pearson correlation between number of pre-installs and device risk score is 6, which the authors use to argue that the decisive factor is the mix of findings rather than raw app count. Devices with the highest scores included the Sony Xperia Z1 and a top ten dominated by Samsung models, while six of the ten lowest-risk devices were 2019-onward models. A score near 100 is defined as indicating a device whose pre-installs almost certainly expose the user to multiple unnoticed trackers, network misconfigurations, dangerous flags, and permissions, whereas a score near 0 indicates a relatively clean factory load (Ozbay et al., 2022).
4. CVSS-derived and time-aware vulnerability scoring
Another RiskInDroid formulation is a mobile-focused enhancement of CVSS intended to improve risk calculation for Android and iOS application vulnerabilities by refining both impact and exploitability (1807.10435). The central criticism is that standard CVSS treats “Partial” impact too coarsely and leaves exploitability static over time. RiskInDroid addresses the first issue by splitting “Partial” into two Android-relevant variants: Partial-Application, with coefficient 0.461, for vulnerabilities confined to an app’s sandbox, and Partial-System, with coefficient 0.515, for vulnerabilities in the Android OS or core services. “Complete” remains 0.660, and “None” remains 0 (1807.10435).
The revised impact equation is
7
where 8. Exploitability is then made time-dependent through a compound Poisson formulation. Let
9
with events corresponding to proof-of-concept release, exploit availability, or patch release. The probability mass function is
0
and exploitability becomes
1
The final RiskInDroid base score is
2
This formulation explicitly encodes exploit pressure as a dynamic process rather than a fixed property (1807.10435).
The empirical motivation comes from a case-control design using vulnerabilities with publicly available exploits as cases and NVD-only vulnerabilities as controls, followed by Android-versus-iOS and OS-versus-app subdivision. The paper reports that roughly 50% of the 44 Android CVEs examined from ExploitDB had previously clustered around a base score of approximately 6.8, with Impact = 6.4 and Exploitability = 8.6. After the RiskInDroid refinement, impact scores spread out; OS flaws moved upward by approximately 0.2–0.3 points, app-only flaws by approximately 0.1–0.2, and base scores redistributed between 5.5 and 8.3 rather than collapsing at 6.8. The same model yields temporal trajectories in which an unpatched OS vulnerability with a proof-of-concept starts at exploitability approximately 8.6, rises to approximately 9.4 once a public exploit appears at month 2, and falls below 7.0 ten months after a patch is shipped (1807.10435).
This version of RiskInDroid differs sharply from APK-level static scanners. Its evidence is not manifest flags or code smells but vulnerability metadata, exploit availability, and patch timing. The underlying decision problem is therefore patch prioritization rather than app vetting.
5. User-facing risk communication and profile-aware risk estimation
RiskInDroid also appears in a user-decision context. The Android marketplace study using Bayesian analysis did not present a full system under that name, but its “Implications for RiskInDroid” are specific: risk communication should occur at app-selection time, use simple padlock-based visual indicators aligned with user mental models, provide early presentation in list view, and leave details available on demand (Momenzadeh et al., 2020). In that experiment, 60 adults used an “Alternate PlayStore” on Nexus 7 tablets across four categories—Flashlight, Weather, Photos, and Games—while the interface displayed privacy grades from PrivacyGrade.org as 1–5 padlocks, with more padlocks denoting lower risk. The per-participant weighted category rating was
3
where 4 is the numerical padlock score and 5 is the download-count weight (Momenzadeh et al., 2020).
The reported results were category dependent. For Flashlight, where privacy and functionality were identical across apps, the posterior for 6 was essentially zero and the 95% HDI lay inside the ROPE. For Weather, the posterior shifted positively and the HDI was approximately 7 padlocks; for Photos, approximately 8. Only 7% of installs were preceded by a click on “View Permissions,” and 78% of participants judged workload minimal. The paper interprets these findings as evidence that simple, positively framed risk indicators can increase risk-averse choices without imposing high mental workload (Momenzadeh et al., 2020). In a RiskInDroid framing, this is a presentation-layer mechanism: it does not infer vulnerabilities directly, but it changes decision quality under risk.
A different user-oriented formulation appears in the large-scale analysis of malicious-app encounters across Android user profiles (Dambra et al., 2023). There, RiskInDroid is associated with profile-specific risk modeling based on app-installation telemetry from approximately 12.2 million devices, reduced after filtering to 8.657 million devices. The core device-level risk metric is the malicious-encounter rate
9
with binary encounter indicator
0
Profiles are defined from counts of user-installed apps by Google Play category, yielding 563 profiles covering 5.4 million devices (Dambra et al., 2023).
The population-level model found strong monotonic associations between certain factors and risk: for example, more than 75 distinct signers versus 1–25 yielded odds ratios of 9.22 for potentially unwanted applications and 7.03 for malware, while more than 75% of apps from alternative markets versus 0–25% yielded 8.21 and 12.03, respectively. Risk was also highly profile dependent. The “Social–Video Players” profile had 1, or 2 the population average; “Entertainment–Game” had 31.50%; “Video Players” had 25.81%; and “Average users” had 8.65%, or 0.60 times the population average. For classification, a Random Forest with 225 trees, maximum depth 20, and 3 improved average accuracy from 52.19% in the whole-population model to 69.84% in per-profile models, with the “Average users” profile rising from 51.90% to 84.35% (Dambra et al., 2023).
Taken together, these two strands show that RiskInDroid can also denote decision support above the code layer: one branch focuses on communicating privacy risk during app choice, and another on adapting risk prediction to heterogeneous user populations. A plausible implication is that RiskInDroid is not limited to vulnerability discovery; it can also serve as a policy engine for personalized warnings, scan frequency, or remediation advice.
6. Related systems, methodological neighbors, and disambiguation
Several adjacent Android security systems illuminate the methodological neighborhood around RiskInDroid without using the name as their primary title. MPDroid, for example, evaluates overprivilege risk by identifying a minimum permission set through collaborative filtering and static analysis (Xiao et al., 2020). It defines extra permissions as 4 and an app-level score
5
with 6 for normal permissions and 7 for dangerous permissions. On 16,343 benign Google Play apps and 524 malicious samples, MPDroid reported improvements over the skewness-based baseline in AUPR, RAR, ARISK, and TRR (Xiao et al., 2020). AUSERA, in turn, organizes Android banking-app assessment as a three-phase static pipeline of sensitive-data tagging, function identification, and weakness detection via taint analysis plus reachability pruning, and it detected 2,157 weaknesses in 693 banking apps with 98.2% precision and average scan time of approximately 1.6 minutes per app (Chen et al., 2018). These systems suggest recurrent design patterns also visible in RiskInDroid: domain-specific evidence extraction, rule- or model-based aggregation, and prioritization-ready outputs.
A similar specialization appears in native-code vulnerability assessment for Android applications, although the paper does not use the RiskInDroid name (Sanna et al., 2024). That methodology scans ELF libraries extracted from APKs, associates them with known products by regex and symbol-table inspection, matches them against a local CVE database, and computes
8
where threat is the CVSS 3.1 exploitability sub-score, impact is the CVSS 3.1 impact sub-score, and vulnerability is a binary indicator of version-plus-function match. In a large-scale analysis of more than 100,000 applications, 38,348 APKs contained at least one .so library, 44,225 ELF files were linked to one of 15 tracked products, and approximately 62% of scorable native-code APKs fell into MEDIUM or HIGH risk (Sanna et al., 2024). This is not RiskInDroid in name, but it reflects the same general movement from raw findings to a compact Android risk label.
RiskInDroid should also be distinguished from the unrelated computer-vision framework DROID, “Driver-centric Risk Object Identification” (Li et al., 2021). DROID analyzes egocentric driving video, builds an Ego–Thing Graph over detected traffic participants, predicts the driver’s imminent Go-versus-Stop response, and assigns causal object risk scores by intervention:
9
where 0 and 1. On an HDD subset, DROID achieved 64.5% Top-1 Precision, 78.2% Top-3 Recall, and mAP 0.58, outperforming several baselines (Li et al., 2021). Despite the lexical similarity, DROID is a driving-risk framework, not an Android security system.
Across the Android literature, the recurring invariant behind RiskInDroid is the compression of multidimensional evidence into a calibrated output that is interpretable enough to guide action. What changes from paper to paper is the evidence domain—Smali code, manifests, trackers, cloud keys, CVE timelines, native libraries, or user-installation behavior—and the decision target—APK vetting, device trust, patch prioritization, or personalized protection.