Papers
Topics
Authors
Recent
Search
2000 character limit reached

CPP4APP: Mobile Contextual Privacy Dataset

Updated 12 July 2026
  • CPP4APP is a mobile context-centric dataset that aligns app UI elements with fine-grained privacy policy segments.
  • It is used in Conflect to evaluate extraction accuracy (up to 94.0%) and to benchmark reductions in processing latency using parallel classification.
  • The dataset employs a 12-category label taxonomy to link UI contexts with specific privacy policy clauses, supporting real-time contextual notice delivery.

Searching arXiv for the Conflect paper and related CPP4APP work. arxiv_search.query({"search_query":"ti:\"Conflect: Designing Reflective Thinking-Based Contextual Privacy Policy for Mobile Applications\" OR all:\"CPP4APP\" OR all:\"Contextual Privacy Policies for Mobile Applications\"","start":0,"max_results":10}) arXiv search results received. Looking for the exact Conflect entry and any arXiv-indexed antecedent on contextual privacy policies for mobile apps. arxiv_search.query({"search_query":"all:\"Conflect\"","start":0,"max_results":10}) CPP4APP, short for Contextual Privacy Policies for Mobile Applications, is a mobile, context-centric dataset introduced by Pan et al. and reused in Conflect as the primary benchmark for evaluating automated extraction of privacy-policy segments relevant to specific on-screen contexts and data types, as well as for measuring the end-to-end latency of a contextual detection-to-notice pipeline on realistic mobile interfaces (Zhang et al., 16 Sep 2025). The dataset was curated to address the gap between app interface contexts—such as fields, buttons, and prompts—and the corresponding fine-grained privacy policy statements required for automated contextual privacy policy generation in mobile applications. Within Conflect, CPP4APP functions as an evaluation resource rather than a training corpus, and the paper does not report a versioning history or updates beyond citing Pan et al. (2024) (Zhang et al., 16 Sep 2025).

1. Provenance and intended scope

CPP4APP originates from Pan et al., A NEW HOPE: contextual privacy policies for mobile applications and an approach toward automated generation (USENIX Security 24). In the Conflect study, the authors explicitly reuse the dataset and do not claim to have created or re-annotated it (Zhang et al., 16 Sep 2025). The reported timeframe is therefore 2024, inferred from that citation, with no additional dataset release details given in the paper.

The dataset’s purpose is tightly coupled to contextual privacy presentation. Rather than treating privacy policies as standalone documents, CPP4APP is designed to support alignment between mobile UI contexts and policy segments. This makes it suitable for evaluating systems that must identify a risk-bearing interface element, retrieve the relevant policy text, and present a just-in-time notice. A plausible implication is that CPP4APP occupies a methodological niche between document-level privacy policy corpora and runtime mobile interface analysis, because its stated role is precisely to bridge those two levels.

The Conflect paper uses CPP4APP in three specific ways: to measure policy segment extraction accuracy, to benchmark pipeline latency from screenshot capture to contextual notice display, and to support human evaluation of generated reflective risk descriptions grounded in dataset-derived policy segments (Zhang et al., 16 Sep 2025).

2. Dataset composition and label taxonomy

In Conflect’s use of CPP4APP, the dataset comprises three reported content types: mobile app user interface contexts or screens, privacy policy text and policy segments aligned to data categories relevant to mobile contexts, and per-segment categorical labels used to map between UI contexts and policy statements (Zhang et al., 16 Sep 2025). The paper does not report the dataset’s full schema, raw document counts, token totals, or language coverage.

The label taxonomy used in Conflect spans 12 common mobile data categories supported by CPP4APP:

  • Location
  • Address
  • Phone
  • Email
  • Birthday
  • Contacts
  • Name
  • Voices
  • Social media
  • Photos
  • Profile
  • Financial info

These categories are described as reflecting typical mobile data practices that a contextual notice would need to cover, including collection, transmission, sharing or disclosure, and purpose (Zhang et al., 16 Sep 2025). In Conflect, a keyword-based mapping adapted from prior work by Pan et al. covers approximately 80% of common policy topics and is used to prompt LLMs to extract, structure, and ground segments in these categories.

The paper does not detail CPP4APP’s original annotation process, including annotator identity, guidelines, tools, or agreement metrics, and it reports no inter-annotator agreement. It also does not disclose corpus-level counts such as the number of apps, screens, policy documents, labeled segments, total words or tokens, segment-length statistics, or class distributions (Zhang et al., 16 Sep 2025). This absence is significant for secondary analysis: it limits direct assessment of representational breadth, possible imbalance, and annotation stability.

3. Integration into the Conflect pipeline

Within Conflect, CPP4APP is not used to train supervised models. Instead, it supports evaluation of a multi-stage pipeline combining OCR, visual classification, language-model-based policy processing, and category-level alignment (Zhang et al., 16 Sep 2025).

The reported preprocessing and formatting sequence begins with policy acquisition and segmentation. GPT-4o retrieves each app’s privacy policy via web search, segments it, and extracts structured data practices using the keyword-based mapping adapted from Pan et al. The extracted segments are organized by the dataset’s data categories, and the content is described as being strictly grounded in the original policy text to minimize hallucinations (Zhang et al., 16 Sep 2025).

Context detection operates on screenshots. PaddleOCR recognizes text on screenshots; GPT-3.5-Turbo classifies the recognized strings into data categories; and a pretrained ResNet classifies icons, which are then mapped into the same categories. Bounding boxes and anchors are preserved so that detected UI elements can be aligned with policy segments. Context–policy alignment then proceeds through shared category labels, after which GPT-4o generates reflective, scenario-based risk descriptions on top of the matched segments (Zhang et al., 16 Sep 2025).

Two runtime strategies are reported. First, outputs such as extracted segments and generated risk descriptions are cached on first app launch to reduce runtime delays. Second, a new analysis is triggered whenever the pixel-wise overlap between consecutive screenshots falls below 80% (Zhang et al., 16 Sep 2025). This suggests that CPP4APP, as operationalized in Conflect, supports not only offline benchmarking but also scheduling decisions for runtime privacy assistance.

4. Evaluation protocol and reported results

The paper evaluates CPP4APP-based policy extraction with standard classification metrics, although only accuracy is reported for policy segment extraction in the technical results (Zhang et al., 16 Sep 2025). The reported average extraction accuracy on CPP4APP is 94.0% across categories.

Selected per-category results reported in the paper are shown below.

Category Extraction accuracy
Location 98.0%
Email 98.0%
Photos 95.9%
Social media 93.9%
Financial info 91.8%
Profile 89.8%

The technical stack behind these results is specified as GPT-4o for retrieval, segmentation, and extraction, with strict grounding in the original policy text. The paper does not report statistical significance tests for technical accuracy; instead, it notes that comparisons are made across datasets and categories (Zhang et al., 16 Sep 2025).

CPP4APP is also used to assess the usefulness of generated reflective descriptions. Human raters gave an average Likert score of 6.4/7 for whether the generated risk description would prompt reflections (Zhang et al., 16 Sep 2025). This result pertains to outputs grounded in CPP4APP-derived policy segments rather than to the dataset as a standalone annotation benchmark.

A stylized illustration in the paper clarifies the intended category–segment linkage: a UI element such as a “Share location” toggle or a location field is mapped to the Location category, which is then linked to grounded policy statements about location collection, sharing with service providers, retention, or recommendation use. Comparable patterns are described for photos, contacts, and financial information. The paper is explicit, however, that these are illustrative rather than raw dataset quotes (Zhang et al., 16 Sep 2025).

5. Latency benchmarking and systems implications

A distinctive feature of CPP4APP in Conflect is that it supports end-to-end latency measurement from screenshot capture to contextual privacy notice display (Zhang et al., 16 Sep 2025). The reported hardware and software environment is a server with 8 vCPUs and 32 GB RAM, using CPU-only execution with no GPU acceleration.

The latency comparison between a serial baseline and the Conflect pipeline is as follows.

Pipeline setting Average latency Dispersion / range
Baseline (serial) 19.78 s SD 14.31 s; min 3.62 s; max 137.57 s
Conflect (parallel classification) 4.35 s SD 0.93 s; min 2.38 s; max 7.71 s

For the Conflect pipeline, the mean component breakdown is also reported: GUI element localization 2.49 s (0.663), GUI element classification 1.84 s (0.679), and matching 0.02 s (0.003) (Zhang et al., 16 Sep 2025). Batch size is not specified, but the paper states that parallelism occurs across detected elements within a screen.

These results position CPP4APP as more than a corpus for textual extraction. In the Conflect setting, it becomes a systems benchmark for contextual privacy delivery under interactive constraints. A plausible implication is that datasets linking UI contexts to policy segments are unusually valuable when evaluation must include both semantic correctness and operational responsiveness.

The Conflect paper situates CPP4APP against several adjacent resources. Relative to OPP-115, Polisis, and general web privacy policy corpora, CPP4APP is described as mobile-first and explicitly contextual, because it targets the alignment between mobile UI elements and policy segments needed for just-in-time contextual privacy policies (Zhang et al., 16 Sep 2025). By contrast, OPP-115 and Polisis are described as focusing on web privacy policies and broad hierarchical annotations detached from mobile runtime contexts.

The paper also reports complementary use of CA4P-483, a fine-grained Chinese software privacy policy dataset, and the MAPP Corpus, a bilingual corpus for which only English policies are used in the reported evaluation. Within that comparative frame, CPP4APP emerges as the benchmark that allows Conflect to measure both policy-segment extraction and real-time pipeline latency in app interaction contexts (Zhang et al., 16 Sep 2025).

Several limitations are explicitly noted. Coverage is incomplete because the keyword-based mapping adapted from Pan et al. covers approximately 80% of common topics, leaving non-standard or novel policy clauses potentially under-covered. The paper does not report counts, splits, language coverage, or class distributions for CPP4APP. Potential class imbalance and modality gaps are not discussed. The screenshot-driven, CPU-bound detection pipeline may also be sensitive to variation in interface complexity, which can affect latency and detection quality (Zhang et al., 16 Sep 2025).

Availability and licensing are likewise not specified in Conflect. No access URL, DOI, or license terms for CPP4APP are provided there. The paper directs readers to Pan et al. (USENIX Security 24) and their supplementary materials or repository for dataset acquisition, governance, and licensing details (Zhang et al., 16 Sep 2025). Reproducibility guidance in Conflect instead concerns the evaluation setup: use the reported 8-vCPU, 32-GB CPU-only environment for latency recreation, apply the same keyword-based mapping and category taxonomy for extraction evaluation, and trigger reanalysis at at most 80% screenshot overlap to mirror pipeline scheduling (Zhang et al., 16 Sep 2025).

In that form, CPP4APP is best understood as a specialized benchmark for contextual privacy-policy alignment in mobile applications: valuable for evaluating extraction fidelity and runtime contextualization, but only partially documented at the corpus level in the presently cited literature.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CPP4APP Dataset.