DataTrust: Trustworthy Data Sharing
- DataTrust is a framework of concepts and architectures that ensure data accuracy, completeness, and integrity through trusted sharing and decentralized governance.
- It integrates cryptographic protocols, remote attestation, and provenance tracking to secure and validate data exchanges across diverse domains such as IoT, AI, and open banking.
- Empirical implementations demonstrate its potential with blockchain, TEE layers, and scalable architectures, while ongoing challenges include standardization and revocation mechanisms.
DataTrust denotes a family of concepts, architectures, and assurance mechanisms concerned with making data sharing, storage, exchange, and downstream use trustworthy. The literature suggests that the term does not yet have a single fixed meaning. In some work, DataTrust is used interchangeably with data trustworthiness from the perspective of a data consumer in inter-organisational sharing; in other work it denotes a fiduciary-style data-sharing arrangement, a decentralized protocol stack for fair data trading, or a data-centric auditing framework that links properties of training data to model trustworthiness indicators (Zimmer et al., 31 Mar 2025, Samaniego, 2022, Su et al., 2020, Wang et al., 2024). Across these usages, recurring concerns include confidence that data are true, complete, accurate, untampered, properly governed, and usable under explicit technical or legal constraints.
1. Conceptual scope and major formulations
The contemporary literature uses DataTrust in several closely related but non-identical senses. Zimmer et al. define it as data trustworthiness from the perspective of a data consumer, adopting the ISO/IEC 15408-1 view of the consumer as a “risk owner” and defining data trustworthiness as “the degree to which a consumer can be confident that shared data are true, complete, accurate and have not been tampered with” (Zimmer et al., 31 Mar 2025). In IoT-oriented work, DataTrust is described as a fiduciary-style data-sharing arrangement in which individual data owners appoint a steward or trustee to collect, manage and exchange IoT-generated data on specified terms, for the benefit of one or more beneficiaries (Samaniego, 2022). In decentralized market systems, DataTrust appears as a protocol objective: BDTF frames it as fair, confidential, integrity-preserving, and decentralized data trading without a centralized broker (Su et al., 2020). In AI assurance, DataTrust becomes a data-centric auditing abstraction that connects training-data quality indicators to evaluation-time trustworthiness indicators (Wang et al., 2024).
| Formulation | Core object | Representative source |
|---|---|---|
| Data trustworthiness | Consumer confidence in shared data | (Zimmer et al., 31 Mar 2025) |
| Fiduciary-style data-sharing arrangement | Trustee-mediated governance of IoT data | (Samaniego, 2022) |
| Decentralized exchange framework | Fair payment and fair data transmission | (Su et al., 2020) |
| Consent-management substrate | Legally enforceable data flow across “worlds” | (Ayappane et al., 2023) |
| Data-centric auditing framework | Correlation between training data and model trustworthiness | (Wang et al., 2024) |
| Trustworthiness index | Weighted aggregation over trust pillars for synthetic data | (Belgodere et al., 2023) |
This multiplicity is not merely terminological. It reflects different problem settings: inter-organisational data sharing, IoT governance, open data markets, synthetic-data auditing, code-model evaluation, and black-box service selection. A plausible implication is that DataTrust is best understood as an umbrella concept whose concrete instantiation depends on whether the primary trust object is the data itself, the party holding the data, the protocol carrying the data, or the model trained on the data.
2. Core properties and trust dimensions
Despite variation in formulation, several trust dimensions recur. In BDTF, four requirements are explicit: fairness, confidentiality, integrity, and decentralization. Fairness means that no party can cheat—either both payment and data are exchanged, or neither is. Confidentiality means that only the intended buyer sees the seller’s data. Integrity means that delivered data must be exactly what the seller originally offered. Decentralization means that no single trusted broker is required (Su et al., 2020). These properties are tailored to data trading, but analogous requirements appear elsewhere.
In provenance-oriented storage models, trustworthiness is decomposed into provenance completeness and integrity. Salgado et al. define an integrity predicate on a dataset as
where is structural consistency, domain constraints, and bit-level correctness. A datum is trustworthy exactly when its provenance is complete and its containing dataset passes integrity:
They further characterize reliability as trust plus availability under expected SLAs (Salgado et al., 2021).
From the consumer perspective, DataTrust centers on completeness, accuracy, integrity, and provenance. Zimmer et al. organize assurance claims around these dimensions and propose that metrics may be combined into a TrustScore of the form
Their conceptual framework includes LoA 0 (Self-asserted), LoA 1 (Automated Controls), LoA 2 (Third-party Audit), and LoA 3 (Formal Certification), with progressively stronger evidence requirements (Zimmer et al., 31 Mar 2025).
In dashboard and data-artifact settings, the barriers to trust are more socio-technical. Sultanum et al. identify six recurring obstacles: Context Dependency, Reliance on Intuition (“Data Smells”), Social Relationships, Fragile Trust Over Time, Ambiguous Definitions, and Untracked Changes. Their “data guards” respond with Overview guards, Details guards, and Community guards, indicating that DataTrust is not only a matter of cryptographic assurance or statistical validation but also of explanation, provenance visibility, and social vetting (Sultanum et al., 2024).
3. Architectural patterns
A major strand of DataTrust research is architectural. BDTF organizes its system into two layers: a blockchain layer and a TEE layer . The blockchain realizes payments, while an Intel SGX enclave implements a “trusted exchange” for fair data transmission. The protocol couples remote attestation, on-chain deposit verification, encrypted off-chain data upload, preview, final payment, and release of the full encrypted data (Su et al., 2020). Its formalization uses a public ledger of blocks 0 and enclave code measurement 1.
A different architectural pattern appears in “smart data” systems. Salgado et al. separate a Physical Persistence Layer, a Transaction Processing Layer, and a Domain-Specific Layer. The persistence layer is append-only and versioned; the transaction layer records provenance through slot interceptors and commit-time serialization; the domain layer exposes entities, roles, and business operations (Salgado et al., 2021). This architecture is designed to preserve provenance, enforce integrity, and support revocation and future distribution.
Consent-centric DataTrust architectures encode legality directly into data flows. The Multiverse framework models an open, distributed network of “worlds” and legal capacities expressed as “role tunnels.” Its core frame is
2
with worlds, data elements, agents, and templates. A legal capacity is the concatenation
3
and stored resources are represented as
4
Access is conditioned on tunnel validity, template constraints, and time-to-live; integrity checks are probabilistically sampled via an access-risk parameter 5 (Ayappane et al., 2023).
The crypto-democracy paradigm pushes DataTrust toward institutional decentralization. Instead of a single physical party, several independent institutions jointly implement a virtual trusted third party—the “Trustworthy”—using secret sharing and MPC. In the proof-of-concept, five universities act as consortium nodes; user data are secret-shared under a 6 threshold scheme, and computations are carried out only if an authorized subset agrees (Gambs et al., 2014). This formulation minimizes single-point trust by distributing custody and computation.
In open banking, TruChain instantiates a three-layer trust service. Layer 1 validates source legitimacy using decentralized identity and verifiable presentations; Layer 2 verifies data authenticity and consistency using cryptographic signing; Layer 3 guarantees tamper-proof storage through the IOTA Tangle, with full records stored off-chain in MongoDB (Rahman et al., 11 Jul 2025). This layered pattern separates source trust, data-object trust, and storage immutability.
PrivTru represents yet another architectural direction: a privacy-by-design data trustee that calculates the minimal amount of information needed from data sources to answer a query, with the stated objective of minimizing information leakage to the data trustee while preserving utility (Gehring et al., 6 Jun 2025).
4. Assurance mechanisms, provenance, and governance
DataTrust systems achieve assurance through heterogeneous but often composable mechanisms. In BDTF, remote attestation establishes trust in enclave code, on-chain evidence establishes payment facts, and the enclave releases data only after verifying payment. The framework also claims non-repudiation because payments and reviews are recorded immutably on 7, and liveness as long as the blockchain continues to confirm transactions (Su et al., 2020).
TruChain uses decentralized identity and credentialing to bind sources to attestable identities. Its Layer 1 challenge-response uses a verifiable presentation signed with Ed25519, while Layer 2 packages financial data as a signed verifiable credential and validates both signature and JSON schema. Layer 3 anchors a Blake2b-256 hash of validated data on the Tangle and stores the full JSON record plus block ID off-chain; query-time re-verification recomputes the hash and compares it to the on-chain value (Rahman et al., 11 Jul 2025).
Provenance capture is the central assurance mechanism in smart data systems. A datum is always written within exactly one transaction, each transaction carries a provenance record 8, and derived objects record both immediate sources and transaction metadata. Because stored objects are immutable, versioned, and linked to predecessors, one can walk back through a provenance DAG, and a later revocation can in principle cascade to all dependents (Salgado et al., 2021).
Consent management frameworks shift emphasis from cryptographic truth to legally tractable flow control. In Multiverse, access is authorized only for agents with an unbroken role tunnel satisfying all incoming relationship constraints. Freshness and revocability are handled through TTL; trusted template authorities and offline vetting are proposed as countermeasures against false template implementations or fake relationships (Ayappane et al., 2023).
Social and organizational assurance also appears in the “data guards” framework. The seven guards include Data and Pipeline Tests, Data Quality Agent, Data and Pipeline Change Alerts, Explanation and Status, Data Traces, Stamp of Approval, and Crowd Wisdom. These mechanisms cover automated checks, narrative disclosures, provenance exploration, and social proof by designated data stewards or user communities (Sultanum et al., 2024). This suggests that DataTrust, in practice, often combines cryptographic, procedural, and social guarantees rather than relying on a single assurance primitive.
5. Quantification and evaluation
Quantitative formulations of DataTrust are domain-specific. In black-box data services, DETECT assigns each service a trust score
9
where Performance aggregates Availability, Task Success Ratio, and Time Efficiency, and DataQuality is modeled as 0, combining data timeliness and database timeliness (Romdhani et al., 2021). This is an operational trust metric for environments in which metadata and internal processing conditions are hidden.
For synthetic data, the auditing framework of “Auditing and Generating Synthetic Data with Controllable Trust Trade-offs” defines a trustworthiness index 1 over trust pillars such as Fidelity, Privacy, Utility, Fairness, and Robustness. After metric alignment and ECDF normalization, each trust dimension is aggregated into 2, and the global index is
3
The framework uses this score for ranking, reporting, and trustworthiness-driven checkpoint selection (Belgodere et al., 2023).
In trustworthy AI data selection, VTruST defines additive value functions for accuracy, fairness, and robustness and allows user-controlled trade-offs through 4. For example,
5
Its subset-selection problem is posed as an online sparse approximation solved by an online version of Orthogonal Matching Pursuit (Das et al., 2024). Here DataTrust is expressed not as a single scalar score on a dataset already given, but as a controllable data-valuation and subset-construction process.
For LLMs for code, DataTrust is formalized as
6
where 7 is a vector of model trustworthiness indicators, 8 is a vector of training-data quality indicators, and 9 is a correlation function between them (Wang et al., 2024). This formulation explicitly treats trust as a relation between evaluation outcomes and training-corpus properties.
Other systems use scalar trust outputs for end users. WebTrust assigns a reliability score from 0 to 1 to each statement and provides a justification; its score is computed by mapping the fine-tuned model output through a sigmoid and affine transform,
2
with hard clipping to 3 (Chandra et al., 5 Jun 2025). In dialog research, VIRATrustData operationalizes trust as a four-class annotation problem—Low Institutional Trust, Low Agent Trust, Neutral Trust, and High Trust—over human-chatbot turns (Friedman et al., 2022).
Taken together, these formulations suggest that quantitative DataTrust currently relies on heterogeneous operationalizations: QoS, provenance, audit metrics, assurance levels, value functions, reliability regression, or classification labels, depending on the object of trust.
6. Application domains, empirical results, and open problems
DataTrust research spans data markets, IoT, open banking, dashboards, synthetic data, code LLMs, inter-organisational data spaces, black-box services, and conversational agents. In BDTF, the system was implemented on Ethereum and Intel SGX. On a ganache local chain with an i7-7700, blockchain throughput peaked at approximately 4 tx/s, enclave response time averaged 5 ms per buyer request, and with 6 confirmations and 7 s block time the overall trade latency was approximately 8 s. Each deposit or payment cost approximately 9 gas 0 Gwei 1 ETH (Su et al., 2020).
TruChain reports a proof-of-concept in Node.js on a local IOTA Tangle testnet. Under concurrency from 2 to 3 connections over 4 s, throughput remained roughly constant; when source verification was skipped per request, end-to-end latency stayed below 5 ms at 6 clients; query latency stayed below 7 ms compared with a Swedbank API baseline of 8 ms to 9 ms; invalid inputs were rejected with a 0 validation rate; CPU peaked under 1 and memory stayed below 2 MiB (Rahman et al., 11 Jul 2025).
WebTrust reports that its fine-tuned Granite-based scorer outperformed rule-based and other small-scale LLM baselines on MAE, RMSE, and 3, and its user study with 4 volunteers reported mean scores of 5 for Ease of Use, 6 for Reliability Score Usefulness, 7 for Trust in Justifications, and 8 for Overall Satisfaction (Chandra et al., 5 Jun 2025). DETECT’s e-health case study showed that service rankings changed systematically as 9 and 0 shifted between performance-centric and data-quality-centric weighting, supporting the feasibility of continuous black-box trust monitoring (Romdhani et al., 2021). VTruST reports better trade-offs than baselines on social, image, and scientific datasets, while VIRATrustData provides a trust-annotated corpus of 1 conversational turns for modeling trust and mistrust in public-health chatbot interactions (Das et al., 2024, Friedman et al., 2022).
The literature also records substantial unresolved issues. Sultanum et al. report a recurring need, but lack of existing standards, for data validation and verification, especially among data consumers (Sultanum et al., 2024). Zimmer et al. describe their Data LoA artifact as early-stage, with specific level definitions, control checklists, and numerical thresholds still to be defined (Zimmer et al., 31 Mar 2025). Salgado et al. note that revocation remains unimplemented and that distribution, consensus, and partition handling are future work (Salgado et al., 2021). Multiverse does not include quantitative performance benchmarks or full formal security proofs (Ayappane et al., 2023). The IoT data trust survey states that implementation details and experiments for its proposed blockchain-based solution are left for further research (Samaniego, 2022). In the code-auditing vision, a central challenge is aligning trustworthiness indicators in evaluation with data-quality indicators in training at scale (Wang et al., 2024).
A plausible synthesis is that DataTrust has evolved from a problem of custody and exchange into a broader research program on verifiable provenance, formally constrained sharing, controllable trust trade-offs, and user-visible assurance. What remains unsettled is not the importance of the problem, but the extent to which these diverse formulations can converge on interoperable definitions, comparable metrics, and deployable assurance levels across domains.