Digital Twin-Based Decision Support
- Digital twin-based decision support systems are virtual replicas that combine real-time data, simulation, and optimization to enhance decision making.
- They employ layered architectures—integrating physical sensors, data analytics, and predictive modeling—to offer descriptive, diagnostic, and prescriptive insights.
- Key applications span urban EV charging, industrial maintenance, and healthcare, demonstrating modular design, uncertainty management, and robust interface integration.
Searching arXiv for papers on digital twin-based decision support systems to ground the article in relevant literature. A digital twin-based decision support system is a decision-oriented digital representation of a physical asset, process, or environment that combines data, models, and computational services to support monitoring, analysis, prediction, optimization, and, in some cases, automated control across a system’s life cycle. Across the literature, the concept appears as a dynamic, continually updated virtual model linked to a physical counterpart through data flows and used to evaluate scenarios, quantify outcomes, and guide decisions before or during real-world execution (Rasheed et al., 2019, Beek et al., 2022). In operational terms, such systems have been instantiated as an agent-based urban charging twin for electric-vehicle infrastructure (Do-Bui-Khanh et al., 21 Oct 2025), a service-based predictive-maintenance twin for Industry 5.0 (Esteves et al., 10 Nov 2025), an operating-room schedule execution twin (Rifi et al., 3 Sep 2025), a construction-phase quality-assurance twin (Islam et al., 18 Feb 2026), and a treatment-optimization clinical twin that combines digital twin simulation, treatment-effect estimation, and reinforcement learning (Qin et al., 16 Jun 2026).
1. Definition and conceptual scope
The literature converges on a view of the digital twin as more than a static model. One formulation describes it as “a virtual representation of a physical asset enabled through data and simulators for real-time prediction, monitoring, control and optimization of the asset for improved decision making throughout the life cycle of the asset and beyond” (Rasheed et al., 2019). Another perspective formalizes the twin as a probabilistic model linked to decision variables and utility, explicitly tying prediction to design and operational choices across the system life cycle (Beek et al., 2022). In applied settings, this general concept becomes domain-specific: a campus EV charging ecosystem represented as infrastructure, vehicles, energy assets, policies, and weather (Do-Bui-Khanh et al., 21 Oct 2025); a patient-specific clinical replica combining rules, machine-learning predictions, and explainability artifacts (Rao et al., 2019); or an executable operating-room model driven by planned or performed schedules (Rifi et al., 3 Sep 2025).
A recurrent distinction in the literature separates a digital model, a digital shadow, and a digital twin. The distinction matters because decision support changes with the degree of coupling. A digital model supports offline analysis without automated synchronization; a digital shadow adds automated physical-to-digital updating; a digital twin adds bidirectional integration and can influence physical operation (Agrawal et al., 2022, Wei et al., 2024). A model-based maturity formulation sharpens this progression into levels from descriptive and analytical to operational, prescriptive, cognitive, and connected cognitive, linking higher maturity to increasing automation, inclusion of physical-to-digital and digital-to-physical flows, and broader system-to-system integration (Wei et al., 2024).
This suggests that “digital twin-based DSS” names not a single architecture but a family of systems whose common denominator is the use of a synchronized or executable virtual representation to support decisions. A plausible implication is that the minimum defining features are not identical across domains: in some applications, such as operating-room analysis, the twin is primarily a scenario-execution backbone (Rifi et al., 3 Sep 2025), whereas in others, such as predictive maintenance or adaptive clinical treatment, it is embedded in a closed loop with online updating, uncertainty handling, and automated action pathways (Esteves et al., 10 Nov 2025, Qin et al., 16 Jun 2026).
2. Core architectural patterns
Despite domain variation, the systems surveyed exhibit a recurring layered structure. One common pattern combines a physical layer, a data layer, a model or simulation layer, a decision or optimization layer, and a visualization or interface layer. In the EV charging infrastructure twin, the physical layer includes the campus road network, parking areas, chargers, PV, wind, and BESS; the data layer combines GIS, field observations, surveys, and weather time series; the model layer is an agent-based model implemented in GAMA; the decision layer embeds metaheuristics; and the dashboard layer exposes controls, animated GIS state, and KPIs (Do-Bui-Khanh et al., 21 Oct 2025). In predictive maintenance, the DT-Create suite is organized as System Environment, Data Collect, Preprocessing, AutoML, Ontology Model, Inference Machine, Autonomous Agent, and Web Interface (Esteves et al., 10 Nov 2025).
A related model-driven architecture derives decision-oriented digital twins from engineering models of the physical twin. That framework separates models of the physical twin, models of the digital twin, and models in the digital twin, then instantiates a runtime architecture composed of a gateway, an engine, a digital shadow, and services such as monitoring, anomaly detection, and cockpit control (Bilal et al., 3 Jul 2026). The methodology associates increasing decision-support capability with increasing automation and connectivity, culminating in prescriptive and cognitive levels where automated digital-to-physical actions become part of the twin itself (Wei et al., 2024).
In simulation-centric deployments, the architecture often takes the form of data acquisition and storage, a simulation model, an analytics or experimentation layer, and a dashboard. The operating-room management system exemplifies this structure: a relational database stores master schedules, case attributes, and durations; FlexSim Healthcare executes the OR twin; an analytics layer manages prospective and retrospective experiments; and the interface presents KPIs and Gantt charts (Rifi et al., 3 Sep 2025). In wastewater treatment, a learned open-loop simulator separates historical state inference from future rollout under prescribed controls and forecasts, thereby supporting operator plan screening over 12–36-hour horizons (Simethy et al., 22 Apr 2026).
Across these examples, the architectural role of the twin is to bind heterogeneous data and executable models into a decision loop. In some cases, the loop is explicitly cyclical: input data and scenario parameters drive simulation, simulation produces KPIs, optimization algorithms update configuration variables, and the loop repeats until near-optimal decisions are found (Do-Bui-Khanh et al., 21 Oct 2025). In others, the loop follows a MAPE-K-like structure in which sensing, analysis, planning, and execution are separated but linked through shared models and knowledge bases (Esteves et al., 10 Nov 2025).
3. Modeling substrates and computational mechanisms
Digital twin-based DSSs are distinguished not only by connectivity but by the kinds of models they embed. Several strands recur: physics-based models, agent-based or discrete-event simulation, probabilistic state estimation, machine learning, semantic reasoning, and optimization. The broad review literature emphasizes that a twin is typically built from “physics-based models,” “data-driven models,” and “hybrid analysis and modeling (HAM),” with data assimilation acting as the mechanism that continually aligns virtual and physical state (Rasheed et al., 2019). A design-oriented perspective similarly characterizes the twin as a hierarchical, multi-model representation equipped with data fusion, Bayesian calibration, uncertainty quantification, and policy optimization (Beek et al., 2022).
At the simulation level, domain-specific abstractions dominate. In EV charging, the twin uses an agent-based model with charging-station agents, energy agents, and vehicle agents, updated every 5 minutes; renewable generation, BESS state, vehicle states, and charging decisions are co-simulated (Do-Bui-Khanh et al., 21 Oct 2025). In operating-room management, the twin is a discrete-event simulation of setup, anesthesia, surgery, reversal, idle, and off-schedule states, supporting both deterministic and stochastic duration models and alternative urgent-case insertion policies (Rifi et al., 3 Sep 2025). In wastewater treatment, the learned simulator uses a controlled continuous-time state-space formulation with typed context encoding, gain-weighted forcing of prescribed and forecast drivers, semigroup-consistent rollouts, and Student- plus hurdle outputs to cope with irregular sampling, missingness, heavy tails, and zero inflation (Simethy et al., 22 Apr 2026).
Clinical decision support introduces additional computational layers. One patient-specific formulation represents the twin as a feature vector, a prediction function, and explanatory functions such as LIME and partial dependence plots, allowing clinicians to interrogate risk predictions and refine rule-based thresholds (Rao et al., 2019). A more advanced adaptive framework treats the patient twin as a learned transition model , combines it with a treatment-outcome model , and trains a BCQ-based policy for sequential treatment optimization under safety constraints (Qin et al., 16 Jun 2026). In that system, the twin is not merely descriptive; it is the environment for offline RL training, a what-if engine for candidate treatment sequences, and a source of uncertainty estimates via model ensembles (Qin et al., 16 Jun 2026, Qin et al., 24 Aug 2025).
Semantic modeling appears where decision support requires contextual reasoning. The DT-Create suite uses OWL 2 ontologies manipulated via OWLReady2, SWRL rules for expert knowledge, and an inference engine that derives alerts, maintenance needs, and failure patterns from sensor data and predictive-model outputs (Esteves et al., 10 Nov 2025). This semantic layer bridges numerical predictions and actionable maintenance decisions, for example by inferring alert codes from combinations of temperature, humidity, and repeated failure patterns (Esteves et al., 10 Nov 2025).
A further methodological development explicitly targets the decision objective of the twin. “DT: Decision-Targeted Digital Twins” argues that minimizing one-step transition error can be suboptimal for ranking policies according to a reward function, and proposes a ranking-aware loss that preserves pairwise policy orderings estimated via fitted Q-evaluation (Amad et al., 24 Jun 2026). This suggests that, for decision support, raw simulation fidelity and decision fidelity may diverge, and that digital twins can be trained directly for policy ranking and reduced decision regret rather than only for trajectory accuracy (Amad et al., 24 Jun 2026).
4. Decision-support functions and workflow types
The surveyed systems support several recurring classes of decisions: descriptive and diagnostic interpretation, predictive scenario evaluation, prescriptive policy comparison, and automated or semi-automated intervention. A practical framework frames these as the progression from description to diagnostic, predictive, and prescriptive analytics, and uses that progression to select an appropriate level of twin sophistication based on value, data and models, performance, and organizational transformation requirements (Agrawal et al., 2022).
In infrastructure planning, a digital twin-based DSS supports infrastructure siting and sizing, policy design, and seasonal energy strategy. The EV charging twin allows users to vary charger counts, PV panels, and policy switches such as gasoline bans, idle fees, relocation, and real-time notifications, then evaluate satisfaction, self-sufficiency, self-consumption, payback, and profits across scenarios (Do-Bui-Khanh et al., 21 Oct 2025). Exhaustive enumeration and embedded metaheuristics are both used, with optimization focusing on the number of 11 kW and 30 kW chargers and the number of solar panels (Do-Bui-Khanh et al., 21 Oct 2025).
In industrial maintenance, decision support operates over alerts, planning, and autonomous actions. DT-Create’s workflow is data acquisition, preprocessing, model training or selection, prediction on new data, ontological enrichment and inference, and human or autonomous decisions (Esteves et al., 10 Nov 2025). Decisions include classifying failures as critical, determining whether intervention is needed, generating alert codes, scheduling maintenance, and sending stop commands via MQTT (Esteves et al., 10 Nov 2025). The system is explicitly described as human-centric: dashboards and mobile interfaces support maintenance managers and operators rather than replacing them (Esteves et al., 10 Nov 2025).
In healthcare, decision-support workflows range from explainable diagnostic support to sequential treatment planning. The liver-disease twin combines DMN rule tables, machine-learning classification, local and global explainability, and threshold refinement to reduce physician-to-physician subjectivity in risk assessment (Rao et al., 2019). The adaptive treatment framework closes a more elaborate loop: observed state is processed by the twin and RL ensemble, actions are filtered by a rule-based safety layer, uncertainty may trigger clinician review, and new outcomes update the models incrementally (Qin et al., 16 Jun 2026). In both cases, the twin supports counterfactual or patient-specific reasoning before intervention (Rao et al., 2019, Qin et al., 16 Jun 2026).
In operations management, digital twins support both prospective and retrospective decision workflows. The operating-room twin evaluates planned schedules prospectively through feasibility checks, deterministic and stochastic performance assessment, robustness to duration variability, resilience to non-elective arrivals, and combinations thereof (Rifi et al., 3 Sep 2025). Retrospectively, it reconstructs performed schedules, checks resource-feasibility under deterministic assumptions, evaluates actual performance, and compares performed and provisional schedules to distinguish weaknesses in the original plan from suboptimal day-of-surgery decisions (Rifi et al., 3 Sep 2025).
In construction QA, the twin supports “structured release or hold decisions before standard-age test results are available” by linking inspections, plant data, embedded sensing, lab tests, and predictive strength models to individual elements (Islam et al., 18 Feb 2026). The resulting element-centric QA state—pending, provisional, released, hold, or non-conforming—supports readiness-based progression decisions while preserving contractual inspection and testing authority (Islam et al., 18 Feb 2026).
A recurring misconception is that a digital twin-based DSS must always imply full automation. The literature does not support that claim. Some systems deliberately remain advisory or planning-oriented, such as the operating-room simulator or conversational household twin (Rifi et al., 3 Sep 2025, Mylonas et al., 30 Jun 2026). Others enable automatic action only after safety checks, uncertainty thresholds, or domain-rule validation, as in predictive maintenance and clinical treatment optimization (Esteves et al., 10 Nov 2025, Qin et al., 16 Jun 2026).
5. Evaluation criteria, metrics, and evidence
Evaluation in digital twin-based DSS research is multi-dimensional. The central axes are usually simulation fidelity, decision quality, uncertainty handling, operational robustness, and practical utility. Different domains instantiate these axes differently, but the pattern is consistent: the twin is not judged solely by predictive error.
In EV charging planning, decision criteria are explicitly scalarized. User satisfaction is defined as the fraction of EV drivers who complete charging among those requesting it, self-sufficiency and self-consumption quantify renewable integration, and payback period captures financial viability. These are combined into an objective function: which is maximized by embedded metaheuristics (Do-Bui-Khanh et al., 21 Oct 2025). Statistical comparisons between scenarios use the Wilcoxon signed-rank test (Do-Bui-Khanh et al., 21 Oct 2025).
In predictive maintenance, model quality is evaluated with Accuracy, AUC, Recall, Precision, F1, Cohen’s Kappa, and MCC, alongside practical feasibility of the service suite (Esteves et al., 10 Nov 2025). Case-specific metrics include a cross-validation mean accuracy of approximately $0.9976$ for textile failure criticality classification and, for heat-treatment furnaces, tuned Decision Tree performance of Accuracy , AUC , Recall , Precision , F1 0, Kappa 1, and MCC 2 (Esteves et al., 10 Nov 2025). These numerical results support the claim that the DSS can generate meaningful predictions and semantically enriched recommendations, though the paper also notes dependence on data preprocessing quality and ontology design (Esteves et al., 10 Nov 2025).
Clinical systems add policy-level and safety-level metrics. The adaptive CDSAS reports mean return, consistency, safety compliance, query rate, response time, throughput, and component performance of the digital twin and outcome model (Qin et al., 16 Jun 2026). In synthetic and TCGA ovarian-cancer settings, the method outperforms computational baselines while maintaining 100% safety compliance with predefined constraints (Qin et al., 16 Jun 2026). The earlier online adaptive variant reports low latency, a low expert query rate, and a bounded-residual twin update rule that stabilizes rollouts under online adaptation (Qin et al., 24 Aug 2025).
For operating-room management, performance indicators include OR utilization, staff overtime, and patient waiting time (Rifi et al., 3 Sep 2025). The case study reports OR utilization of 3 and staff overtime of 4 across compared urgent-case insertion rules, revealing that structural constraints, not merely rule choice, drove underperformance relative to targets (Rifi et al., 3 Sep 2025). In wastewater treatment, the learned simulator is evaluated by RMSE and CRPS at long horizon; at 5, it achieves RMSE 6 and CRPS 7, reducing RMSE by 8 relative to Neural CDE baselines (Simethy et al., 22 Apr 2026). Importantly, the paper complements forecast metrics with operator-relevant case studies such as plan ranking and robustness to context-only sensor outages (Simethy et al., 22 Apr 2026).
At the interface layer, the conversational household twin evaluates the reliability of the decision-support frontend itself. On 45 prompts, the agentic interface achieves 100% schema conformance, 96.1% field-level F1, 90.4% value accuracy, and a 95.6% end-to-end simulation success rate (Mylonas et al., 30 Jun 2026). This shifts evaluation beyond the underlying twin to the human-access path through which decisions are formulated.
The decision-targeted DT literature further argues that evaluation should include policy ranking and decision regret. DT9 reports consistent improvements in policy ranking and reductions in decision regret relative to conventional one-step-loss DT training, while maintaining “a good level of raw simulation fidelity” (Amad et al., 24 Jun 2026). This suggests that decision-support evaluation should explicitly test the twin’s usefulness for selecting among alternatives, not only its ability to reproduce observed trajectories.
6. Cross-domain applications, transferability, and limitations
The application range of digital twin-based DSS is broad: urban mobility and energy (Do-Bui-Khanh et al., 21 Oct 2025), industrial maintenance (Esteves et al., 10 Nov 2025), hospital operations (Rifi et al., 3 Sep 2025), disaster management (Dogan et al., 2021), civil-infrastructure QA (Islam et al., 18 Feb 2026), residential energy (Mylonas et al., 30 Jun 2026), wastewater operations (Simethy et al., 22 Apr 2026), clinical diagnosis and treatment (Rao et al., 2019, Qin et al., 16 Jun 2026), and Industry 4.0 systems engineering (Bilal et al., 3 Jul 2026, Wei et al., 2024). The common structural pattern is the integration of a virtual representation with data-driven state updating or scenario specification, followed by predictive, comparative, or prescriptive analysis.
Transferability depends heavily on modularity. The EV charging twin emphasizes generic agent classes and GIS-driven modularity, making it straightforward to swap spatial data and scale from campus to district or city (Do-Bui-Khanh et al., 21 Oct 2025). The model-based Industry 4.0 framework similarly treats the digital twin as a configurable product line instantiated from physical-twin models and service choices (Bilal et al., 3 Jul 2026). DT-Create frames its contribution as a service suite for specifying digital twins in predictive maintenance, implying reuse across domains through intelligent processing, semantic enrichment, and self-adaptation (Esteves et al., 10 Nov 2025).
Several limitations recur. One is data quality and representativeness. Predictive maintenance depends on balanced and well-preprocessed datasets (Esteves et al., 10 Nov 2025). Construction QA faces heterogeneous, incompatible systems and non-standardized element identifiers (Islam et al., 18 Feb 2026). Wastewater simulation must tolerate 43% missingness and irregular 1–20 minute sampling (Simethy et al., 22 Apr 2026). Clinical systems confront retrospective bias, residual confounding, and lack of prospective trials (Qin et al., 16 Jun 2026). These issues constrain how strongly one can trust the decision recommendations.
A second recurring limitation is the gap between planning twins and live operational twins. The EV charging model uses realistic historical and site-specific data but is “not yet fed live operational data for real-time control” (Do-Bui-Khanh et al., 21 Oct 2025). The residential energy HDT is scenario-executable and physics-based, but it is not a continuously calibrated operational twin (Mylonas et al., 30 Jun 2026). The disaster-management twin is explicitly presented as a proposal and proof of concept, with live IoT ingestion and broad cloud deployment left for future work (Dogan et al., 2021). This suggests that many current “digital twin-based DSSs” are still closer to high-fidelity scenario laboratories than to fully online cyber-physical controllers.
A third limitation concerns organizational and contractual integration. The digitalization framework for DT practice emphasizes that selecting the right sophistication level requires aligning value, transformation, data and models, and performance; otherwise, projects risk strategic misalignment and adoption failure (Agrawal et al., 2022). Construction QA makes a related point in contractual terms: predictive models can support provisional release or hold decisions, but they do not replace established acceptance procedures (Islam et al., 18 Feb 2026). In many settings, therefore, the decisive challenge is not only model fidelity but governance, workflow embedding, and role allocation.
7. Research directions and emerging design principles
Several design principles emerge repeatedly. First, the twin should be built around the decision purpose, not around technology labels. A widely cited concern is that prediction, simulation, AI, and machine learning are often rebranded as necessary DT constituents when they may or may not be appropriate for the value sought (Agrawal et al., 2022). Purpose-driven selection is therefore foundational: description may suffice in some settings, while others justify predictive or prescriptive sophistication (Agrawal et al., 2022, Wei et al., 2024).
Second, combining models with semantics and decision logic appears increasingly important. DT-Create’s integration of AutoML, ontologies, SWRL inference, and an Autonomous Agent shows one pathway from raw data to explainable maintenance action (Esteves et al., 10 Nov 2025). Clinical systems similarly couple rule-based safety, treatment-effect models, uncertainty-driven referral, and simulation-based treatment comparison (Qin et al., 16 Jun 2026). This suggests that the most mature digital twin-based DSSs are not only simulators but composite reasoning systems.
Third, uncertainty handling is becoming a first-class design requirement. Ensemble dynamics models and Q-networks support uncertainty-driven clinician queries in adaptive treatment recommendation (Qin et al., 16 Jun 2026, Qin et al., 24 Aug 2025). Continuous-time state-space simulation in wastewater uses distributional outputs and skill-vs-persistence analysis to determine practical decision horizons (Simethy et al., 22 Apr 2026). Decision-targeted training further implies that, when model capacity is limited or data are imperfect, uncertainty should be treated with respect to decision quality, not just prediction intervals (Amad et al., 24 Jun 2026).
Fourth, interface design increasingly matters. The conversational HDT demonstrates that a physics-based twin can be made accessible through a two-tier agentic interface while preserving schema conformance and operational reliability (Mylonas et al., 30 Jun 2026). This suggests that natural-language access layers may become an important mediation mechanism between sophisticated twins and non-specialist stakeholders, provided the simulation backend remains authoritative and deterministic post-processing protects numerical integrity (Mylonas et al., 30 Jun 2026).
Finally, a plausible implication is that future digital twin-based DSSs will increasingly converge toward hybrid stacks: domain models or mechanistic simulators for structural validity, data-driven components for adaptation and speed, uncertainty-aware policy evaluation, and interfaces that support both expert interpretation and automated action. The literature already points in this direction through hybrid analysis and modeling (Rasheed et al., 2019), lifecycle-wide probabilistic decision framing (Beek et al., 2022), embedded optimization and simulation (Do-Bui-Khanh et al., 21 Oct 2025), and decision-targeted training objectives (Amad et al., 24 Jun 2026).