Published 16 Feb 2026 in cs.DB, cs.AI, and cs.AR | (2602.14699v1)
Abstract: This paper envisions a quantum database (Qute) that treats quantum computation as a first-class execution option. Unlike prior simulation-based methods that either run quantum algorithms on classical machines or adapt existing databases for quantum simulation, Qute instead (i) compiles an extended form of SQL into gate-efficient quantum circuits, (ii) employs a hybrid optimizer to dynamically select between quantum and classical execution plans, (iii) introduces selective quantum indexing, and (iv) designs fidelity-preserving storage to mitigate current qubit constraints. We also present a three-stage evolution roadmap toward quantum-native database. Finally, by deploying Qute on a real quantum processor (origin_wukong), we show that it outperforms a classical baseline at scale, and we release an open-source prototype at https://github.com/weAIDB/Qute.
The paper proposes an end-to-end quantum-native database architecture that integrates SQL-to-circuit compilation, quantum indexing, storage, optimization, and hybrid classical fallbacks.
Qute models quantum operators by latency, success probability, and approximation error, enabling noise-aware planning that adapts shots, circuit variants, or execution domains within declared budgets.
Prototype measurements on a 72-qubit QPU support cost-model extrapolations predicting a crossover near 2³⁰ rows for unindexed filtering, although large-scale performance remains unexecuted and transactional semantics remain unresolved.
Qute is a vision paper that proposes an end-to-end architecture for a quantum-native database, treating quantum computation as a first-class execution substrate rather than an isolated accelerator. The authors—affiliated with Shanghai Jiao Tong University, Tsinghua University, NUS, Microsoft, and HKBU—argue that prior work falls into two disconnected camps: applying quantum algorithms to classical databases (query optimization, index selection, approximate query processing) and using database abstractions to manage quantum workloads. Qute instead integrates quantum execution across parsing, planning, indexing, storage, and optimization.
Motivation and design principles
The paper targets workloads where classical systems struggle: full-table scans over unindexed data, high-dimensional similarity joins, and aggregation over large qualifying sets. Four challenges motivate the design: reformulating SQL operators into quantum primitives (superposition, interference), planning across noise-sensitive operators with fundamentally different cost characteristics, operating within a limited qubit budget that permits loading only small data fractions, and storing quantum-active attributes in fidelity-preserving formats. The stated design principles are to minimize quantum data loading, exploit amplitude-level parallelism instead of tuple-at-a-time processing, adopt fidelity-aware optimization, and maintain hybrid executability so classical paths remain viable under hardware noise.
System architecture
Qute comprises five components. The quantum SQL compiler lowers extended SQL through a three-tier LLVM-style IR (logical IR, quantum-extended IR, physical circuit IR) into gate-efficient circuits, maintaining a dynamic catalog of rewrite rules for joins and multi-predicate filters. The hybrid plan optimizer models each quantum operator as a tuple (Tq,Pq,εq) capturing latency, success probability, and approximation error, and generates plans embedding both quantum and classical realizations of eligible operators. Quantum-accelerable operators offload costly operations such as similarity joins onto circuits encoding computation in probability amplitudes. Quantum-aware indexing performs selective probing before allocating qubits. Finally, the storage engine stores quantum-active attributes as compressed tensor networks acting as logical pages, exposing fidelity-aware primitives (LOAD with error bound ϵ, bounded SAMPLE, REFRESH) while keeping classical metadata separately for recoverability.
The paper contrasts this state-centric storage with prior page-based designs, though it concedes implicitly that destructive measurement and the no-cloning theorem break conventional assumptions about caching and data reuse—an issue the paper acknowledges but does not resolve.
Three designs are detailed. For filtering, row identifiers form the search space; compound predicates are compiled into ancilla-assisted phase-marking oracles, followed by classical reconciliation (deduplication and rechecking) to guarantee correctness—a deliberate concession that quantum results alone are not trusted. For similarity join, vectors are encoded via parameterized rotations and inner products estimated through a CSWAP-based SWAP test, where the ancilla measurement probability ϵ2 yields the overlap without dimension-wise traversal. For aggregation, SUM is recast as probability estimation: each row controls a Good flag qubit via a rotation proportional to ϵ3, and amplitude estimation recovers ϵ4. Notably, the aggregation formulation assumes coherent value access (e.g., QRAM); on NISQ hardware without QRAM, data-loading cost would dominate, a dependency the paper states explicitly rather than hiding.
Selective quantum indexing
Because existing devices can hold only a small fraction of data in superposition, Qute proposes probing one-dimensional quantum Bϵ5-trees on selected predicate dimensions, taking the smallest candidate set (ϵ6), and branching dynamically: if ϵ7, classical post-filtering on remaining dimensions costs ϵ8; otherwise, the system escalates to a quantum multi-divided KD-tree that jointly indexes multiple dimensions. Disjunctive predicates are handled by independently probing each dimension and combining superpositions. This avoids the exponential ϵ9 cost of naive multi-dimensional range trees when selectivity is favorable, but the benefit is workload-dependent—the paper does not characterize the distribution of O(N)0 under which escalation dominates.
Noise-aware hybrid optimizer
The optimizer's estimation model grounds all three profile components in circuit structure and calibrated hardware parameters. Execution time sums per-layer gate durations plus control overhead after topology-aware routing, making SWAP-induced depth inflation explicit. Success probability composes multiplicatively across layers as O(N)1, so deeper or poorly routed circuits are pruned as fragile. Expected runtime blends quantum and classical paths, O(N)2, guaranteeing correctness through deterministic fallbacks. Cross-domain coordination costs (data movement, binding, reconciliation) are included in the objective, and deferred binding lets the runtime dispatcher choose domains based on calibration snapshots and queue delays. Runtime adaptation—increasing shots, switching circuit variants, falling back classically—operates strictly under declared error and latency budgets.
Preliminary results
A minimal prototype supporting filtering was deployed on the origin_wukong QPU (72 physical qubits, noisy), built on QPanda3 v1.0 with 2000 shots per experiment. Evaluation used a synthetic dataset of O(N)3 tuples with predicate selectivity bounded at 2%. Because classical runtimes depend heavily on system tuning, comparison used cost models on both sides: measured Qute latencies up to O(N)4 agree closely with the model's extrapolation (Grover error within ±8.0%), validating the cost model as an extrapolation instrument. The headline result is a crossover at approximately O(N)5: below this scale the classical baseline wins on constant overheads; above it, Qute is projected to outperform classical unindexed filtering. Two caveats deserve emphasis: the large-scale results are model-extrapolated rather than executed, and the comparison is against an unindexed classical scan—the regime where Grover's advantage is largest.
Limitations and open questions
The paper is candid about several constraints. NISQ hardware limits qubit counts, coherence, and connectivity; destructive measurement and no-cloning undermine conventional storage, caching, and transaction semantics. No transactional model exists for probabilistic, measurement-disturbing operations—consistency, traceability, and trust in hybrid results remain unsolved. The evaluation conflates measured and modeled performance, and the claimed advantage holds only for extreme-scale, unindexed filtering. The three-stage roadmap (quantum-assisted co-processing → quantum-centric execution with hybrid indexes → fully quantum-native storage under fault tolerance) is aspirational; stages two and three presuppose advances in quantum memory and fault-tolerant hardware that do not currently exist. Open questions include whether the O(N)6 profiles remain accurate under correlated noise, how index-selectivity distributions affect the KD-tree escalation threshold in practice, and what transactional semantics can support hybrid quantum-classical consistency.
Conclusion
Qute contributes a coherent end-to-end blueprint for quantum-native data management: SQL-to-circuit compilation via a tiered IR, stochastic cost modeling with classical fallbacks, selective quantum indexing, amplitude-encoded operators, and tensor-network storage. Its empirical claim—a crossover advantage beyond roughly O(N)7 rows for unindexed filtering—is supported by real-QPU measurements at small scale plus validated cost-model extrapolation, not by end-to-end execution at scale. The paper's principal value lies in its system-level integration and honest treatment of noise, fallback, and approximation as first-class optimization concerns, leaving hardware scaling and transactional semantics as the decisive open problems.