What Happens When the Model Eats the Stack? Rethinking the Research Agenda for Data Agents to Withstand the Bitter Lesson
Abstract: The bitter lesson poses an existential question for the data systems community, whereby LLMs trained end-to-end are rapidly internalizing new capabilities that previously required carefully engineered data agents. Guided by empirical insights, we argue that as models continue to improve, many proposed system layers designed to compensate for model limitations on a given task will increasingly be subsumed by the model itself. We instead identify enduring research opportunities, which lie in supporting data agents across many queries with curated contextual information about the data environment, which we call persistent semantic context. We find that these context layers demonstrate strong promise for improving data agent performance, but they also raise significant system challenges. Thus, a key requirement for future data systems will lie in natively serving persistent semantic contexts as a first-class abstraction in order to enable capable data agents working over huge, complex knowledge corpora. Towards this vision, we outline exciting new research opportunities, including designing efficient context data structures, storage methods, compression techniques, and semantic consistency protocols, to ensure integrity and correctness of the stored contextual knowledge.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper asks an important question:
As artificial intelligence models become smarter, what kinds of computer systems will still be needed?
The authors study data agents. These are AI programs that can answer questions about large collections of data. For example, a data agent might be asked:
“Why did sales fall in Europe last quarter?”
To answer, it may need to find the right tables, combine information, write computer code, check its calculations, and explain the result.
The paper argues that many carefully designed tools used by data agents may become less important as LLMs, or LLMs, improve. However, one problem will remain: AI agents still need reliable knowledge about the particular data environment they are working in.
The authors call this reusable knowledge persistent semantic context.
2. What questions did the researchers ask?
The paper focuses on four main questions:
- Are general-purpose coding agents becoming better than specially designed data agents?
- Do newer AI models solve data problems more accurately and efficiently?
- What kinds of mistakes do advanced agents still make?
- Can stored information about a company’s data help agents answer future questions?
The researchers were especially interested in whether AI systems still need many human-designed steps, or whether stronger models can figure out the steps themselves.
3. How did the researchers investigate this?
Comparing different kinds of AI agents
The researchers compared two types of systems:
- General coding agents: AI systems that receive a task and decide for themselves how to solve it, often by writing and running code.
- Human-designed data agents: Systems built with specific instructions and fixed procedures created by people.
This is similar to comparing:
- A student who follows a very detailed recipe, and
- A student who understands the goal and chooses their own method.
Testing different models
They tested several generations of LLMs, including:
o3GPT-5GPT-5.6 Sol
The paper says these models were tested across 2025 and 2026.
Using two benchmarks
A benchmark is a standard test used to compare computer systems fairly. The researchers used:
- TAG-Bench, which asks agents to combine calculations, database information, and general knowledge.
- Data-Agent Benchmark (DAB), which tests more complicated analysis tasks involving information spread across different business datasets.
The agents were judged on:
- Accuracy: Did they produce the correct answer?
- Efficiency: How many steps and how many words, or “tokens,” did they use?
- Failure types: What went wrong when they made mistakes?
Testing stored context
The researchers also created extra information for the AI to use. This information described things such as:
- What tables exist
- What different columns mean
- Which datasets are trustworthy
- How tables are connected
- What earlier users discovered
This stored information was called persistent semantic context.
An everyday analogy is a student keeping a well-organized notebook about a complicated school project. Instead of rediscovering every fact each time, the student can check the notebook before starting a new task.
4. What did the researchers find?
General coding agents became stronger than specially designed agents
With older or less capable models, a carefully designed data agent sometimes performed better. However, as the models improved, the general coding agents became more accurate and efficient.
The authors suggest that newer models can increasingly perform tasks such as:
- Planning
- Writing code
- Using tools
- Finding and fixing errors
- Checking their own answers
Because of this, fixed systems designed by humans may eventually limit the AI rather than help it.
Newer models used fewer steps
The newer coding agents needed fewer turns, or back-and-forth actions, to complete a task.
On the DAB benchmark, the newest model used about four times fewer turns than the older o3 model. It also improved its results by more than 35 percentage points and used fewer tokens.
This is important because fewer steps can mean:
- Lower cost
- Faster answers
- Fewer chances for something to go wrong
The result challenges the idea that future AI agents will always create huge numbers of inefficient database requests.
The biggest remaining problem was understanding the data environment
Even the strongest agents still made many mistakes because they misunderstood the data around them.
More than 60% of the remaining errors were related to environmental knowledge. For example, the agent might:
- Misunderstand what the question means
- Choose the wrong table
- Use the wrong column or measurement
- Connect two tables using the wrong identifier
- Misunderstand a company’s special definitions
This is different from simply making a calculation mistake. The agent may know how to calculate correctly but still use the wrong information.
For example, imagine asking:
“How many customers renewed their contracts?”
The agent might calculate perfectly but accidentally use a table about new customers instead of renewing customers. The arithmetic would be correct, but the answer would still be wrong.
Stored context improved performance
Giving the agent useful information in advance helped it perform better.
One accuracy-focused context improved performance by 19 percentage points compared with giving the agent no stored context.
A schema-focused context also helped the agent spend fewer turns exploring the database. In other words, it already knew more about where information was located.
However, there was a trade-off: some contexts became large, expensive, and time-consuming to create. The schema-focused context took more than 3,400 seconds to build and occupied much more space than the other contexts.
The experiments were also fairly small, using only 12 example tasks and 12 datasets. Real companies may have thousands of tables and terabytes of data, so creating and updating the context could be much harder in practice.
5. What does “persistent semantic context” mean?
The term can be understood by separating it into three parts:
- Persistent: It is saved and can be used again later.
- Semantic: It describes the meaning of information, not just its location.
- Context: It gives background knowledge that helps the AI understand a task.
For example, a context file might say:
- “This table contains official monthly revenue.”
- “The Finance department uses this definition of profit.”
- “Customer ID in Table A matches Account Number in Table B.”
- “For European sales, use the tax-adjusted revenue column.”
- “This older dataset should not be used for current reports.”
This information could be written as text files, stored in a knowledge graph, or organized in other searchable forms.
6. Why is this important?
The paper’s main message is that better AI models may handle more of the general reasoning and coding work by themselves. Human engineers may not need to build as many rigid step-by-step systems.
But AI models cannot automatically know every organization’s private and changing rules. They may not know:
- Which dataset is considered official
- What a company means by “active customer”
- How departments define important measurements
- Which tables contain current information
- How different systems refer to the same person or product
This knowledge is usually outside the model’s original training data. It must be provided separately and kept up to date.
7. What future research does the paper suggest?
The authors argue that future database systems should treat persistent semantic context as an important part of the system, rather than as an ordinary note or prompt.
They identify several challenges:
Keeping information correct
Databases, company rules, and documentation change over time. If the stored context is not updated, the AI could confidently use outdated information.
Researchers need ways to notice when a change in one table or definition affects other stored explanations.
Choosing how widely updates should spread
Some information may belong only to one user. Other information may apply to a whole team or an entire company.
A future system might need different levels of context:
- Personal context for one user
- Team context for a project
- Company-wide context for official rules
- Public context from regulations or outside sources
Storing information efficiently
Context might include millions of words from tables, documents, previous questions, and agent actions. An AI cannot read all of this every time.
Systems will need smart ways to:
- Search the most relevant information
- Store frequently used knowledge
- Remove repeated information
- Compress large collections of text
- Decide what information is no longer useful
Balancing speed, cost, and accuracy
Building a detailed context can take a lot of time and computer power. A short context may be cheap but miss important details. A large context may be more accurate but slow and expensive.
Finding the right balance will be a major research problem.
Conclusion
This paper argues that the future of data agents will not mainly depend on adding more complicated, hand-built instructions. As LLMs improve, general-purpose agents may learn to plan, code, debug, and check their work without as much human-designed scaffolding.
The harder problem is helping agents understand the specific world they are working in. A company’s data has its own names, rules, connections, and history. Saving this knowledge in a reusable form—persistent semantic context—can make agents more accurate and faster.
The possible impact is significant. If researchers can build context systems that are reliable, efficient, and easy to update, people may be able to ask complex questions about business, science, medicine, or government data using ordinary language. However, these systems must be carefully maintained, because outdated or incorrect context could cause an AI to give confident but misleading answers.
Knowledge Gaps
Knowledge Gaps, Limitations, and Open Questions
- The evaluation covers only two benchmarks,
TAG-BenchandDAB, so it remains unclear whether the reported superiority of general coding agents transfers to other domains, database architectures, modalities, enterprise environments, and real-world workloads. - The experiments use only three model versions from a single model ecosystem, limiting conclusions about whether the observed trends reflect general model scaling or provider-specific behavior.
- The study does not report statistical significance, confidence intervals, variance across repeated runs, or sensitivity to stochastic agent behavior, making the robustness of the performance differences uncertain.
- The comparison with
Agentar-Scale-SQLis not fully reproducible because its inference code is unavailable and the authors modify its procedure with a custom refinement step. - The human-designed baselines and the general coding agents may not be matched for prompt design, tool access, execution budgets, context windows, or engineering effort, leaving the fairness of the comparison unresolved.
- The paper evaluates primarily low-reasoning configurations and states that other reasoning levels show consistent trends, but it does not provide detailed results quantifying how reasoning budgets affect accuracy, latency, cost, and context use.
- The failure taxonomy is assigned by
GPT-5.6 Sol, creating potential evaluator bias and uncertainty about the reliability, reproducibility, and inter-rater validity of the root-cause classifications. - The paper does not establish whether schema exploration is truly the dominant bottleneck in production data-agent workloads, where network latency, permissions, data quality, query execution, and human approval may contribute substantially to total cost.
- The proposed notion of persistent semantic context lacks a precise formal definition, including its representation, provenance, dependency structure, validity conditions, and interface with the underlying data system.
- The experiments construct context from only 12 sample trajectories and 12 datasets, so the reported improvements may not generalize to larger, more heterogeneous, or less representative collections of historical traces.
- The context experiments do not isolate the effects of additional information from those of prompt length, formatting, model-generated instructions, or accidental leakage of benchmark-specific knowledge.
- The study does not test whether context generated from historical traces remains effective under distribution shift, newly introduced schemas, unseen business questions, changed data distributions, or unfamiliar organizational conventions.
- The contexts are appended to the prompt as static text; the paper does not evaluate retrieval, selective context loading, hierarchical memory, structured context, or context-window management strategies.
- The reported accuracy gains are not evaluated against the full lifecycle cost of context construction, storage, maintenance, retrieval, and invalidation over a realistic multi-query workload.
- The paper provides no break-even analysis showing how many queries are required for persistent context to offset its offline construction and maintenance costs.
- The scalability claims for terabyte-scale environments and thousands of tables are speculative; no experiments measure construction time, storage growth, retrieval latency, or update cost as environment size increases.
- The proposed context-generation methods do not include systematic comparisons with database catalogs, knowledge graphs, vector retrieval, materialized views, metadata systems, or hybrid representations under identical workloads.
- The tradeoff between context size and task performance is not characterized, including whether larger contexts improve coverage or instead increase distraction, prompt interference, latency, and hallucination.
- The paper does not define reliable metrics for semantic context quality, such as factuality, coverage, provenance completeness, freshness, contradiction rate, retrieval utility, or downstream task risk.
- It remains unresolved how an agent or system should detect and correct hallucinated, ambiguous, obsolete, or internally contradictory statements in generated context.
- The paper does not investigate provenance mechanisms that link each semantic claim to source schemas, documentation, queries, data samples, policies, or human decisions.
- No experiments evaluate whether agents can distinguish authoritative context from lower-confidence observations, conflicting user preferences, outdated traces, or speculative model-generated summaries.
- The proposed semantic consistency problem lacks operational guarantees: the paper does not specify how semantic dependencies should be represented, detected, validated, or proven correct after an underlying change.
- The appropriate consistency model for different applications is left undefined, including how to select among strong, eventual, query-triggered, or multi-version semantic consistency based on risk and workload requirements.
- The paper does not address conflict resolution when user, team, enterprise, regulatory, and external contexts provide incompatible definitions or recommendations.
- The effects of concurrent updates, access control, privacy policies, and authorization changes on persistent context are not studied.
- The paper does not examine how to prevent sensitive information, proprietary business logic, personal data, or secrets from being copied into broadly accessible agent context.
- No mechanism is proposed for auditing who created, modified, retrieved, or relied upon a semantic context artifact, which is important for regulated or high-impact data-agent applications.
- The relative merits of incremental maintenance versus complete context regeneration are discussed conceptually but not evaluated experimentally across update frequency, dependency density, environment size, and workload characteristics.
- The paper does not establish whether semantic dependencies can be represented explicitly enough to support reliable incremental maintenance, especially for implicit relationships encoded in natural language.
- The impact of stale context on downstream decisions is not quantified, including the severity, detectability, and reversibility of errors caused by outdated semantic knowledge.
- The paper does not evaluate context versioning, rollback, temporal queries, or reproducibility of agent results under historical semantic states.
- Compression is identified as an open problem, but no experiments compare textual summarization, retrieval-based pruning, structured compression, learned memory, or lossless representations for persistent context.
- It remains unknown which information can safely be discarded during context compression and how compression affects rare queries, edge cases, join correctness, and long-horizon task execution.
- The interaction between persistent context and model pretraining or test-time adaptation is not examined, including whether context duplicates parametric knowledge or creates harmful conflicts with it.
- The paper does not evaluate smaller, open-weight, fine-tuned, or multimodal models, leaving the accessibility and generality of the proposed research agenda uncertain.
- The relationship between improved single-query efficiency and total system throughput under concurrent multi-user workloads is not measured.
- The study does not assess whether persistent context changes agent behavior in ways that reduce exploration too aggressively, causing agents to miss relevant data sources or fail to verify inherited assumptions.
- The paper does not investigate human oversight requirements, including how users should inspect, edit, approve, correct, or override model-authored semantic context.
- No deployment study measures long-term context evolution, maintenance burden, failure accumulation, or user trust over extended periods of real usage.
- The claim that hand-engineered agent layers will be subsumed by increasingly capable models remains a forecast rather than a causally tested conclusion; the paper does not identify which system components are likely to remain valuable under future model and workload changes.
Practical Applications
Immediate Applications
- Enterprise natural-language data analysis (Software, business intelligence, finance, operations)
- Potential tools/workflows: conversational BI assistants, automated quarterly business reviews, finance variance analysis, supply-chain investigations, and self-service SQL/Python generation.
- Dependencies: reliable tool execution, read-only or carefully scoped database access, robust validation, data permissions, and human review for high-impact decisions. Performance may degrade when schemas, metric definitions, or join relationships are undocumented.
- Offline generation of semantic data catalogs (Data engineering, knowledge management)
- Potential products: agent-readable data catalogs, semantic schema registries, automated “how to use this dataset” documentation, and context packages attached to data products.
- Dependencies: representative historical traces, access to authoritative metadata, safeguards against hallucinated or obsolete descriptions, and periodic review by data owners.
- Context-augmented coding agents for analytics and data operations (Software engineering, data platforms)
- Potential workflows: automatic query repair, data-pipeline debugging, migration assistance, notebook generation, and documentation-linked code generation.
- Dependencies: context must be selectively retrieved rather than indiscriminately appended to prompts; otherwise, large or inaccurate context can increase latency, cost, or over-reliance on incorrect guidance.
- Automated semantic failure diagnosis (Data quality, observability, platform engineering)
- Potential tools: agent observability dashboards, benchmark suites, automatic trace labeling, root-cause alerts, and targeted regression tests.
- Dependencies: reliable trace capture and evaluation metrics that distinguish syntactic correctness from semantic correctness.
- Agent-assisted data governance and metric discovery (Governance, compliance, finance, public administration)
- Potential workflows: governed self-service analytics, metric certification, audit preparation, and policy-aware report generation.
- Dependencies: explicit ownership, versioning, access control, provenance, and human approval. The context should not be treated as authoritative unless its source and freshness are verifiable.
- Academic research assistants for complex data environments (Academia, computational science, social science)
- Dependencies: provenance preservation, reproducibility requirements, domain-specific validation, and protection of confidential or sensitive research data.
- Personal and small-business data assistants (Daily life, education, small-business operations)
- Dependencies: secure local or private-cloud deployment, accurate synchronization, simple correction mechanisms, and clear disclosure that generated answers may be wrong.
Long-Term Applications
- Persistent semantic context as a native database abstraction (Database systems, cloud infrastructure)
- Potential products: “context databases,” agent-native warehouses, semantic indexes, and context-aware query engines.
- Dependencies: scalable representations, workload-aware retrieval, interoperability across database systems, and evidence that maintenance costs are lower than repeated online exploration.
- Semantic consistency and dependency management (Database research, enterprise governance, regulated sectors)
- Potential workflows: automatically flagging contexts affected by a renamed field, propagating a new regulatory definition, or preventing an agent from using an obsolete metric.
- Dependencies: the ability to represent implicit semantic dependencies, define acceptable freshness guarantees, and reconcile conflicting user, team, enterprise, and external contexts.
- Versioned and scope-aware organizational memory (Enterprise software, collaboration, policy)
- Potential tools: version-controlled semantic workspaces, policy-aware context registries, and collaborative organizational-memory platforms.
- Dependencies: identity and authorization systems, conflict-resolution policies, auditability, and controls preventing private information from leaking across scopes.
- Semantic query optimization and agent planning (Data systems, software infrastructure)
- Potential products: semantic indexes, context-aware query planners, retrieval-augmented SQL engines, and agent-specific materialized views.
- Dependencies: accurate workload models, safe integration with query optimizers, and safeguards against recommendations that are semantically plausible but computationally or statistically invalid.
- Compression and lifecycle management for machine-readable organizational memory (AI infrastructure, storage systems)
- Potential tools: semantic compilers, context summarizers, long-term memory stores, and hot/cold context tiers.
- Dependencies: compression must preserve information needed for downstream decisions; lossy summaries require provenance, confidence scores, and mechanisms for recovering the underlying evidence.
- High-assurance agents for healthcare, finance, energy, and public services (Regulated and safety-critical sectors)
- Dependencies: strong semantic consistency, traceable evidence, privacy protection, domain validation, formal or procedural verification, and human authorization. The paper’s results alone do not establish suitability for autonomous decisions in these domains.
- Robotics and embodied systems with persistent environmental knowledge (Robotics, manufacturing, logistics)
- Dependencies: integration with sensor data and real-time state estimation, safe update mechanisms, uncertainty handling, and guarantees that stale context cannot cause physical harm.
- Adaptive educational and scientific knowledge environments (Education, research, knowledge management)
- Dependencies: privacy and consent, protection against reinforcing incorrect knowledge, teacher or researcher oversight, and mechanisms for distinguishing authoritative content from generated suggestions.
- Policy and regulatory infrastructure for agent-readable data environments (Public policy, standards, compliance)
- Dependencies: cross-vendor interoperability, measurable definitions of semantic consistency, sector-specific risk thresholds, and legal clarity regarding responsibility for stale or incorrect context.
Glossary
- Agentic capabilities: Abilities of an AI system to autonomously plan and perform multi-step actions. “agentic capabilities are undergoing a dramatic transformation”
- Agentic speculation: A workload pattern involving many exploratory or inefficient queries issued by an AI agent. “a workload characterized by sheer scale and inefficiency of agent-issued queries”
- Amortization: Distribution of a one-time computational or storage cost across multiple later operations. “its associated costs can be amortized across many user queries”
- Authoritative data source: A data source treated as the official or most trustworthy basis for a fact or measurement. “identifying authoritative datasets across multiple systems”
- Bitter Lesson: The principle that scalable general-purpose computation and learning tend to outperform manually engineered domain-specific methods. “This trend reflects the Bitter Lesson”
- Canonical metric: A formally accepted definition of a measurement used consistently across an organization. “canonical metric definitions”
- Compression: Reduction of the storage or processing requirements of information while attempting to preserve its useful content. “Future systems will require efficient techniques to consolidate accumulated knowledge”
- Context layer: A system component that stores contextual knowledge for use by AI agents. “persistent semantic contexts introduce significant system overheads”
- Contextual knowledge: Information about the environment, data, conventions, or task history relevant to an agent’s operation. “Contextual knowledge and understanding over large, complex data environments remains a substantial challenge”
- Data agent: An AI system designed to answer questions or perform analyses over structured or enterprise data. “future data systems will need to natively serve persistent semantic contexts as a first-class abstraction”
- Data profiling: Examination of datasets to discover their structure, contents, quality, and statistical properties. “steps spent on schema exploration or data profiling remain a substantial relative cost”
- Data structure: An organized representation used to store and retrieve information efficiently. “One central open question is: what data structure should represent semantic context?”
- Declarative: Describing the desired result or behavior without specifying every execution step. “natural language or semantic knowledge contained within the context layer”
- End-to-end model training: Training a model jointly across the complete processing pipeline rather than engineering separate task-specific components. “general methods that scale computation (e.g., end-to-end model training) ultimately displace hand-engineered domain knowledge”
- Entity relationship: A semantic association between identifiable objects or records in a data system. “entity relationships”
- Eventual consistency: A consistency model in which updates may propagate asynchronously but replicas converge over time. “collaborative analytics may tolerate eventual semantic consistency”
- Execution-guided refinement: Iteratively improving generated queries or programs by using the results of their execution. “We use its published SQL-generation prompt and add a custom execution-guided refinement step”
- First-class abstraction: A concept directly and explicitly supported by a system’s programming or storage model. “natively serving persistent semantic contexts as a first-class abstraction”
- Garbage collection: Automatic identification and reclamation of storage occupied by information no longer considered useful. “motivating new garbage collection and compression policies”
- Generalization: The ability of a model or system to perform successfully on cases beyond those directly used during development. “their performance and generalization will likely surpass that of data agents”
- Holistic regeneration: Rebuilding an entire representation from its underlying sources instead of updating only changed portions. “increasingly capable long-context models may make holistic regeneration an attractive alternative”
- Incremental view maintenance: Updating a derived database representation by propagating source changes rather than recomputing it completely. “This resembles incremental view maintenance”
- Inference: The process of using a trained model to produce outputs for new inputs. “whose full inference code is not publicly available”
- Join key: An attribute used to match records across tables in a relational operation. “incorrectly using entities, join keys, and identifiers”
- Knowledge graph: A graph-based representation of entities, relationships, and associated information. “structured knowledge graphs and vector indexes”
- Long-context model: A model capable of processing unusually large amounts of input context in one interaction. “increasingly capable long-context models”
- Materialized view: A stored result of a query that can be reused instead of recomputed. “indexes, materialized views, and caches”
- Multi-agent system: A system in which multiple interacting agents jointly perform a task. “DeepEye is a multi-agent system”
- Non-parametric knowledge: Explicit knowledge supplied to a model externally rather than encoded in its learned parameters. “an increasingly important requirement for more capable data agents will lie in non-parametric contextual information”
- Parametric knowledge: Information encoded within a model’s learned parameters during training. “what the model does not acquire during training as parametric knowledge”
- Persistent semantic context: A durable, language-based representation of knowledge about a data environment that can be reused across tasks. “We refer to this contextual knowledge as persistent semantic context”
- Physical design: Decisions about how data or metadata are organized and stored to optimize system performance. “The physical organization of semantic context presents new opportunities”
- Prompt: Input text or instructions provided to a LLM to guide its output. “appending the generated context to the user prompt”
- Query-agnostic context: Context that is useful across many queries rather than tailored to one specific query. “agents are likely to rely on both query-agnostic context”
- Query-triggered consistency: A consistency strategy that revalidates stored information only when it is accessed by a query or agent. “systems may adopt query-triggered semantic consistency”
- Relational database: A database that organizes data into tables connected through defined relationships. “queries that require combining exact computation, semantic reasoning, and world knowledge over relational databases”
- Semantic consistency: The property that stored semantic knowledge remains correct and current relative to the underlying environment. “ensuring that the natural language or semantic knowledge contained within the context layer is up-to-date and correct”
- Semantic dependency: A meaning-based relationship in which one piece of contextual knowledge depends on another. “changes may propagate through implicit semantic dependencies rather than explicit relational dependencies”
- Semantic reasoning: Drawing conclusions based on meaning, concepts, and relationships rather than only literal data operations. “queries that require combining exact computation, semantic reasoning, and world knowledge”
- Schema exploration: Investigating the structure, fields, relationships, and metadata of available data sources. “the relative cost of schema exploration has increased over successive model generations”
- Schema knowledge: Information about the organization and attributes of data sources. “capturing schema knowledge to reduce online exploration”
- Subagent: An auxiliary agent launched by a primary agent to perform part of a larger task. “forking subagents when paired with a terminal, file system and execution loop”
- Token efficiency: The amount of useful work or accuracy achieved relative to the number of language-model tokens consumed. “general coding agents are rapidly improving both accuracy and token efficiency”
- Vector index: An index that enables efficient retrieval of items based on similarity between numerical vector representations. “structured knowledge graphs and vector indexes”
- Workload-aware strategy: A system strategy tailored to the characteristics and access patterns of expected tasks. “developing workload-aware strategies for semantic maintenance”





