Papers
Topics
Authors
Recent
Search
2000 character limit reached

What Happens When the Model Eats the Stack? Rethinking the Research Agenda for Data Agents to Withstand the Bitter Lesson

Published 2 Sep 2026 in cs.DB | (2609.03141v1)

Abstract: The bitter lesson poses an existential question for the data systems community, whereby LLMs trained end-to-end are rapidly internalizing new capabilities that previously required carefully engineered data agents. Guided by empirical insights, we argue that as models continue to improve, many proposed system layers designed to compensate for model limitations on a given task will increasingly be subsumed by the model itself. We instead identify enduring research opportunities, which lie in supporting data agents across many queries with curated contextual information about the data environment, which we call persistent semantic context. We find that these context layers demonstrate strong promise for improving data agent performance, but they also raise significant system challenges. Thus, a key requirement for future data systems will lie in natively serving persistent semantic contexts as a first-class abstraction in order to enable capable data agents working over huge, complex knowledge corpora. Towards this vision, we outline exciting new research opportunities, including designing efficient context data structures, storage methods, compression techniques, and semantic consistency protocols, to ensure integrity and correctness of the stored contextual knowledge.

Summary

  • this paper find sustained system evolution for persistent semantic context allow more accurate data analysis agonist the model’s parametric competence.
  • empirical evaluation shows coding agent accuracy increased by 19 percentage points when using a simple persistent context strategies.
  • data agents fail due to semantic misinterpretation, highlighting the need for reusable semantic context despite significant build and storage costs.

Empirical motivation and central thesis

The paper examines which components of data-agent systems are likely to remain valuable as foundation models acquire increasingly strong capabilities for planning, code generation, tool use, debugging, and verification. Its organizing premise is Sutton’s “bitter lesson”: methods that scale general computation tend to displace domain-specific, hand-engineered procedures. Applied to data systems, this premise raises a direct question: if general-purpose models increasingly internalize the capabilities encoded in agent scaffolds, which systems abstractions will continue to matter?

The authors’ answer is persistent semantic context. They distinguish between capabilities that can be supplied by the model’s parametric competence and knowledge that is specific to a changing data environment. The latter includes schema interpretation, authoritative-source selection, join relationships, organization-specific metric definitions, undocumented conventions, historical discoveries, and user preferences. Such knowledge is external to model parameters, changes as the environment evolves, and is repeatedly rediscovered by agents unless it is stored and reused.

The paper therefore advances two related claims. First, increasingly capable general coding agents can subsume much of the functionality provided by human-designed data-agent pipelines. Second, the persistent systems problem is not primarily the optimization of individual agent trajectories, but the construction, serving, maintenance, and validation of reusable semantic knowledge across many queries. This argument extends earlier data-agent work such as TAG (Biswal et al., 2024), while placing greater emphasis on cross-query environmental knowledge than on query-local orchestration.

Experimental design

The empirical analysis compares general coding agents with human-designed data agents across two benchmarks. TAG-Bench combines exact computation, semantic reasoning, and world knowledge over relational databases (Biswal et al., 2024). DAB evaluates multi-step analytical tasks over fragmented enterprise data (Ma et al., 21 Mar 2026). The evaluation spans three model generations—o3, GPT-5, and GPT-5.6 Sol—and uses a Codex-style coding-agent harness for the general agents.

The human-designed baselines differ by benchmark. On TAG-Bench, the authors evaluate an Agentar-Scale-SQL-inspired system based on its published SQL-generation prompt and add an execution-guided refinement step. On DAB, they use DeepEye, a multi-agent workflow system designed for multi-turn data-analysis tasks [Li_2026]. This setup permits a comparison between a relatively general agent loop and systems that encode explicit assumptions about decomposition, SQL generation, workflow construction, and repair.

The comparison is informative but not fully symmetric. The paper explicitly states that the complete Agentar-Scale-SQL inference implementation is unavailable, so its baseline is an approximation rather than a reproduction of the published system. The results should consequently be interpreted as evidence about the relative trajectory of general coding agents and a representative human-designed pipeline, not as a definitive ranking of all possible data-agent architectures.

Figure 1

Figure 1

Figure 1

Figure 1: Performance evolution of general coding agents and human-designed data agents on TAG-Bench and DAB across successive model generations.

General coding agents increasingly subsume task-specific pipelines

The principal empirical result is that general coding agents improve more rapidly than the human-designed data-agent baselines. With the weaker o3 model, the human-designed Agentar-inspired system performs better than the coding agent on TAG-Bench and is more token-efficient. With subsequent model generations, however, the relationship reverses. GPT-5.6 Sol’s general coding agent achieves substantially higher accuracy than the human-designed agents on both TAG-Bench and DAB while also improving token efficiency.

This reversal supports the paper’s strongest claim: as models internalize planning, code synthesis, debugging, and validation, rigid agent architectures can become liabilities because they impose human priors that restrict generalization. The authors do not claim that all scaffolding is immediately obsolete. Rather, their longitudinal comparison suggests that scaffolding designed to compensate for model weaknesses loses relative value as those weaknesses are reduced by model scaling.

On TAG-Bench, the GPT-5.6 Sol coding agent reaches accuracy comparable to an oracle baseline consisting of expert-written queries executed with the LOTUS runtime (Patel et al., 2024). On DAB, the GPT-5.6 Sol coding agent improves by more than 35 percentage points over the o3 coding agent while also achieving more than a 2×2\times improvement in token efficiency. These gains arise from a general-purpose harness without benchmark-specific task decomposition or data-specific engineering.

The implication is methodological as well as architectural. Data-agent research that evaluates only a fixed model generation may overestimate the long-term value of specialized orchestration. A system improvement that compensates for current deficiencies in planning or error recovery may be outperformed by a later model that learns those behaviors directly. Evaluations should therefore measure not only absolute performance but also how system components interact with model capability over successive generations.

Efficiency gains change the bottleneck

The paper analyzes DAB trajectories by decomposing agent turns into schema exploration, analytical computation, semantic querying, verification, and recovery. The number of turns required per query declines sharply across model generations. In particular, GPT-5.6 Sol requires approximately four times fewer turns per task than o3.

Figure 2

Figure 2: Turns per query for coding agents across model generations on DAB, decomposed by action type.

The reduction is concentrated in analytical and semantic querying, as well as verification and recovery. This result directly challenges the premise that future data systems will primarily be bottlenecked by “agentic speculation”—large volumes of inefficient, exploratory queries issued by increasingly autonomous agents. The measured trajectory instead indicates that improved models formulate better actions, recover from fewer failures, and reach correct answers with fewer interactions.

This finding does not imply that query volume or execution cost is irrelevant. It establishes a narrower point: the measured per-task reasoning workload is declining across the evaluated model generations, so data-system designs premised on an indefinite increase in speculative agent-issued queries may be miscalibrated. The result also shifts attention toward the residual cost that does not decline at the same rate: understanding the external data environment.

Although total turns decrease, schema exploration and profiling remain a substantial relative component of GPT-5.6 Sol’s workload. As computation, planning, and debugging become more efficient, environmental acquisition accounts for a larger fraction of the remaining effort. The authors thus identify a changing bottleneck: not how to make the model reason through more query variants, but how to provide it with accurate knowledge of the relevant data environment before and during execution.

Failure analysis: environmental knowledge dominates residual errors

The authors classify failed coding-agent trajectories into five root-cause clusters:

Cluster Failure type
C1 Misinterpreted task meaning or invalid proxy
C2 Wrong data source, table, field, or metric
C3 Entity, join-key, or identifier mismatch
C4 Incorrect filtering, aggregation, ranking, or output logic
C5 Runtime, tooling, or command-execution failure

For GPT-5.6 Sol, more than 60% of failures arise from semantic misinterpretation. These include misunderstanding the task, selecting an invalid proxy, choosing the wrong table or metric, and mishandling entities, join keys, or identifiers. By contrast, execution and tooling failures decline as model capability improves.

The distinction between C5 and the semantic clusters is central to the paper’s argument. C5 errors are plausibly addressable through better general model competence: improved command construction, code generation, debugging, and recovery. C1–C3 errors depend on facts about a particular environment. A model can be highly capable at reasoning in the abstract while still lacking knowledge of which dataset is authoritative, whether a field represents forecast or actual revenue, how an organization defines an entity, or which identifier supports a valid join.

The implication is that further improvements to generic execution competence may yield diminishing returns on these residual failures. The paper instead motivates an external knowledge layer that records environmental semantics and makes them available across tasks. This diagnosis is consistent with the broader distinction between parametric and non-parametric memory in retrieval-augmented generation (Lewis et al., 2020), but the proposed setting is more operationally demanding: the memory must represent a live, heterogeneous, organization-specific data environment rather than a relatively static document corpus.

Persistent semantic context

Persistent semantic context is defined as an explicit, language-based representation of knowledge about a data environment. It is constructed offline, accessed during online task execution, and amortized across multiple queries. Candidate contents include schema descriptions, canonical metric definitions, authoritative data sources, entity relationships, historical query discoveries, organizational conventions, workflow guidance, and user-specific preferences.

The paper evaluates this idea on held-out DAB tasks using four context-generation strategies: Self-Curated contexts optimized for accuracy, latency, or schema knowledge, and a GEPA-generated accuracy context (Agrawal et al., 25 Jul 2025). Each context is produced offline by GPT-5.6 Sol from 12 development trajectories and then appended to the agent’s prompt during evaluation.

The results establish that even simple agent-authored context can materially affect performance. The Self-Curated Accuracy context improves accuracy by 19 percentage points over the no-context baseline. The Schema context substantially reduces turns devoted to schema exploration, demonstrating that persistent knowledge can amortize repeated environmental discovery. However, the Schema context also causes a modest accuracy decline, which the authors attribute to possible over-reliance on the curated artifact and overfitting to the small development sample.

Figure 3

Figure 3

Figure 3: Coding-agent performance and execution composition under different persistent-context strategies.

The results expose a multidimensional trade-off rather than a uniformly beneficial intervention. Context optimized for accuracy need not minimize latency; context optimized for latency need not preserve accuracy; and context optimized for schema acquisition can reduce exploration while introducing stale or overgeneralized assumptions. Persistent context should therefore not be treated as a monolithic prompt augmentation. It is better understood as a materialized semantic state with workload-dependent objectives and physical-design choices.

The reported construction costs make this distinction especially important:

Context strategy Build time Build cost Context size
Self-Curated Accuracy 164.9 s $1.09 2.52 KB
Self-Curated Latency 154.9 s $1.13 2.25 KB
Self-Curated Schema 3,413.6 s $9.60 162.14 KB
GEPA Accuracy 1,359.8 s $12.16 7.55 KB

Even in the small experimental setting, schema-oriented context construction requires more than 3,400 seconds and produces a 162.14 KB artifact. The experiment uses only 12 development trajectories and 12 data environments, whereas a production enterprise may contain terabytes of data, thousands of tables, heterogeneous storage systems, evolving documentation, and long-horizon traces. The paper therefore does not present context construction as a cheap replacement for online exploration. Its result is more precise: offline construction can improve accuracy and reduce repeated exploration, but the resulting build, storage, update, and validation costs become first-class systems problems.

A database research agenda for semantic state

The proposed agenda treats persistent semantic context as a first-class data-system abstraction rather than as an incidental prompt file. This framing separates the work from conventional prompt engineering. The relevant system must construct, index, retrieve, update, compress, validate, and govern semantic artifacts whose content may be partly natural language and whose dependencies may be implicit.

Semantic consistency

The most distinctive systems problem is semantic consistency. Traditional consistency models operate over explicit objects and operations: rows, records, writes, reads, and transactions. Persistent semantic context contains natural-language statements and inferred relationships. A schema change, documentation update, revised business definition, or statistical shift can invalidate context entries whose dependencies are not explicitly represented.

The paper identifies several dimensions of this problem. Consistency scope may be user-specific, project-wide, enterprise-wide, or external. A user’s preferred data source can coexist with an enterprise canonical metric, while external regulatory guidance may impose another scope. Systems must determine which agents observe an update, how updates propagate across overlapping scopes, and how conflicting semantic views are represented.

Consistency models also require an application-dependent spectrum. Regulatory reporting may require strong semantic consistency, whereas exploratory analytics may tolerate eventual consistency. Query-triggered validation offers a lazy alternative in which context is revalidated when accessed. The appropriate choice depends on the trade-off among freshness, accuracy, availability, latency, and maintenance cost.

Maintenance methods present a further design choice. Incremental semantic maintenance could revalidate only artifacts affected by a detected change, analogous to incremental view maintenance, but semantic dependencies are often implicit and heterogeneous. Holistic regeneration avoids explicit dependency tracking but may be computationally expensive and can erase useful stable knowledge. The paper leaves open the problem of workload-aware maintenance granularity.

Representation, physical design, and compression

The paper does not prescribe a single representation. Free-form text files are simple and editable but have weak structural guarantees. Knowledge graphs expose entities and relations but require schema design and dependency maintenance. Vector indexes support semantic retrieval but may obscure provenance, freshness, and exact dependency structure. A practical system may combine these representations in the same way that database systems combine indexes, materialized views, and caches.

Physical design must distinguish query-agnostic context from query-aware context. Schema descriptions, business terminology, and organizational documentation can be reused broadly. Intermediate reasoning, execution traces, and task-specific discoveries may be useful only for particular workloads. Systems must decide what belongs in persistent storage, what should be cached, how memory should be partitioned across fast short-term representations and compact long-term storage, and how retrieval should be conditioned on query scope.

Compression is unavoidable for long-lived context layers. A mature context store may contain millions of tokens from schemas, documentation, historical queries, execution logs, and agent-generated observations. Such a corpus cannot be supplied wholesale to every model invocation. The resulting lifecycle problem includes deciding what to retain, merge, summarize, compress, or discard while preserving information relevant to future queries. Existing work on KV-cache compression addresses session-level working memory (Zweiger et al., 18 Feb 2026), but the paper emphasizes that persistent semantic memory introduces a different retention and correctness problem.

Limitations and open questions

The empirical claims are bounded by the evaluation design. The study uses two benchmarks, three model generations, and a particular coding-agent harness. The model names and benchmark environments do not establish that the same trends hold for other models, modalities, database engines, enterprise workloads, or longer-lived deployments. The Agentar comparison is approximate because the complete published inference pipeline is unavailable, and the context experiments use only 12 development trajectories. That small sample makes the observed gains informative but also creates a substantial risk of context overfitting.

The persistent-context evaluation appends generated artifacts to the prompt, so it does not isolate the benefits of semantic organization from the costs and limitations of prompt injection. A context layer implemented through retrieval, structured dependency tracking, provenance-aware storage, or model-side memory may exhibit different accuracy and latency behavior. The study also reports context size and construction cost but does not provide a complete end-to-end accounting of update, invalidation, retrieval, validation, and storage costs over a long operational horizon.

Several technical questions consequently remain open. How should semantic dependencies be represented so that a changed schema or metric definition triggers complete but selective invalidation? What correctness guarantees can be provided for model-generated context? How should conflicting user, team, enterprise, and external semantic views be reconciled? Under what conditions does retrieval outperform a compact global context, and when does additional context induce distraction or over-reliance? Finally, the paper does not establish whether persistent context can retain its accuracy advantage under continuous schema evolution and distribution shift rather than a fixed held-out evaluation.

Conclusion

The paper presents a longitudinal argument about the changing division of labor between foundation models and data systems. General coding agents increasingly outperform human-designed data-agent pipelines and require fewer turns for individual tasks, weakening the case for system layers whose main purpose is to compensate for deficiencies in generic planning, code generation, debugging, or verification.

The enduring difficulty identified by the study is environmental knowledge. Agents continue to fail because they misunderstand data semantics, select inappropriate sources, and use incorrect entity or join relationships. Persistent semantic context addresses this gap by externalizing reusable knowledge about a changing data environment. Its demonstrated benefits—most notably a 19-point accuracy improvement for an accuracy-oriented context and large reductions in schema exploration—are accompanied by substantial construction and maintenance costs.

The paper’s central systems proposal is therefore not simply to add more context to prompts, but to make semantic context a managed, auditable, physically designed, and consistency-aware data abstraction. The unresolved issue is how to provide this abstraction with database-level guarantees despite the implicit, heterogeneous, and model-generated nature of its contents.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper asks an important question:

As artificial intelligence models become smarter, what kinds of computer systems will still be needed?

The authors study data agents. These are AI programs that can answer questions about large collections of data. For example, a data agent might be asked:

“Why did sales fall in Europe last quarter?”

To answer, it may need to find the right tables, combine information, write computer code, check its calculations, and explain the result.

The paper argues that many carefully designed tools used by data agents may become less important as LLMs, or LLMs, improve. However, one problem will remain: AI agents still need reliable knowledge about the particular data environment they are working in.

The authors call this reusable knowledge persistent semantic context.

2. What questions did the researchers ask?

The paper focuses on four main questions:

  1. Are general-purpose coding agents becoming better than specially designed data agents?
  2. Do newer AI models solve data problems more accurately and efficiently?
  3. What kinds of mistakes do advanced agents still make?
  4. Can stored information about a company’s data help agents answer future questions?

The researchers were especially interested in whether AI systems still need many human-designed steps, or whether stronger models can figure out the steps themselves.

3. How did the researchers investigate this?

Comparing different kinds of AI agents

The researchers compared two types of systems:

  • General coding agents: AI systems that receive a task and decide for themselves how to solve it, often by writing and running code.
  • Human-designed data agents: Systems built with specific instructions and fixed procedures created by people.

This is similar to comparing:

  • A student who follows a very detailed recipe, and
  • A student who understands the goal and chooses their own method.

Testing different models

They tested several generations of LLMs, including:

  • o3
  • GPT-5
  • GPT-5.6 Sol

The paper says these models were tested across 2025 and 2026.

Using two benchmarks

A benchmark is a standard test used to compare computer systems fairly. The researchers used:

  • TAG-Bench, which asks agents to combine calculations, database information, and general knowledge.
  • Data-Agent Benchmark (DAB), which tests more complicated analysis tasks involving information spread across different business datasets.

The agents were judged on:

  • Accuracy: Did they produce the correct answer?
  • Efficiency: How many steps and how many words, or “tokens,” did they use?
  • Failure types: What went wrong when they made mistakes?

Testing stored context

The researchers also created extra information for the AI to use. This information described things such as:

  • What tables exist
  • What different columns mean
  • Which datasets are trustworthy
  • How tables are connected
  • What earlier users discovered

This stored information was called persistent semantic context.

An everyday analogy is a student keeping a well-organized notebook about a complicated school project. Instead of rediscovering every fact each time, the student can check the notebook before starting a new task.

4. What did the researchers find?

General coding agents became stronger than specially designed agents

With older or less capable models, a carefully designed data agent sometimes performed better. However, as the models improved, the general coding agents became more accurate and efficient.

The authors suggest that newer models can increasingly perform tasks such as:

  • Planning
  • Writing code
  • Using tools
  • Finding and fixing errors
  • Checking their own answers

Because of this, fixed systems designed by humans may eventually limit the AI rather than help it.

Newer models used fewer steps

The newer coding agents needed fewer turns, or back-and-forth actions, to complete a task.

On the DAB benchmark, the newest model used about four times fewer turns than the older o3 model. It also improved its results by more than 35 percentage points and used fewer tokens.

This is important because fewer steps can mean:

  • Lower cost
  • Faster answers
  • Fewer chances for something to go wrong

The result challenges the idea that future AI agents will always create huge numbers of inefficient database requests.

The biggest remaining problem was understanding the data environment

Even the strongest agents still made many mistakes because they misunderstood the data around them.

More than 60% of the remaining errors were related to environmental knowledge. For example, the agent might:

  • Misunderstand what the question means
  • Choose the wrong table
  • Use the wrong column or measurement
  • Connect two tables using the wrong identifier
  • Misunderstand a company’s special definitions

This is different from simply making a calculation mistake. The agent may know how to calculate correctly but still use the wrong information.

For example, imagine asking:

“How many customers renewed their contracts?”

The agent might calculate perfectly but accidentally use a table about new customers instead of renewing customers. The arithmetic would be correct, but the answer would still be wrong.

Stored context improved performance

Giving the agent useful information in advance helped it perform better.

One accuracy-focused context improved performance by 19 percentage points compared with giving the agent no stored context.

A schema-focused context also helped the agent spend fewer turns exploring the database. In other words, it already knew more about where information was located.

However, there was a trade-off: some contexts became large, expensive, and time-consuming to create. The schema-focused context took more than 3,400 seconds to build and occupied much more space than the other contexts.

The experiments were also fairly small, using only 12 example tasks and 12 datasets. Real companies may have thousands of tables and terabytes of data, so creating and updating the context could be much harder in practice.

5. What does “persistent semantic context” mean?

The term can be understood by separating it into three parts:

  • Persistent: It is saved and can be used again later.
  • Semantic: It describes the meaning of information, not just its location.
  • Context: It gives background knowledge that helps the AI understand a task.

For example, a context file might say:

  • “This table contains official monthly revenue.”
  • “The Finance department uses this definition of profit.”
  • “Customer ID in Table A matches Account Number in Table B.”
  • “For European sales, use the tax-adjusted revenue column.”
  • “This older dataset should not be used for current reports.”

This information could be written as text files, stored in a knowledge graph, or organized in other searchable forms.

6. Why is this important?

The paper’s main message is that better AI models may handle more of the general reasoning and coding work by themselves. Human engineers may not need to build as many rigid step-by-step systems.

But AI models cannot automatically know every organization’s private and changing rules. They may not know:

  • Which dataset is considered official
  • What a company means by “active customer”
  • How departments define important measurements
  • Which tables contain current information
  • How different systems refer to the same person or product

This knowledge is usually outside the model’s original training data. It must be provided separately and kept up to date.

7. What future research does the paper suggest?

The authors argue that future database systems should treat persistent semantic context as an important part of the system, rather than as an ordinary note or prompt.

They identify several challenges:

Keeping information correct

Databases, company rules, and documentation change over time. If the stored context is not updated, the AI could confidently use outdated information.

Researchers need ways to notice when a change in one table or definition affects other stored explanations.

Choosing how widely updates should spread

Some information may belong only to one user. Other information may apply to a whole team or an entire company.

A future system might need different levels of context:

  • Personal context for one user
  • Team context for a project
  • Company-wide context for official rules
  • Public context from regulations or outside sources

Storing information efficiently

Context might include millions of words from tables, documents, previous questions, and agent actions. An AI cannot read all of this every time.

Systems will need smart ways to:

  • Search the most relevant information
  • Store frequently used knowledge
  • Remove repeated information
  • Compress large collections of text
  • Decide what information is no longer useful

Balancing speed, cost, and accuracy

Building a detailed context can take a lot of time and computer power. A short context may be cheap but miss important details. A large context may be more accurate but slow and expensive.

Finding the right balance will be a major research problem.

Conclusion

This paper argues that the future of data agents will not mainly depend on adding more complicated, hand-built instructions. As LLMs improve, general-purpose agents may learn to plan, code, debug, and check their work without as much human-designed scaffolding.

The harder problem is helping agents understand the specific world they are working in. A company’s data has its own names, rules, connections, and history. Saving this knowledge in a reusable form—persistent semantic context—can make agents more accurate and faster.

The possible impact is significant. If researchers can build context systems that are reliable, efficient, and easy to update, people may be able to ask complex questions about business, science, medicine, or government data using ordinary language. However, these systems must be carefully maintained, because outdated or incorrect context could cause an AI to give confident but misleading answers.

Knowledge Gaps

Knowledge Gaps, Limitations, and Open Questions

  • The evaluation covers only two benchmarks, TAG-Bench and DAB, so it remains unclear whether the reported superiority of general coding agents transfers to other domains, database architectures, modalities, enterprise environments, and real-world workloads.
  • The experiments use only three model versions from a single model ecosystem, limiting conclusions about whether the observed trends reflect general model scaling or provider-specific behavior.
  • The study does not report statistical significance, confidence intervals, variance across repeated runs, or sensitivity to stochastic agent behavior, making the robustness of the performance differences uncertain.
  • The comparison with Agentar-Scale-SQL is not fully reproducible because its inference code is unavailable and the authors modify its procedure with a custom refinement step.
  • The human-designed baselines and the general coding agents may not be matched for prompt design, tool access, execution budgets, context windows, or engineering effort, leaving the fairness of the comparison unresolved.
  • The paper evaluates primarily low-reasoning configurations and states that other reasoning levels show consistent trends, but it does not provide detailed results quantifying how reasoning budgets affect accuracy, latency, cost, and context use.
  • The failure taxonomy is assigned by GPT-5.6 Sol, creating potential evaluator bias and uncertainty about the reliability, reproducibility, and inter-rater validity of the root-cause classifications.
  • The paper does not establish whether schema exploration is truly the dominant bottleneck in production data-agent workloads, where network latency, permissions, data quality, query execution, and human approval may contribute substantially to total cost.
  • The proposed notion of persistent semantic context lacks a precise formal definition, including its representation, provenance, dependency structure, validity conditions, and interface with the underlying data system.
  • The experiments construct context from only 12 sample trajectories and 12 datasets, so the reported improvements may not generalize to larger, more heterogeneous, or less representative collections of historical traces.
  • The context experiments do not isolate the effects of additional information from those of prompt length, formatting, model-generated instructions, or accidental leakage of benchmark-specific knowledge.
  • The study does not test whether context generated from historical traces remains effective under distribution shift, newly introduced schemas, unseen business questions, changed data distributions, or unfamiliar organizational conventions.
  • The contexts are appended to the prompt as static text; the paper does not evaluate retrieval, selective context loading, hierarchical memory, structured context, or context-window management strategies.
  • The reported accuracy gains are not evaluated against the full lifecycle cost of context construction, storage, maintenance, retrieval, and invalidation over a realistic multi-query workload.
  • The paper provides no break-even analysis showing how many queries are required for persistent context to offset its offline construction and maintenance costs.
  • The scalability claims for terabyte-scale environments and thousands of tables are speculative; no experiments measure construction time, storage growth, retrieval latency, or update cost as environment size increases.
  • The proposed context-generation methods do not include systematic comparisons with database catalogs, knowledge graphs, vector retrieval, materialized views, metadata systems, or hybrid representations under identical workloads.
  • The tradeoff between context size and task performance is not characterized, including whether larger contexts improve coverage or instead increase distraction, prompt interference, latency, and hallucination.
  • The paper does not define reliable metrics for semantic context quality, such as factuality, coverage, provenance completeness, freshness, contradiction rate, retrieval utility, or downstream task risk.
  • It remains unresolved how an agent or system should detect and correct hallucinated, ambiguous, obsolete, or internally contradictory statements in generated context.
  • The paper does not investigate provenance mechanisms that link each semantic claim to source schemas, documentation, queries, data samples, policies, or human decisions.
  • No experiments evaluate whether agents can distinguish authoritative context from lower-confidence observations, conflicting user preferences, outdated traces, or speculative model-generated summaries.
  • The proposed semantic consistency problem lacks operational guarantees: the paper does not specify how semantic dependencies should be represented, detected, validated, or proven correct after an underlying change.
  • The appropriate consistency model for different applications is left undefined, including how to select among strong, eventual, query-triggered, or multi-version semantic consistency based on risk and workload requirements.
  • The paper does not address conflict resolution when user, team, enterprise, regulatory, and external contexts provide incompatible definitions or recommendations.
  • The effects of concurrent updates, access control, privacy policies, and authorization changes on persistent context are not studied.
  • The paper does not examine how to prevent sensitive information, proprietary business logic, personal data, or secrets from being copied into broadly accessible agent context.
  • No mechanism is proposed for auditing who created, modified, retrieved, or relied upon a semantic context artifact, which is important for regulated or high-impact data-agent applications.
  • The relative merits of incremental maintenance versus complete context regeneration are discussed conceptually but not evaluated experimentally across update frequency, dependency density, environment size, and workload characteristics.
  • The paper does not establish whether semantic dependencies can be represented explicitly enough to support reliable incremental maintenance, especially for implicit relationships encoded in natural language.
  • The impact of stale context on downstream decisions is not quantified, including the severity, detectability, and reversibility of errors caused by outdated semantic knowledge.
  • The paper does not evaluate context versioning, rollback, temporal queries, or reproducibility of agent results under historical semantic states.
  • Compression is identified as an open problem, but no experiments compare textual summarization, retrieval-based pruning, structured compression, learned memory, or lossless representations for persistent context.
  • It remains unknown which information can safely be discarded during context compression and how compression affects rare queries, edge cases, join correctness, and long-horizon task execution.
  • The interaction between persistent context and model pretraining or test-time adaptation is not examined, including whether context duplicates parametric knowledge or creates harmful conflicts with it.
  • The paper does not evaluate smaller, open-weight, fine-tuned, or multimodal models, leaving the accessibility and generality of the proposed research agenda uncertain.
  • The relationship between improved single-query efficiency and total system throughput under concurrent multi-user workloads is not measured.
  • The study does not assess whether persistent context changes agent behavior in ways that reduce exploration too aggressively, causing agents to miss relevant data sources or fail to verify inherited assumptions.
  • The paper does not investigate human oversight requirements, including how users should inspect, edit, approve, correct, or override model-authored semantic context.
  • No deployment study measures long-term context evolution, maintenance burden, failure accumulation, or user trust over extended periods of real usage.
  • The claim that hand-engineered agent layers will be subsumed by increasingly capable models remains a forecast rather than a causally tested conclusion; the paper does not identify which system components are likely to remain valuable under future model and workload changes.

Practical Applications

Immediate Applications

  • Enterprise natural-language data analysis (Software, business intelligence, finance, operations)
    • Potential tools/workflows: conversational BI assistants, automated quarterly business reviews, finance variance analysis, supply-chain investigations, and self-service SQL/Python generation.
    • Dependencies: reliable tool execution, read-only or carefully scoped database access, robust validation, data permissions, and human review for high-impact decisions. Performance may degrade when schemas, metric definitions, or join relationships are undocumented.
  • Offline generation of semantic data catalogs (Data engineering, knowledge management)
    • Potential products: agent-readable data catalogs, semantic schema registries, automated “how to use this dataset” documentation, and context packages attached to data products.
    • Dependencies: representative historical traces, access to authoritative metadata, safeguards against hallucinated or obsolete descriptions, and periodic review by data owners.
  • Context-augmented coding agents for analytics and data operations (Software engineering, data platforms)
    • Potential workflows: automatic query repair, data-pipeline debugging, migration assistance, notebook generation, and documentation-linked code generation.
    • Dependencies: context must be selectively retrieved rather than indiscriminately appended to prompts; otherwise, large or inaccurate context can increase latency, cost, or over-reliance on incorrect guidance.
  • Automated semantic failure diagnosis (Data quality, observability, platform engineering)
    • Potential tools: agent observability dashboards, benchmark suites, automatic trace labeling, root-cause alerts, and targeted regression tests.
    • Dependencies: reliable trace capture and evaluation metrics that distinguish syntactic correctness from semantic correctness.
  • Agent-assisted data governance and metric discovery (Governance, compliance, finance, public administration)
    • Potential workflows: governed self-service analytics, metric certification, audit preparation, and policy-aware report generation.
    • Dependencies: explicit ownership, versioning, access control, provenance, and human approval. The context should not be treated as authoritative unless its source and freshness are verifiable.
  • Academic research assistants for complex data environments (Academia, computational science, social science)
    • Dependencies: provenance preservation, reproducibility requirements, domain-specific validation, and protection of confidential or sensitive research data.
  • Personal and small-business data assistants (Daily life, education, small-business operations)
    • Dependencies: secure local or private-cloud deployment, accurate synchronization, simple correction mechanisms, and clear disclosure that generated answers may be wrong.

Long-Term Applications

  • Persistent semantic context as a native database abstraction (Database systems, cloud infrastructure)
    • Potential products: “context databases,” agent-native warehouses, semantic indexes, and context-aware query engines.
    • Dependencies: scalable representations, workload-aware retrieval, interoperability across database systems, and evidence that maintenance costs are lower than repeated online exploration.
  • Semantic consistency and dependency management (Database research, enterprise governance, regulated sectors)
    • Potential workflows: automatically flagging contexts affected by a renamed field, propagating a new regulatory definition, or preventing an agent from using an obsolete metric.
    • Dependencies: the ability to represent implicit semantic dependencies, define acceptable freshness guarantees, and reconcile conflicting user, team, enterprise, and external contexts.
  • Versioned and scope-aware organizational memory (Enterprise software, collaboration, policy)
    • Potential tools: version-controlled semantic workspaces, policy-aware context registries, and collaborative organizational-memory platforms.
    • Dependencies: identity and authorization systems, conflict-resolution policies, auditability, and controls preventing private information from leaking across scopes.
  • Semantic query optimization and agent planning (Data systems, software infrastructure)
    • Potential products: semantic indexes, context-aware query planners, retrieval-augmented SQL engines, and agent-specific materialized views.
    • Dependencies: accurate workload models, safe integration with query optimizers, and safeguards against recommendations that are semantically plausible but computationally or statistically invalid.
  • Compression and lifecycle management for machine-readable organizational memory (AI infrastructure, storage systems)
    • Potential tools: semantic compilers, context summarizers, long-term memory stores, and hot/cold context tiers.
    • Dependencies: compression must preserve information needed for downstream decisions; lossy summaries require provenance, confidence scores, and mechanisms for recovering the underlying evidence.
  • High-assurance agents for healthcare, finance, energy, and public services (Regulated and safety-critical sectors)
    • Dependencies: strong semantic consistency, traceable evidence, privacy protection, domain validation, formal or procedural verification, and human authorization. The paper’s results alone do not establish suitability for autonomous decisions in these domains.
  • Robotics and embodied systems with persistent environmental knowledge (Robotics, manufacturing, logistics)
    • Dependencies: integration with sensor data and real-time state estimation, safe update mechanisms, uncertainty handling, and guarantees that stale context cannot cause physical harm.
  • Adaptive educational and scientific knowledge environments (Education, research, knowledge management)
    • Dependencies: privacy and consent, protection against reinforcing incorrect knowledge, teacher or researcher oversight, and mechanisms for distinguishing authoritative content from generated suggestions.
  • Policy and regulatory infrastructure for agent-readable data environments (Public policy, standards, compliance)
    • Dependencies: cross-vendor interoperability, measurable definitions of semantic consistency, sector-specific risk thresholds, and legal clarity regarding responsibility for stale or incorrect context.

Glossary

  • Agentic capabilities: Abilities of an AI system to autonomously plan and perform multi-step actions. “agentic capabilities are undergoing a dramatic transformation”
  • Agentic speculation: A workload pattern involving many exploratory or inefficient queries issued by an AI agent. “a workload characterized by sheer scale and inefficiency of agent-issued queries”
  • Amortization: Distribution of a one-time computational or storage cost across multiple later operations. “its associated costs can be amortized across many user queries”
  • Authoritative data source: A data source treated as the official or most trustworthy basis for a fact or measurement. “identifying authoritative datasets across multiple systems”
  • Bitter Lesson: The principle that scalable general-purpose computation and learning tend to outperform manually engineered domain-specific methods. “This trend reflects the Bitter Lesson”
  • Canonical metric: A formally accepted definition of a measurement used consistently across an organization. “canonical metric definitions”
  • Compression: Reduction of the storage or processing requirements of information while attempting to preserve its useful content. “Future systems will require efficient techniques to consolidate accumulated knowledge”
  • Context layer: A system component that stores contextual knowledge for use by AI agents. “persistent semantic contexts introduce significant system overheads”
  • Contextual knowledge: Information about the environment, data, conventions, or task history relevant to an agent’s operation. “Contextual knowledge and understanding over large, complex data environments remains a substantial challenge”
  • Data agent: An AI system designed to answer questions or perform analyses over structured or enterprise data. “future data systems will need to natively serve persistent semantic contexts as a first-class abstraction”
  • Data profiling: Examination of datasets to discover their structure, contents, quality, and statistical properties. “steps spent on schema exploration or data profiling remain a substantial relative cost”
  • Data structure: An organized representation used to store and retrieve information efficiently. “One central open question is: what data structure should represent semantic context?”
  • Declarative: Describing the desired result or behavior without specifying every execution step. “natural language or semantic knowledge contained within the context layer”
  • End-to-end model training: Training a model jointly across the complete processing pipeline rather than engineering separate task-specific components. “general methods that scale computation (e.g., end-to-end model training) ultimately displace hand-engineered domain knowledge”
  • Entity relationship: A semantic association between identifiable objects or records in a data system. “entity relationships”
  • Eventual consistency: A consistency model in which updates may propagate asynchronously but replicas converge over time. “collaborative analytics may tolerate eventual semantic consistency”
  • Execution-guided refinement: Iteratively improving generated queries or programs by using the results of their execution. “We use its published SQL-generation prompt and add a custom execution-guided refinement step”
  • First-class abstraction: A concept directly and explicitly supported by a system’s programming or storage model. “natively serving persistent semantic contexts as a first-class abstraction”
  • Garbage collection: Automatic identification and reclamation of storage occupied by information no longer considered useful. “motivating new garbage collection and compression policies”
  • Generalization: The ability of a model or system to perform successfully on cases beyond those directly used during development. “their performance and generalization will likely surpass that of data agents”
  • Holistic regeneration: Rebuilding an entire representation from its underlying sources instead of updating only changed portions. “increasingly capable long-context models may make holistic regeneration an attractive alternative”
  • Incremental view maintenance: Updating a derived database representation by propagating source changes rather than recomputing it completely. “This resembles incremental view maintenance”
  • Inference: The process of using a trained model to produce outputs for new inputs. “whose full inference code is not publicly available”
  • Join key: An attribute used to match records across tables in a relational operation. “incorrectly using entities, join keys, and identifiers”
  • Knowledge graph: A graph-based representation of entities, relationships, and associated information. “structured knowledge graphs and vector indexes”
  • Long-context model: A model capable of processing unusually large amounts of input context in one interaction. “increasingly capable long-context models”
  • Materialized view: A stored result of a query that can be reused instead of recomputed. “indexes, materialized views, and caches”
  • Multi-agent system: A system in which multiple interacting agents jointly perform a task. “DeepEye is a multi-agent system”
  • Non-parametric knowledge: Explicit knowledge supplied to a model externally rather than encoded in its learned parameters. “an increasingly important requirement for more capable data agents will lie in non-parametric contextual information”
  • Parametric knowledge: Information encoded within a model’s learned parameters during training. “what the model does not acquire during training as parametric knowledge”
  • Persistent semantic context: A durable, language-based representation of knowledge about a data environment that can be reused across tasks. “We refer to this contextual knowledge as persistent semantic context”
  • Physical design: Decisions about how data or metadata are organized and stored to optimize system performance. “The physical organization of semantic context presents new opportunities”
  • Prompt: Input text or instructions provided to a LLM to guide its output. “appending the generated context to the user prompt”
  • Query-agnostic context: Context that is useful across many queries rather than tailored to one specific query. “agents are likely to rely on both query-agnostic context”
  • Query-triggered consistency: A consistency strategy that revalidates stored information only when it is accessed by a query or agent. “systems may adopt query-triggered semantic consistency”
  • Relational database: A database that organizes data into tables connected through defined relationships. “queries that require combining exact computation, semantic reasoning, and world knowledge over relational databases”
  • Semantic consistency: The property that stored semantic knowledge remains correct and current relative to the underlying environment. “ensuring that the natural language or semantic knowledge contained within the context layer is up-to-date and correct”
  • Semantic dependency: A meaning-based relationship in which one piece of contextual knowledge depends on another. “changes may propagate through implicit semantic dependencies rather than explicit relational dependencies”
  • Semantic reasoning: Drawing conclusions based on meaning, concepts, and relationships rather than only literal data operations. “queries that require combining exact computation, semantic reasoning, and world knowledge”
  • Schema exploration: Investigating the structure, fields, relationships, and metadata of available data sources. “the relative cost of schema exploration has increased over successive model generations”
  • Schema knowledge: Information about the organization and attributes of data sources. “capturing schema knowledge to reduce online exploration”
  • Subagent: An auxiliary agent launched by a primary agent to perform part of a larger task. “forking subagents when paired with a terminal, file system and execution loop”
  • Token efficiency: The amount of useful work or accuracy achieved relative to the number of language-model tokens consumed. “general coding agents are rapidly improving both accuracy and token efficiency”
  • Vector index: An index that enables efficient retrieval of items based on similarity between numerical vector representations. “structured knowledge graphs and vector indexes”
  • Workload-aware strategy: A system strategy tailored to the characteristics and access patterns of expected tasks. “developing workload-aware strategies for semantic maintenance”

Tweets

Sign up for free to view the 11 tweets with 170 likes about this paper.