September 21, 2026 in data science
The Harness Is the Bottleneck
Engineering Enterprise Data for AI Agents
SHARE: PRINT ARTICLE:
https://doi.org/10.1287/LYTX.2026.03.13
An agent harness is everything that sits between a prompt and production. It covers the orchestration loop, the tool and retrieval interfaces, context assembly and memory, the guardrails on action, and the observability that makes behavior auditable. Most organizations treat the harness as plumbing and the model as the source of capability. Published evaluation data points to the opposite allocation of effort.
Gartner has forecast that more than 40% of agentic AI projects will be canceled by the end of 2027, attributing this to escalating costs, unclear business value, and inadequate risk controls.1 Model capability does not appear on that list. The binding constraint sits in the layer between the model and the enterprise data estate, and that layer is a data engineering problem.
The Evidence Is on the Leaderboard
Spider 2.0 is the most useful public probe of this claim. It comprises 632 text-to-SQL workflow problems drawn from production enterprise systems, with databases that often exceed 1,000 columns across several SQL dialects.2 On release, GPT-4o solved 10.1% of these tasks against 86.6% on the far simpler Spider 1.0 schemas. The model was identical. The data environment was not.
What has happened there since is more instructive. Figure 1 shows published execution accuracy on the fixed 547-task Spider 2.0-Snow split for five base models, each evaluated under several independently built harnesses. Holding the model constant, the harness alone accounts for swings of roughly 8 to 52 percentage points. The open-weights model gpt-oss-120B reaches 65.5 under its strongest harness, more than double the best published result for o1-preview on the same tasks.

Figure 1. Weakest and strongest published harness for each fixed base model on the 547-task Spider 2.0-Snow split. The model is constant within each row, so the span is attributable to the harness and the data context it supplies. Source: Spider 2.0 leaderboard, July 2026.
Three Mechanisms Explain the Spread
First, a schema is not a semantic model. A benchmark by researchers at data.world measured GPT-4 zero-shot question answering against an enterprise insurance schema at 16.7% accuracy. Posing the same questions over a knowledge graph representation of that database lifted accuracy to 54.2%, and adding ontology-based query checking with automated repair lifted it to 72.55%.3 The model was untouched at every step. What changed was whether join paths, entity identity, and metric definitions were declared, rather than inferred.
Second, context is a budget rather than a container. Chroma's evaluation of 18 frontier models found that accuracy degrades non-uniformly as input length grows, often well before the advertised window is full.4 Serializing a 3,000-column DDL into a prompt is therefore not a neutral act, because it purchases recall at the cost of attention.
Third, reliability compounds downward. The tau-bench authors introduced pass^k, the probability that an agent solves the same task on all k independent trials. A GPT-4o function-calling agent with better than 60% single-trial success fell below 25% at k=8.5 Enterprise agents run the same task thousands of times, which makes single-trial accuracy the wrong acceptance criterion.
A Reference Architecture
Figure 2 presents a reference architecture that separates the concerns. The harness owns planning, routing, and action guardrails. A context interface plane exposes governed data through narrow typed entry points. A governed metadata plane, versioned in the same repository as transformation code, holds semantic models, catalog entries, lineage, contracts, and access policy. The physical plane keeps computation where the data already lives.
A request resolves downward through those planes. The router selects a narrow typed interface instead of a raw connection, the semantic layer resolves the metric deterministically, and the policy plane applies row and column filters under the end user's propagated identity rather than a shared service account. The agent never loads schema it did not need, and it never sees a row the requester could not have queried unaided. Failure then surfaces as an explicit error rather than a plausible wrong number, which is what matters when output reaches an auditor.

Figure 2. Reference architecture for a data-grounded agent harness. Semantics, lineage, and policy resolve beneath the model, keeping context small and the model replaceable. The right-hand loop is the DataOps control cycle in which agent traces become evidence that promotes or blocks a semantic model change.
Four Moves That Pay for Themselves
Build a thin semantic contract rather than boiling the ocean. The company dbt Labs re-ran a 2023 benchmark in 2026 against current frontier models. Unmodeled text-to-SQL reached 64.5%, while a semantic layer over the same warehouse was correct 100% of the time for questions inside its coverage. Adding three models to close the gaps raised text-to-SQL to between 84% and 90%, and the semantic layer to between 98% and 100%.6 That study also found model tier and reasoning effort changed results very little, while modeling quality changed them a great deal. Start with metrics carrying financial or regulatory consequence, then version them alongside the pipelines that produce them.
Treat the catalog as a retrieval index rather than a wiki. Column descriptions, glossary terms, sample values, and certification status should be artifacts the agent retrieves on demand – the way a large tool registry is searched rather than loaded in full. This is progressive disclosure applied to metadata, and it keeps schema context proportional to the question.
Emit lineage as a runtime signal. Column-level lineage captured through a vendor-neutral standard, such as the OpenLineage specification governed by the LF AI and Data Foundation, serves two purposes at inference time. It lets the harness annotate an answer with the provenance and freshness of every field involved, and it gives the platform team a blast-radius query the moment an upstream contract breaks. Provenance is also what makes an answer defensible after the fact: when a number reaches an auditor or a customer, lineage is the difference between showing how it was derived and asking them to trust it. Captured as runtime events rather than parsed from static SQL, that same lineage records what the pipeline actually did on each run, so a definition that drifts becomes visible as a change in the graph rather than a silent divergence downstream.
Push computation down rather than through. Anthropic reported that presenting tool servers as code APIs inside a sandbox – so the agent filters and aggregates before anything returns to the model – cut one workflow from roughly 150,000 tokens to 2,000.7 Aggregate inside Snowflake, Amazon Redshift, or Databricks and return only the result set. Run evaluation sweeps and offline scoring as batch jobs on managed compute such as Amazon SageMaker, where they are cheap enough to run on every merge.
The Payoff is Optionality
Once semantics, policy, and lineage resolve below the model, the model becomes a swappable component rather than an architectural commitment. Deterministic, high-volume steps can route to smaller or open-weights models from the Qwen, DeepSeek, and gpt-oss families, while ambiguous planning and recovery stay with frontier models from Anthropic, OpenAI, Kimi, GLM, Google, and their peers. The dbt result implies exactly this, since two frontier models produced identical accuracy once the semantic layer did the work.
Instrument accordingly. Track pass^k rather than single-trial accuracy, since an agent that succeeds 90% of the time on one attempt but has to be right across every run in a batch is a different reliability story than the single-trial number suggests. Track tokens per resolved task rather than tokens per call, because a harness that answers in one clean pass is often cheaper than one that loops through many small calls, even when each call looks inexpensive. And track the share of agent queries resolving through governed semantic definitions rather than improvised SQL. That last measure is the honest indicator of how much of an agent program rests on engineering and how much still rests on luck.
References
1. Gartner, 2025, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027,” June 25, https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027.
2. Fangyu, L., et al., 2025, “Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows,” ICLR 2025 (Oral), arxiv.org/abs/2411.07763; leaderboard at spider2-sql.github.io.
3. Sequeda, J. F., Allemang, D., Jacob, B., 2024, “A Benchmark to Understand the Role of Knowledge Graphs on LLM Accuracy for Question Answering on Enterprise SQL Databases,” GRADES-NDA ’24, ACM, doi.org/10.1145/3661304.3661901; Allemang, D., Sequeda, J., 2024, “Increasing the LLM Accuracy for Question Answering: Ontologies to the Rescue!” arxiv.org/abs/2405.11706.
4. Kelly, H., Troynikov, A., Huber, J., 2025, “Context Rot: How Increasing Input Tokens Impacts LLM Performance,” Chroma Technical Report, July, research.trychroma.com/context-rot.
5. Shunyu, Y., Shinn, N., Razavi, P., Narasimhan, K., 2024, “τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains,” arxiv.org/abs/2406.12045.
6. Ganz, J., Perigaud, B., 2026, “Semantic Layer vs. Text-to-SQL: 2026 Benchmark Update,” dbt Developer Blog, April 7, docs.getdbt.com/blog; code at github.com/dbt-labs/dbt-llm-sl-bench.
7. Jones, A., Kelly, C., 2025, “Code Execution With MCP: Building More Efficient Agents,” Anthropic Engineering, https://www.anthropic.com/engineering/code-execution-with-mcp.
Anjanava Biswas is a Principal Architect in Applied AI at ServiceNow Inc. He has more than 19 years of experience across enterprise architecture, cloud systems, applied machine learning, and data engineering, including six years as a Senior Generative AI Specialist Architect at AWS. He co-authored Building Agentic AI Systems, is an IEEE member, and is a fellow of the IET (UK), BCS (UK), and IETE (India).