September 3, 2026 in Agentic AI

The Missing Triad

Science, Service, and Metrics in Durable Generative and Agentic AI Systems

SHARE: PRINT ARTICLE:print this page https://doi.org/10.1287/orms.2026.03.06

robot with gears coming off of head

Most organizations miss the essence of GenAI systems. They continue to treat GenAI as a conventional IT product that is designed, built, tested, deployed, and maintained through operational models. But the real advantage of GenAI systems is that they are science-enabled, continuously adaptive, behaviorally complex service systems. Treating them otherwise produces a predictable outcome: stalled pilots, production failures, and low adoption rates. As organizations shift from single-turn generative tools toward autonomous, multi-step agentic AI, the same immaturity is now compounding at the orchestration layer. 

The Scope of the Problem 

Failure rates for AI projects are still very high. Gartner reports that only 7% of CFOs see meaningful return from AI in finance.1 A recent MIT study found that 95% of GenAI pilots yield no measurable value and never scale into production.2 This trend persists despite record investment. The pattern is intensifying for agentic AI; Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear value, and inadequate risk controls.

Traditional explanations – unclear objectives, weak data pipelines, culture gaps, change-management challenges, fragmented systems, poor planning – are common stumbling blocks across IT products generally, not unique to AI. What differentiates failing GenAI systems is the absence of scientific rigor and adaptive service design.

Without scientific foundations, organizations build fragile systems that cannot evolve, self-correct, or survive changing conditions.

Science-Enabled Products Behave Differently 

A traditional IT system encodes human-designed logic; its behavior is stable unless the code changes. A science-enabled AI system behaves differently; it encodes patterns, relationships, and evolving statistical beliefs derived from data. These beliefs shift as the environment shifts, so validity depends on continuous experimentation, measurement, recalibration, and evidence-gathering. Such systems must sense, learn, coordinate activity, and reconfigure resources into new competencies. They must shift from “build once, deploy many” to “deploy, motor, learn, and evolve,” something that many organizations fail to do. 

“Gartner predicts that more than 40% of agentic AI projects will be canceled by the end of 2027. 

This shift transforms the entire life cycle. Data are not static input, but an ongoing stream of observations, and assumptions require regular testing rather than one-time validation. Models function as living scientific constructs rather than monolithic artifacts, and deployment marks an entry into continuous evolution rather than a final milestone. For agentic AI, this shift extends further: a decision is not a single inference but a trajectory of reasoning, tool calls, and hand-offs that must itself be observed and validated. 

The lesson is clear: if the scientific life cycle is not embedded into engineering processes, the product life cycle collapses.

The Service Dimension: Where Products Break 

GenAI is not simply a product; it’s a service that interacts with people, context, tools, and evolving behavioral patterns, and its health depends on several interlocking capabilities. These include value co-creation with the user, real-time data ingestion, and ongoing labeling and feedback, which occur alongside continuous and iterative model retraining and effective drift detection and mitigation. Human-centered validation loops and rigorous safety, trust, and interpretability evaluation are equally essential, as are rapid operational escalation processes grounded in validations and metrics. 

Systems lacking these operational muscles degrade predictably: performance drops, errors multiply, trust declines, and adoption collapses.

The Metric Framework: Without Measurement, There Is No Science

Effective GenAI requires a measurement architecture connecting scientific integrity, operational reliability, customer experience, and business value. Traditional IT metrics – throughput, latency, and uptime – are insufficient. GenAI requires multidimensional measurement. As GenAI evolves into agentic AI – planning, calling tools, and handing work to other agents –the framework must extend further still, from grading a single output to grading the full decision trajectory that produced it.

Scientific Accuracy and Product Launch Metrics

These metrics ensure the system’s core intelligence remains valid. 

Science model functionality confirms that no constraints are being violated, such as under-bias or over-bias, among other considerations. Output accuracy is captured through prediction accuracy metrics, such as the F1 score, which measures shared-word overlap with ground truth, and BLEU, which assesses how closely generated text matches a reference, along with other context-dependent measures.
  
Model reproducibility reflects the system’s ability to replicate results under equivalent conditions, whereas uncertainty estimation and calibration assess the quality of the model’s confidence scores. Data completeness, freshness, and representativeness remain essential to avoiding bias and drift. Risk and safety metrics address hateful or unfair content, sexual or violent content, self-harm-related content, direct and indirect jailbreak attempts, and protected material.  

Data are not static input, but an ongoing stream of observations, and assumptions require regular testing rather than one-time validation.” 

Overall quality metrics capture coherence, fluency, groundedness, relevance, and similarity, whereas experimentation throughput tracks the number and quality of A/B tests or model versions. For agentic systems, trajectory and process metrics evaluate whether the sequence of reasoning steps and tool calls was sound, not merely whether the final answer was correct. 

Without these metrics, the scientific core decays. 

Service, Adoption, Value Co-Creation, and Customer Experience Metrics 

These metrics reflect user impact and real-world performance.  

Value metrics capture the gains created and pains relieved for the user, such as time or cost saved and the quality of experience, whereas adoption rate and feature usage track how many users engage with the capability. Task success rate measures the proportion of user tasks completed successfully via the system, or “out-tasking.” User trust and interpretability reflect how users perceive and understand system output.  

Long-term retention or behavioral change serves as an indicator of sustained engagement, whereas feedback-to-deployment latency captures how long it takes to incorporate user feedback or new data into the model. Finally, model degradation rate tracks how quickly accuracy or trust decline without intervention. 

Without these, the system becomes irrelevant.

Complex Adaptive Service System Metrics

Measuring complex adaptive service systems requires traditional performance metrics alongside specialized metrics capturing the system’s ability to adapt, self-organize, and respond to change, spanning resilience and adaptivity, operational performance, and emergent behavior.

Resilience and Adaptivity Metrics

Given the “complex adaptive” nature of these systems, metrics measuring disruption-handling, adaptation, and recovery are essential. Robustness reflects the system’s ability to maintain function under expected conditions and disruptions, whereas recovery rate, or rapidity, captures the speed of return to a stable or new operational state after a performance dip. Recovery level, or adaptive capacity, describes the steady-state performance achieved after disruption, which may exceed the prior level in a phenomenon known as antifragility. Settling time measures how long it takes for performance to stabilize after an adaptation or load change. 

The frequency and volatility ratio of adaptive strategies indicates how often, and how many, unique, adaptive strategies are used, reflecting both responsiveness and diversity, whereas adaptive cost captures the resources required to implement and maintain adaptivity, such as the minimum structural adaptive cost.  

Mean time to repair rounds out this set, measuring the average time needed to fix a failure and restore operational status.

Operational Performance Metrics

These metrics assess process efficiency and effectiveness. 

Response time, or average wait time, measures how long it takes to acknowledge or begin processing a request, whereas resolution time captures the average time needed to fully resolve a request. Throughput, or processing rate and service volume, reflects the total requests handled within a given timeframe, and error rate tracks the frequency of failures or incorrect outcomes, serving as a key indicator of application health.  

Resource utilization rate measures CPU, memory, or personnel usage to aid capacity planning, whereas first contact resolution rate captures the percentage of issues resolved on the first interaction. Availability measures the percentage of time the system is operational and accessible. 

Adaptation-Specific and Emergent Behavior Metrics 

These metrics capture the system’s ability to self-organize, learn, and adapt. The Adaptivity Distribution Index measures how widely adaptive features are distributed across components, whereas the frequency of adaptive strategies tracks how often an adaptive strategy executes over time.  

The Volatility Ratio reflects the number of unique strategies used, indicating the breadth of adaptation. The Domain-Specific Adaptivity Index captures the share of application-specific factors influencing adaptation.  

Service thinking reframes the AI capability as an evolving offering rather than a delivered product.” 

Customer satisfaction ratings offer a survey-based measure of user sentiment, whereas the Customer Effort Score reflects the ease of use or issue resolution, with lower scores indicating a better experience. Perceived usefulness and appropriateness of adaptations provide a qualitative assessment of the system’s adaptive behavior from the user’s perspective. 

For complex adaptive service systems, evaluation should prioritize flexibility over rigid, fixed metrics and redefine success as the capacity for learning and adaptation amid changing conditions. A mixed-methods approach – quantitative and qualitative – is often necessary to capture emergent behavior. These measures determine whether a system evolves or collapses.

Agent Behavior Tracing: One More Metric Family for Agentic Systems

As GenAI matures into agentic AI – systems that plan, choose tools, call APIs, and hand off work to other agents – an additional metric family emerges: agent behavior tracing and trajectory observability. Visibility into how an agent reaches an answer, including its reasoning loops, tool selection, memory look-ups, and state transitions, requires stateful orchestration frameworks paired with specialized agent-observability platforms.

The orchestration framework dictates how a trace is structured. The observability engine visualizes the resulting trajectory, which is the chronological tree of intermediate actions. Frameworks such as LangGraph, Microsoft’s Agent Framework, the OpenAI and Claude Agent SDKs, CrewAI, and PydanticAI/LlamaIndex Workflows each force the agent to emit a distinct span per micro-decision rather than running as a black-box script, recording node-level routing, hand-off telemetry, tool arguments, delegation, and schema changes, respectively. 

A complementary layer of pure-play observability platforms (LangSmith, Arize Phoenix, Galileo, Langfuse, and AgentOps) makes traces reviewable, replaying prompts and tool calls and surfacing failures such as infinite loops, role drift, or mid-workflow hallucination rather than only final-answer error.6 

Two practices are converging into a de facto standard for evaluating agentic systems in production:

  • Golden Trajectories and Tool-Contract Tests: A small set of critical paths (typically five to ten) is asserted against in full – tools called, order, retrieved context, and final state – rather than grading only the final answer. A correct answer via the wrong tool call counts as a latent incident. Each tool or MCP end point is treated as an external dependency with its own contract; stale data, malformed responses, and timeouts are injected deliberately to confirm the agent degrades safely.7
  • Invariant Checks: These are deterministic guardrails, including price floors and ceilings, regulatory and tax constraints, and safety boundaries, that the agent can never violate regardless of its reasoning, wrapped around the non-deterministic core.8

This argues for a layered evaluation harness rather than a single gate: component-level, trajectory, counterfactual/backtested, and invariant checks, followed by bounded online deployment with monitoring. Aggregate A/B or offline accuracy alone is insufficient for agentic systems, because those methods assume a fixed answer key and independent, stationary units – assumptions multi-step agents routinely violate – and can mask a rare catastrophic action occurring in a small fraction of runs. 

At minimum, every GenAI initiative that involves agentic behavior should log four trace layers: 

  • Cognitive/Planning: the system instructions and chain-of-thought behind each decision
  • Tool Execution: the exact arguments sent to each tool and the raw response received
  • State and Memory Transitions: what changed in working memory or conversation state
  • Hand-off Logistics: which agent or sub-agent received control, and why 

Without these, an organization can confirm that an agentic system produced a correct answer once, without ever knowing whether it will do so reliably, safely, or for defensible reasons the next time. By tracking across these dimensions, an organization gains visibility into where the system is succeeding, and where it isn’t. These metrics connect business outcomes (customer experience and product adoption) with scientific foundations (model accuracy and data lifecycle) and operational health (service performance). After two decades studying adaptive systems across industries, it is clear that the enterprises that succeed measure adaptation as seriously as they measure accuracy.9

The Bottom Line

GenAI projects fail not because the technology is inadequate, but because the supporting ecosystem is immature. A successful initiative requires a science-enabled product architecture, a service-oriented operational model, and a metric framework aligning scientific validity, product adoption, and operational reliability. Where autonomous agents are involved, that framework is incomplete without trace-level behavior observability.

The shift from traditional IT to science-enabled products requires a hybrid mindset combining product management, service thinking, operational rigor, and scientific method. Service thinking reframes the AI capability as an evolving offering rather than a delivered product. Success is measured by sustained improvement, value co-creation, and user trust rather than initial performance alone. A culture of continuous measurement is essential. Without metrics tied to scientific accuracy, data health, and service performance, organizations lose visibility into decline. 

In essence, the science life cycle becomes the product life cycle, and the service loop becomes the innovation loop. For agentic systems, the trace becomes the audit trail. By embracing this triad of science, service, and metrics, organizations can move generative AI from fascination to transformation.

 

References 

1.  Kiel, M. 2025, “5 Steps CFOs Can Take to Maximize ROI From AI Initiatives,” Gartner, https://www.gartner.com/en/articles/how-cfos-can-maximize-roi-from-ai

2. Challapally, A., Pease, C., Raskar, R., and Chari, P., 2025, “The GenAI Divide: State of AI in Business 2025,” https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf

3. Gartner, Inc., 2025, “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027,” https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027.

4. Demirkan, H., and Dal, B., 2016, “Product or Service? Get smart: Digital business innovation,” Analytics, INFORMS, https://pubsonline.informs.org/do/10.1287/LYTX.2016.01.02/full/

5. Future AGI 2026, “The Definitive Guide to AI Agent Evaluation,” https://futureagi.substack.com/p/the-definitive-guide-to-ai-agent.  

6. Vongthongsri, K, 2026, “LLM Agent Evaluation Metrics in 2026,” Confident AI, https://www.confident-ai.com/blog/llm-agent-evaluation-complete-guide

7. Braintrust, 2026, “Agent Observability: The Complete Guide for 2026,” https://www.braintrust.dev/articles/agent-observability-complete-guide-2026.  

8. Kinzer, K., 2026, "AI Observability Tools," Comet, https://www.comet.com/site/blog/ai-observability-tools/

9. Demirkan, H., and Dal, B. 2014, “Why do so many analytics projects fail?” Analytics, INFORMS,  https://pubsonline.informs.org/do/10.1287/LYTX.2014.04.02/full/

Haluk Demirkan
([email protected])

SHARE:

Keywords:
INFORMS site uses cookies to store information on your computer. Some are essential to make our site work; Others help us improve the user experience. By using this site, you consent to the placement of these cookies. Please read our Privacy Statement to learn more.