
Episode #74
How to Increase Agent Accuracy & Reduce Cost with Context
What happens when an enterprise data agent asks you: “What is our churn this quarter?” The AI agent has access to the warehouse, the tables, and the query engine. It carries out the necessary calculations. However, the answer it provides turns out to be inaccurate. Why? It’s not because the model is inadequate. It’s because it didn’t know which definition of the term 'churn' the business had in mind, which table was out of date, or which join resulted in customers being counted twice. When you apply this situation to all the questions your enterprise directs at AI, you begin to understand why so many agent projects come to a halt at the demo stage. In this episode of the Don't Panic, It's Just Data podcast , host Shubhangi Dua, Podcast Producer and B2B journalist at EM360Tech , sits down with Suresh Srinivas, the CEO and Co-Founder of Collate and also a Co-Founder of the open-source project OpenMetadata . They talk about the context layer and the reasons why AI agents fail when it comes to enterprise data, even though the underlying models are continually getting better. He explains why the solution must be an open context layer rather than a proprietary one. Also Read: OpenAI's self-service AI data agent, built on OpenMetadata What OpenAI's internal data agent shows about scale Srinivas cites a use case to Dua to explain how the solution can be deployed at scale. He talks about OpenAI's own internal data agent, Kepler , which is used by more than 3,500 people and carries out reasoning across about 70,000 datasets and over 600 petabytes of data, as proof that the issue is not particular to Collate's customers. OpenAI developed a seven-layered context system based on OpenMetadata, and is one of its largest users as well as a major open source contributor. Srinivas maintains that even though there are seven layers, they all come down to three basic elements: the context of the data (what exists), the ontology and semantics (what it means, expressed in business terms such as "revenue" or "churn"), and memory (the corrections and feedback that the agent learns from). Srinivas states that after OpenAI had invested in the context layer, query performance on Kepler decreased from 22 minutes to 90 seconds, representing a roughly 16-fold improvement, together with improvements in token efficiency. Anthropic too has published its own research on the importance of context layers for data agents, which shows that there is convergence among the leading research labs, not merely a point made by Collate. Key Takeaways McKinsey: fewer than 10 per cent of enterprises have scaled AI agents to value 80 per cent of enterprises cite data limitations as the scaling barrier On the Spider 2.0 benchmark, the top models achieved only 59 per cent accuracy when tested on enterprise data. Collate's context layer raised accuracy from 59 per cent to 94 per cent on the same benchmark That's an 86% reduction in wrong answers from context alone, not a bigger model Token spend fell 75 per cent once a context layer was added Database queries per question dropped from 190 to 27, an 86 per cent reduction in compute OpenAI's Kepler agent spans 70,000 datasets and 600+ petabytes, built on OpenMetadata OpenAI's query performance improved from 22 minutes to 90 seconds with context Context layer reduces to three primitives: data context, ontology/semantics, memory OpenMetadata has 15,000+ community members and 4,000 production deployments Collate curates context per persona rather than serving the full context to every agent Collate's AI Governance Studio audits what agents access and the risk they pose Chapters 00:00 Introduction to AI and Data Context Challenges 02:16 The Root Cause of Data Problems: Missing Context 04:20 Evolving Definition of Context in AI and Data 06:52 Why Data Limitations Still Hinder AI Scaling 07:22 The Impact of Context on AI Model Accuracy and Cost 09:21 How Providing Context Improves AI Performance 11:02 Verifying Data Accuracy in a Constantly Changing Enterprise 12:25 AI Agents and Continuous Data Quality Management 14:25 OpenAI's Use of Context Layers for Better AI Performance 16:03 Avoiding Noise and Cost in Context Management 18:47 The Role of Persona-Based Context Curation 19:45 Open Source as a Foundation for Context Layers 21:37 Accountability for Incorrect or Stale Context 22:24 Collaborative Role of Data Teams and AI Agents in Context Management 24:01 Guardrails and Deterministic Answers in Enterprise AI 25:53 Reducing Token Consumption with Context Layers 28:47 Key Takeaway: Building a Robust Context Layer for AI 30:37 Upcoming Industry Discussions and Challenges at Big Data London 32:28 The Future of AI Agents and the Importance of Context in Data Strategy 33:12 Closing Remarks and Next Steps for AI and Data Leaders @CollateData @enterprisemanagement360 #AIagents #EnterpriseAI #DataGovernance #OpenMetadata #ContextLayer #AIaccuracy #AgenticAI #CIO #DataStrategy #EM360Tech #DontPanicItsJustData #opensourcesoftware #opensourceai AI agents, enterprise AI, CIO, CTO, IT leadership, data governance, AI governance, context layer, OpenMetadata, data strategy, agentic AI, digital transformation, tech podcast, enterprise technology






