Why Enterprise AI Agents Fail Beyond the Model With Databricks

Why Enterprise AI Agents Fail Beyond the Model With Databricks

Why do businesses replace the AI model when the failure may have started somewhere else entirely?

In this episode of Tech Talks Daily, I speak with Richard Shaw, Technology General Manager for Databricks in the UK and Ireland. Richard leads the field engineering organization that works closely with customers on data and AI problems, giving him a practical view of what happens when promising agentic AI projects meet production workloads.

Richard argues that the model often receives the blame because it is the most visible part of the system. The actual fault may come from stale data, missing business context, inconsistent permissions, an unsuccessful tool call, or another point in the workflow. Replacing the model before tracing the request from start to finish can recreate the same problem in a new place. This is why lineage, end-to-end tracing, and continuous evaluation matter once an agent moves beyond a controlled pilot.

We discuss what a production-readiness rehearsal should include. Richard recommends realistic data, realistic user volumes, unauthorized requests, ambiguous questions, failed tool calls, and tests of what the agent should refuse to do. Teams also need agreed standards for quality, security, cost, and auditability, along with a clear decision about which actions an agent can complete independently and where a person must review or approve the result.

The conversation also looks at model choice and infrastructure cost. Richard believes the strongest test is performance on the organization's actual work rather than a benchmark leaderboard. A frontier model may suit complex reasoning, while a smaller or open-weight model may perform routine extraction or classification at a lower cost. Access policies, observability, and spend controls need to remain consistent as those model choices change.

Conversational analytics creates another governance challenge. Databricks customers such as Virgin Atlantic and Repsol are using natural-language tools to make company data easier for employees to question. Richard says wider access should preserve existing permissions, ownership, definitions, and lineage. An answer becomes far more useful when the user can see where it came from and which team owns the information behind it.

We also cover the boundary between historical analytical data and fast operational workloads. Richard describes how Databricks positions the lakehouse for broad enterprise context and Lakebase for immediate reads and writes, such as updating an account, placing an order, or storing agent memory, while keeping both connected to a common data and governance base.

Are companies ready to trace and test the whole AI workflow, or are too many treating the model as both the hero and the culprit? Listen to the episode and share your thoughts with me.

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(2000)

Turning AI Investment Into Measurable Business Value With HP

Turning AI Investment Into Measurable Business Value With HP

How can businesses turn growing investment in AI infrastructure, cloud capacity and devices into outcomes that employees, customers and finance teams can actually measure? In this episode of Tech Talk...

15 Sep 29min

Turning Two Hours of Document Work Into Eight Minutes With Templafy

Turning Two Hours of Document Work Into Eight Minutes With Templafy

What could your team achieve if a document that previously took two hours could be created in eight minutes? In this episode of Tech Talks Daily, I speak with Oskar Konstantyner, Chief Product Officer...

14 Sep 32min

When AI Gets Your Brand Wrong With Bluefish AI

When AI Gets Your Brand Wrong With Bluefish AI

What happens when an AI system gives a customer the wrong price, misrepresents a product, or recommends a competitor using outdated information? In this episode of Tech Talks Daily, I speak with Alex ...

13 Sep 32min

What Does AI Say About Your Brand With DOM

What Does AI Say About Your Brand With DOM

What happens when an AI assistant becomes the first place a potential customer learns about your company, but the answer it provides is inaccurate, outdated, or influenced by a competitor? In this epi...

12 Sep 23min

Modernizing the Power Grid for AI and Rising Demand With SAP

Modernizing the Power Grid for AI and Rising Demand With SAP

Can electricity grids built for an earlier era support AI data centers, expanding manufacturing, electric vehicles, severe weather, and rising customer expectations at the same time? In this episode o...

11 Sep 27min

Turning AI Agents Into Revenue Workflows With Outreach

Turning AI Agents Into Revenue Workflows With Outreach

What happens when an AI system moves beyond recommending the next sales action and begins running a connected revenue workflow? In this episode of Tech Talks Daily, I speak with Abhijit Mitra, CEO of ...

10 Sep 24min

Building a Faster Specialty Insurance Market With Accelerant

Building a Faster Specialty Insurance Market With Accelerant

What would happen if specialty insurance underwriters and risk capital providers could work from the same timely, detailed information? In this episode of Tech Talks Daily, I speak with Jeff Radke, CE...

9 Sep 28min

Populært innen Politikk og nyheter

giver-og-gjengen-vg
aftenpodden
bt-dokumentar-2
aftenpodden-usa
forklart
stopp-verden
fotballpodden-2
popradet
nokon-ma-ga
det-store-bildet
rss-espen-lee-usensurert
hanna-de-heldige
rss-gukild-johaug
dine-penger-pengeradet
rss-ness
aftenbla-bla
unitedno
saken
frokostshowet-pa-p5
lydartikler-fra-aftenposten