Evals, error analysis, and better prompts: A systematic approach to improving your AI products | Hamel Husain (ML engineer)
How I AI13 Loka 2025

Evals, error analysis, and better prompts: A systematic approach to improving your AI products | Hamel Husain (ML engineer)

Hamel Husain, an AI consultant and educator, shares his systematic approach to improving AI product quality through error analysis, evaluation frameworks, and prompt engineering. In this episode, he demonstrates how product teams can move beyond “vibe checking” their AI systems to implement data-driven quality improvement processes that identify and fix the most common errors. Using real examples from client work with Nurture Boss (an AI assistant for property managers), Hamel walks through practical techniques that product managers can implement immediately to dramatically improve their AI products.


What you’ll learn:

1. A step-by-step error analysis framework that helps identify and categorize the most common AI failures in your product

2. How to create custom annotation systems that make reviewing AI conversations faster and more insightful

3. Why binary evaluations (pass/fail) are more useful than arbitrary quality scores for measuring AI performance

4. Techniques for validating your LLM judges to ensure they align with human quality expectations

5. A practical approach to prioritizing fixes based on frequency counting rather than intuition

6. Why looking at real user conversations (not just ideal test cases) is critical for understanding AI product failures

7. How to build a comprehensive quality system that spans from manual review to automated evaluation

Brought to you by:

GoFundMe Giving Funds—One account. Zero hassle: https://gofundme.com/howiai

Persona—Trusted identity verification for any use case: https://withpersona.com/lp/howiai

Where to find Hamel Husain:

Website: https://hamel.dev/

Twitter: https://twitter.com/HamelHusain

Course: https://maven.com/parlance-labs/evals

GitHub: https://github.com/hamelsmu

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

In this episode, we cover:

(00:00) Introduction to Hamel Husain

(03:05) The fundamentals: why data analysis is critical for AI products

(06:58) Understanding traces and examining real user interactions

(13:35) Error analysis: a systematic approach to finding AI failures

(17:40) Creating custom annotation systems for faster review

(22:23) The impact of this process

(25:15) Different types of evaluations

(29:30) LLM-as-a-Judge

(33:58) Improving prompts and system instructions

(38:15) Analyzing agent workflows

(40:38) Hamel’s personal AI tools and workflows

(48:02) Lighting round and final thoughts

Tools referenced:

• Claude: https://claude.ai/

• Braintrust: https://www.braintrust.dev/docs/start

• Phoenix: https://phoenix.arize.com/

• AI Studio: https://aistudio.google.com/

• ChatGPT: https://chat.openai.com/

• Gemini: https://gemini.google.com/

Other references:

• Who Validates the Validators? Aligning LLM-Assisted Evaluation of LLM Outputs with Human Preferences: https://dl.acm.org/doi/10.1145/3654777.3676450

• Nurture Boss: https://nurtureboss.io

• Rechat: https://rechat.com/

• Your AI Product Needs Evals: https://hamel.dev/blog/posts/evals/

• A Field Guide to Rapidly Improving AI Products: https://hamel.dev/blog/posts/field-guide/

• Creating a LLM-as-a-Judge That Drives Business Results: https://hamel.dev/blog/posts/llm-judge/

• Lenny’s List on Maven: https://maven.com/lenny

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Jaksot(69)

Zapier’s CEO shares his personal AI stack | Wade Foster

Zapier’s CEO shares his personal AI stack | Wade Foster

Wade Foster is the co-founder and CEO of Zapier. In this episode, Wade shows how he uses meeting transcripts, Zapier agents, and even Grok to analyze company culture, evaluate interview candidates, an...

5 Tammi 41min

How Webflow’s CPO built an AI chief of staff to manage her calendar, prep for meetings, and drive AI adoption | Rachel Wolan

How Webflow’s CPO built an AI chief of staff to manage her calendar, prep for meetings, and drive AI adoption | Rachel Wolan

Rachel Wolan, the chief product officer at Webflow, has embraced AI not just as a product leader but as a hands-on builder. A coder since age 16, Rachel has returned to her technical roots by creating...

29 Joulu 202543min

How to get your whole team excited about AI (and actually using it) | Brian Greenbaum (product designer at Pendo)

How to get your whole team excited about AI (and actually using it) | Brian Greenbaum (product designer at Pendo)

Brian Greenbaum is a Senior Staff Product Designer at Pendo who led a company-wide AI transformation after a personal epiphany while on paternity leave. After experiencing the power of AI coding tools...

22 Joulu 202547min

How Zapier’s EA built an army of AI interns to automate meeting prep, strengthen team culture, and scale internal alignment | Cortney Hickey

How Zapier’s EA built an army of AI interns to automate meeting prep, strengthen team culture, and scale internal alignment | Cortney Hickey

Cortney Hickey is the executive assistant to the CEO at Zapier, where she’s leveraging AI to transform traditional EA responsibilities into scalable, organization-wide systems. In this episode, she de...

15 Joulu 202544min

ChatGPT agent mode: The “little helper” that transformed recruiting, crafted user personas, and solved parking nightmares | Michal Peled (Honeybook)

ChatGPT agent mode: The “little helper” that transformed recruiting, crafted user personas, and solved parking nightmares | Michal Peled (Honeybook)

Michal Peled is a Technical Operations Engineer at HoneyBook who specializes in building internal tools and automations that eliminate friction for teams. In this episode, Michal demonstrates three pr...

8 Joulu 202558min

Gemini 3 vs. Claude Opus 4.5 vs. GPT-5.1 Codex: Which AI model is the best designer?

Gemini 3 vs. Claude Opus 4.5 vs. GPT-5.1 Codex: Which AI model is the best designer?

I put three cutting-edge AI models to the test in a head-to-head design competition. Using the exact same prompt, I challenged Google’s Gemini 3, Anthropic’s Opus 4.5, and OpenAI’s Codex 5.1 to redesi...

3 Joulu 202525min

“PMs who use AI will replace those who don’t”: Google’s AI product lead on the new PM toolkit | Marily Nika

“PMs who use AI will replace those who don’t”: Google’s AI product lead on the new PM toolkit | Marily Nika

Marily Nika, AI Product Lead at Google and founder of the AI Product Academy, demonstrates how product managers can leverage AI tools to dramatically accelerate their workflow. Using a smart-fridge co...

1 Joulu 202540min

How to create your own AI performance coach: Optimizing your unique nutrition, recovery, and injury management needs | Lucas Werthein (Cactus)

How to create your own AI performance coach: Optimizing your unique nutrition, recovery, and injury management needs | Lucas Werthein (Cactus)

Lucas Werthein, the COO and co-founder of Cactus, shares how he built a personalized AI wellness coach using ChatGPT to optimize his athletic performance while managing past injuries. After multiple s...

24 Marras 202551min