Testing AI: Engineering Confidence in Non-Deterministic Systems with Jason Arbon

Testing AI: Engineering Confidence in Non-Deterministic Systems with Jason Arbon

In this episode 600 of the TestGuild Automation Podcast, Joe Colantonio talks with Jason Arbon, founder of Testers.ai, Jank.AI and IcebergQA and author of the new book Testing AI: Engineering Confidence in Non-Deterministic Systems.

Take Our 2027 Survey Now: https://testgld.link/27data

Jason makes a case most testers have not heard yet. Coding is being absorbed by AI. Specification work is thinning out. Product, development, and test roles are converging into one. And when the music stops, the only seat left belongs to the person who can look at what the machine produced and make an evidence backed call on whether it ships. He calls that confidence engineering, and he argues it is not a rebrand of QA. It is what QA was always supposed to be.

Along the way, Joe and Jason get into the containment problem and why alignment, not lockdown, is now the real safety goal. They dig into why testing cost scales quadratically, meaning ten times more generated code creates roughly a hundred times more testing demand. Jason also pushes back hard on skeptics of agentic testing, pointing out that almost nobody has run the obvious experiment of testing a site themselves for a week and comparing their results against what AI finds.

You will also hear Jason's most practical piece of advice in the whole conversation. If you are not running the same suite five times against the same build and looking at the actual results, not just flake, you are not testing seriously in an AI world.

Plus a detour into grokking, the Chinese Room, and Geoffrey Hinton, because it would not be a Jason Arbon episode without one.

Listen up!

Denne episoden er hentet fra en åpen RSS-feed og er ikke publisert av Podme. Den kan derfor inneholde annonser.

Episoder(605)

Playwright With AI: How to Automate Tests Without Shipping AI Slop with Andrew Knight

Playwright With AI: How to Automate Tests Without Shipping AI Slop with Andrew Knight

Your developers just got supercharged by AI coding agents. Your test coverage did not. So how do you keep quality high when product code is shipping faster than any test team can follow? In this episo...

4 Aug 38min

Test Automation Won't Save Your QA Career, but These Skills Will with Keith Klain

Test Automation Won't Save Your QA Career, but These Skills Will with Keith Klain

Keith Klain has spent 25+ years leading enterprise quality programs in financial services and is one of the testing industry's most respected voices. In this episode, he joins Joe to discuss his Brows...

28 Jul 45min

AI Testing Strategy: Stop Being a Cost Center, Start Protecting Revenue with Nandini Srinivasan

AI Testing Strategy: Stop Being a Cost Center, Start Protecting Revenue with Nandini Srinivasan

Nandini Srinivasan has spent 25 years in the quality industry and leads a global QA organization of over 150 engineers across the US, Canada, India, and Pakistan. In this episode, she breaks down exac...

22 Jul 36min

How to Move from Prompt Engineering to Harness Engineering in Testing with Matt Wynne

How to Move from Prompt Engineering to Harness Engineering in Testing with Matt Wynne

Matt Wynne, co-creator of Cucumber and BDD practitioner, joins Joe for the first time in over a decade to talk about what two years inside a Silicon Valley AI startup taught him about the future of so...

14 Jul 39min

Agentic Engineering for Testers: How to Automate Your Way to the Top with Amit Rawat

Agentic Engineering for Testers: How to Automate Your Way to the Top with Amit Rawat

Amit Rawat is an agentic engineer who spent two decades in QA before shifting fully into building AI agents. He's the creator of PromptWright, a desktop tool that turns natural language prompts into a...

7 Jul 43min

How to Test Any API Without Documentation with Liudas Jankauskas

How to Test Any API Without Documentation with Liudas Jankauskas

Most API testing stops at the happy path. The problem is that the bugs that actually hurt you in production are sitting in everything many testers skip, like the boundary values, the oversized payload...

30 Jun 23min

Your AI Code Review Is Lying to You (Here's the Fix) with Evan Marshall

Your AI Code Review Is Lying to You (Here's the Fix) with Evan Marshall

Your AI code review tools read the diff. They stare at your code. But they never actually run it. So the bugs that only show up at runtime, the broken user flows, the bad query plan, the duplicate sub...

23 Jun 36min

Populært innen Fakta

fastlegen
dine-penger-pengeradet
relasjonspodden-med-dora-thorhallsdottir-kjersti-idem
foreldreradet
treningspodden
jakt-og-fiskepodden
rss-kunsten-a-leve
mikkels-paskenotter
rss-strid-de-norske-borgerkrigene
hverdagspsyken
sinnsyn
fryktlos
gravid-uke-for-uke
rss-sarbar-med-lotte-erik
rss-var-forste-kaffe
kvallm
tomprat-med-gunnar-tjomlid
rss-impressions-2
rss-bak-luftfarten
laringsmiljo-i-skole-og-barnehage-uis-podkast