Claude Opus 5 review: this model is brilliant (but annoying)
How I AI24 Heinä

Claude Opus 5 review: this model is brilliant (but annoying)

I’m tired of new models. Every week there’s a new benchmark, a new frontier intelligence claim, a new thing to test. But here we are, because Opus 5 just dropped and I’ve had real hands-on time with it, so you’re getting the honest version.


This is my full Opus 5 review: personality analysis, live benchmark results from my 7-model How I AI eval, and an actual verdict on whether I’m swapping it in. Spoiler: the answer surprised me.


What you’ll learn:

  1. Why I think we’ve hit an intelligence overhang and what that means for which model variables actually matter now
  2. How Opus 5’s “neurotic” personality showed up in real coding sessions, including a merge conflict it refused to touch
  3. What I learned from asking both Opus 5 and GPT‑5.6 Sol “who’s smarter, you or me?”
  4. Where Opus 5, GPT‑5.6 Sol, Sonnet 5, and Gemini 3.1 Pro actually landed on the HIA benchmark leaderboard
  5. The one use case where Opus 5 earned straight 5s from me
  6. My actual plan for using Opus 5 going forward

In this episode, I cover:

(00:00) Opus 5 is here

(03:15) First impressions

(06:12) Opus 5 vs. GPT‑5.6 Sol personality comparison

(14:39) Claude Slop: the verbosity problem and why it makes my blood boil

(16:55) How the How I AI benchmark works (7 models, 6 tasks, blind scoring)

(18:30) Live benchmark results: the leaderboard reveal

(23:25) My verdict and how I’ll actually use Opus 5

Tools referenced:

• Claude Opus 5:

• Anthropic blog: https://www.anthropic.com/news

• GPT‑5.6 Sol: https://openai.com/index/previewing-gpt-5-6-sol/

• Sonnet 5: https://www.anthropic.com/news/claude-sonnet-5

• Gemini 3.1 Pro: https://deepmind.google/models/gemini/pro/

Where to find Claire Vo:

ChatPRD: https://www.chatprd.ai/

Website: https://clairevo.com/

LinkedIn: https://www.linkedin.com/in/clairevo/

X: https://x.com/clairevo

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email jordan@penname.co.

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(96)

Computer & browser use in Codex (5 real examples)

Computer & browser use in Codex (5 real examples)

Today I’m walking you through one of my absolute favorite AI features right now: browser and computer use via Codex (the ChatGPT desktop app). I use this every single day, personally and professionall...

22 Heinä 0s

How the founder of Morning Brew built a Claude content machine that never runs out of ideas and never sounds like slop | Alex Lieberman

How the founder of Morning Brew built a Claude content machine that never runs out of ideas and never sounds like slop | Alex Lieberman

Alex Lieberman co-founded Morning Brew in college and grew it into one of the most-read business newsletters in the world before selling it to Business Insider. Now he’s the co-founder and co-managing...

20 Heinä 0s

This solo builder runs 24/7 local AI on his own hardware | Alex Finn

This solo builder runs 24/7 local AI on his own hardware | Alex Finn

Alex Finn is an AI builder, YouTuber, and the creator of Vibe Code Academy, a community for people learning to build with AI tools. He runs one of the most ambitious local AI setups I’ve come across: ...

13 Heinä 35min

GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark

GPT-5.6 Sol vs. Claude Fable: Why OpenAI’s new model crushes my benchmark

GPT-5.6 Sol is back, and I ran it through my full How I AI vibe benchmark against GPT-5.6 Terra, Luna, Claude Fable 5, and Sonnet 5 across five categories: PRDs, prototypes, wireframes, debugging, and...

9 Heinä 36min

What a harness is and how to build one with Claude Agent SDK

What a harness is and how to build one with Claude Agent SDK

Everybody is saying, “It’s not the model, it’s the harness,” but almost nobody stops to explain what a harness actually is. So I did. I built one live on the show: a Sentry bug-debugging harness for m...

8 Heinä 24min

How I run autonomous coding agents from my phone with OpenAI Symphony + Linear | Alessio Fanelli (Kernel Labs)

How I run autonomous coding agents from my phone with OpenAI Symphony + Linear | Alessio Fanelli (Kernel Labs)

Alessio Fanelli, founder of Kernel Labs and co-host of Latent Space podcast, walks us through two very different AI workflows: (1) a fully autonomous coding setup using OpenAI Symphony + Linear, where...

6 Heinä 35min

Sonnet 5 review: I ran 64 generations to find out if it's worth it

Sonnet 5 review: I ran 64 generations to find out if it's worth it

I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat o...

30 Kesä 25min