GPT 5.4 vs Gemini: Benchmarks, Codex, Excel

GPT 5.4 vs Gemini: Benchmarks, Codex, Excel

Beth Lyons and Andy Halliday open the show with a focused breakdown of GPT-5.4, framing it less as a universal leap and more as a strong advance in white-collar knowledge work and real-world task performance. Much of the conversation compares GPT-5.4 with Gemini 3.1 Pro Preview, Claude models, Codex, and other systems across benchmarks like GPT-Val, coding, long-context reasoning, hallucination resistance, and visual reasoning, with repeated emphasis that users still need to pick models based on the actual job to be done. Beth also shares a practical complaint about Gemini hallucinating around silent screen recordings and uses that to argue for a more dependable “colleague layer” in agentic systems. Later, Karl Yeh joins to talk through hands-on experience with GPT-5.4 in Codex, comparisons with Claude in Excel and Gemini in Sheets, and where the new release feels genuinely useful in day-to-day work.


Key Points Discussed


00:00:18 Welcome and setup for a GPT-5.4-focused episode

00:02:47 GPT-Val and white-collar knowledge work framing

00:08:51 Benchmark comparison across GPT-5.4, Claude, Gemini, and others

00:16:26 Gemini strengths in video and visual reasoning

00:18:05 Beth’s Gemini transcription / hallucination workflow example

00:23:54 “Then we’ll move to more news” and handoff to Karl Yeh

00:24:24 Karl Yeh on real-world use cases over benchmarks

00:55:30 Closing recommendations: try GPT-5.4, use Codex, newsletter and community plug


The Daily AI Show Co Hosts: Beth Lyons, Andy Halliday, Karl Yeh

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(874)

Who Should You Trust to Teach You AI?

Who Should You Trust to Teach You AI?

The episode opened with Perplexity Deep Research suddenly behaving very differently from the product Brian had used for months. Instead of detailed research, it returned short answers, mixed old conve...

25 Elo 57min

Is the Backlash Against AI Data Centers Justified?

Is the Backlash Against AI Data Centers Justified?

The episode opened with a fact-check of claims defending the current AI data center buildout. Brian compared arguments about electricity prices, taxes and water use against research he had gathered, w...

24 Elo 1h

The Synthetic Anchor Conundrum

The Synthetic Anchor Conundrum

Mirage’s AI news experiment points to a version of media that does not need a studio, a broadcast schedule, or a human anchor reading from a desk. A channel can appear in a day. It can label synthetic...

22 Elo 29min

Should We Rebuild Work Around AI?

Should We Rebuild Work Around AI?

The episode opened with a practical warning for people building AI systems: timestamps and time zones can quietly break databases, automations and search tools. That led into Slack Code, a new collabo...

21 Elo 1h 7min

Is Grok Bot the Best AI Work Assistant?

Is Grok Bot the Best AI Work Assistant?

The episode opened with a hands-on comparison of Grokbot, Codex and Claude. Gareth found Grokbot strong for delegation, organization and everyday work, but weaker on difficult problem solving. The dis...

20 Elo 1h 3min

Do We Need to Rethink What Work Is?

Do We Need to Rethink What Work Is?

The episode opened with Apple Vision Pro being used to map a house while running Ethernet cable, letting a worker see marked locations through floors and walls. That led to a wider discussion about di...

19 Elo 1h

Are Custom GPTs Reaching the End?

Are Custom GPTs Reaching the End?

The episode opened with a practical example of how quickly AI coding agents are moving beyond software. Someone used Claude to write a Mac driver for an old Windows-only HP printer, leading to a wider...

18 Elo 50min

Are AI Harnesses the New AI Wrappers?

Are AI Harnesses the New AI Wrappers?

The episode opened with the reported Stripe acquisition of OpenRouter at a $7 billion valuation and questions about how OpenRouter’s business model supports that price. The conversation expanded into ...

17 Elo 58min