
When Will AI Models Blackmail You, and Why?
In the last few days Anthropic have released an impressive honest account of how all models blackmail, no matter what goal they have, and despite prompt warnings, and other preventions. But do these m...
24 Kesä 202526min

Apple’s ‘AI Can’t Reason’ Claim Seen By 13M+, What You Need to Know
What to make of those headlines that AI can’t reason, seen by tens of millions? I cover the paper in layman’s terms, what it means and doesn’t mean, and what’s next. Thanks to Storyblocks for sponso...
12 Kesä 202514min

AI Accelerates: New Gemini Model + AI Unemployment Stories Analysed
There’s a new best language model, so let’s go through the up and downs of Gemini 2.5 Pro 06-05. Record-breaking common-sense, but dumb mistakes remain. And it’s not even their best model, which remai...
6 Kesä 202516min

Claude 4: Full 120 Page Breakdown … Is it the Best New Model?
Not only did I get early access and ran my own tests, as per the title I read both the 120 page Claude 4 Opus and Claude 4 Sonnet System Card, and 25 page report on ASL-3 being triggered, plus the 2 h...
22 Touko 202519min

Google Takes No Prisoners Amid Torrent of AI Announcements
Google just announced at least 12 things that are each worthy of a video, but here are the top I/O highlights. From Veo 3 to Deep Research now being useable, Deep Think breaking records to Gemini Diff...
21 Touko 202517min

AI Improves at Self-improving
AlphaEvolve is not the first system to exhibit self-improvement, but it may be the most impressive yet. AI is literally improving the hardware, architectures, data and training methods of AI itself. A...
19 Touko 202517min

o3 breaks (some) records, but AI becomes pay-to-win
A green card, o3 vs Gemini 2.5, 6 Benchmarks and a whole bunch of my thoughts on what on earth is happening in AI, from here to 2030. Plus, how AI is becoming pay-to-win, and why. Crazy times, 14 mins...
25 Huhti 202514min

o3 and o4-mini - they’re great, but easy to over-hype
Critical analysis of the two most powerful new models behind ChatGPT, o3 and o4-mini. Not just the system cards, benchmarks, and my own tests, but some you may not have seen before. Yes, they can whip...
16 Huhti 202514min



















