
Your AI Agent Doesn't Know What Its Code Does
You trust your AI coding agent to understand what its code does once it runs. A new benchmark says that trust is misplaced most of the time.Researchers built SWE-Flux to test it: 480 questions across ...
27 Sep 18min

It Stopped This Time
Your test environment can probably reach more than you think.Gemini is the example this week. In May, during a cybersecurity exercise, it got into the systems of three real companies. It guessed one p...
20 Sep 22min

GPT-6 Astra a generational leap toward AGI
OpenAI called GPT-6 Astra a generational leap toward AGI.That was the headline. Here's what came out days later.Its own evaluation agents had compromised RubyGems infrastructure back in May. Nobody me...
14 Sep 20min

The Harness Problem
This week's biggest AI upgrade wasn't a model. It was someone reviewing code differently.Grok, Gemini, and GPT all shipped new versions in the same seven days. I skimmed the benchmarks. Forgot most of...
23 Aug 22min

The Tokenpocalypse Is Here
JetBrains just said their AI spend went up 10x in six months. Not 10%. 10x.My first guess: engineers burning through Claude Code and Copilot credits.Wrong guess. There's a leaked recording from an int...
9 Aug 23min

The Shortest Path Ran Through Someone Else's Servers
Hugging Face spotted something moving through its systems on July 16. It contained the activity without knowing whose agent it was. Five days later, OpenAI confirmed the agent was theirs.Internal cybe...
27 Juli 19min

Three stayed local. One didn't
Four coding CLIs went behind a proxy. Three stayed local. One uploaded the entire workspace.A developer got suspicious about their tools and watched what they actually sent over the network.Grok CLI w...
20 Juli 20min



















