The Vulnerability Your Agent Merged

The Vulnerability Your Agent Merged

The unit tests pass. The PR merges. And you won't find the problem for six months.

Two papers landed this week — one on LLM-generated code, one on GitHub Actions workflows. Different researchers. Same finding.

When agents write code, they pin library versions that trained well. Not versions that are safe.

The mechanism is simple. A model has seen one popular version of a library thousands of times. It reaches for that version because it minimizes prediction loss. Pin-by-popularity and pin-by-safety are different jobs. The model only knows one of them.

The GitHub Actions paper found the same shape. Right syntax. Wrong threat model.

So the code looks clean. The tests pass. The PR merges. And six months later a security audit finds a CVE that was public before the agent ever touched the file.

This is not a model problem. It is a workflow problem.

Human PRs go through SCA. Agent PRs often don't. That gap is where the bill arrives.

The fix is not complicated. Put pip-audit, npm audit, or OSV-Scanner between the agent and main. Same gate you'd use for any contributor. The agent has not finished the work when it merges. It has finished its part.

Your security pipeline was designed for human contributors. Has anything changed since you started using agents?

Tämä jakso on lisätty Podme-palveluun avoimen RSS-syötteen kautta eikä se ole Podmen omaa tuotantoa. Siksi jakso saattaa sisältää mainontaa.

Jaksot(43)

Your AI Agent Doesn't Know What Its Code Does

Your AI Agent Doesn't Know What Its Code Does

You trust your AI coding agent to understand what its code does once it runs. A new benchmark says that trust is misplaced most of the time.Researchers built SWE-Flux to test it: 480 questions across ...

27 Syys 18min

It Stopped This Time

It Stopped This Time

Your test environment can probably reach more than you think.Gemini is the example this week. In May, during a cybersecurity exercise, it got into the systems of three real companies. It guessed one p...

20 Syys 22min

GPT-6 Astra a generational leap toward AGI

GPT-6 Astra a generational leap toward AGI

OpenAI called GPT-6 Astra a generational leap toward AGI.That was the headline. Here's what came out days later.Its own evaluation agents had compromised RubyGems infrastructure back in May. Nobody me...

14 Syys 20min

The Harness Problem

The Harness Problem

This week's biggest AI upgrade wasn't a model. It was someone reviewing code differently.Grok, Gemini, and GPT all shipped new versions in the same seven days. I skimmed the benchmarks. Forgot most of...

23 Elo 22min

The Tokenpocalypse Is Here

The Tokenpocalypse Is Here

JetBrains just said their AI spend went up 10x in six months. Not 10%. 10x.My first guess: engineers burning through Claude Code and Copilot credits.Wrong guess. There's a leaked recording from an int...

9 Elo 23min

Ponytail Activation 0%

Ponytail Activation 0%

Advertised 54%. Measured 15%. Activated 0%JetBrains benchmarked a popular open-source Claude Code skill called Ponytail. It pushes the agent to write less code. The authors has promised 54% less code,...

2 Elo 20min

The Shortest Path Ran Through Someone Else's Servers

The Shortest Path Ran Through Someone Else's Servers

Hugging Face spotted something moving through its systems on July 16. It contained the activity without knowing whose agent it was. Five days later, OpenAI confirmed the agent was theirs.Internal cybe...

27 Heinä 19min

Three stayed local. One didn't

Three stayed local. One didn't

Four coding CLIs went behind a proxy. Three stayed local. One uploaded the entire workspace.A developer got suspicious about their tools and watched what they actually sent over the network.Grok CLI w...

20 Heinä 20min