Aardvark Agent Security: Scaling Defense and Finding 92% of Code Vulnerabilities with GPT-5

Aardvark Agent Security: Scaling Defense and Finding 92% of Code Vulnerabilities with GPT-5

Join us as we explore Aardvark, OpenAI’s groundbreaking agentic security researcher, now available in private beta. Powered by GPT-5, Aardvark is an autonomous agent designed to help developers and security teams discover and fix security vulnerabilities at scale.

Software security is one of the most critical and challenging frontiers in technology. With over 40,000 CVEs reported in 2024 alone, and estimates showing that around 1.2% of commits introduce bugs, software vulnerabilities pose a systemic risk to infrastructure and society. Aardvark is working to tip this balance in favor of defenders, representing a new, defender-first model that delivers continuous protection as code evolves.

Unlike traditional program analysis techniques like fuzzing, Aardvark uses LLM-powered reasoning and tool-use to understand code behavior and identify vulnerabilities. It approaches security like a human researcher would: reading code, running tests, analyzing findings, and using tools.

Aardvark operates through a multi-stage pipeline to identify, explain, and fix issues:

  1. Analysis: It begins by producing a threat model based on the project’s security objectives.
  2. Commit scanning: It continuously monitors and inspects commit-level changes against the entire repository, identifying vulnerabilities and explaining them step-by-step.
  3. Validation: It attempts to trigger the potential vulnerability in an isolated, sandboxed environment to confirm its exploitability and ensure accurate insights.
  4. Patching: Aardvark integrates with OpenAI Codex to generate and scan a patch, which is then attached to the finding for efficient human review.

The results are significant: in benchmark testing on "golden" repositories, Aardvark identified 92% of known and synthetically-introduced vulnerabilities. It also uncovers other issues, such as logic flaws, incomplete fixes, and privacy concerns. Aardvark integrates seamlessly with existing workflows and has already surfaced meaningful vulnerabilities within OpenAI's internal codebases and external alpha partners.

Furthermore, Aardvark has already been applied to open-source projects, contributing to the security of the ecosystem and resulting in the responsible disclosure of numerous vulnerabilities—ten of which have received CVE identifiers. By catching vulnerabilities early and offering clear fixes, Aardvark helps strengthen security without slowing innovation.

Tune in to understand how this new breakthrough in AI and security research is expanding access to security expertise.

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(167)

AT&T's AI Automation Push: WIRED Reports More Job Cuts, Retiring Copper Landlines, and a Leaner Telecom Network—What the Company's Transformation Means for Workers and Customers

AT&T's AI Automation Push: WIRED Reports More Job Cuts, Retiring Copper Landlines, and a Leaner Telecom Network—What the Company's Transformation Means for Workers and Customers

AT&T is rebuilding its telecom business for the AI era, and the shift could mean fewer jobs, less copper infrastructure, and a very different network. In this episode of The Daily AI Chat, we unpack W...

23 Sep 20min

GPT-6 Sol and Luna Launch: OpenAI Says Its New AI Models Cut API Costs in Half and Make Fewer Mistakes as Competition With Anthropic Heats Up Across ChatGPT and Codex

GPT-6 Sol and Luna Launch: OpenAI Says Its New AI Models Cut API Costs in Half and Make Fewer Mistakes as Competition With Anthropic Heats Up Across ChatGPT and Codex

OpenAI has expanded its GPT-6 lineup with Sol and Luna. The company says the new models cost half as much through the API as their GPT-5.6 counterparts while making fewer factual and coding mistakes. ...

22 Sep 17min

AI Data Center Backlash Is Growing: Why Pennsylvania Residents, Unions, and Environmental Groups Are Fighting New Construction as the AI Infrastructure Boom Accelerates

AI Data Center Backlash Is Growing: Why Pennsylvania Residents, Unions, and Environmental Groups Are Fighting New Construction as the AI Infrastructure Boom Accelerates

AI companies are racing to build the data centers that power bigger models and new products. But the communities asked to host that infrastructure are raising questions about electricity costs, water,...

22 Sep 17min

Nscale’s $103 Billion AI Contract Backlog Meets Wall Street: The IPO Test for Microsoft and Anthropic Dependence, Financing Risk, Data Centers, and the Neocloud Boom

Nscale’s $103 Billion AI Contract Backlog Meets Wall Street: The IPO Test for Microsoft and Anthropic Dependence, Financing Risk, Data Centers, and the Neocloud Boom

Nscale says it has more than $103 billion in contracts. Now the British AI cloud company is preparing to go public, and its filing exposes a question investors across the AI boom can no longer avoid: ...

22 Sep 10min

AI Speaks to the Other 3 Billion: Inside the Gates Foundation Coalition Fixing the Global Language Data Gap With Anthropic, Google and OpenAI—and Why It Matters

AI Speaks to the Other 3 Billion: Inside the Gates Foundation Coalition Fixing the Global Language Data Gap With Anthropic, Google and OpenAI—and Why It Matters

Artificial intelligence can write essays, generate code and answer questions in seconds—but for billions of people, it still struggles to understand the language they actually speak. The Gates Foundat...

21 Sep 21min

UN Scientists Warn AI Agents Are Outpacing Traditional Safeguards: Inside the Global Push for Precaution, Governance and Controls Before Autonomous Systems Scale

UN Scientists Warn AI Agents Are Outpacing Traditional Safeguards: Inside the Global Push for Precaution, Governance and Controls Before Autonomous Systems Scale

AI agents are rapidly moving beyond simple question-and-answer tools. They can plan, use software, communicate with other systems, and take actions with limited human supervision. Now a new United Nat...

21 Sep 18min

Only 7,000 Humanoid Robots Sold Worldwide: The Reality Behind the AI Robotics Boom, China’s Ambitions, Factory Pilots and the Race to 1.2 Million by 2030

Only 7,000 Humanoid Robots Sold Worldwide: The Reality Behind the AI Robotics Boom, China’s Ambitions, Factory Pilots and the Race to 1.2 Million by 2030

Humanoid robots can run, box, dance, carry parts, and dominate technology demonstrations—but how many are actually being sold and put to work? In this episode of The Daily AI Chat, we examine a striki...

21 Sep 20min

Big Tech’s Hidden $300 Billion AI Debt Bet: How Off-Balance-Sheet Guarantees Are Financing the Data-Center Boom—and What Investors Should Watch, Explained

Big Tech’s Hidden $300 Billion AI Debt Bet: How Off-Balance-Sheet Guarantees Are Financing the Data-Center Boom—and What Investors Should Watch, Explained

Big Tech’s artificial-intelligence spending boom may be even larger—and more financially complex—than corporate balance sheets suggest. In this episode of The Daily AI Chat, we unpack Financial Times ...

20 Sep 20min

Populärt inom Politik & nyheter

svenska-fall
aftonbladet-krim
fordomspodden
p3-krim
rss-krimstad
flashback-forever
rss-vad-fan-hande
aftonbladet-daily
rss-sanning-konsekvens
rss-krimreportrarna
svd-ledarredaktionen
spar
politiken
rss-flodet
en-runda-till
rss-frandfors-horna
motiv
omni-podd
rss-expressen-dok
rss-klubbland-en-podd-mest-om-frolunda