Absolute AppSec podcast

Episode 333 - LLM Patching Flaws, AI Code Regressions, Bug Bounty Economy

0:00
NaN:NaN:NaN
Rewind 15 seconds
Fast Forward 15 seconds
Sponsored by GuardSquare (guardsquare.com), the discussion of Episode 333 opens with an analysis of a 1Password academic paper evaluating how frontier LLMs perform at autonomous vulnerability patching. The research indicates that LLMs successfully generate functional, side-effect-free patches only 26% of the time, often introducing new security flaws, breaking application behavior, or hallucinating fixes due to a lack of environmental context and "correctness collapse". The hosts critique the industry push toward auto-remediation, arguing that automated patch generation fails to address root causes like noisy tooling or organizational culture issues, and they emphasize that human domain expertise remains necessary for reliable patching. Turning to real-world AI security risks, the episode examines a Snowflake vulnerability where an AI coding tool (GitHub Copilot Autofix) regressed a GitHub Actions workflow into an unauthenticated Remote Code Execution (RCE) flaw via command injection, which was subsequently discovered and validated within five days by Wiz's automated "Red Agent". Finally, the hosts cover Dark Reading reporting on how the AI-driven "vulnpocalypse" is repricing the bug bounty economy. As automated scanning harnesses double report volumes, companies face budget constraints that reduce payout amounts per finding, forcing organizations to narrow program scopes toward high-priority assets.

More episodes from "Absolute AppSec"