In agentic engineering, engineers hand tasks to coding agents that write code, run it, and fix their own mistakes in fast loops. Bartosz Ocytko recently mapped where agentic engineering is heading, from a surge of commits on GitHub to bottlenecks in verification, open source, and compute. Most security tooling and processes weren’t built for that volume of change. While reading his piece, I kept coming back to 1 question: what does this shift mean for application security (AppSec)?
Ocytko argues that agentic engineering will reward organizations with strong engineering foundations: those that can verify, coordinate, and safely absorb machine-generated change. I think AppSec is part of that foundation. AppSec teams that already build tooling, run large-scale remediation, and treat risk quantitatively are well placed to make agentic engineering safe at scale. Attackers raise the stakes for strong security engineering, too: Google’s Threat Intelligence Group reports that AI models lower the barrier for adversaries to develop exploits. In May 2026, the group reported the first threat actor it believes used a zero-day exploit developed with AI.
The same agents that create the volume can also help with problems that AppSec teams couldn’t solve at human scale: too many findings, too few security engineers, and feedback that arrived long after the code was written. I see 3 shifts that decide which AppSec teams keep up:
- From detection to mitigation: prevent findings inside the agent’s loop and fix the backlog at fleet scale.
- From writing policies to building tools: ship security tooling into the context that coding agents work in.
- From gatekeeping to partnership: manage security risk together with site reliability engineering (SRE), platform, and cloud engineering teams, the way they already manage reliability.
Findings Pile Up Faster Than Teams Fix Them
Most mature organizations already have good visibility into known vulnerability classes. Their scanners cover infrastructure, cloud platforms, application code, and continuous integration and delivery (CI/CD) pipelines. AI adds even more findings: HackerOne argues that AI-powered vulnerability discovery has improved by an order of magnitude, while remediation hasn’t kept up.
Backlog data shows how far remediation lags behind. Edgescan, a vulnerability management vendor, analyzed its 2024 scanning and penetration testing results in its Vulnerability Statistics Report 2025. At enterprises with more than 1,000 employees, 45.4% of the vulnerabilities discovered within 12 months were still open. High and critical findings with the highest predicted exploit probability (an Exploit Prediction Scoring System score above 0.7) took 115.7 days on average to fix, slightly longer than the 109.4 days for those least likely to be exploited. Every finding that nobody fixes grows the backlog and extends the exposure window, the time during which an attacker can exploit the vulnerability in production.
Coding agents make this worse. They generate code at machine speed, so new findings appear faster. Lars Janssen calls the growing gap between how fast teams generate code and how fast they can check it verification debt. If remediation still depends on engineers picking up tickets between feature work, the backlog keeps growing. Risk only goes down when teams fix findings faster than scanners discover new ones.
That requires AppSec teams to move from “we found it, you fix it” to “we prevent it, and we fix what slips through”.
Move Security Feedback Into the Agent Loop
In the traditional loop, a scanner reports an issue in the pipeline, someone files a ticket, and an engineer fixes it when there’s time. In the agentic loop, the scan results go straight back to the coding agent, which rewrites the code in the same session until the checks pass. Many findings never reach a ticket because the agent fixes them before anyone commits the code. GitHub’s Copilot Autofix suggests fixes directly in the pull request, which comes close to this loop. During its 2024 public beta, developers fixed new code scanning alerts in pull requests in a median of 28 minutes, compared with 1.5 hours without Autofix.
Today’s security gates weren’t built for this loop. When a coding agent opens a pull request (PR) in seconds, a security gate that takes minutes becomes the slowest step. Binary gates also break down at this volume: “block on medium severity” blocks everything, and “ignore everything below critical” leaves real risks open. The risk budgets I describe further down give teams a better decision rule.
Some findings will still slip through, and most organizations already carry a large backlog. Fixing them 1 ticket at a time doesn’t scale. In campaign-based remediation, AppSec targets a whole vulnerability class at once, such as an outdated dependency, an infrastructure as code (IaC) misconfiguration, or an insecure pattern repeated across hundreds of services.
Spotify’s fleetshift is a good mental model: it makes fleet-wide code changes by running a Docker image against each repository. Spotify reports that fleetshift created more than 270,000 PRs in 2022 and merged 77% of them automatically, and that it patched 80% of its production backend services within 9 hours of the Log4j disclosure.
With coding agents, the AppSec team designs the campaign, and agents open, test, and verify the remediation PRs across the fleet. GitHub already supports part of this: in security campaigns, teams can assign alerts to the Copilot cloud agent, which opens pull requests with fixes.
Build and Ship Security Tooling
In this model, AppSec engineers build and ship security tooling and contribute remediation code directly. Their goal is infrastructure that makes coding agents produce secure code by default.
Security Skills and Machine-Readable Policy
Skills are packaged instructions that coding agents load when a task needs them, and they follow an open specification. Project CodeGuard from the Coalition for Secure AI (CoSAI) publishes general security skills for cryptography, input validation, authentication and authorization, and supply chain security. Organization-specific skills go further and encode internal policies (“use the central auth SDK, not raw JWT handling”), approved libraries, and deployment constraints. If the organization distributes skills centrally, AppSec can publish a skill update when a new vulnerability class appears, and every agent in the organization has it the same day.
The security policy itself needs the same treatment, because a coding agent can’t do much with a 50-page PDF. AppSec teams can turn it into agent instructions, a searchable vector database, or policy-as-code rules that agents check while they work. The agent then verifies compliance mid-task without waiting for a human reviewer, and AppSec’s work moves from reviewing individual changes to maintaining the policies and risk models that agents enforce.
Local Scanners and Review Agents
Security scanning can’t live only in version control platforms and CI pipelines anymore. Coding agents need static application security testing (SAST), dynamic application security testing (DAST), and reviewers based on large language models (LLMs) on the engineer’s machine, so they can check their own output before they push. AppSec teams package these tools as command-line interfaces (CLIs) or Model Context Protocol (MCP) servers and manage their configuration and policies centrally. MCP gives agents a standard way to call tools. Engineering teams then get scanning that works without setup.
Dedicated review agents combine security skills with local tooling. They check code against organization-specific policies and report issues in a form the coding agent can act on directly. roborev points in this direction: it’s a local daemon that reviews each commit in the background through a git hook and reports every finding with a severity and a file and line location before the code reaches a PR.
A shared review ledger ties these checks together. Every source, whether static analyzer, dynamic scanner, or LLM-based reviewer, writes to the same structured log, typically in the Static Analysis Results Interchange Format (SARIF). The coding agent reads the ledger after each iteration and knows which checks passed, which failed, and what to fix next. LLM-based reviewers can catch context-dependent issues that rule-based scanners miss, but they don’t return the same result on every run, so the ledger needs both.
Sandboxes and Short-Lived Credentials
Coding agents are most useful when they can run the code they write: they implement a change, run the tests, read the failure, and try again. That makes every agent task an untrusted workload running arbitrary code, often on an engineer’s laptop. For AppSec, this adds up to a large internal remote code execution (RCE) surface, and securing it touches workstation setup as well as engineering practice.
Sandboxing and prompt injection defense go together. A coding agent reads external input such as dependency metadata, issue descriptions, PR comments, and documentation, and each of these can carry a prompt injection that tells the agent to do something else. A sandbox limits the damage when that happens, and the common options differ in isolation strength and cost:
- Lightweight virtual machines (micro-VMs) such as Firecracker and Docker Sandboxes give each task a disposable environment with its own kernel.
- OS-level sandboxes are lighter. nono applies irrevocable, kernel-enforced filesystem allow-lists through Landlock on Linux and Seatbelt on macOS, and Sandvault runs agents as a separate macOS user inside
sandbox-exec. The nono authors point out that OS-level sandboxing doesn’t give the memory isolation of a virtual machine. - Containers share the host kernel, so escaping 1 container puts the attacker on the host.
AppSec picks the isolation boundary that fits each workload’s trust level and, together with platform teams, ships the sandbox setup as the default configuration engineers get out of the box.
Credentials need the same care. Coding agents should never hold persistent identity and access management (IAM) roles or long-lived API keys. An identity broker issues short-lived, task-scoped credentials on demand and revokes them when the task completes. AppSec can build the broker together with platform and IAM teams and hide it behind the agent’s tooling, so agents exchange tokens automatically and never fall back on whatever credential is at hand.
Beyond internal credentials, agents increasingly act on behalf of users in third-party services. That needs delegation and consent infrastructure with narrow, time-limited permissions the user can revoke at any moment. Without it, teams end up storing tokens in agent contexts or hard-coding service accounts, and both fail as soon as an attacker compromises an agent or the agent misreads a prompt.
Threat Modeling and Offensive Testing
Traditional scanners catch known vulnerabilities listed as Common Vulnerabilities and Exposures (CVEs), common misconfigurations, and obvious injection points. They miss logic flaws such as broken access control across multi-step workflows, business logic bypasses, and race conditions. To catch design-level flaws, AppSec teams have to encode their domain knowledge into agents.
Threat modeling, where a team maps how an attacker could abuse a system’s design, usually happens in periodic workshops. Jack Naglieri built threat models with AI agents and MCP and argues that agents remove the weeks of meetings and documentation review that kept threat modeling a once-a-year exercise. Naglieri’s example comes from detection engineering, but the approach carries over to applications. AppSec teams can build agent setups that regularly review applications and repositories, analyze access control patterns, data flow boundaries, and trust assumptions, and report drift from the organization’s security posture.
Agentic penetration testing (pentesting) complements this, and AppSec can offer it to engineering teams as an internal service. These agents explore the application’s business logic and look for the flaws a human pentester would target: authorization boundaries, multi-step attack paths, and privilege escalation. To build them, AppSec teams turn their offensive expertise into repeatable attack strategies. Public setups today still keep a human in the loop. In AI-assisted web pentesting with Claude Code and Burp MCP, for example, the agent suggests the next test and the pentester decides whether to run it.
Partner With SRE and Platform Teams
Can a system ever truly be considered reliable if it isn’t fundamentally secure? Or can it be considered secure if it’s unreliable?
โ Building Secure and Reliable Systems
In SRE, the product owner sets a service level objective (SLO) such as 99.9% availability, and the remaining 0.1% becomes the service’s error budget, which the team can spend on outages and risky releases. AppSec can join that model with a risk budget. The risk budget sets how long a high-risk, reachable vulnerability may stay open in production before it burns error budget, just like downtime does. AppSec provides the framework and the tooling, and the product owner sets the target, just as for the SLO.
A risk budget helps with a problem that agentic engineering makes worse. Coding agents generate change faster than people can triage the findings, and without a quantitative target every finding competes for the same attention. With a risk budget, teams fix what burns the budget and defer what doesn’t. The conversation moves from “fix all the bugs” to “manage the exposure window”, and security becomes an operational concern that AppSec shares with the teams running the service.
Teams can only use a risk budget if they can tell real exposure from noise, and that requires reachability analysis: tracing the call graph from the vulnerable code to the application’s entry points. A CVE in dead code doesn’t burn budget, but the same CVE behind a public API does. At the CI gate, teams then block only changes that introduce reachable risk and stop failing builds on raw finding counts. In production, a monitoring layer tracks how long reachable findings stay open and burns error budget accordingly.
Platform and cloud engineering teams already ship the paved roads that engineering teams build on: base images, service templates, IaC modules, identity primitives, and networking defaults. AppSec teams get the most reach by building security into those golden paths. A hardened base image or a secure-by-default service template then carries security to every service that uses it, including the code that agents write. Platform engineering owns the templates and modules, and AppSec owns the security defaults inside them. I cover this model in more detail in Shift Down.
Where to Start
None of this has to happen at once. A team that wants to start can work through these steps in order:
- Sandbox the coding agents that engineers already run, and replace their long-lived credentials with short-lived, task-scoped ones.
- Package the existing scanners as CLIs or MCP servers with central configuration, and make them write SARIF.
- Publish an organization-specific security skill that encodes the internal policies that matter most.
- Pick 1 vulnerability class from the backlog and fix it across the fleet as a campaign.
- Agree on a risk budget for reachable vulnerabilities with 1 service owner and their SRE team.
The sandbox comes first because the exposure exists as soon as an engineer runs an agent with access to their machine.
