Dr Anton Chuvakin Blog (Original)

Security writing since the mid-2000s. New posts are mirrored from Anton on Security on Medium.

The 15-Minute Patch, Reverse-Engineered: What Had to Be True?

In Part 2 of the Patch Sound Barrier series, I shared a thought experiment from a talk in May 2026: imagine any vulnerability in your environment can be patched within 15 minutes of release. Now reverse-engineer that reality. What had to be true?

Gemini imagining this blog

Usman and I revisited this exercise in our “basics at AI scale” post. Readers asked: “How do I actually run this exercise with my team?”

So here’s how. Two ground rules first:

  • This is not a 15-minute patch SLA. If you take this post and write “all vulnerabilities remediated within 15 minutes” into a policy, I will find you… I utterly hated “30 days flat” policies as an analyst and “15 minutes flat” is the same mistake with more zeros. The number is a probe, not a target.
  • You will never get there across your whole environment. Legacy systems guarantee it. The point is the gap between the fantasy and your reality, because that gap is the most honest map of your program you will ever get. The Patch Sound Barrier is real; this exercise tells you where yours is and what it is made of.

Why 15 minutes, and not “fast”

“Fast” allows cheating. Some consider 90 days “fast” because they started at an annual cadence. “Patch faster” is generic advice everyone agrees with, but rarely acts on (or: barely acts on). An aggressive, specific number forces you to identify which pipeline steps physically cannot fit.

The Cloudflare quote from Part 2 summarizes this: demanding faster patching without redesigning the process leads to skipped steps. If regression testing takes a day, a two-hour SLA just skips testing. This exercise targets pipeline design, not raw speed. A 15-minute window leaves room only for automated steps. Everything else itemizes your sound barrier.

The 15 minutes, minute by minute

Assume a patch drops at T+0. At each stage, ask:

  1. What must be true for this stage to fit the timeframe?
  2. What is current reality, and what is the gap?

Details:

Process machinery 1

and

Process machinery 2

A 15-minute timeline leaves zero room for manual deliberation. In this model, humans define policies, build automation, and handle (rare!) exceptions. They do not manually approve or execute patches.

The five things the exercise refuses to let you skip

  1. Inventory as a live graph. If identifying affected systems takes hours, downstream automation cannot compensate. First move: measure how long it takes to locate a specific library or version today. That duration is your baseline gap. Most organizations are shocked. Log4j much?
  2. Rapid automated testing. Manual testing blocks rapid remediation. First move: time how long it takes to get from “patch available” to “confident it won’t break prod” on core systems, then work out what it would take to make that 5 minutes.
  3. Replaceable infrastructure. Rapid patching relies on redeploying updated templates rather than patching running instances. First move: determine your “replaceability ratio”: the percentage of systems that can be automatically rebuilt from source fast enough to fit the deploy window.
  4. Pipeline-level governance. Per-change approval gates prevent rapid execution. Approve the policy and the pipeline, not each patch. First move: transition low-risk change classes to pre-approved automated pipelines.
  5. Rehearsed rollback. Teams patch slowly because they fear breaking prod, and they fear it because they’ve never practiced un-breaking it. First move: execute and time an intentional rollback on a non-critical production system.

The cheat code: change the verb from “patched” to “non-exploitable”

For legacy systems, OT networks, or appliances you can’t modify, 15-minute patching is fiction. So change the verb: make the vulnerability non-exploitable or harmless within 15 minutes. That’s the Part 1 move: assume you can’t patch, then decide what you’ll do instead.

(source)

The difference (reminder for most, new for some...):

  • Patching eliminates the flaw. It requires vendor fixes, testing, deployment, and mutable target systems, factors largely outside your immediate operational control.
  • Mitigation reduces exposure risk. It relies on controls you own (network segmentation, identity policies, access gateways … at times an ax) without modifying target code, so actions can be pre-engineered and pre-approved easier.

The 15-minute mitigation exercise tests four capabilities:

  1. Can you isolate network paths immediately via automated policy?
  2. Can you revoke or scope down credentials in minutes?
  3. Do you know the blast radius well enough to isolate a component without breaking the business?
  4. Can you push a protection rule via API in minutes? (See API or Die.)

Your patch clock partly belongs to the vendor (partially). Your mitigation clock is all yours. For legacy, that’s the only 15-minute clock that matters.

But a segmentation diagram is not segmentation, and a kill switch nobody has pulled is a hypothesis. As I argued in Survival of the Basics, a mitigation counts only if it is a tested, operational control.

How to actually run this with your team

Book 90 minutes, a whiteboard, and the right people: platform/infra, an app owner, whoever runs change management, and QA. Security alone will find the gaps, but won’t own any of them.

  1. Select three specific target systems: modern (containerized service), median (standard application server), and legacy (high-risk core system).
  2. Map the clock for each system: document stage durations and identify the gaps.
  3. Identify the primary bottleneck stage.
  4. Define a single corrective move for that bottleneck: one bottleneck, fixed end to end (same logic as Baby ASO).
  5. Apply the mitigation model to legacy targets: establish rapid containment controls where patching is impossible.
  6. Track stage durations quarterly: progress shows up as shorter stages on your critical systems.

No time for a workshop? Pick a widely used component (a core library or agent) and time how long your team takes to find every active instance with high confidence. That’s your T+0 to T+1 baseline.

So, will you get to 15 minutes?

Containerized microservices can get there (yes, really, but provided some moons are aligned just right). Standard application servers can perhaps go from weeks to hours. Legacy systems need rapid mitigation (segmentation, credential scoping, containment) rather than patching.

What do you think? Is this totally crazy or what?

Related posts and resources:


The 15-Minute Patch, Reverse-Engineered: What Had to Be True? was originally published in Anton on Security on Medium, where people are continuing the conversation by highlighting and responding to this story.


Originally published at Medium.