Back to Home
#AI#Agents#Automation

I Replaced a Process, Not a Person

Toni Nowak
I Replaced a Process, Not a Person

I didn't automate a security analyst out of a job. I automated the 15 hours a week they lose to vulnerability triage. Here's how the agent is built — and the four places it tries to do something stupid.

The headline has the frame backwards

"AI will take your job" is the wrong unit. Jobs aren't monoliths an agent swallows whole. Jobs are bundles of processes, and most of those processes are toil.

In my consulting work I've watched security teams burn 15+ hours a week on vulnerability triage — reading scanner output, chasing CVE details, deciding what actually matters — before a single fix gets written. Nobody took that job because they love triage. They took it for the hard parts: judgment, architecture, the weird incident at 2 a.m.

So I didn't build an agent to be a security engineer. I built one to eat the triage. It runs on a schedule against my own infrastructure: it scans, ranks, proposes fixes, applies the safe ones, and hands me the rest. Below is the anatomy — and, more usefully, the four places it tries to do something stupid and how I stop it.

Pick a process, not a job

The actual skill in using agents isn't prompting. It's scoping — carving a well-shaped process out of a fuzzy job. Get this wrong and no model will save you.

A process is a good agent target when it looks like this:

Good target for an agentBad target
High-volume, repetitiveOne-off, bespoke
Has a checkable "right answer"Ambiguous, taste-driven
Output a human can verify fastOutput nobody can audit
MeasurableJudged on vibes
Low blast radius, or gateableIrreversible by default

Vulnerability triage scores well on every row. "Be our security architect" scores badly on all of them. Point an agent at the second and you get a confident, plausible, unaccountable mess. Point it at the first and you get your week back.

The shape of an agent that does real work

The mistake is one giant prompt that "handles security." Real work needs a small pipeline of narrow roles, each doing one job and handing structured output to the next — exactly the triage → research → remediation split I've written about before.

In practice that's: a scanner (Trivy, Grype) produces findings; a local model (Devstral via LM Studio) ranks them and drops the noise; Claude CLI drafts the actual fix; a scheduler (systemd) runs the loop; and every step writes to a log you can read later. Cheap reasoning stays local, expensive reasoning goes to a frontier model, and nothing is a black box.

But the architecture isn't the point. The control flow is:

# The shape that matters: propose -> gate -> act. Never act -> explain.
findings = scan()                       # what's vulnerable
plan     = triage(findings)             # rank, drop false positives

for fix in plan.high_confidence:        # the boring 80%
    if low_risk(fix) and tests_pass(sandbox(fix)):
        apply(fix); log(fix)            # automatic, fully audited
    else:
        queue_for_human(fix)            # the risky 20%

notify(summary)                         # what I did, what I left for you

An agent that acts and then explains is a liability. An agent that proposes, gets gated, then acts is a colleague. The whole design lives in that ordering.

The four places it tries to do something stupid

This is where the engineering actually is. Every one of these is a real failure mode, not a hypothetical.

1. It hallucinates a fix — or returns garbage. Local models especially will hand you confident, malformed output. So model output is untrusted input. My agent parses the model's JSON defensively and rejects anything that doesn't validate, rather than acting on a half-formed plan. If the structure is wrong, the fix doesn't happen.

2. It wants to act on a false positive. Scanners over-report. An eager agent will happily "fix" things that were never broken, churning your systems for nothing. The gate is severity plus verification: confirm the vulnerability is real and reachable before anything touches it.

3. A "fix" breaks three other things. "Just update the library" sounds simple until the maintainer changed the API and two services fall over. Nothing auto-applies without passing tests in a sandbox first, and every change has a rollback. Low-risk patches go through automatically; anything with blast radius waits for a human.

4. It has more access than it needs. An autonomous agent with broad permissions is the most dangerous thing in your stack — it runs with your rights, and a single bad instruction (or a prompt injection in something it reads) becomes a real action. Least privilege, isolation, and a log of every command aren't optional. Residency and local models don't save you here; access control does.

The pattern under all four: the agent proposes, a gate decides, and a human owns anything risky. Automate the boring 80%. Keep eyes on the 20%.

What stays human

Augmentation isn't a softer word for replacement. It's a division of labour.

The agent owns the toil: the reading, the ranking, the boilerplate fix, the paperwork. The human owns what the agent can't be trusted with — judgment on edge cases, the high-blast-radius approvals, the system design, and the accountability when something breaks. Someone is still responsible. That doesn't move to the model.

The analyst whose triage I automated didn't lose a job. They lost the part of the job that was making them quit. Same person, reviewing instead of grinding.

What this generalizes to

Vulnerability management is just the cleanest example — high-volume, measurable, gateable. The same recipe transfers to any process shaped like it:

  • incident response triage
  • first-pass code review
  • infrastructure provisioning
  • data-pipeline babysitting
  • support-ticket escalation

Find a process that's costing someone 15 hours a week. Give it the propose-gate-act treatment. Keep the human on the risky 20%. Repeat.

The bottom line

The question was never "will an agent replace you." It's "which of your weeks will you stop wasting" — and "are you the one designing the agent, or the one being designed around."

An agent didn't replace the analyst. It gave them their Thursday back.

Key takeaways

  • Agents don't replace jobs; they replace processes. The skill is scoping the right one.
  • Good targets are high-volume, verifiable, measurable, and gateable. Bad targets are taste-driven and irreversible.
  • Build a pipeline of narrow roles, not one mega-prompt. Cheap reasoning local, expensive reasoning frontier.
  • The safety lives in the ordering: propose → gate → act, never act → explain.
  • Treat model output as untrusted, test before applying, keep rollbacks, and enforce least privilege.
  • Automate the boring 80%; keep a human — and the accountability — on the 20%.

I'm a solutions architect specializing in AI/ML, secure infrastructure, and agent automation. Through AI-Flow, I help organizations put AI agents to work on real processes — without handing over the keys.

#AIAgents #Automation #DevSecOps #AI #HumanInTheLoop #AgentEngineering