Skip to content
kiprozess
Processes & automation2 min

Agents in the mid-market: what really happens after the pilot

Almost every company is running an AI pilot by now. Few make it into day-to-day operations. Why that is — and what the exceptions do differently.

The pilot is the easy part. A team builds an agent in three weeks that pulls quotes out of email, the result looks impressive, and leadership is convinced. Six months later the system still runs on the laptop of the person who built it.

This pattern repeats so reliably that it can no longer be coincidence. The jump from prototype to production rarely fails because of model quality. It fails because of everything missing around the model.

The pilot measures the wrong thing

A prototype proves a task is solvable in principle. Production demands something else: that it is solvable five days out of five, on bad inputs, when nobody is watching. Between 85 % and 99 % accuracy lies not fine-tuning but a different project altogether.

  • Who corrects errors when the agent is wrong — and would anyone notice?
  • What happens when the API goes down? Does the process continue or does the department stop?
  • Who is accountable for the output towards customers, auditors and works councils?
  • How can decisions be reconstructed after the fact?

What the exceptions do differently

Projects that reach production share three traits. First, they automate a process that was already documented. Second, they start with a human in the approval step and only remove them once the numbers allow it. Third, they log everything from day one: every decision, every input, every correction.

That sounds unspectacular, and that is exactly the point. The difference between an impressive pilot and a productive system almost never lies in model choice. It lies in whether someone was willing to build the boring part.

A realistic timeline

  1. Weeks 1–2: observe the process and break it into steps without talking about technology.
  2. Weeks 3–5: prototype on real historical data, not on examples.
  3. Weeks 6–10: shadow mode — the agent decides alongside, but nobody acts on it.
  4. From week 11: partial release for clearly bounded cases, everything else stays manual.

Three months sounds long for something the demo video finished in three weeks. It is the price of a result that survives the first vacation of the person responsible.

ShareLinkedInX

Author

Florian AlbertFounder & editor

Has spent over a decade at the intersection of performance marketing and automation, building AI systems for mid-market companies and agencies.

FA
Back to articles

Keep reading