In short. Rewiring one workflow takes ninety days, and forty of them produce no automation. They produce the baseline you will be judged against and the exception log that tells the agent what the job actually is. Ford proved the moving assembly line on a single component before touching the rest of the plant. Start where he started.
Ford did not rewire the plant. He rewired one component.
In April 1913, the flywheel magneto at Ford's Highland Park plant was built the way almost everything was built at the time. One worker, one complete unit, around twenty minutes.
Ford's team broke that single sub-assembly into twenty-nine operations along a moving belt, twenty-nine men, each fitting a few parts and pushing the flywheel to the next. Assembly time fell to thirteen minutes, then to five.
Only after that did the idea travel. To the engine, then the chassis, then the whole car. The moving assembly line that reorganized manufacturing for a century started as one department's experiment on one component, run long enough to produce a number nobody could argue with.
Call it the magneto principle. You do not need a plan for the second workflow. You need the first one to be provable.
Ninety days is enough for exactly that. Not a transformation. One workflow, rewired properly, with a number at the end of it.

Days 1 to 15: choose it, and write down what you will never allow
Three questions pick the workflow. Is volume rising while the cost of handling it rises in step? Is the decision repeatable while the inputs arrive messy? Is there someone whose week goes into assembling information rather than judging it? Prior authorization qualifies. Supplier qualification qualifies. Your monthly board pack almost certainly does not.
Then do the unglamorous part. Write down what the workflow costs today: cycle time, touches per case, exception rate, rework. This is the number you will be judged against in December, and it cannot be reconstructed later.
Name the reviewer now. Not when the agent is ready. Now, while it is still cheap to argue about. The gap that stops most operations is not technical readiness but the moment everyone realises no one will put their name against an automated decision. Settle it in week two and it never becomes a blocker in week twelve.
Days 16 to 40: instrument the workflow while humans are still doing it
This is the phase everyone skips, and skipping it is why most pilots cannot be defended afterwards.
For four weeks, change nothing about who does the work. Log everything: every decision, every exception, every escalation, every case that took three times longer than it should and the reason why. Your team keeps working exactly as before.
Two things come out of it. The first is a defensible before-and-after, because a baseline measured after you started building is not a baseline. The second is more valuable. That exception log is the specification. The twentieth edge case is never in the process document and always in somebody's head, and four weeks of honest logging is how it gets out.
Days 41 to 70: build it in shadow, where being wrong is free
Now wire the agent into live systems: the ERP, the payer portal, the document store. Real records, not a pasted sample, because the gap between a demo and production is almost entirely a question of where the data comes from.
Then run it in shadow. The agent decides. A person decides. You compare the two, every day, on real cases that are still being handled by the person. Nothing the agent produces reaches a customer, a payer or a supplier.
Shadow mode is the cheapest accuracy data you will ever get, and it turns the eventual go-live conversation from a matter of confidence into a matter of evidence. By day seventy you will know precisely which case types the agent handles better than your team, which it handles worse, and where the line between them sits.
Days 71 to 90: open the gate one notch, not all the way
Promote the narrowest slice where shadow accuracy was highest. If the agent was reliable on standard cases and shaky on anything involving a prior denial, standard cases go live and everything else escalates to the named reviewer. That is not timidity. It is how the audit trail stays clean enough to be worth having.
Then publish the comparison against your day-one baseline, including the parts that came in below what you hoped. An operation that only reports its wins teaches its people to hide the losses, and you will need those people to trust the next ninety days.
Four ways this goes wrong
You pick a workflow with no baseline. If nobody can say what it costs today, nobody can say what changed. Pick a different workflow.
You skip instrumentation to save a month. You will save the month and lose the argument, because you will have a working agent and no evidence.
Scope widens around day forty-five. Someone senior sees the shadow results and asks whether it could also handle a second workflow. It could, later. Ford did the engine after the magneto, not alongside it.
Accuracy becomes the only metric. An agent can be accurate and still slow, expensive, or impossible to audit. Measure cycle time, escalation rate and cost per case as well.
Ninety days from today is 23 November
That is one budget cycle, one quarter, and roughly the same stretch Ford's team spent on a single component before anyone at Highland Park had heard the phrase assembly line.
Your operation will handle several thousand cases between now and then either way. The only question is whether, at the end of it, you have a workflow you can prove got better, or another quarter of the same numbers and a longer list of candidates.
Which workflow would you choose, and what would you have to measure this week to make the comparison honest?
Next week: who is accountable when the agent is right and the outcome is wrong.
Recent blogs
Secure your agents
We’d love to chat with you about how your team can secure and govern Ai agents everywhere







