In short. Your teams have built agents in Claude, in Copilot, in whatever was to hand, and most of them work. Ask what one costs per case, what it decided last Tuesday and why, or who last changed its instructions, and the room goes quiet. The agents are not the problem. The layer above them is missing, and adding it does not mean rebuilding any of them.
This is what success looks like early, and it is worth understanding why
Agents get built where the work is, by the people closest to it, in whatever environment they already had access to. That is not shadow IT and it is not a governance failure. It is the correct response to a capable tool arriving in the hands of people with real problems.
The trouble is what happens next. Six agents becomes twenty. Someone tunes a prompt on a Friday and nobody records it. An agent that was accurate in March drifts in July because the underlying model was upgraded and nothing was re-tested. Spend arrives as one invoice with no way to attribute it to a workflow, so finance treats the whole thing as an experiment rather than an operating cost.
Call it the invisible fleet. Each agent is individually fine. Collectively they are unmanaged, and you cannot govern what you cannot see.
Four questions decide whether you have a fleet or a collection
The question
What did this agent cost per case last month?
What did it decide, and can you reconstruct why?
Who last changed its instructions, and was that reviewed?
What happens when the model underneath is upgraded?
Why it decides things
Until spend attaches to a workflow, no agent has a defensible business case
A log that cannot reconstruct reasoning is storage. Evidence is what survives a dispute
Prompt and skill changes alter behavior as much as code does, and are usually ungoverned
Behavior can move without anyone changing anything. Somebody has to notice
An organization that can answer all four has a governed operation. One that cannot has a promising collection of experiments, whatever the demo looked like.
The layer you add, and the thing you do not touch
The instinct at this point is to consolidate: pick one enterprise ai platform, rebuild everything on it, retire the rest. It is a reasonable instinct and it is usually the wrong move. Rebuilding working agents costs months, breaks the trust of the teams who built them, and delivers no new capability at the end of it.
The alternative is to leave the agents where they are and add the layer above them.

Your agents keep running where they were built.
Instructions Manager takes the prompts and skills out of wherever they currently live and puts them under version control, with review before a change goes live and the approved version served at runtime. Nobody redeploys an agent to change how it behaves, and every change has an author and a date. This is also where your domain knowledge accumulates as an asset you own rather than as text pasted into someone's editor.
Guardrails run policy as code before and after the model: PHI and PII handling, which tools an agent may call, token budgets, off-topic refusal. Before the action, not as an alert afterwards, because a control that reports a breach is not the same as one that prevents it.
elsai observe gives you traces, cost, latency and token consumption per project, per workflow and per agent, on OpenTelemetry. It is what turns the four questions above from awkward into routine.
Workflow is where agents and human gates get composed into something that resembles an operation, with a named reviewer at each point where the cost of being wrong exceeds the cost of a pause.
Agents built with elsai Agentkit run under the same layer. So do agents you built somewhere else. The control plane does not care who wrote the agent, which is the entire point.
What this looked like in practice
One organization we work with had gone a long way with Claude. Real agents, real usage, built by capable people, spread across teams. What they did not have was a single place to see any of it.
We put a control plane over the top without disturbing a single existing agent. Nothing was rebuilt and nothing was migrated. Once the instructions were organized and the traces were flowing, two things happened. The agents performed better, because inconsistent instructions had been quietly costing accuracy. And the team could finally see their token capital, their accuracy and their false-positive rate as operating numbers rather than as impressions.
The visibility did not arrive after the governance work. It was the governance work.
Where to start, and it is smaller than you think
Inventory first. One page listing every agent you know of, who owns it, what it touches and roughly what it costs. The gaps in that page are the finding.
Then take the two or three agents that touch regulated data or customer-facing decisions and put those under the layer first. Instructions under version control, a guardrail on the data they may handle, traces flowing. Leave the rest alone until that works.
Ninety days is enough for that, and it does not require anyone to stop building.
If your model provider shipped an upgrade tomorrow, who in your organization would know whether your agents still behave the same way?
Next in this series: why policy that runs before the action is the only kind that counts. And if the operational question on your mind is what this looks like on a real workflow, that is covered in Your Prior Authorization Queue Has a Deadline Your EHR Doesn't Know About.
Recent blogs
Secure your agents
We’d love to chat with you about how your team can secure and govern Ai agents everywhere







