Agent guide
Running agents in production
An agent that works in a demo can fail quietly in production: looping, calling the wrong tool, or costing more per task than a person. Running agents well means seeing every step, measuring outcomes, and rolling out in stages.
SeeEvery step traced
MeasureTask success, cost
Roll outRead only first
ReviewWeekly at first
Roll out in phases
| Phase | What the agent may do | Move on when |
|---|---|---|
| 1. Shadow | Suggests actions; people do them | Suggestions are right most of the time |
| 2. Read and draft | Reads systems, drafts tickets and replies for people to send | Drafts need few edits |
| 3. Low-risk actions | Performs reversible, low-risk changes on its own | No harmful mistakes over an agreed period |
| 4. Approved actions | Prepares high-risk changes for one-click approval | Approvers trust the preparation |
What to measure
- Task success: did the request actually get resolved, judged by the requester or a reviewer.
- Steps and tool calls per task: rising numbers often mean the agent is confused.
- Escalation rate: how often it hands over to a person, and why.
- Cost per task: model usage and tool calls, compared with doing it by hand.
- Blocked actions: attempts stopped by guardrails, each one worth a look.
Typical tools
| Area | Common choices |
|---|---|
| Agent frameworks | LangGraph, Microsoft Semantic Kernel, vendor agent SDKs |
| Tool connections | MCP servers |
| Tracing | Langfuse, OpenTelemetry |
| Evaluation | Recorded test tasks re-run on every change |
More agent guides: MCP architecture · Guardrails and approvals · Agent FAQ · Use case: IT service desk agent
Thinking about AI agents?
Tell us the process you would like to automate and the systems it touches. We will come back with a plain view of what an agent could safely do, what should stay with people, and how to start small.