Home / Agentic AI / Running agents
Agent guide

Running agents in production

An agent that works in a demo can fail quietly in production: looping, calling the wrong tool, or costing more per task than a person. Running agents well means seeing every step, measuring outcomes, and rolling out in stages.

SeeEvery step traced
MeasureTask success, cost
Roll outRead only first
ReviewWeekly at first

Roll out in phases

PhaseWhat the agent may doMove on when
1. ShadowSuggests actions; people do themSuggestions are right most of the time
2. Read and draftReads systems, drafts tickets and replies for people to sendDrafts need few edits
3. Low-risk actionsPerforms reversible, low-risk changes on its ownNo harmful mistakes over an agreed period
4. Approved actionsPrepares high-risk changes for one-click approvalApprovers trust the preparation

What to measure

  • Task success: did the request actually get resolved, judged by the requester or a reviewer.
  • Steps and tool calls per task: rising numbers often mean the agent is confused.
  • Escalation rate: how often it hands over to a person, and why.
  • Cost per task: model usage and tool calls, compared with doing it by hand.
  • Blocked actions: attempts stopped by guardrails, each one worth a look.

Typical tools

AreaCommon choices
Agent frameworksLangGraph, Microsoft Semantic Kernel, vendor agent SDKs
Tool connectionsMCP servers
TracingLangfuse, OpenTelemetry
EvaluationRecorded test tasks re-run on every change

Thinking about AI agents?

Tell us the process you would like to automate and the systems it touches. We will come back with a plain view of what an agent could safely do, what should stay with people, and how to start small.