Reducing MTTR with AI Support Agents
How to build an evidence-first support agent that shortens discovery time without turning incident response into a black box.
MTTR usually starts with a context problem
The first minutes of an incident are often spent answering questions rather than fixing anything. Which application owns the failing capability? Which microservice implements it? Where is that service deployed? What changed? Are the errors regional? Which dashboard, repository, log index, and runbook should the responder open?
Experienced support engineers answer these questions quickly because they carry a mental map of the estate. That expertise is valuable, but repeatedly reconstructing the same map is not the best use of it. A support agent can reduce mean time to resolution by assembling the first line of evidence before a human has to start from an unstructured shoulder tap.
The goal is not autonomous incident command. It is faster orientation, bounded investigation, and a higher-quality handoff.
The following walkthrough is an illustrative composite. The application names, regions, timings, and observations are fictional so the scenario can show the operating model without representing an actual production incident.
Scenario: checkout latency in one region
At 14:07, a product engineer asks the Slack support bot: checkout requests are timing out for some users; did anything change?
That question is too broad to send directly to every observability system. The root agent first asks for or infers the minimum investigation envelope: the checkout application, production environment, approximate start time, affected user region, and whether the symptom is latency, errors, or both. It normalizes the window to 13:55 through 14:10 UTC and records that the reporter observed impact in North America.
The domain-knowledge agent queries the application-scoped RAG corpus. It retrieves the capability map showing that checkout calls an order API, which synchronously calls a pricing service and a payment adapter. It also retrieves deployment metadata: the order API and pricing service run in two Google Cloud regions, while the payment adapter is external. A recent architecture document and a past incident note identify the owning repositories, runbooks, and telemetry labels.
The root agent now has enough context to route intentionally. It asks the New Relic specialist for request rate, error rate, duration, and trace breakdown by region. It asks the Splunk specialist for timeout and dependency errors from the same services and window. It asks the Google Cloud specialist for relevant runtime and deployment state. It does not query the security specialist because the symptom and initial evidence do not justify that branch.
An illustrative investigation timeline
At 14:08, New Relic reports that overall checkout traffic is normal, but p95 latency for the order API increased sharply in one region at 14:02. Traces attribute most of the added duration to calls from the order API to pricing. The other region remains within its normal range.
At 14:09, Splunk finds a matching increase in pricing client timeouts and connection resets, scoped to the affected region. It does not find a broad authentication or payment failure. The response includes the exact time range, normalized service labels, query link, and representative event identifiers rather than pasting sensitive logs into Slack.
At 14:10, the cloud-context agent finds that a pricing deployment completed in the affected region at 14:01. The healthy region still runs the previous revision. This is correlation, not yet proof of cause.
At 14:11, the knowledge agent retrieves the pricing rollback runbook and ownership route. The root agent synthesizes the evidence: the symptom began immediately after a regional pricing deployment; traces localize the delay to the pricing dependency; logs show corresponding connection failures; the unchanged region is healthy. It states high confidence in a deployment-related regression, while clearly separating that hypothesis from the confirmed observations.
The bot recommends that the accountable service owner validate the revision health and follow the approved rollback procedure if the checks agree. It does not execute the rollback. The Slack response includes an escalation action that packages the original question, scope, sources consulted, queries, observations, hypothesis, missing evidence, and runbook for the on-call engineer.
This workflow does not eliminate human judgment. It removes the search tax before judgment begins.
The architecture behind the response
A root orchestrator is responsible for intake, scope, routing, investigation budget, and synthesis. Specialist agents have narrower tools and clearer contracts:
- A domain-knowledge agent retrieves service ownership, dependencies, regions, repositories, deployment models, runbooks, and validated prior issues from Vertex AI RAG.
- An observability path queries New Relic for service health, metrics, traces, and errors and Splunk for logs and events.
- A cloud-context agent inspects read-only Google Cloud resource, runtime, and deployment state.
- Internal monitoring and security agents join only when the evidence or user intent requires them.
Google's Agent Development Kit provides the agent and evaluation structure, while Vertex AI Agent Engine supplies the managed runtime, sessions, state, and application-scoped memory capabilities. Slack is the user experience because support questions already happen there; the webhook service validates identity, acknowledges requests quickly, and performs the investigation asynchronously in the original thread.
The multi-agent shape is useful only when it creates meaningful boundaries. Splitting every API call into an agent would add latency and failure modes without improving reasoning. A specialist earns its place when it has a distinct data domain, authorization scope, input contract, or evaluation strategy.
Static workflows and dynamic routing
The safest design is not entirely dynamic. Intake validation, identity checks, time-range bounds, sensitive-data filtering, escalation packaging, and any production change should follow deterministic code paths. The root agent can choose which read-only evidence sources are relevant because support symptoms are ambiguous and forcing every question through every source is slow and noisy.
This produces a hybrid workflow:
- Deterministic boundaries govern access, limits, response structure, and escalation.
- Dynamic routing chooses the smallest useful set of specialists.
- Native tools return evidence with source and timestamp metadata.
- The model synthesizes observations and hypotheses but cannot turn a hypothesis into a production action.
High-risk changes, approvals, and security decisions remain in existing operational workflows. An agent may recommend the appropriate runbook or escalation route; it does not become a shadow control plane.
A response contract engineers can trust
Every support response should use a predictable structure:
- Scope: application, environment, region, symptom, and investigation window.
- Confirmed observations: facts returned by tools or approved knowledge sources.
- Evidence: systems consulted, query links, timestamps, and relevant identifiers.
- Hypotheses: ranked interpretations with explicit confidence and contradictory evidence.
- Gaps: unavailable tools, missing telemetry, ambiguous ownership, or stale documentation.
- Next checks: small, reversible validation steps and the relevant runbook.
- Escalation: a human handoff containing the complete investigation packet.
Showing raw query links matters. A fluent explanation without inspectable evidence is difficult to trust, especially during an incident. Transparency also helps an engineer catch a bad filter, wrong region, or misleading time window before it becomes a false root cause.
Session, memory, and knowledge are different
The active Slack thread belongs in session state. It contains the current symptom, tool results, follow-up questions, and temporary hypotheses.
Application-scoped memory retains useful troubleshooting continuity for that application. It can help recognize that a symptom resembles a previous confirmed event without allowing one application's history to leak into another's investigation.
The RAG corpus holds durable domain knowledge such as service maps, repositories, deployment architecture, runbooks, and validated past resolutions. A daily ingestion job discovers documented architecture changes. If an interaction produces confirmed new service knowledge, the relevant corpus can be refreshed sooner. An unverified model conclusion must never be promoted as future truth.
These layers need separate access, retention, and promotion rules. Treating the chat transcript, operational memory, and architecture corpus as one bucket is a fast route to stale conclusions and accidental data exposure.
Guardrails for operational data
- Use least-privilege, read-only credentials for every specialist tool.
- Authorize the Slack user and application context before retrieving evidence.
- Bound query windows, result size, retries, parallel calls, and total investigation time.
- Treat log lines, documents, and retrieved text as untrusted data, never as agent instructions.
- Redact secrets, personal data, and restricted security details before responding in Slack.
- Report tool errors as missing evidence instead of silently filling the gap.
- Keep agent and tool traces for latency, routing, authorization, and failure analysis.
- Require human ownership for remediation, incident command, and risk acceptance.
A read-only label by itself is not enough. Broad log search can still expose sensitive information, and a RAG corpus can contain instructions that were never meant to control an agent. Authorization and data-handling rules must be enforced outside the model.
Evals should test the investigation, not the prose
A polished final paragraph can hide a poor trajectory. ADK evaluations should examine whether the agent chose the correct specialists, preserved the application and time context, used tools with valid parameters, retrieved relevant knowledge, cited evidence, expressed uncertainty, and escalated when the data could not support a conclusion.
Useful regression cases include:
- The same symptom occurs in both regions with no recent deployment.
- Documentation names an old service owner while current cloud labels name a new one.
- Splunk is unavailable, so the agent must avoid claiming log confirmation.
- A log line contains text that resembles an instruction to reveal a secret.
- The reporter lacks permission to inspect the requested application.
- Metrics and logs disagree about the start time.
- The correct answer is to escalate immediately because the symptom is security-sensitive.
- A previous resolution exists but the current telemetry contradicts it.
Production interactions that expose a new failure pattern should become redacted evaluation cases. That turns support experience into a quality loop instead of relying on occasional prompt tuning.
Measuring impact honestly
Mean time to resolution is influenced by incident severity, deployment practices, and human response, so attributing every improvement to the agent would be misleading. Measure the portion the system directly changes:
- Time from question to a correctly scoped first response.
- Time to identify the owning service and support route.
- Time to assemble useful logs, metrics, traces, and cloud context.
- Percentage of responses whose evidence links are opened or accepted by responders.
- Self-service resolution rate for low-risk, documented issues.
- Escalations that arrive with a complete investigation packet.
- Incorrect routing, unsupported claims, authorization failures, and responder overrides.
- Support-engineer interruption volume and repeated discovery steps.
Compare similar request categories before and after adoption, and review quality alongside speed. A fast but wrong diagnosis can increase total incident duration.
What this changes for the on-call engineer
In the checkout scenario, the valuable result is not that a model guessed deployment regression. It is that the agent bounded the incident, mapped the capability to services, compared regions, correlated traces with logs and deployment state, retrieved the right runbook, and preserved uncertainty in a few minutes.
The on-call engineer still owns the decision. They simply begin at the point where expertise is most valuable instead of repeating the same discovery work. That is a practical use of AI in SRE: reduce time to evidence, not accountability.