Infrastructure Assist Agent
A production multi-agent support system that combines application knowledge, logs, metrics, cloud state, and security signals to provide first-line troubleshooting and evidence-rich human escalation in Slack.
Problem
Platform and application engineers were repeatedly interrupted by shoulder taps for first-line production support. Answering even a familiar question meant reconstructing application ownership, service dependencies, deployment regions, code locations, previous incidents, logs, metrics, and current cloud state across several systems.
The same discovery work was repeated before an engineer could determine whether the issue had a known resolution, required deeper investigation, or belonged with another support team.
What I built
I built a production support agent with Google's Agent Development Kit and deployed it to Vertex AI Agent Engine. A root agent coordinates specialized sub-agents that retrieve application knowledge and inspect read-only operational data from Splunk, New Relic, Google Cloud, and internal monitoring and security tools.
The system is delivered through Slack. Consumers describe a production symptom in a familiar support workflow, receive an evidence-linked troubleshooting response, and can escalate the investigation to a support engineer from every bot response when the agent cannot reach a well-supported resolution.
System architecture
The agent performs first-line investigation with read-only tools. It does not execute remediation or replace the accountable support engineer.
Root and specialist agents
The root agent turns an open-ended support question into a bounded investigation. It establishes application, environment, region, and time-window context; chooses the specialists needed for the symptom; and combines their results into one response.
- The domain-knowledge agent queries Vertex AI RAG Engine for service responsibilities, application behavior, deployment documentation, code locations, runbooks, and relevant previous incidents.
- The Splunk agent runs scoped searches for logs and events associated with the identified services and time window.
- The New Relic agent reviews metrics, traces, service health, errors, and reliability signals.
- The Google Cloud agent gathers current resource and runtime context for the affected application.
- Internal monitoring and security agents contribute approved reliability or security evidence when the investigation requires it.
Each tool is read-only, bounded to the requested scope, and returns timestamps and source context that the root agent can preserve in its response.
Knowledge, session, and memory model
Vertex AI RAG Engine holds application and infrastructure domain knowledge: which microservice owns a capability, where it runs, how it is deployed, where its code lives, and which past incidents or runbooks may apply.
The RAG corpus is updated after an interaction when the investigation produces confirmed new service knowledge. A daily ingestion process also discovers documented architecture changes and refreshes the relevant application content. Only validated knowledge and confirmed resolutions are promoted; an agent hypothesis is not written back as future truth.
Agent Engine session state contains the active Slack investigation. Memory Bank is scoped per application and retains relevant troubleshooting history across interactions. This separation lets the agent recognize recurring application patterns without mixing unrelated application context or treating transient conversation state as durable architecture knowledge.
Investigation response contract
Every response is structured to help a consumer act and to make escalation efficient:
- Application, environment, region, and investigation time window.
- Confirmed observations from retrieved knowledge and operational tools.
- Evidence links, query timestamps, and the systems consulted.
- Leading hypotheses with uncertainty rather than unsupported root-cause claims.
- Missing or unavailable evidence and tools that did not return successfully.
- Recommended validation or troubleshooting steps.
- A Slack action to escalate the complete investigation to a support engineer.
When a security signal requires restricted handling, the response avoids exposing sensitive detail in a broadly visible Slack channel and directs the escalation to the appropriate support path.
Slack escalation workflow
The webhook service validates the Slack request, acknowledges it quickly, and runs the multi-tool investigation asynchronously in the original thread. Request identifiers make Slack retries idempotent so one question does not start duplicate investigations.
If the agent cannot reach a sufficiently grounded resolution, the user can escalate from the response. The handoff includes the original symptom, normalized incident context, specialists invoked, searches performed, evidence links, observations, hypotheses, missing data, and the conversation summary. Support engineers begin with the first-pass investigation already assembled instead of repeating discovery.
Security and operational guardrails
- Slack identity and application context are evaluated before querying operational systems.
- Specialist agents use least-privilege, read-only credentials and approved query boundaries.
- Log ranges, result sizes, retries, tool calls, and total investigation time are bounded.
- Retrieved documents and telemetry are treated as untrusted evidence rather than agent instructions.
- Sensitive log fields and security details are filtered before they are returned to Slack.
- Tool failures are reported as missing evidence instead of being silently converted into confident conclusions.
- Agent and tool traces support investigation of latency, failures, routing decisions, and unexpected behavior.
Evaluation strategy
I built ADK evaluations around the full investigation trajectory, not only the final wording. The evaluation set checks whether the root agent selects the right specialists, passes the right application and time context, retrieves relevant knowledge, uses operational tools correctly, cites evidence, communicates uncertainty, and escalates when the available data cannot support a resolution.
Regression scenarios include stale or conflicting documentation, incomplete telemetry, tool outages, ambiguous symptoms, unauthorized requests, prompt-like text inside logs or documents, and incidents where human escalation is the correct outcome. Difficult production interactions can be converted into new evaluation cases as the support surface grows.
Design tradeoffs
Specialized agents make tool boundaries and responsibilities clearer, but they add orchestration latency and cost. The root agent therefore routes only to relevant specialists and operates within a bounded investigation budget.
Application-scoped memory improves continuity, but troubleshooting history can contain stale conclusions or sensitive evidence. Session state, durable application memory, and validated RAG knowledge are kept as separate layers so each can have appropriate access, retention, and promotion rules.
The agent is intentionally a diagnostic and escalation system rather than an autonomous remediation service. This limits automation depth, but keeps production changes with accountable engineers and established operational controls.
Impact
The system converted repeated shoulder taps into Slack-based self-service first-line support. Questions that previously required an engineer to search application documentation, Splunk, New Relic, Google Cloud, and internal systems now begin with an agent-led investigation across those same sources.
Consumers receive faster initial evidence and relevant next steps. When human expertise is still required, the engineer receives an assembled investigation instead of starting from an unstructured message. This reduces support interruption and repeated discovery while preserving a clear path to accountable human support.