Build an Incident Triage Agent Architecture
Design a multi-agent incident-triage workflow that enriches incidents with monitoring evidence, assigns work to the right teams, and keeps high-risk decisions reviewable by humans.
Before You Build: What This Tutorial Verifies
This tutorial explains a verified architecture for AI-assisted incident triage based on the Triangle research described in Microsoft Research publications. It does not claim that the supplied sources verify a particular Claude model, Anthropic SDK version, FastAPI application, SQLite schema, cloud integration, or production deployment. Those are implementation choices that must be validated separately against the provider and framework documentation before code is written.
The verified problem is operationally important. Incident triage means quickly and accurately assigning an incident to the appropriate team so that mitigation can begin. In a large cloud-service environment, poor triage can increase Time to Engage, or TTE. A longer TTE can delay mitigation and damage service quality and customer trust. The difficulty is not limited to labeling a ticket: engineers must inspect incident details, use several tools, collaborate with multiple teams, and change a decision when new evidence appears.
Triangle addresses this difficulty with a multi-agent design intended to emulate collaborative reasoning among effective human expert teams. A particularly important verified capability is Team Manager information enrichment. The Team Manager extracts a relevant time range and component names from the incident, queries the monitoring database associated with its team, and summarizes discussions using both the incident and related Monitor Logs.
That distinction governs the design below. The agent is not an unconstrained chatbot and should not be described as an autonomous operator. It is an evidence-enrichment and routing workflow. Its output should help an authorized responder decide which team needs to engage, what evidence supports that assignment, and where uncertainty remains.
Prerequisites
- A clearly defined incident record containing the problem description and any available component or time information.
- Access to the monitoring data that the responsible team is authorized to query.
- A documented team directory or routing catalogue that maps components and services to responsible teams.
- An agreed human review process for uncertain, high-impact, or conflicting triage recommendations.
- A method for recording the incident, evidence retrieved, proposed assignment, reviewer decision, and subsequent changes.
The sources establish the operational need and the information-enrichment pattern, but they do not prescribe a particular programming language, web framework, database, model provider, authentication system, or observability vendor. Keep those decisions explicit in your own design documentation rather than presenting them as properties of Triangle.
Step 1: Define the Triage Objective and the Assignment Contract
Begin with the outcome that the workflow must support: assign an incident to the appropriate team quickly and accurately enough to reduce avoidable delay. Do not begin by asking which model or framework to use. A model call is only one part of a triage system, while the assignment contract determines what information the rest of the system must preserve.
At minimum, define the following logical fields:
- Incident identity: a stable identifier and the original incident text.
- Observed scope: the named service, component, or system area, when available.
- Relevant time range: the reported start time and any bounded investigation interval that can be derived from the incident.
- Candidate teams: teams that could plausibly investigate or mitigate the issue.
- Evidence: the incident facts and the monitoring records used to support the recommendation.
- Recommendation: the proposed destination team and the reasoning that connects evidence to the assignment.
- Uncertainty: missing, contradictory, stale, or ambiguous information.
- Review state: whether the result is proposed, accepted, rejected, or awaiting clarification.
The important design choice is to separate observed facts from the recommendation. An incident may state that requests are failing, while monitoring records may show several affected components. The system should preserve both inputs instead of flattening them into an unexplained label. This makes later review possible and prevents a generated summary from becoming the only surviving representation of the evidence.
Define what “appropriate team” means for your organization. It may mean the team that owns the affected component, the team responsible for the relevant monitoring signal, or the team currently able to mitigate the failure. These are not always the same. If the organization has no routing policy, the agent cannot reliably invent one; it can only expose the ambiguity for a human decision.
Step 2: Separate the Multi-Agent Responsibilities
Triangle is relevant because it treats incident triage as a collaborative reasoning problem rather than a single classification step. Use that idea to assign narrow responsibilities to separate logical roles. The exact number and names of agents are implementation decisions; the division of responsibilities is the more important architectural principle.
A useful design begins with an incident interpretation role. It reads the incident...
Continue Reading
Log in for free to read the rest of this article and access exclusive AI tools.
Log in / Register