Alert investigation — metrics
Load this file when investigating a metric-based alert (aggregated metric threshold, custom metric anomaly, composite alert).
What the alert context carries
Metric alerts carry:
alertID,alertName,alertValue— threshold and valuegroup,groupValue— the dimension (if grouped)query— the metric querythresholdWindow— hold durationmetricsLink— deep linktimeRange— alert window
Investigation shape
- Replay the alert's metric query using the
query-aggregationstool, sameproduct_type,query, and time window as the alert. Confirm you see the same value that breached. - Expand the time window just enough to see onset — if the threshold was crossed at T, pull
query-aggregationsfromT - 2 * thresholdWindowtoTto see the rise. - Drill into the underlying records. A metric alert on
errorspoints to specific error groups; onsessionsto specific sessions; ontracesto specific traces. Pull the relevant record-level data for the top contributing dimension. - Check external triggers. Deploy correlation, flag flip, traffic surge, dependency issue.
What goes in the diagnosis
- What triggered — alert name, threshold, alertValue, and (if grouped) which dimension value crossed it.
- Likely cause — specific underlying records (error group IDs, trace IDs, session IDs) that drove the metric up. Quote one or two.
- Scope — affected count, time-window pattern, whether this is a sharp spike or a drift.
- Next steps — concrete actions per the root cause found; if the metric reflects underlying behavior that needs fixing, point to the record-level remediation.
Common mistakes
- Only reporting the metric value without drilling into record-level evidence. The metric is a summary; the cause lives in the records.
- Mis-attributing a spike to a specific cause without correlating with a deploy/flag/dependency timeline.
- Querying too-wide a time range and washing out the signal. Stay near the alert window.