Building a CloudWatch Omni Dashboard
Use this when building, saving, reading back, or fixing a CloudWatch Omni dashboard
— "build me a dashboard for checkout", "show what's most critical for this
service", "lay this out", "which chart for this signal?", "why is that panel
empty?", "why does that field have no effect after save?", "why was my save rejected?" —
and for the dashboard API itself
(CreateOmniDashboard and its siblings, aws cloudwatchomni *-omni-dashboard*).
A CloudWatch Omni dashboard is a set of panels laid out on a fixed 60-column
grid, most of them charts over CloudWatch Omni queries. The dashboard body is
JSON — a panels[] array (plus an optional scope) that the Omni UI renders. It
belongs to a Space and is stored, as a JSON string, in the body field of a
dashboard resource you create and update through the API (section 7). The
deliverable is therefore two things: a valid, grounded body, and the API call that
saves it.
Contents
- Omni dashboard vs. CloudWatch dashboard
- Decide what the dashboard shows
- Ground every panel query
- Author the panel bodies
- Lay out on the 60-column grid
- Troubleshooting
- API reference
- Appendix — Archetype templates
Two rules hold across every step:
- Never chart a metric, field, or label you have not confirmed exists. An invented name renders an empty panel, which on a saved dashboard reads like an outage — worse than no panel. Ground first (section 3); drop what you cannot ground.
- Query text follows the CloudWatch Omni query dialect. The SQL rules for logs and traces live in sql-logs-traces.md; metric semantics and PromQL selectors live in promql-metrics.md. This reference owns the dashboard body — the panel schema, chart choices, grid math — and the dashboard API. It does not restate the dialects.
ALWAYS STATE — how a dashboard body fails. It fails in several ways, and only one of them errors — the rest fail silently, so call the right one out whenever you author, explain, or debug a body (full detail in "API save semantics", section 4):
- A structural or key error is REJECTED at save (HTTP 400
ValidationException). The root and panel objects are schema-validated (additionalPropertiesoff), so this covers an unknown or misspelled key there (a top-levelrefreshInterval, a per-panelid, a misspelledtitle), a badvariantenum, a missing or non-arraypanels, a panel missing its requiredtypeorlayout, and an out-of-bounds coordinate FIELD (x/ybelow 0,wbelow 1). The save does NOT succeed — none of these is silently dropped. Validate the structure and every root/panel key against section 4. - A bad VALUE saves with HTTP 200, then blanks at render. The schema does not check
cross-field invariants, and the only enum VALUE it validates is
variant(a bad one is rejected at 400 — previous bullet);visualization, paneltype, andalertsstatevalues are unchecked. An unknownvisualizationVALUE blanks THAT panel while its siblings still render. Everything else it does not catch blanks the WHOLE dashboard (every sibling included): the cross-fieldx + w > 60overflow (per-field bounds are validated, their sum is not), aw > 60, a badhunit (80vh), an unrecognised paneltypeVALUE, a badalertsstate, or anow-30s/now-1ytime bound. Either way there is no error, and a200is not proof of renderability; nothing enforces thex + w <= 60sum. - A wrong key INSIDE
configis accepted and silently ignored.config(and its nested objects likechartOptions) is permissive: a misspelled or unknown key there is stored, comes back on read, and never takes effect. This — not root/panel keys — is where "a field I set doesn't take effect, with no error" actually happens. queryLanguagemust be set explicitly per panel; a mismatch fails silently. Set eachexplorepanel'sconfig.queryLanguageto'sql'or'promql'yourself — it is not inferred from the query text. (Metrics →promql; log/trace content →sql.)
Work sections 2–5 in order — each depends on the one before — then save with the API in section 7.
1. Omni dashboard vs. CloudWatch dashboard
Both are called "a CloudWatch dashboard" in ordinary speech, but they are different resources with different bodies, different APIs, and different query languages:
| CloudWatch Omni dashboard (this file) | CloudWatch dashboard | |
|---|---|---|
| Lives in | An Omni Space (spaceId on every call) |
The AWS account/Region |
| API / CLI | CreateOmniDashboard etc. — aws cloudwatchomni … |
PutDashboard etc. — aws cloudwatch … |
| Body | { "panels": [...], "scope": {...} } — panel type, config, layout {x,y,w,h} |
{ "widgets": [...] } — widget type, properties, x/y/width/height |
| Grid | 60 columns; y/h in pixels |
24 columns; all four in grid units |
| Panel queries | Omni SQL (logs/traces) and PromQL (metrics), one query per panel | CloudWatch metric definitions, Metrics Insights, Logs Insights |
| Identity | dashboardId (system-minted UUID); the name is metadata |
The dashboard name |
The two bodies are not interchangeable — a CloudWatch widgets[] body pasted into
CreateOmniDashboard is rejected, and an Omni panels[] body means nothing to
PutDashboard. CloudWatch dashboards are documented in
../cloudwatch/dashboards.md.
Confirm Omni is enabled before acting. The Omni dashboard API is only reachable
once a Space exists in the target Region. Before you create, update, or delete a
dashboard, probe with aws cloudwatchomni list-spaces in that Region. If there is
no Space, say so plainly and route the same need to the CloudWatch path — that is the
correct fallback, not a wrong answer. A knowledge question ("does Omni have a
dashboard API?", "what's the panel schema?") is answered from this file directly;
do not gate it behind the probe. See concepts.md for what a Space is.
Facts you MUST surface when building or saving an Omni dashboard
Author and ground the panels, and surface the items below that are relevant to the task —
each is a save-time or render-time trap the API does not fully catch (only structural,
root/panel key, coordinate-field, and variant errors are loud; the rest are silent).
When the task is choosing a panel's chart type, surface the visualization and panel-shape
items; when saving or reading a dashboard back, surface the save-boundary behavior for
their config. Each item points to the section that details it; state what fits the task,
not a re-paste of every detail.
- Pick a
visualizationvalue from §4's supported set deliberately, there is no safe default. Any value outside that set, plus a missing or empty value, saves200and silently blanks that one panel at render, with no fallback that draws. visualizationandqueryLanguageare both REQUIRED on everyexplore(data) panel, and aqueryLanguagemismatch is a silent blank rather than an error. They areconfigfields of theexplorepanel type only. The non-explore panel types (such asmarkdown,divider,alerts,spans, andapplication-map) carry neither. Set eachexplorepanel'sconfig.queryLanguagetosqlorpromqlyourself. It is never inferred from the query text (§4 preamble, §4 "explore", §4 "API save semantics").- The two panel shapes authors most often mis-pick: a single scalar aggregate →
number, nottable; a top-N ranking (ORDER BY … LIMIT/topk) →table, notline(§4, "Visualizations"). - The
number-tileCOUNT(*)product bug. ACOUNT(*)-shaped scalar in anumberpanel renders the result-set ROW COUNT (usually1), not the count value; put theCOUNT(*)in atablepanel instead, and for any other scalar aggregate make it the first selected column so thenumbertile reads the value rather than the row count (§4, "Visualizations", "numbertile trap"). - The save boundary — where the API is loud vs silent. SILENT (saves HTTP 200,
fails only at render): an unknown key INSIDE
configis stored, round-trips on read, and is silently ignored; and a cross-fieldx + w > 60overflow blanks the whole canvas. LOUD (REJECTED at save with HTTP 400ValidationException): a structural, top-level or per-panel key, per-field coordinate-BOUND (negativex/y,w < 1), orvariantviolation. AndGetOmniDashboardreturns the body byte-for-byte (§4, "API save semantics").
2. Decide what the dashboard shows
Pick the panels, order them into sections, and choose a visualization per signal.
Start focused — a tight dashboard beats an exhaustive one
- The default is a FOCUSED ~6–8 panels: roughly 4 KPI tiles + 2–3 golden-signal trends + at most one breakdown table. A single-signal dashboard is smaller still (3–6 panels). Don't emit dozens of panels for a vague ask.
- But build the COMPLETE core — never collapse to one row. A real dashboard reads as summary → trend → breakdown (below): a KPI row AND at least one trend chart AND a breakdown. If some metric-based panels cannot ground, fill those tiers from traces/logs (section 3) rather than shipping a lone KPI row.
- Cap the KPI row at 4 tiles (
numberpanels). Four is also the natural grid fit —15 × 4 = 60columns — so a fifth tile forces an awkward wrap. - Always include a saturation panel in the core when the grounded catalog has one — backlog / queue depth, consumer lag / liveness, or circuit-open / write-denied. This is the incident-visibility panel; the broader "infra / USE health" sweep below remains a deferrable follow-up, but the single saturation signal belongs in the core.
- Build the core, then offer depth as a follow-up "add section." Defer these rather than placing them up front: per-service latency, error / status-code breakdown, infra / USE health, and log signals. Name the one the user most likely wants next instead of pre-building all of them.
Ordering principle — headline first, detail last
Every dashboard reads top-to-bottom as summary → trend → breakdown → raw:
- KPI row — a few
numberpanels with the headline figures (availability, error rate, p99 latency, request rate). The fast "is it healthy?" read. - Trend charts —
line/area/bartime series for the golden signals. - Ranked breakdown — a
table(top-N by the primary signal) so the worst offenders surface. - Raw / drill-down — logs, spans, alerts, or an application map. Optional depth, offered as a follow-up rather than built by default.
Open a logically distinct group with a divider, and open an on-call / incident
dashboard with a short markdown runbook panel.
If the ask names a resource type, start from a template
When the request names a known resource TYPE — an EC2 fleet, a Lambda function, a Kubernetes/EKS workload, an RDS/database instance, an ECS/web service, or an API Gateway API — don't classify from scratch. Start from that archetype's ready panel set in the Appendix, then ground and lay it out.
The panel content for a resource type — which signals, how each is read
(gauge vs. counter, max for a 0/1 check, min for a credit balance), which exist
only under a configuration (an installed agent, a burstable instance type, Container
Insights), and the identity label to group by — is owned by the per-service catalog in
promql-metrics.md (its "Service-health question — facts you
MUST surface" section and section 5). When the ask is which signals or panels a named
resource type needs, state those catalog facts as the proposed panel set; this file then
supplies the body and the grid. Do not decide the content by enumerating resources
through the service's control plane.
Recipes (for a general intent)
- Golden-signals service health (the default when the ask is vague). For
"build a dashboard for
<service>" / "how is<service>doing?". Use RED (rate/errors/duration) for a request-serving service, USE (utilization, saturation, errors) for an infrastructure resource. CORE (~6–8): a KPI row (availability %, error rate, p99 latency, request rate); a request-rateline; an errorsbar(stacked) orarea; a latencyline(p50/p90/p99) when a real percentile source exists (section 3); and a top-Ntable. Build each family from the service's OWN grounded metric first, falling back to a generic OTel metric, then a span / log derivation, only when no higher tier grounds it. - Errors-focused. A headline error-rate
number; a stackedbar/areaof the status-code or fault/error breakdown over time; atableof top error sources. Offer a recent-error log panel and — if the space has alerts — analertspanel filtered toCRITICAL(orWARNING) as a follow-up. - Latency-focused. A p99
number; alineof p50/p90/p99 together; atableof the slowest operations. If the service exposes no percentile-capable source, chart the average with an honest title — never a fake p99. - Throughput-focused. A request/message-rate
number; alineover time grouped by the identity label — read the aggregation off the counter's temporality (sum_over_timefor a delta Sum,rate()/increase()for a cumulative Sum), not a uniformrate(); for a queue, pair inbound vs. outbound counters and add the backlog gauges (depth, age). - Saturation / capacity. Utilization gauges (
solidgaugefor a bounded % against a threshold,linefor a trend): CPU, memory, disk/queue depth, credit balances; atableof the resources closest to their limit. An empty series often means "not applicable," not zero. - Incident review (a fixed past window). Open with a
markdownsummary, set a fixedscope.timeRange(absolute start and end) on the dashboard so every panel covers the incident, then golden-signal charts plus aspanspanel — or adashboardpanel withsource.id: "trace-details"— for the offending path. Find the offending trace with a span query per sql-logs-traces.md and pin itstraceIdin that panel'ssource.params.traceConfig. - Multi-service overview (a fleet or whole space). One compact row per service
— a
numberKPI plus a smallline— under adividerper service, or a single widetableranking all services by health. Prefer atableortopk(...)over alinewith more than ~25 series.
One query per explore panel. To compare two signals, emit two side-by-side
panels (section 5), never one overlaid panel. If the intent is genuinely one
number, a single number panel is a complete answer — don't pad it.
3. Ground every panel query
Ground each data panel's query against live telemetry before you write the body. The query languages themselves are deferred to sql-logs-traces.md and promql-metrics.md; this section is the discipline.
The grounding loop, per data panel:
- Discover what exists. Run discovery queries against the Space — don't guess
service names, field names, or metric names.
- Logs and traces:
EXPLAIN (ANALYZE_FIELDS)orSELECT@record… LIMIT 10, narrowed to the service and telemetry type (see Schema Discovery in sql-logs-traces.md). Interactive discovery queries need a`@timestamp`bound — only the final panel query is windowless (below). - Metrics: enumerate the real metric names and their identity labels per promql-metrics.md — which metrics a service emits, each metric's instrument type (Sum / Gauge / Histogram) and temporality (delta / cumulative), and the labels that scope it.
- Follow a service's dependencies (the services it calls, from span data) to find neighbors worth a panel.
- Logs and traces:
- List the real names. For each signal a panel will chart, list the actual
metric / field / label names for that data set and signal kind (LOGS, TRACES,
METRICS). Use those names exactly — don't shorten, alias, pluralize, or
invent them. Nouns in the request ("throttles", "5xx", "timeouts") describe
what to aggregate, not field names — map them to a real grounded name.
Metric names are emitter/OTel-specific — they are NOT Prometheus
conventions. Do not assume
requests_total,http_requests_total,*_total,*_seconds,*_count, or*_bucket; a PromQL selector MUST name a metric from the grounded list verbatim. If the list lacks the metric you wanted, it does not exist here — use a fallback surface (step 4), never a guessed name. Pick the SOURCE per signal by a strict 3-tier priority — reach for a lower tier only when no higher tier grounds the signal: (1) the service's OWN domain metric it emits (a business counter / gauge / histogram) FIRST; (2) a generic OTel metric (e.g.Count/Errors) as a fallback; (3) a trace-span / log derivation (spandurationNanopercentiles, log counts) LAST. - Pick the surface and set
queryLanguageto match. Metrics → PromQL; log or trace content → SQL. Metrics execute as NATIVE PromQL (queryLanguage: "promql"), NEVER SQL — there is nometrics.defaultSQL table; a SQLFROMis onlydefault,logs.default, ortraces.default. Set the panel'squeryLanguageexplicitly — it is not inferred from the query text, and a mismatch is a silent failure. - Fall back to a lower tier before dropping. The same RED signal can come from more than one place: request rate, error rate, and latency can be derived from TRACES/spans (SQL) — counts, error status, duration percentiles — and error/volume counts from LOGS (SQL), not only from METRICS. If the metric you wanted is not grounded, re-express the SAME panel from a lower-tier surface rather than dropping it. A service with sparse metrics but rich traces/logs should still yield a FULL dashboard.
- Self-check the draft. Scan every field / metric / label reference in the
drafted
queryStringagainst the grounded lists. If a reference cannot be confirmed, re-ground; if it still cannot — on any surface — drop that panel rather than shipping an invented query. Optionally run the final windowless query once with a temporary time bound added, to confirm it returns rows.
Bind panels to the grounded metric identity. Vended AWS metrics carry an
instrumentation scope (@instrumentation.@name="cloudwatch.aws/<service>") plus a
datapoint attribute (e.g. FunctionName); OTLP-native signals carry
@resource.* labels. Select the metric by name — use an {__name__="<exact>"}
matcher for a name that starts with a digit or contains %, ., or spaces (e.g.
4xxErrors, traces.span.metrics.*); a plain non-dotted name may be used bare.
A bare Prometheus job= / service= selector is an ungrounded reference — drop
panels that use one. The per-service scope + datapoint-attribute matrix lives in
promql-metrics.md.
Pick the aggregation from the metric's type. Read it off the grounded
instrument type and temporality — sum_over_time(<m>[w]) for a DELTA Sum,
rate() / increase() for a CUMULATIVE Sum, avg() / max() for a Gauge, and
histogram_quantile(0.99, rate(<base>[5m])) for a Histogram (use the BASE metric
name — NO _bucket, NO by (le)). Never label an average as a percentile, and
never chart a counter as a raw value.
Two grounding cautions:
- Keep every panel query WINDOWLESS. Never embed a time window in the
queryString— no`@timestamp`bound, noNOW() - INTERVAL, no PromQL range that pins an absolute time. The panel (or dashboard)scopedrives the time range; a hard-coded window fights the time picker and freezes the panel. (This is the one place a dashboard query differs from an interactive one, which requires a time bound — so don't copy an interactive query verbatim.) - An empty series is often correct. A signal that only exists under a specific configuration (an agent installed, request metrics enabled, a burstable instance) returns nothing when that condition is absent. Confirm a metric is expected to have data before making it a headline panel.
4. Author the panel bodies
This is the round-tripping JSON schema — the format a user hand-edits in the
source editor and the one the UI persists. Every field here survives save → reopen.
In a body you author now, a key that is not in the schema at the root or panel
level is rejected at save (see "API save semantics"), not dropped; a key inside
config is accepted but ignored. Only a pre-existing legacy body (stored before
the schema was enforced) has non-schema keys dropped or rewritten on read. The whole
object is what you serialize into the API's body string (section 7).
Top-level shape
{
"panels": [ /* Panel[] — required, may be empty */ ],
"scope": { "timeRange": { "start": "now-3h", "end": "now" } } // optional
}panelsis the only required key. Panels sit on a fixed 60-column grid; the column count is not a field and is not configurable.scope.timeRangeis the window the whole dashboard opens under. Omit it to inherit whatever range is already active.
There is no
views, nolayout: { columns: 60 }, no per-panelid, norefreshInterval, and no top-levelowner. The root and panel objects are schema-validated withadditionalPropertiesoff, so sending one of these is not silently dropped — it fails the save with aValidationException(see "API save semantics"). Author new bodies in thepanels[]form only. A body authored under the current schema round-trips exactly:GetOmniDashboardreturns itsbodybyte-for-byte (stored verbatim as an opaque string; whitespace and key order preserved). Only a pre-existing legacy body is upgraded on read (see the carve-out above), so those do not round-trip byte-for-byte.
Panel
{
"type": "explore", // required — selects the panel and its config shape
"layout": { "x": 0, "y": 0, "w": 30, "h": 320 }, // required — rejected at save if absent (section 5)
"title": "Read IOPS", // optional
"description": "…", // optional, one line
"config": { /* shape depends on type */ },
"scope": { "timeRange": { "start": "now-7d", "end": "now" } }, // optional per-panel override
"variant": "transparent" // optional — drop the card frame (headings, spacers, prose)
}Both type and layout are required — a panel missing either is rejected at save
with a ValidationException (panels[N].type / panels[N].layout required). The
renderer does not reflow overlaps (section 5), so give every panel a deliberate
layout to keep placement deterministic. An untitled panel takes
a heading from its content. variant accepts "transparent" (drops the card frame)
or "default" (explicit no-op) — omit the key for the default card frame. variant
is a panel-level enum, so any other value is rejected at save with a
ValidationException (see "API save semantics" below), not silently blanked; other
variant values you read back in product-authored bodies are private — do not author
them.
Scope and time range
scope.timeRange has the same shape at the dashboard and panel level; a panel's
scope overrides the dashboard's for that panel only (e.g. a 7-day baseline beside a
1-hour detail view).
"timeRange": { "start": "now-1h", "end": "now" } // rolling window
"timeRange": { "start": "2026-01-01T00:00:00Z", "end": "now" } // growing window (since incident)
"timeRange": { "start": "2026-01-01T00:00:00Z",
"end": "2026-02-01T00:00:00Z" } // fixed windowEach bound is independently an ISO-8601 instant, the literal now, or a relative
now-<N><unit>. Units: m minutes, h hours, d days, w weeks, mo months.
There is no seconds or years unit and no rounding syntax — use an absolute instant
for a calendar boundary. The scope drives a query's time range; never embed a
window in a queryString (section 3).
Panel types
alerts · explore · markdown · divider · spans · application-map ·
dashboards-list · investigation-list · dashboard · agent-kpi-strip ·
recent-error-traces-table · agent-health-table · agent-playground. Author
only these thirteen — this is the public authoring contract. dashboard and
agent-playground are full-bleed surfaces (public and hand-authorable, but not
offered for free composition; a dashboard panel is always x: 0, w: 60). There
is no navigation type. alert-detail and trace-details are withdrawn: a
saved body that already holds one still reads back, but do not author them in a
new body — see below for the dashboard-panel replacement. Other product-authored
types you read back are private; leave them as they are and do not author new ones.
A missing type is rejected at save (panels[N].type required, HTTP 400). An
unrecognised type VALUE is not caught by the schema — it saves with HTTP 200 and
blanks the canvas at render (see "API save semantics" below).
explore — charts and query results (the workhorse)
"config": {
"queryString": "sum by (FunctionName) (rate({__name__=\"Invocations\", \"@instrumentation.@name\"=\"cloudwatch.aws/lambda\"}[5m]))",
"queryLanguage": "promql", // 'sql' | 'promql' — REQUIRED; SET IT EXPLICITLY; not inferred; a mismatch silently blanks the panel
"visualization": "line", // REQUIRED; one of the nine below; missing or unknown silently blanks the panel (no 'table' fallback)
"autoRun": true, // run on open — set this on any data panel
"promqlQueryOptions": { "step": 300 }, // PromQL step in SECONDS (default 60)
"chartOptions": { /* appearance — see below */ }
}- Always set
queryLanguageexplicitly — it is not inferred from the query text; a mismatch is a silent failure. The telemetry source is the query'sFROMclause (SQL) or metric selector (PromQL); there is no separate source field. - Always set
autoRun: trueon a panel meant to show data without a click — otherwise it opens showing its query, not its result. - Match
stepto the range vector. A[5m]window sampled at the default 60s over-samples 5×; setpromqlQueryOptions.stepto300(SECONDS). inputMode. Omit for a chart tile (the default). Set"inputMode": "editor"only when authoring a full Explore-surface panel — that is the one accepted value, and it is load-bearing on read (a saved Explore surface withinputModedropped reopens as a tile).- The
@-labels are double-quoted PromQL selectors, so they are JSON-escaped (\") once more inside thequeryString; and because the whole body is itself passed to the API as a JSON string, it is escaped a third time on the wire — build the body as an object and let your JSON library serialize it (section 7) rather than hand-escaping.
markdown — formatted text
"config": { "content": "## On call\n\n1. Check error rate.\n2. Open the failing trace." }Pair with "h": "auto" and "variant": "transparent" for prose that reads as
part of the canvas.
divider — collapsible section heading
"config": { "title": "Service health", "collapsed": false }Always full width (x: 0, w: 60); give it a small fixed height (40). Groups the
panels beneath it, down to the next divider, into a collapsible section.
alerts
{ "type": "alerts", "config": { "state": "CRITICAL" } } // OK | WARNING | CRITICAL | NODATA (the Omni alert states); omit for all. NOT the classic CloudWatch alarm states `ALARM` / `INSUFFICIENT_DATA` (and `ALERT` is not a state at all) — an unknown state value silently blanks the canvasOne alert's detail is not a panel type of its own any more — alert-detail is
withdrawn (readable, not authorable). Author a dashboard panel instead:
{ "type": "dashboard",
"source": { "kind": "system", "id": "alert-details", "version": 1,
"params": { "alertName": "checkout-error-rate-high", "alertId": "6c89…" } }, // alertName required; alertId optional
"layout": { "x": 0, "y": 0, "w": 60, "h": "auto" } }The required parameter is the alert's name, but alert names are not unique
within a space — two alerts can share one name. The stable identity is the
alertId (see alerts.md); pass it too so the panel is unambiguous.
Resolve a name to its alertIds with ListAlerts (filterCriteria.names) first,
and rename one via UpdateAlert if two collide.
application-map / spans
Both take a public config; either also renders correctly with no config and
picks up the dashboard's time range.
{ "type": "application-map", "config": { "focusServiceName": "checkout" } } // opens focused on one service; a one-time opening intent the console drops on its next save
{ "type": "spans", "config": { "lens": "application", // application | agent (default application)
"grain": "traces", // application lens: traces | spans (default traces)
"agentGrain": "traces", // agent lens: traces | sessions (default traces)
"agentName": "…" } } // agent lens only: scope to one agent; absent = allAn absent or unrecognised selector falls back to its default. Everything else — service/operation/status/duration/trace-id filters, camera, node selection, rail state — is per-user console state, not saved.
dashboards-list
"config": {} — a list of the Space's saved dashboards.
one trace's detail — a dashboard panel, not trace-details
trace-details is a withdrawn panel type: a saved body that holds one still
reads back, but do not author it in a new body — like any unrecognised type it
saves with HTTP 200 and blanks the canvas. Author a dashboard panel whose
source.id is trace-details:
{ "type": "dashboard",
"source": { "kind": "system", "id": "trace-details", "version": 1,
"params": { "traceConfig": {
"traceId": "…", // required
"startTime": 1730000000000, "endTime": 1730000600000,
"initialMode": "waterfall", // waterfall | graph | flame | raw (`flamegraph` is a legacy alias read as `graph`)
"focusSpanId": "…" } } }, // optional — deep-link to a specific span
"layout": { "x": 0, "y": 0, "w": 60, "h": "auto" } }startTime/endTime are epoch milliseconds in traceConfig (not the panel
scope), fixed at author time. A pinned trace is useful only while it is still
retained — prefer a spans panel for a durable dashboard.
Visualizations — exactly nine legal values
line · area · bar · scatter · pie · table · number · solidgauge ·
heatmap.
stacked-bar,single-metric, anddonutare not valid — usebarwithchartOptions.plotOptions.style.barOptions.stacked: true(there is noplotOptions.stacked; see Chart options),number, andpierespectively.visualizationis REQUIRED and there is no fallback default that renders: a missing value, an unknown value such asstacked-bar, or the empty string""saves with HTTP 200 but silently blanks the panel at render.
number tile trap (known product bug). A number panel driven by a
SELECT count(*) AS n (or similar COUNT(*)) query currently renders the ROW
COUNT of the result set — usually 1 when the query aggregates to one row —
instead of the value in the aliased column. Author the query so the aggregate
value is the first column of the first row (e.g. SELECT <aggregate_expression>),
and reserve COUNT(*)-shaped queries for table panels until the bug is fixed.
Pick by signal kind:
| Signal | Visualization |
|---|---|
| Time series, ≤ ~25 series, comparable units | line |
| Additive series / discrete buckets (2xx/4xx/5xx, per-AZ) | bar (stacked) or area |
| Correlation point cloud (no connecting line) | scatter |
| One number, with a delta | number |
| Resources ranked by a metric | table (default sort: descending on the primary signal) |
| Genuine part-of-whole | pie |
| A bounded ratio against thresholds (availability %) | solidgauge |
Two shapes authors most often get wrong — check them explicitly: a scalar
aggregate (one number, no by / no GROUP BY over time) should be a number
tile, not a table; a top-N ranking (ORDER BY … LIMIT / topk(...)) should be
a table, not a line.
Anti-patterns: no line past ~25 series (use topk(...) or a table); don't
stack a bar whose series go negative; don't put two units on one chart — split
into two w: 30 panels. heatmap does not yet ingest histogram-bucket output —
for a latency distribution use line with a histogram_quantile.
Chart options
chartOptions sets appearance and is a discriminated union keyed on view, which
must match the panel's visualization (when both are present, view wins — so
setting only visualization is the simpler, recommended form).
"chartOptions": {
"view": "line",
"title": { "text": "Requests/sec", "show": true },
"plotOptions": {
"legend": { "show": true, "position": "bottom" },
"yAxis": [ { "min": 0, "title": "req/s" } ] // TUPLE of 1 or 2 axes (2nd = right-hand)
}
}- Cartesian (
line/area/bar/scatter):legend,stacking,xAxis, andyAxisas a tuple of one or two axes (the second is the right-hand axis). Usetype: 'datetime'for a time axis (not'time'). solidgauge:yAxisis an object{ min, max }(not a tuple), plusplotBands: [{ from, to, color }].heatmap:xAxis/yAxiswithcategories, andcolorScale.table:hiddenColumns,summaryColumns(MIN/MAX/SUM/AVG),layout(horizontal/vertical),stickySummary,showTimeSeriesData,formatJson(true|'raw'|'raw-single-line'). Column order is not a chart option — the console does not save it.numberandpietake little beyond the base fields.- Stacking depends on the view. For
line/area, setplotOptions.stacking: "normal"(or"percent") — an enum, not a boolean, keyed directly onplotOptions; it wins over the equivalentplotOptions.style.lineOptions.stacked: true. Abarpanel does not readstackingat all — the only stacking field the bar renderer reads isplotOptions.style.barOptions.stacked: true. There is noplotOptions.stackedfield; a flag placed there is an unknownconfig-level key — it saves without error but is ignored at render (it survives read-back yet never takes effect), so the chart will not stack.
API save semantics — what happens to a body on read/write
The save API schema-validates the body's structure, its root/panel keys, each
coordinate FIELD's bounds, and the variant enum — but not cross-field invariants or
the other enum values (visualization, panel type, alerts state). A body
fails in three distinct ways, and only the first errors:
- Rejected at save (HTTP 400
ValidationException). The root and panel objects haveadditionalPropertiesoff, so an unknown or misspelled key there — a top-levelrefreshInterval/owner/layout, a per-panelid, a misspelledtitle— fails withdoes not satisfy the 'additionalProperties' constraint. The schema also rejects: a badvariantvalue (enumconstraint); a missing or non-arraypanels(required/type); a panel missing its requiredtypeorlayout(required); and an out-of-bounds coordinate FIELD —layout.xorlayout.ybelow 0, orlayout.wbelow 1 (minimum). The dashboard is not created or updated. - Accepted then blank (HTTP 200). A value the schema does NOT check persists with a
200and adashboardId, then fails at render — the blast radius depends on the kind: an unknownvisualizationVALUE blanks that panel while its siblings render; everything else the schema does not catch — an unrecognised paneltypeVALUE, a badalertsstate, the cross-fieldx + w > 60overflow (each field is in bounds, but their sum is not validated), aw > 60(no maximum onw), a badhunit (80vh, a fraction), or anow-30s/now-1ybound (no seconds/years unit) — blanks the WHOLE dashboard (valid siblings included). A200proves only that the body parsed and passed the schema, not that it renders. - Accepted then ignored (HTTP 200, renders fine minus that key). An unknown or
misspelled key inside
config(or its nested objects likechartOptions) is stored, comes back on read, and simply never takes effect — this is the one place the "a field I set doesn't take effect, with no error" behavior lives (e.g. a Britishvisualisation, or aplotOptions.stackedthat should beplotOptions.stackingfor line/area orplotOptions.style.<view>Options.stackedfor bar).
Consequences:
- Do not rely on a 200 to confirm renderability. Validate the body against this
reference before the save call; a
200+ adashboardIdproves only that the JSON parsed and passed the schema (structure, keys, and per-field bounds), not that it renders. - Structural / key / field-bound /
varianterrors are loud; other enum-value and cross-field errors are silent. The schema catches missing structure, unknown keys, per-field coordinate bounds, and thevariantenum with aValidationException; it does NOT catch an unknownvisualizationvalue — which blanks that one panel — or a bad paneltype/alertsstate, thex + w <= 60sum, and badhunits — which blank the whole canvas — all only at render. layoutis required on every panel — a panel with nolayoutis rejected at save (panels[N].layoutrequired), not drawn full-width.- Read-back is byte-for-byte for a current-schema body. The
bodyis stored verbatim as an opaque string;GetOmniDashboardreturns it exactly as written (whitespace and key order preserved), so it round-trips exactly. (Only a pre-existing legacy body is upgraded on read and so does not round-trip; section 6.) queryLanguagemismatches are silent failures, not errors. Set it perexplorepanel.- Dashboard names (the API's
name, not part of the body) must be 1–256 chars matching^[a-zA-Z0-9_.@~()-]+$. Nothing is trimmed: surrounding whitespace, spaces,/,:,#,+, accented letters and emoji all fail the pattern, and 257+ chars fails the length check. The body itself must be 1–1,048,576 bytes once serialized.
5. Lay out on the 60-column grid
Assign each panel's layout: { x, y, w, h } so panels tile cleanly with no overlap
and no overflow. This is the single most error-prone part of a body.
ALWAYS STATE when laying out a grid — two different failure modes. (1) A per-field bound violation is rejected at save (
ValidationException, HTTP 400): a negativexory, or awbelow 1, trips the schema'sminimumconstraint and the save does not succeed. (2) The cross-fieldx + w > 60overflow, aw > 60, and a badh(a fraction or avh/%unit) are NOT validated: they save with HTTP 200 and adashboardIdand then render the entire dashboard — every otherwise-valid sibling panel included — as a blank page, with no error. Nothing enforces thex + w <= 60sum, and the renderer does not reflow or auto-arrange overlaps. Keepx + w <= 60on every panel yourself; a200is not proof the layout renders.
The two-unit rule (memorize this)
x/w and y/h do not share a unit:
| Field | Unit | Rule |
|---|---|---|
x |
grid column | integer 0–59; x: 0 is the left edge. A negative x is rejected at save (minimum, HTTP 400) |
w |
grid columns | integer, w >= 1 is enforced at save (w: 0 is rejected, minimum, HTTP 400); there is no enforced maximum (w: 61 saves). x + w must not exceed 60, but that sum is not enforced — an overflow saves with HTTP 200 and blanks the canvas |
y |
pixels | integer pixels from the top, not a row index; a negative y is rejected at save (minimum, HTTP 400) |
h |
pixels or 'auto' |
integer pixels within roughly 100–2000 (divider 40 is the sanctioned sub-100 row), or the literal 'auto'; never fractional and never a relative CSS unit (vh / % / vw) — those are not rejected at save, they save with HTTP 200 and blank the canvas at render |
Mixing the units (treating y as a row index, or w as pixels) is the most common
authoring bug. Prefer 'auto' for content-sized panels (markdown, lists, nested
dashboard panels) and a fixed pixel height where the content needs one; a relative
CSS unit such as '80vh' is an unknown value that blanks the canvas.
Width conventions
Every row MUST match one of the width-per-count patterns below. Author the correct shape yourself — nothing re-shapes it for you.
| Row count | Width per panel (w) |
Use for |
|---|---|---|
| 1 panel | 60 |
full-width strip (hero chart, table, status strip, divider, application-map, dashboard) |
| 2 panels | 30 each |
two halves side-by-side (most time-series pairs) |
| 3 panels | 20 each |
three thirds (read/write/idle, 2xx/4xx/5xx) |
| 4 panels | 15 each |
four quarters (KPI row) |
Hard rules — all MUST:
- A row's widths sum to EXACTLY 60, with
xas the cumulative sum from 0 (0, 30;0, 20, 40;0, 15, 30, 45). - Every panel in a row shares the same
yAND the sameh. - Never more than 4 panels in one row. A 5th starts a new row below: 5 panels = 3 then 2; 6 = 3+3; 7 = 4+3. Five across does not fit the 60-column grid and the overflow lands on top of its neighbours.
- No mixed-width rows except an intentional hero-supporting layout, which uses
TWO rows: row 1 is one
w: 60hero, row 2 isw: 30×2 orw: 20×3 supporting panels — never side-by-side with the hero. - No heterogeneous chart/table rows. A
linechart and atableon the same row read badly — split them into two rows (chart row above, table row below).
Stacking rows with y
y is the pixel offset from the top and it orders rows. Give successive rows
increasing y equal to the running sum of prior row heights:
- KPI row (
h: 120) aty: 0 - first chart row (
h: 320) aty: 120 - next chart row at
y: 440, and so on.
When a row uses 'auto' heights, still give the next row a larger y; the renderer
resolves final positions from each row's height, so approximate but
monotonically-increasing y values are fine.
Height guidance by panel kind
| Panel kind | Suggested h |
|---|---|
number KPI tile |
120 (short — one figure) |
Time-series (line/area/bar/scatter) |
~320 |
table |
~320–400, taller for many rows |
pie / solidgauge |
~280–320 (roughly square reads best) |
divider |
40 (always full width, x: 0, w: 60) |
dashboard (nested system dashboard) |
'auto' (always full width, x: 0, w: 60) |
markdown runbook |
'auto' (pair with variant: 'transparent') |
application-map / spans |
full width, tall (~480) |
Layout shortcuts
- Full-width strip —
{ x: 0, y: <row>, w: 60, h: <h> }. - Two halves —
{ x: 0, …, w: 30 }and{ x: 30, …, w: 30 }, samey. - Three thirds —
x: 0,x: 20,x: 40, eachw: 20, samey. - KPI row of four then a wide chart — four tiles at
y: 0,w: 15,x: 0/15/30/45,h: 120; one chart aty: 120,x: 0,w: 60,h: 320. - Divider-led section — a full-width
divider(h: 40) at the row'sy, then the section's panels at a largerybeneath it.
Overflow and overlap
- A panel that overflows the grid (
x + w > 60) is never reflowed into what you meant — it saves with HTTP 200 and blanks the whole canvas (section 4, API save semantics). Keepx + w <= 60on every panel yourself. - The renderer does not reflow overlaps. Two panels sharing a
ymust not share columns — give them non-overlapping[x, x+w)ranges. Follow the width-per-count patterns above and give each row a largeryand you will never overlap.
6. Troubleshooting
A panel rendered empty
Work down this list in order; the first match is usually the cause.
- An ungrounded name. A metric, field, or label that does not exist returns
nothing, not an error. Re-run discovery (section 3) for exactly the names in
the
queryString; a*_total/*_secondsPrometheus-convention name, or a barejob=/service=selector, is the CloudWatch tell. queryLanguagemismatch. A PromQL selector under"queryLanguage": "sql"(or SQL underpromql) fails silently. Check the language against the text.- A time window embedded in the query. A
`@timestamp`bound, aNOW() - INTERVAL, or an absolute PromQL range pins the panel to a window the time picker no longer covers. Remove it; thescopedrives the range. - The panel's or dashboard's
scopeexcludes the data. A fixed incident-window scope, or a per-panel override, can legitimately show nothing outside its range. autoRunis nottrue. The panel opens showing its query, not its result — it looks empty until clicked.- The metric selector's identity is wrong. A vended metric without its
@instrumentation.@namescope, or the wrong datapoint attribute (FunctionNamevs.ApiName), matches nothing. Compare against the archetype identity forms and promql-metrics.md. - The series is genuinely empty — and that is correct. Conditional signals
(
IteratorAgefor a non-stream consumer,RunningTaskCountwithout Container Insights, burst-credit metrics on a non-burstable instance) return nothing when the condition is absent. Say so, and either drop the panel or retitle it. - A
stepfar coarser than the window. With a very largepromqlQueryOptions.stepover a short scope there may be no evaluation point.
A field has no effect, or the canvas is blank
- The field is not in the schema, or its name is misspelled. Where it lands
decides the symptom. A misspelled or unknown key inside
config(e.g.plotOptions.stacked, which should beplotOptions.stacking: "normal"orplotOptions.style.<view>Options.stacked) saves without error and is ignored at render — that is the "field has no effect" case. A wrong key at the root or panel level (layout.columns, per-panelid,refreshInterval,owner,views) does NOT reach this state — it is rejected at save with aValidationException(section 4), so its symptom is a failed save, not a silently-ignored field. inputModewas omitted from an Explore-surface panel. It reopens as a tile.- The whole canvas is blank after a 200. Nothing was rejected — the save returned
HTTP 200 — but the body carries a value the schema does not check that breaks the
whole render: a cross-field
x + w > 60overflow or aw > 60, a badh(fractional or avhunit), an unrecognised paneltypeVALUE, a badalertsstate, atimeRangebound with a seconds/years unit, asourceon a non-dashboardpanel, or adashboardpanel without one. No sibling renders until it is fixed — validate each panel against section 4. (An unknownvisualizationVALUE instead blanks only its own panel; see the next bullet. A missing/non-arraypanels, a panel missingtypeorlayout, or a negativex/yorw: 0is rejected at save with aValidationException, not this silent-blank symptom.) - A panel is blank but the rest render. A missing or invalid
visualizationVALUE (stacked-bar,donut,single-metric,""; there is notablefallback) blanks that panel alone while its siblings render. - A pre-existing legacy body was rewritten on read. In bodies stored before the
schema was enforced,
viewsbecamepanelsand old relative ranges became{ start, end }on read — that is the read-time upgrade of grandfathered bodies, not data loss. It does NOT apply to a body you author now: sendingviews(or any unknown root/panel key) in a new body is rejected at save with aValidationException(see "API save semantics"), not upgraded.
Reading back a body (pasted into chat, or fetched with GetOmniDashboard)
- Get the JSON. From the API,
omniDashboard.bodyis a JSON string — parse it before inspecting it. From a paste, strip any surrounding prose. - Check the top level.
panelsmust be an array; note anyscope.timeRange(a fixed absolute window means an incident dashboard). - Walk each panel and describe it in the user's terms — what it shows, from
which telemetry, over what range — while checking, for each:
typeis one of the thirteen authorable types;layoutuses the two units correctly and stays inside 60 columns; rows shareyandh; forexplore,queryLanguagematches the query text,visualizationis one of the nine, the query is windowless, and every name in it is plausible for the grounded catalog;chartOptions.viewmatchesvisualization; no unknown keys. - Explain the two silent failures (an unknown key inside
configthat is stored but ignored at render;queryLanguagenot inferred) whenever you find an instance of either, because the user will not have seen an error. An unknown root- or panel-level key cannot appear in a read-back body from a current-schema save — it would have been rejected at write with a ValidationException. - Do not narrate the schema when the user asked what the dashboard shows — lead with the panels' meaning; surface the defects as a short list after.
7. API reference
Five operations manage dashboards. They are exposed by the CloudWatch Omni service
(endpoint prefix cloudwatch-omni, SigV4 signing name cloudwatch — see
programmatic-access.md) and reachable from the AWS CLI as
aws cloudwatchomni <kebab-case-operation>, from every AWS SDK, or through an AWS
MCP aws___call_aws tool with the same operation and parameter names.
Common to all five:
- Space-scoped.
spaceId(the Space's UUID) is required on every call. There is no account-wide list of dashboards — you list per Space. - The identity is
dashboardId, a system-minted lowercase UUID (^[0-9a-f]{8}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{4}-[0-9a-f]{12}$). The caller-suppliednameis metadata, is not unique, and is not part of the id — renaming never changes the id or the ARN. - ARN:
arn:<partition>:cloudwatch:<region>:<account-id>:omni-dashboard/<dashboardId>— thecloudwatchvendor namespace with anomni-dashboard/resource type, the same convention as alerts. Neither the Space nor the name appears in it. bodyis a string, 1–1,048,576 bytes, containing the serialized JSON from section 4. The API schema-validates the root and per-panel objects on write (additionalPropertiesoff) and stores the string verbatim, returning a current-schema body byte-for-byte on read (only pre-existing legacy bodies are upgraded on read; section 6). An APIValidationExceptiononbodycan therefore be a length/type problem OR a body-schema one — an unknown root/panel key, a badvariantenum, a missing/non-arraypanels, a panel missingtype/layout, or an out-of-bounds coordinate field (negativex/y,w: 0) (section 4). What is NOT enforced at the API — a bad enum VALUE (visualization, paneltype), thex + w > 60sum, or unknownconfigkeys — surfaces only when the UI renders.name: 1–256 chars,^[a-zA-Z0-9_.@~()-]+$(section 4's rule).description: 1–1024 chars, optional.- Errors every operation can return:
ValidationException,AccessDeniedException,InternalServerException,ThrottlingException. Per-operation additions are noted below. - Authorization is the Space's access grants, exactly as for a console user. A
correctly-signed call returning
AccessDeniedExceptionusually means the caller has no grant on the Space — check that before suspecting the request. The IAM action names are the operation names (cloudwatch:CreateOmniDashboard, etc.).
CreateOmniDashboard
Creates a dashboard in a Space. Returns the full omniDashboard — read
dashboardId and arn from it; you do not need a follow-up GetOmniDashboard.
| Input | Required | Notes |
|---|---|---|
spaceId |
✔ | The Space to create the dashboard in |
name |
✔ | Display name, 1–256 chars, ^[a-zA-Z0-9_.@~()-]+$. Not unique |
body |
✔ | The serialized { "panels": [...] } JSON, 1–1,048,576 bytes |
description |
1–1024 chars | |
tags |
Key/value map | |
clientToken |
Idempotency token — a retry with the same token returns the original result instead of creating a duplicate. SDKs and the CLI fill it automatically |
Errors beyond the common set: ServiceQuotaExceededException (the Space's
dashboard limit), ResourceNotFoundException (unknown Space), ConflictException.
# body.json holds the panels[] object from section 4; file:// passes its text as the string value
aws cloudwatchomni create-omni-dashboard \
--region us-east-1 \
--space-id 3f2c1e9a-7b4d-4c0e-9a1b-2d3e4f5a6b7c \
--name checkout-service-health \
--description "RED golden signals for checkout" \
--body file://body.jsonBuild body.json with a JSON library from the panel objects — do not hand-escape
the nested \" in PromQL selectors a second time; the CLI/SDK serializes the string
for the wire.
GetOmniDashboard
Retrieves one dashboard including its body. This is the only way to read a saved body — list results do not carry it.
| Input | Required |
|---|---|
spaceId |
✔ |
dashboardId |
✔ |
Returns omniDashboard: dashboardId, arn, name, body (a JSON string —
parse it), createdBy (the creating principal), description, tags, createdAt,
updatedAt. Read-only. Errors beyond the common set: ResourceNotFoundException.
aws cloudwatchomni get-omni-dashboard --region us-east-1 \
--space-id <spaceId> --dashboard-id <dashboardId> \
--query 'omniDashboard.body' --output text > body.jsonListOmniDashboards
Enumerates the dashboards in a Space, optionally filtered by name prefix, with pagination.
| Input | Notes |
|---|---|
spaceId |
Required |
namePrefix |
1–256 chars, same character set as name; matches names that start with it |
maxResults |
1–100 per page |
nextToken |
Pagination token |
Read the summaries from items and page with nextToken — items is the only
list member the response carries. Each entry is an
OmniDashboardSummary: dashboardId, arn, name, createdBy, description,
tags, createdAt, updatedAt — no body. Because names are not unique, a
namePrefix (or an exact-name match on the client side) can return several
dashboards; this is the usual way to resolve a name to a dashboardId before
Get/Update/Delete. Follow nextToken rather than treating the first page as
the complete set. Errors beyond the common set: ResourceNotFoundException
(unknown Space).
aws cloudwatchomni list-omni-dashboards --region us-east-1 \
--space-id <spaceId> --name-prefix checkout- --max-results 50UpdateOmniDashboard
Modifies a dashboard in place, addressed by dashboardId. PATCH semantics:
only the fields you send change; an omitted field is left unchanged. Returns the
full updated omniDashboard, so you can confirm the result without a second
call. @idempotent.
| Input | Required | Notes |
|---|---|---|
spaceId |
✔ | |
dashboardId |
✔ | The dashboard to update |
body |
Full replacement body, 1–1,048,576 bytes. There is no partial-panel patch — send the complete panels[] |
|
name |
Renaming is supported; the id and ARN do not change | |
description |
1–1024 chars |
tags cannot be changed through this operation. Errors beyond the common set:
ResourceNotFoundException, ConflictException, ServiceQuotaExceededException.
To edit one panel: GetOmniDashboard → parse body → change the panel →
serialize → UpdateOmniDashboard with the whole body. Because a body is replaced,
not merged, an update built from a stale copy silently discards someone else's
edits — fetch immediately before you write.
Confirm before updating. Draft the exact change, state that body replaces the
whole panel set, and confirm with the user before calling.
aws cloudwatchomni update-omni-dashboard --region us-east-1 \
--space-id <spaceId> --dashboard-id <dashboardId> \
--body file://body.jsonDeleteOmniDashboard
Permanently removes a dashboard. Returns an empty response.
| Input | Required |
|---|---|
spaceId |
✔ |
dashboardId |
✔ |
Idempotent — deleting a dashboard that has already been removed succeeds
without error (the operation does not return ResourceNotFoundException), so a
successful response does not prove the dashboard existed. Deletion cannot be undone;
if the body might be wanted later, GetOmniDashboard and keep it first.
Confirm before deleting. This is destructive and irreversible. Name the specific dashboard (id and name) and confirm with the user before calling — never delete on inference, and because a success does not tell you whether you deleted the one you meant, confirm the target up front rather than checking after.
aws cloudwatchomni delete-omni-dashboard --region us-east-1 \
--space-id <spaceId> --dashboard-id <dashboardId>End-to-end: build, save, verify
- Probe:
aws cloudwatchomni list-spaces— pick thespaceId(section 1). - Plan, ground, author, lay out (sections 2–5) → a
panels[]object. - Serialize it to
body.json;create-omni-dashboardwith a validname. - Read back with
get-omni-dashboard, parsebody, and run the read-back checks in section 6 — the API schema-checks only root/panel keys and per-field bounds, so this is where aconfig-level slip (bad enum value, unknown key) would surface. - Report the
dashboardId,arn, and a one-line description of each panel.
CLI says cloudwatchomni is not a valid choice. The local AWS CLI predates
Omni's service model; upgrade it. The failure is client-side argument parsing and
says nothing about whether Omni is enabled. aws cloudwatch <omni-operation>
will never work at any version — CloudWatch is a different service that
shares only the signing name.
8. Appendix — Archetype templates
When the ask names a known resource TYPE, start from its ready panel set below,
then ground every query (section 3) and lay it out (section 5). Build the
CORE subset by default and offer the optional-depth panels as a follow-up "add
section." Look up each panel's real metric and reading (gauge vs. counter, derived
formulas, whether a real percentile source exists) in
promql-metrics.md; the query strings are TEMPLATES — the
identity form is canonical, but the metric name and statistic MUST be grounded. If
a signal cannot be grounded, drop that panel and re-sequence y so no blank row
is left.
Every template binds to the @-label identity (section 3) — vended AWS metrics on
@instrumentation.@name="cloudwatch.aws/<svc>" + datapoint attribute, OTLP /
span-RED on @resource.*. Set queryLanguage explicitly, autoRun: true on every
data panel, and promqlQueryOptions.step to match the range vector; keep every
query windowless. … in a template stands for the identity labels shown in that
archetype's first row.
EC2 fleet
Infrastructure resource — USE / saturation; scope cloudwatch.aws/ec2 by
InstanceId (the CWAgent guest-OS metrics share the identity). Reading rules from the
catalog: the vended EC2 metrics reach this surface as delta exponential histograms
(CPUUtilization, StatusCheckFailed*, CPUCreditBalance) and delta Sums
(NetworkIn/NetworkOut), so a max is histogram_quantile(1, …), a min
histogram_quantile(0, …), an average histogram_sum(…) / histogram_count(…), and a
network volume sum_over_time(…[5m]) — bare max()/avg() and rate() render empty
panels here. Every StatusCheckFailed* metric is a per-period 0/1 flag read as a max
(an average hides a failed check), and _System (AWS hardware) is a different fault from
_Instance (the guest OS or network) — show them side by side; CPUCreditBalance exists
only on burstable (T-family) instances, and near zero at baseline CPU means the instance is
being throttled; mem_used_percent / disk_used_percent exist only where the CloudWatch
agent is installed (an empty panel means "no agent", so ground it and drop it if absent);
DiskRead* / DiskWrite* are instance-store only — EBS volume I/O is the AWS/EBS
family keyed by VolumeId, a separate row if the fleet's disks matter. __name__ takes no
regex, so a panel that shows several check types is one selector per name joined with or
and tagged with label_replace. S below stands for "@instrumentation.@name"="cloudwatch.aws/ec2".
CORE (~8 panels — the default build; every query below returns series on an enriched Space, except the two configuration-gated tiles noted inline):
| Panel | Viz | Query template | Layout |
|---|---|---|---|
| Instances failing any check | number |
count(histogram_quantile(1, {__name__="StatusCheckFailed", S}) == 1) or vector(0) — count() over an empty match is itself empty, so without or vector(0) a healthy fleet renders blank; with it the tile reads 0 both when the fleet is healthy and when the selector matches no instance, so confirm the bare selector returns series before trusting a 0 |
{x:0,y:0,w:15,h:120} |
| Peak CPU % | number |
max(histogram_quantile(1, {__name__="CPUUtilization", S})) |
{x:15,y:0,w:15,h:120} |
| Lowest CPU credits | number |
min(histogram_quantile(0, {__name__="CPUCreditBalance", S})) — burstable instances only; drop if empty |
{x:30,y:0,w:15,h:120} |
| Peak memory % | number |
max({__name__="mem_used_percent", S}) — CloudWatch agent only, a plain gauge (switch to histogram_quantile(1, …) only if its __type__ is ExponentialHistogram); drop if empty |
{x:45,y:0,w:15,h:120} |
| CPU % by instance | line |
histogram_quantile(1, {__name__="CPUUtilization", S}) — one series per InstanceId |
{x:0,y:120,w:30,h:320} |
| Status checks by instance and type | line |
label_replace(histogram_quantile(1, {__name__="StatusCheckFailed_System", S}), "check", "System", "", "") or label_replace(histogram_quantile(1, {__name__="StatusCheckFailed_Instance", S}), "check", "Instance", "", "") or label_replace(histogram_quantile(1, {__name__="StatusCheckFailed_AttachedEBS", S}), "check", "AttachedEBS", "", "") — one 0/1 series per instance and check type, never averaged |
{x:30,y:120,w:30,h:320} |
| Network in/out by instance | line |
label_replace(sum by (InstanceId) (sum_over_time({__name__="NetworkIn", S}[5m])), "direction", "in", "", "") or label_replace(sum by (InstanceId) (sum_over_time({__name__="NetworkOut", S}[5m])), "direction", "out", "", "") — bytes per 5 minutes; append / 300 for bytes per second |
{x:0,y:440,w:60,h:320} |
| Hottest instances | table |
topk(10, histogram_quantile(1, {__name__="CPUUtilization", S})) |
{x:0,y:760,w:60,h:360} |
Optional depth (follow-up "infra / USE health"): CPUCreditBalance by instance
(histogram_quantile(0, …), line); disk_used_percent by instance (CWAgent); the EBS
volume side (VolumeQueueLength, VolumeAvgReadLatency / VolumeAvgWriteLatency on
cloudwatch.aws/ebs by VolumeId — check their __type__ first). Append two-up and
re-sequence y.
Lambda
Serverless function — RED-shaped; scope cloudwatch.aws/lambda by FunctionName.
Duration is a gauge with no percentile label; the error-rate tile is the derived
Errors / Invocations.
CORE (~8 panels — the default build):
| Panel | Viz | Query template | Layout |
|---|---|---|---|
| Invocations/s | number |
sum(rate({__name__="Invocations", "@instrumentation.@name"="cloudwatch.aws/lambda", FunctionName="<fn>"}[5m])) |
{x:0,y:0,w:15,h:120} |
| Error rate | number |
sum(rate({__name__="Errors", …, FunctionName="<fn>"}[5m])) / sum(rate({__name__="Invocations", …, FunctionName="<fn>"}[5m])) |
{x:15,y:0,w:15,h:120} |
| Throttles/s | number |
sum(rate({__name__="Throttles", …, FunctionName="<fn>"}[5m])) |
{x:30,y:0,w:15,h:120} |
| Concurrency | number |
max({__name__="ConcurrentExecutions", …, FunctionName="<fn>"}) |
{x:45,y:0,w:15,h:120} |
| Invocation rate | line |
sum by (FunctionName) (rate({__name__="Invocations", …, FunctionName="<fn>"}[5m])) |
{x:0,y:120,w:30,h:320} |
| Errors/s | line |
sum by (FunctionName) (rate({__name__="Errors", …, FunctionName="<fn>"}[5m])) |
{x:30,y:120,w:30,h:320} |
| Duration (ms) | line |
{__name__="Duration", …, FunctionName="<fn>"} |
{x:0,y:440,w:60,h:320} |
| Recent errors | table (SQL) |
scope to /aws/lambda/<fn> with a grounded log-group filter; queryLanguage:'sql' |
{x:0,y:760,w:60,h:360} |
Optional depth (follow-up): Throttles line; Concurrent executions line;
Stream-consumer lag line (IteratorAge — only for a stream/event-source
consumer, empty otherwise). Append two-up (w:30,h:320) and re-sequence y.
Kubernetes/EKS
OTLP-native — no vended scope. RED from the span metrics grouped by
@resource.service.name; USE from the OTLP Kubernetes resource metrics grouped
by @resource.k8s.*. The span-metric names, span-status attribute, and k8s.*
resource-metric names are ingestion-path-dependent — ground them against the live
catalog (section 3).
CORE (~8 panels — the default build):
| Panel | Viz | Query template | Layout |
|---|---|---|---|
| Request rate | number |
sum(rate({__name__="traces.span.metrics.calls", "@resource.service.name"="<svc>"}[5m])) |
{x:0,y:0,w:15,h:120} |
| Error rate | number |
ratio of …calls, "@status.code"="ERROR" to all calls |
{x:15,y:0,w:15,h:120} |
| p99 latency | number |
histogram_quantile(0.99, rate({__name__="traces.span.metrics.duration", …}[5m])) |
{x:30,y:0,w:15,h:120} |
| Running pods | number |
count({"k8s.pod.phase", "@resource.k8s.namespace.name"="<ns>", "@resource.k8s.deployment.name"="<deploy>"}) |
{x:45,y:0,w:15,h:120} |
| Request rate | line |
sum by ("@resource.service.name") (rate({__name__="traces.span.metrics.calls", …}[5m])) |
{x:0,y:120,w:30,h:320} |
| Errors by status | bar |
sum by ("@status.code") (rate({__name__="traces.span.metrics.calls", …}[5m])) |
{x:30,y:120,w:30,h:320} |
| Latency p50/p90/p99 | line |
histogram_quantile(0.99, rate({__name__="traces.span.metrics.duration", …}[5m])) — one series per quantile (0.50, 0.90, 0.99) |
{x:0,y:440,w:60,h:320} |
| Top services by errors | table |
topk(10, sum by ("@resource.service.name") (rate({__name__="traces.span.metrics.calls", "@status.code"="ERROR"}[5m]))) |
{x:0,y:760,w:60,h:360} |
Optional depth (follow-up "infra / USE health"): Pod CPU usage, Pod memory
usage, Container restarts — grouped by @resource.k8s.* (k8s.pod.cpu.usage,
k8s.pod.memory.usage, increase(k8s.container.restarts[15m])). Append two-up and
re-sequence y.
RDS/database
Infrastructure resource — USE/saturation; scope cloudwatch.aws/rds by
DBInstanceIdentifier. Reading rules: IOPS is already per-second → never rate();
latency is in seconds → × 1000 for ms; mind FreeStorageSpace / FreeableMemory
polarity (higher is healthier).
CORE (~7 panels — the default build):
| Panel | Viz | Query template | Layout |
|---|---|---|---|
| CPU % | number |
max({__name__="CPUUtilization", "@instrumentation.@name"="cloudwatch.aws/rds", DBInstanceIdentifier="<db>"}) |
{x:0,y:0,w:15,h:120} |
| Connections | number |
max({__name__="DatabaseConnections", …, DBInstanceIdentifier="<db>"}) |
{x:15,y:0,w:15,h:120} |
| Free storage | number |
min({__name__="FreeStorageSpace", …, DBInstanceIdentifier="<db>"}) |
{x:30,y:0,w:15,h:120} |
| Read latency (ms) | number |
max({__name__="ReadLatency", …, DBInstanceIdentifier="<db>"}) * 1000 |
{x:45,y:0,w:15,h:120} |
| CPU % | line |
{__name__="CPUUtilization", …, DBInstanceIdentifier="<db>"} |
{x:0,y:120,w:30,h:320} |
| Connections | line |
{__name__="DatabaseConnections", …, DBInstanceIdentifier="<db>"} |
{x:30,y:120,w:30,h:320} |
| Top instances by CPU | table |
topk(10, max by (DBInstanceIdentifier) ({__name__="CPUUtilization", …})) |
{x:0,y:440,w:60,h:360} |
Optional depth (follow-up "infra / USE health"): Read IOPS line; Read latency
(ms) line; add the WriteIOPS / WriteLatency counterparts for a full
read+write view. Append two-up and re-sequence y.
ECS/web service
Cross-signal — USE from the ECS service tier combined with RED from the
load balancer in front (or from Application Signals). Identity + reading rules: ECS
scopes by ClusterName + a non-empty ServiceName; ALB TargetResponseTime is a
gauge in seconds with no percentile label; the target status-code metrics carry the
_Count suffix.
CORE (~7 panels — the default build; this archetype has no ranked table):
| Panel | Viz | Query template | Layout |
|---|---|---|---|
| Request rate | number |
sum(rate({__name__="RequestCount", "@instrumentation.@name"="cloudwatch.aws/applicationelb", LoadBalancer="<lb>"}[5m])) |
{x:0,y:0,w:15,h:120} |
| 5xx rate | number |
sum(rate({__name__="HTTPCode_Target_5XX_Count", …, LoadBalancer="<lb>"}[5m])) |
{x:15,y:0,w:15,h:120} |
| Peak response (ms) | number |
max({__name__="TargetResponseTime", …, LoadBalancer="<lb>"}) * 1000 |
{x:30,y:0,w:15,h:120} |
| Service CPU % | number |
max({__name__="CPUUtilization", "@instrumentation.@name"="cloudwatch.aws/ecs", ClusterName="<cluster>", ServiceName="<svc>"}) |
{x:45,y:0,w:15,h:120} |
| Request rate | line |
sum(rate({__name__="RequestCount", …, LoadBalancer="<lb>"}[5m])) |
{x:0,y:120,w:30,h:320} |
| Status codes | bar |
sum by (__name__) (rate({__name__=~"HTTPCode_Target_..._Count", …, LoadBalancer="<lb>"}[5m])) |
{x:30,y:120,w:30,h:320} |
| Response time (ms) | line |
{__name__="TargetResponseTime", …, LoadBalancer="<lb>"} * 1000 |
{x:0,y:440,w:60,h:320} |
Optional depth (follow-up "infra / USE health"): Memory util % line; Running
tasks line (RunningTaskCount — needs ECS Container Insights or OTel enrichment,
empty otherwise). If the space runs Application Signals, swap the RED panels for
Service/Fault/Error and the availability formula (1 − Fault/Total) × 100.
Append two-up and re-sequence y.
API Gateway
Request-serving — RED. Identity split: REST v1 on ApiName with 4XXError /
5XXError, HTTP v2 on ApiId with 4xx / 5xx; gateway overhead is the formula
Latency − IntegrationLatency.
CORE (~8 panels — the default build):
| Panel | Viz | Query template | Layout |
|---|---|---|---|
| Request count/s | number |
sum(rate({__name__="Count", "@instrumentation.@name"="cloudwatch.aws/apigateway", ApiName="<api>"}[5m])) |
{x:0,y:0,w:15,h:120} |
| 5xx rate | number |
sum(rate({__name__="5XXError", …, ApiName="<api>"}[5m])) |
{x:15,y:0,w:15,h:120} |
| 4xx rate | number |
sum(rate({__name__="4XXError", …, ApiName="<api>"}[5m])) |
{x:30,y:0,w:15,h:120} |
| Peak latency (ms) | number |
max({__name__="Latency", …, ApiName="<api>"}) |
{x:45,y:0,w:15,h:120} |
| Request count | line |
sum(rate({__name__="Count", …, ApiName="<api>"}[5m])) |
{x:0,y:120,w:30,h:320} |
| Errors 4xx/5xx | bar |
sum by (__name__) (rate({__name__=~"[45]XXError", …, ApiName="<api>"}[5m])) |
{x:30,y:120,w:30,h:320} |
| Latency (ms) | line |
{__name__="Latency", …, ApiName="<api>"} |
{x:0,y:440,w:60,h:320} |
| Top resources by 5xx | table |
topk(10, sum by (Resource) (rate({__name__="5XXError", …, ApiName="<api>"}[5m]))) |
{x:0,y:760,w:60,h:360} |
Optional depth (follow-up "per-service latency"): Integration latency (ms)
line (IntegrationLatency) — chart it beside Latency so the gap reads as
gateway overhead. Append two-up and re-sequence y.