---
name: dashboard-builder
description: Build useful monitoring dashboards for Grafana, SigNoz, and similar tools. Use this skill to turn metrics into clear views that help operators find problems and take action.
origin: ECC direct-port adaptation
version: "1.0.0"
---

# Dashboard Builder

Use this skill when the task is to build a dashboard that people can use during real work.

Do not try to show every metric. Build the dashboard to answer these questions:

- Is the system healthy?
- Where is the slow point?
- What changed?
- What should someone do next?

## When to Use

Use this skill for tasks like:

- "Build a Kafka monitoring dashboard."
- "Create a Grafana dashboard for Elasticsearch."
- "Make a SigNoz dashboard for this service."
- "Turn this metric list into a useful operations dashboard."

## Rules

- Start with operator questions, not page layout.
- Do not add every metric you can find.
- Keep health, traffic, speed, and resource panels in clear groups.
- Give every panel a clear title and unit.
- Add limits only when they have a known meaning.
- Do not guess metric names, labels, units, or safe limits.
- Do not use fake data in a final dashboard.
- Do not add charts that lead to the same action.
- Do not add analytics, telemetry, or new data calls.
- Keep secrets, tokens, host names, and private IDs out of the file.

## Workflow

### 1. Learn the Use Case

Write down:

- Who will use the dashboard
- What system it covers
- What action the user can take
- What time range matters
- Which team owns each problem

If this is not known, state your assumptions.

### 2. Define Operator Questions

Group questions into these areas:

- Health and uptime
- Speed and wait time
- Traffic and work volume
- Resource use and full limits
- Risks that are special to the service

Each panel must answer one clear question.

### 3. Check the Target Format

Inspect an existing local dashboard or template first.

Confirm:

- File and JSON shape
- Query language
- Data source name or ID
- Time range and refresh rate
- Variable and filter rules
- Panel types that the tool supports

Do not copy private values from another dashboard.

If no example exists, use the smallest valid format for the target tool. Mark any value that the user must fill in.

### 4. Pick Metrics

Add a metric only if it helps the user:

- See a problem
- Find the cause
- Judge the size of the problem
- Know what action to take

For each metric, check:

- Exact name
- Labels and filters
- Unit
- Normal range
- Missing data behavior
- Whether it is a count, rate, gauge, or total

Use rates for counters. Show percentiles for wait time when they exist. Show both load and limit for resources.

### 5. Build the Layout

Place the most important facts at the top.

Use this order when it fits:

1. Main health state
2. Error rate and request rate
3. Wait time
4. Busy or full resources
5. Service-specific risks
6. Detail panels for finding the cause

Group panels by question. Keep related panels next to each other.

### 6. Write Clear Panels

Each panel needs:

- A title that asks or answers a question
- A correct unit
- A useful note or short help text
- A legend with clear names
- A query that handles the chosen time range
- A clear state for no data

Use color with care. Red should mean that action is needed. Do not use red only because a value is high.

### 7. Set Limits

Use known service goals, alert rules, or system limits.

If no safe limit is known:

- Do not invent one.
- Show the raw trend.
- Add a note that the limit must be set.
- Ask for the service goal if user input is allowed.

### 8. Handle Edge Cases

Plan for:

- No data
- Late data
- A new service with little history
- A counter that resets
- A label with too many values
- A divide-by-zero result
- One host hiding a bad group result
- An average hiding a slow request
- A changed metric name
- A missing data source
- A dashboard used across many regions or teams

Keep filters safe. Set useful default values. Avoid queries that can return a huge number of lines.

### 9. Check the Result

Before delivery, verify:

- The file is valid for the target tool.
- Every query uses real metric names.
- Units match the data.
- Filters work.
- No panel is empty due to a bad label.
- No two panels give the same answer.
- The top row shows health at a glance.
- Each bad state points to a next step.
- No secret or private value is present.

## Concrete Example

Task: Build a Grafana dashboard for a web API.

Operator questions:

- Is the API up?
- Are users seeing errors?
- Are requests slow?
- Is traffic rising?
- Are CPU or memory limits close?

Suggested panels:

1. **Is the API up?**
   - Show the share of successful health checks.
   - Unit: percent.
   - No data: show "No health data."

2. **Are users seeing errors?**
   - Show the error request rate divided by the full request rate.
   - Split by route only if the route list is small.
   - Unit: percent.

3. **Are requests slow?**
   - Show p50, p95, and p99 wait time.
   - Unit: milliseconds.
   - Do not use only the average.

4. **How much traffic is there?**
   - Show requests per second.
   - Split by service or region when useful.

5. **Are resources close to full?**
   - Show CPU use against CPU limit.
   - Show memory use against memory limit.
   - Unit: percent.

6. **What should the user check next?**
   - Add detail panels for the worst route, host, region, or error type.
   - Link only to local views that already exist.

The first row should show uptime, errors, wait time, and traffic. Resource and detail panels should come after it.