All skills
aws avatar

/amazon-opensearch-service

@04f39cf

Guides migration, provisioning, search, log-analytics, trace-analytics, and Agentic AI Assistant workflows for Amazon OpenSearch Service and Serverless across six capabilities — migration (Solr/ES/self-managed into AOS/AOSS, schema/query translation, sizing, cutover); provisioning (domain + AOSS lifecycle, upgrades, FGAC, monitoring); search (vector / semantic / hybrid / RAG with Bedrock); log-analytics (PPL, OSI, anomaly detection, Dashboards); trace-analytics (OTel spans, service maps, Data Prepper); ai-assistant (natural language data exploration, incident investigation, root cause analysis). Triggers on OpenSearch, AOS, AOSS, Elasticsearch, Solr, vector/k-NN/semantic/hybrid search, RAG, log analytics, PPL, trace analytics, ISM, FAISS, HNSW, Migration Assistant, UltraWarm, OR1, query my data, analyze logs, investigate errors, root cause analysis.

Use this Skill: https://skilld.dev/gh/aws/agent-toolkit-for-aws/amazon-opensearch-service

This session only. Nothing lands on disk.

referencestrace-analytics-trace-queries.md

≈3.6k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Trace-analytics capability — entry point and query templates

The AWS MCP server is recommended for executing these commands but is not required — all steps use standard awscurl / AWS CLI syntax.

This file is the entry point for the trace-analytics capability. It covers distributed traces with OpenTelemetry — span queries, service maps, latency analysis (p50/p95/p99), error rate by service, and root-cause via parent/child spans.

When to use this capability

SKILL.md routes here when the user is working with distributed traces on AOS / AOSS. Concrete triggers:

  • Phrases: "trace analytics", "service map", "otel", "distributed traces", "span query", "otel-v1-apm-span-"*, "Data Prepper", "latency p99"
  • Tasks: query trace spans, build service maps, ingest traces (OTel collector → Data Prepper / OSI), troubleshoot trace pipeline or query issues

All trace-analytics files (capability index)

User need File
Span queries (PPL on otel-v1-apm-span-*) this file
Trace ingestion (OTel collector → Data Prepper / OSI) trace-analytics-trace-ingestion.md
Troubleshoot trace pipeline or queries trace-analytics-troubleshooting.md

Cross-cutting refs you may also load: security.md, personas.md (observability-engineer).

Cross-capability handoff

Query discipline (MUST follow)

For the full PPL command and function catalog, see observability-ppl-reference.md.

  • Unknown command → upstream grammar. If a PPL command or function is not in observability-ppl-reference.md, or an emitted query fails with a syntax error, you MUST fetch the raw upstream doc from opensearch-project/sql under docs/user/ppl/ before answering. You MUST NOT invent PPL syntax, because OpenSearch PPL has version-specific commands that do not exist in other query languages.
  • Verify or disclose. When the domain/collection endpoint is reachable, you MUST validate every emitted query against _plugins/_ppl; on 0 rows, fall back to _plugins/_ppl/_explain to confirm the plan and surface the empty result; on error, fix and re-validate. When no endpoint is reachable, you MUST state that the query is unverified, because an untested query can silently reference a non-existent field.

Data Plane Access with awscurl

All queries below use the PPL API at /_plugins/_ppl. Use awscurl for SigV4-authenticated requests:

Credentials & monitoring: source the SigV4 signing credentials from an IAM role (instance profile, ECS task role, or SSO session), not static access keys. Enable CloudTrail for management-plane audit and OpenSearch audit logs for data-plane access (to attribute PPL queries to callers), and alarm on _plugins/_ppl error rates.

Base Command (AOS)

awscurl --service es --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/_plugins/_ppl" \
  -H 'Content-Type: application/json' \
  -d '{"query": "<PPL_QUERY>"}'

Base Command (AOSS)

awscurl --service aoss --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/_plugins/_ppl" \
  -H 'Content-Type: application/json' \
  -d '{"query": "<PPL_QUERY>"}'

Prerequisites: pip install awscurl, AWS credentials configured via aws configure or environment variables.

Verifying Trace Indices

awscurl --service es --region $AWS_REGION \
  "$OPENSEARCH_ENDPOINT/_cat/indices/otel-v1-apm-*?v&h=index,health,docs.count,store.size"

Sampling Recent Spans

awscurl --service es --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/otel-v1-apm-span-*/_search" \
  -H 'Content-Type: application/json' \
  -d '{"size": 5, "sort": [{"startTime": "desc"}], "query": {"match_all": {}}}'

Trace Index Key Fields

Field Type Description
traceId keyword Unique 128-bit trace identifier
spanId keyword Unique 64-bit span identifier
parentSpanId keyword Parent span ID (empty for root spans)
serviceName keyword Service that produced the span
name keyword Span operation name
kind keyword Span kind (SPAN_KIND_SERVER, SPAN_KIND_CLIENT, SPAN_KIND_INTERNAL, SPAN_KIND_PRODUCER, SPAN_KIND_CONSUMER)
startTime date Span start timestamp
endTime date Span end timestamp
durationInNanos long Span duration in nanoseconds
status.code integer 0=Unset, 1=Ok, 2=Error
attributes.gen_ai.operation.name keyword GenAI operation type
attributes.gen_ai.agent.name keyword Agent name
attributes.gen_ai.agent.id keyword Agent identifier
attributes.gen_ai.request.model keyword Requested model
attributes.gen_ai.usage.input_tokens long Input token count
attributes.gen_ai.usage.output_tokens long Output token count
attributes.gen_ai.tool.name keyword Tool name
attributes.gen_ai.tool.call.id keyword Tool call identifier
attributes.gen_ai.tool.call.arguments text Tool call arguments (JSON)
attributes.gen_ai.tool.call.result text Tool call result (JSON)
attributes.gen_ai.conversation.id keyword Conversation identifier
attributes.error_type keyword Error type category
events.attributes.exception.type keyword Exception class/type
events.attributes.exception.message text Exception message
events.attributes.exception.stacktrace text Exception stacktrace

GenAI Operation Types

Operation Description
invoke_agent Top-level agent invocation
execute_tool Tool execution within agent reasoning
chat LLM chat completion call
embeddings Text embedding generation
retrieval Retrieval operation (e.g., RAG)
create_agent Agent creation/initialization
text_completion Text completion (non-chat)
generate_content Generic content generation

PPL Query Templates

Usage: Replace <PPL_QUERY> in the base command above with any query below. Example:

awscurl --service es --region us-east-1 \
  -X POST "https://my-domain.us-east-1.es.amazonaws.com/_plugins/_ppl" \
  -H 'Content-Type: application/json' \
  -d '{"query": "source=otel-v1-apm-span-* | where `attributes.gen_ai.operation.name` = '\''invoke_agent'\'' | head 20"}'

Agent Invocation Spans

source = otel-v1-apm-span-* | where `attributes.gen_ai.operation.name` = 'invoke_agent' | fields traceId, spanId, `attributes.gen_ai.agent.name`, `attributes.gen_ai.request.model`, durationInNanos, startTime | sort - startTime | head 20

Tool Execution Spans

source = otel-v1-apm-span-* | where `attributes.gen_ai.operation.name` = 'execute_tool' | fields traceId, spanId, `attributes.gen_ai.tool.name`, durationInNanos, startTime | sort - startTime | head 20

Slow Spans

Default threshold: 5 seconds (5,000,000,000 nanoseconds). Adjust as needed.

source = otel-v1-apm-span-* | where durationInNanos > 5000000000 | fields traceId, spanId, serviceName, name, durationInNanos, startTime | sort - durationInNanos | head 20

Error Spans

status.code = 2 means ERROR in OTel:

source = otel-v1-apm-span-* | where `status.code` = 2 | fields traceId, spanId, serviceName, name, `status.code`, startTime | sort - startTime | head 20

Token Usage by Model

source = otel-v1-apm-span-* | where `attributes.gen_ai.usage.input_tokens` > 0 | stats sum(`attributes.gen_ai.usage.input_tokens`) as total_input, sum(`attributes.gen_ai.usage.output_tokens`) as total_output by `attributes.gen_ai.request.model`

Token Usage by Agent

source = otel-v1-apm-span-* | where `attributes.gen_ai.usage.input_tokens` > 0 | stats sum(`attributes.gen_ai.usage.input_tokens`) as total_input, sum(`attributes.gen_ai.usage.output_tokens`) as total_output by `attributes.gen_ai.agent.name`

Service Operations Listing

source = otel-v1-apm-span-* | stats count() by serviceName, `attributes.gen_ai.operation.name`

Trace Tree Reconstruction

source = otel-v1-apm-span-* | where traceId = '<TRACE_ID>' | fields traceId, spanId, parentSpanId, serviceName, name, startTime, endTime, durationInNanos, `status.code` | sort startTime

Root Span Identification

source = otel-v1-apm-span-* | where traceId = '<TRACE_ID>' AND parentSpanId = '' | fields traceId, spanId, serviceName, name, durationInNanos, startTime, endTime

Spans with Exceptions

source = otel-v1-apm-span-* | where `status.code` = 2 | fields traceId, spanId, serviceName, name, `events.attributes.exception.type`, `events.attributes.exception.message`, `attributes.error_type`, startTime | sort - startTime | head 20

Conversation Tracking

source = otel-v1-apm-span-* | where `attributes.gen_ai.conversation.id` != '' | stats count() as turns, sum(`attributes.gen_ai.usage.input_tokens`) as total_input_tokens, sum(`attributes.gen_ai.usage.output_tokens`) as total_output_tokens by `attributes.gen_ai.conversation.id`

Tool Call Inspection

source = otel-v1-apm-span-* | where `attributes.gen_ai.operation.name` = 'execute_tool' | fields traceId, spanId, `attributes.gen_ai.tool.name`, `attributes.gen_ai.tool.call.id`, `attributes.gen_ai.tool.call.arguments`, `attributes.gen_ai.tool.call.result`, durationInNanos, startTime | sort - startTime | head 20

Service Map Queries

Important: In otel-v2-apm-service-map-*, sourceNode and targetNode are nested struct objects with keyAttributes.name for the service name — not flat strings.

Service Topology

source = otel-v2-apm-service-map-* | dedup nodeConnectionHash | fields sourceNode, targetNode, sourceOperation, targetOperation

Remote Service Identification with coalesce()

Different OTel instrumentation libraries use different attributes. Use coalesce() to check multiple fields:

source = otel-v1-apm-span-* | where serviceName = 'frontend' | where kind = 'SPAN_KIND_CLIENT' | eval _remoteService = coalesce(`attributes.net.peer.name`, `attributes.server.address`, `attributes.rpc.service`, `attributes.db.system`, `attributes.gen_ai.system`, 'unknown') | stats count() as calls by _remoteService | sort - calls

Query DSL Examples (awscurl)

For complex aggregations that PPL doesn't support well, use Query DSL with awscurl:

Latency Percentiles by Service

awscurl --service es --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/otel-v1-apm-span-*/_search" \
  -H 'Content-Type: application/json' \
  -d '{
  "size": 0,
  "query": {"range": {"startTime": {"gte": "now-1h"}}},
  "aggs": {
    "by_service": {
      "terms": {"field": "serviceName", "size": 20},
      "aggs": {
        "latency_percentiles": {
          "percentiles": {
            "field": "durationInNanos",
            "percents": [50, 90, 95, 99]
          }
        }
      }
    }
  }
}'

Error Rate by Service

awscurl --service es --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/otel-v1-apm-span-*/_search" \
  -H 'Content-Type: application/json' \
  -d '{
  "size": 0,
  "query": {"range": {"startTime": {"gte": "now-1h"}}},
  "aggs": {
    "by_service": {
      "terms": {"field": "serviceName", "size": 20},
      "aggs": {
        "total": {"value_count": {"field": "spanId"}},
        "errors": {
          "filter": {"term": {"status.code": 2}},
          "aggs": {
            "count": {"value_count": {"field": "spanId"}}
          }
        }
      }
    }
  }
}'

Throughput Over Time

awscurl --service es --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/otel-v1-apm-span-*/_search" \
  -H 'Content-Type: application/json' \
  -d '{
  "size": 0,
  "query": {"range": {"startTime": {"gte": "now-1h"}}},
  "aggs": {
    "over_time": {
      "date_histogram": {
        "field": "startTime",
        "fixed_interval": "5m"
      },
      "aggs": {
        "by_service": {
          "terms": {"field": "serviceName", "size": 10}
        }
      }
    }
  }
}'

Slow Operations (P99 > 1s)

awscurl --service es --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/otel-v1-apm-span-*/_search" \
  -H 'Content-Type: application/json' \
  -d '{
  "size": 0,
  "aggs": {
    "by_operation": {
      "terms": {"field": "name", "size": 50},
      "aggs": {
        "p99_latency": {
          "percentiles": {
            "field": "durationInNanos",
            "percents": [99]
          }
        },
        "high_latency": {
          "bucket_selector": {
            "buckets_path": {"p99": "p99_latency.99"},
            "script": "params.p99 > 1000000000"
          }
        }
      }
    }
  }
}'

Find Spans by Service (DSL)

awscurl --service es --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/otel-v1-apm-span-*/_search" \
  -H 'Content-Type: application/json' \
  -d '{
  "query": {
    "bool": {
      "must": [
        {"term": {"serviceName": "ORDER_SERVICE"}},
        {"range": {"startTime": {"gte": "now-1h"}}}
      ]
    }
  },
  "sort": [{"startTime": "desc"}],
  "size": 20
}'

Get Full Trace by ID (DSL)

awscurl --service es --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/otel-v1-apm-span-*/_search" \
  -H 'Content-Type: application/json' \
  -d '{
  "query": {"term": {"traceId": "TRACE_ID_HERE"}},
  "sort": [{"startTime": "asc"}],
  "size": 100
}'

Service Map (DSL)

awscurl --service es --region $AWS_REGION \
  -X POST "$OPENSEARCH_ENDPOINT/otel-v2-apm-service-map-*/_search" \
  -H 'Content-Type: application/json' \
  -d '{"size": 200, "query": {"match_all": {}}}'

Source: SKILL.md on GitHub

No alerts28d3 checks · Risk SAFE
  • Gen Agent Trust Hub28d

    This skill is a highly structured and security-conscious guide for managing Amazon OpenSearch Service and Serverless. It provides comprehensive instructions for migrations, provisioning, and analytics while strictly adhering to AWS security best practices, such as using SigV4 signing, IAM least-privilege, and AWS Secrets Manager for credential handling.

  • Socket28d

    No alerts

  • Snyk28d

    Risk: LOW · No issues

Signed by skilld at 04f39cf. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 months ago
metadata
{
  "version": "2"
}

README badge

README badge for aws/agent-toolkit-for-aws/amazon-opensearch-service