All skills
aws avatar

/amazon-opensearch-service

@04f39cf

Guides migration, provisioning, search, log-analytics, trace-analytics, and Agentic AI Assistant workflows for Amazon OpenSearch Service and Serverless across six capabilities — migration (Solr/ES/self-managed into AOS/AOSS, schema/query translation, sizing, cutover); provisioning (domain + AOSS lifecycle, upgrades, FGAC, monitoring); search (vector / semantic / hybrid / RAG with Bedrock); log-analytics (PPL, OSI, anomaly detection, Dashboards); trace-analytics (OTel spans, service maps, Data Prepper); ai-assistant (natural language data exploration, incident investigation, root cause analysis). Triggers on OpenSearch, AOS, AOSS, Elasticsearch, Solr, vector/k-NN/semantic/hybrid search, RAG, log analytics, PPL, trace analytics, ISM, FAISS, HNSW, Migration Assistant, UltraWarm, OR1, query my data, analyze logs, investigate errors, root cause analysis.

Use this Skill: https://skilld.dev/gh/aws/agent-toolkit-for-aws/amazon-opensearch-service

This session only. Nothing lands on disk.

referencesprovisioning-monitoring.md

≈874 tokens on demand. Your agent reads this file only when SKILL.md points to it.

CloudWatch Monitoring for AOS

Note: Enable OpenSearch application logs (index slow logs, search slow logs, error logs, audit logs) and configure CloudTrail for API-level auditing. Store logs in encrypted CloudWatch Logs groups (specify --kms-key-id at log group creation: aws logs create-log-group --log-group-name /aws/opensearch/my-domain --kms-key-id arn:aws:kms:<region>:<account>:key/<key-id>).

Key Metrics to Monitor

Metric Threshold Action
CPUUtilization > 80% sustained Scale up instance type or add nodes
JVMMemoryPressure > 80% Increase instance size; check for large aggregations
ClusterStatus.red = 1 Immediate: check for unassigned shards
ClusterStatus.yellow = 1 Investigate: replica shards not allocated
FreeStorageSpace < 20 GB (adjust based on provisioned storage) Add EBS capacity or migrate old indices to UltraWarm
SearchLatency > 500ms p99 Optimize queries; consider adding data nodes
IndexingLatency > 100ms p99 Check bulk queue; scale indexing capacity
ThreadpoolSearchRejected > 0 Search queue full; scale or throttle clients

Creating CloudWatch Alarms

Cluster Health (Red)

aws cloudwatch put-metric-alarm --alarm-name aos-cluster-red \
  --namespace AWS/ES --metric-name ClusterStatus.red \
  --dimensions Name=DomainName,Value=my-domain Name=ClientId,Value=<account-id> \
  --statistic Maximum --period 60 --evaluation-periods 1 \
  --threshold 1 --comparison-operator GreaterThanOrEqualToThreshold \
  --alarm-actions arn:aws:sns:<region>:<account>:my-alerts

REQUIRED: SNS topics receiving CloudWatch alarms MUST have KMS encryption enabled. CloudWatch alarm notifications may contain cluster status, metric values, and other sensitive operational data. Enable encryption when creating the topic:

aws sns create-topic --name my-alerts \
  --attributes KmsMasterKeyId=alias/aws/sns

For existing topics: aws sns set-topic-attributes --topic-arn <arn> --attribute-name KmsMasterKeyId --attribute-value alias/aws/sns Verify all SNS subscription recipients belong to authorized personnel before deploying alarms.

JVM Memory Pressure

aws cloudwatch put-metric-alarm --alarm-name aos-jvm-pressure \
  --namespace AWS/ES --metric-name JVMMemoryPressure \
  --dimensions Name=DomainName,Value=my-domain Name=ClientId,Value=<account-id> \
  --statistic Maximum --period 300 --evaluation-periods 3 \
  --threshold 80 --comparison-operator GreaterThanOrEqualToThreshold \
  --alarm-actions arn:aws:sns:<region>:<account>:my-alerts

Free Storage Space

aws cloudwatch put-metric-alarm --alarm-name aos-low-storage \
  --namespace AWS/ES --metric-name FreeStorageSpace \
  --dimensions Name=DomainName,Value=my-domain Name=ClientId,Value=<account-id> \
  --statistic Minimum --period 300 --evaluation-periods 1 \
  --threshold 20480 --comparison-operator LessThanOrEqualToThreshold \
  --alarm-actions arn:aws:sns:<region>:<account>:my-alerts

Recommended Alarm Set

For production domains, create alarms for:

  1. ClusterStatus.red (immediate)
  2. ClusterStatus.yellow (sustained 15 min)
  3. JVMMemoryPressure > 80% (sustained 15 min)
  4. CPUUtilization > 80% (sustained 15 min)
  5. FreeStorageSpace < 20 GB (immediate; adjust based on provisioned storage)
  6. ThreadpoolSearchRejected > 0 (sum over 5 min)
  7. AutomatedSnapshotFailure > 0 (immediate)

Source: SKILL.md on GitHub

No alerts28d3 checks · Risk SAFE
  • Gen Agent Trust Hub28d

    This skill is a highly structured and security-conscious guide for managing Amazon OpenSearch Service and Serverless. It provides comprehensive instructions for migrations, provisioning, and analytics while strictly adhering to AWS security best practices, such as using SigV4 signing, IAM least-privilege, and AWS Secrets Manager for credential handling.

  • Socket28d

    No alerts

  • Snyk28d

    Risk: LOW · No issues

Signed by skilld at 04f39cf. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 months ago
metadata
{
  "version": "2"
}

README badge

README badge for aws/agent-toolkit-for-aws/amazon-opensearch-service