All skills
aws avatar

/amazon-opensearch-service

@04f39cf

Guides migration, provisioning, search, log-analytics, trace-analytics, and Agentic AI Assistant workflows for Amazon OpenSearch Service and Serverless across six capabilities — migration (Solr/ES/self-managed into AOS/AOSS, schema/query translation, sizing, cutover); provisioning (domain + AOSS lifecycle, upgrades, FGAC, monitoring); search (vector / semantic / hybrid / RAG with Bedrock); log-analytics (PPL, OSI, anomaly detection, Dashboards); trace-analytics (OTel spans, service maps, Data Prepper); ai-assistant (natural language data exploration, incident investigation, root cause analysis). Triggers on OpenSearch, AOS, AOSS, Elasticsearch, Solr, vector/k-NN/semantic/hybrid search, RAG, log analytics, PPL, trace analytics, ISM, FAISS, HNSW, Migration Assistant, UltraWarm, OR1, query my data, analyze logs, investigate errors, root cause analysis.

Use this Skill: https://skilld.dev/gh/aws/agent-toolkit-for-aws/amazon-opensearch-service

This session only. Nothing lands on disk.

referencesprovisioning-serverless-deploy-search.md

≈1.7k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Amazon OpenSearch Serverless — Deploy Search Configuration

The AWS MCP server is recommended for executing these commands but is not required — all steps use standard AWS CLI syntax.

Deploy indices, ML models, and pipelines to a provisioned serverless collection.

Route by Strategy

  • Neural Sparse → Neural Sparse Path
  • Dense Vector or Hybrid → Dense Vector Path
  • BM25 → BM25 Path

Neural Sparse Path (Automatic Semantic Enrichment)

Create index with automatic enrichment via AWS API:

POST /opensearchserverless/CreateIndex
{
  "id": "<collection-id>",
  "indexName": "<index-name>",
  "indexSchema": {
    "mappings": {
      "properties": {
        "<text-field>": {
          "type": "text",
          "semantic_enrichment": {
            "status": "ENABLED",
            "language_options": "english"
          }
        }
      }
    }
  }
}

Note: Use aws opensearchserverless create-index for this operation (or call_aws opensearchserverless create-index if the AWS MCP server is available). The semantic_enrichment configuration is specified in the index schema.

  • language_options: "english" or "multi-lingual"
  • System automatically deploys sparse model and creates ingest/search pipelines
  • Standard match queries are automatically rewritten to neural sparse queries
  • No manual model or pipeline management required

Dense Vector Path

1. Create IAM Role for Bedrock

# Both aws:SourceAccount and aws:SourceArn conditions are required to prevent
# confused-deputy: ArnLike narrows trust to a specific AOSS collection so
# other collections in the same account can't assume this role.
aws iam create-role --role-name opensearch-bedrock-role \
  --assume-role-policy-document '{
    "Version":"2012-10-17",
    "Statement":[{
      "Effect":"Allow",
      "Principal":{"Service":"ml.opensearchservice.amazonaws.com"},
      "Action":"sts:AssumeRole",
      "Condition":{
        "StringEquals":{"aws:SourceAccount":"<account>"},
        "ArnLike":     {"aws:SourceArn":    "arn:aws:aoss:<region>:<account>:collection/<collection-id>"}
      }
    }]
  }'

aws iam put-role-policy --role-name opensearch-bedrock-role \
  --policy-name BedrockInvokePolicy \
  --policy-document '{"Version":"2012-10-17","Statement":[{"Effect":"Allow","Action":"bedrock:InvokeModel","Resource":"arn:aws:bedrock:<region>::foundation-model/amazon.titan-embed-text-v2:0"}]}'

2. Create ML Connector

POST <collection-endpoint>/_plugins/_ml/connectors/_create
{
  "name": "Amazon Bedrock Titan Embedding V2",
  "version": 1,
  "protocol": "aws_sigv4",
  "parameters": { "region": "<aws-region>", "service_name": "bedrock" },
  "credential": { "roleArn": "<iam_role_arn>" },
  "actions": [{
    "action_type": "predict",
    "method": "POST",
    "url": "https://bedrock-runtime.<aws-region>.amazonaws.com/model/amazon.titan-embed-text-v2:0/invoke",
    "headers": { "content-type": "application/json", "x-amz-content-sha256": "required" },
    "request_body": "{ \"inputText\": \"${parameters.inputText}\" }",
    "pre_process_function": "connector.pre_process.bedrock.embedding",
    "post_process_function": "connector.post_process.bedrock.embedding"
  }]
}

3. Register and Deploy Model

POST <collection-endpoint>/_plugins/_ml/model_groups/_register
{ "name": "bedrock_embedding_models", "description": "Bedrock embedding model group" }

POST <collection-endpoint>/_plugins/_ml/models/_register
{
  "name": "bedrock-titan-embed-v2",
  "function_name": "remote",
  "model_group_id": "<model_group_id>",
  "connector_id": "<connector_id>"
}

POST <collection-endpoint>/_plugins/_ml/models/<model-id>/_deploy

Test: POST /_plugins/_ml/models/<model-id>/_predict with {"parameters": {"inputText": "hello world"}}. Verify 1024-dim embeddings.

4. Create Ingest Pipeline

PUT <collection-endpoint>/_ingest/pipeline/bedrock-embedding-pipeline
{
  "processors": [{
    "text_embedding": {
      "model_id": "<model_id>",
      "field_map": { "<text-field>": "<vector-field>" }
    }
  }]
}

5. Create Index

PUT <collection-endpoint>/<index-name>
{
  "settings": {
    "index.knn": true,
    "index.default_pipeline": "bedrock-embedding-pipeline"
  },
  "mappings": {
    "properties": {
      "<text-field>": { "type": "text" },
      "<vector-field>": {
        "type": "knn_vector",
        "dimension": 1024
      }
    }
  }
}

Note: Omit engine/mode unless you have confirmed support for your collection generation — NextGen rejects them (Field parameter 'engine' is not supported) and auto-selects the engine. NextGen support can change over time, so rather than treating this as a fixed prohibition, prefer omitting these fields and, if you need to set them, attempt creation and handle any validation error (same pattern used for OCU values in provisioning-serverless-provision.md). space_type is optional (defaults to L2); to use another metric, add it at the field level, e.g. "space_type": "cosinesimil".

6. Search Pipeline (hybrid only)

PUT <collection-endpoint>/_search/pipeline/hybrid-search-pipeline
{
  "phase_results_processors": [{
    "normalization-processor": {
      "normalization": { "technique": "min_max" },
      "combination": { "technique": "arithmetic_mean", "parameters": { "weights": [0.3, 0.7] } }
    }
  }]
}

BM25 Path

Create index with text mappings:

PUT <collection-endpoint>/<index-name>
{ "mappings": { "properties": { "<text-field>": { "type": "text" } } } }

Index Sample Documents & Test

After index creation (all paths):

  1. Index test documents to verify setup
  2. Test search queries:
    • Neural Sparse: standard match queries (auto-rewritten)
    • Dense Vector: neural query with model_id
    • BM25: standard match queries

Next Step (optional)

Security Considerations

  • Encryption in transit: all data-plane calls (index creation, indexing, search) MUST use HTTPS. Verify the collection endpoint starts with https://, because an unencrypted request exposes documents and queries in transit.
  • Encryption at rest: indexed embeddings and document content are sensitive; AOSS encrypts at rest by default, and compliance workloads should use a customer-managed KMS key on the encryption policy (see provisioning-serverless-provision.md).
  • Least-privilege connector role: scope the ML model connector role to the minimum data-access actions and the specific model ARN it invokes, because a broad connector role is a standing path into the data plane.

Source: SKILL.md on GitHub

No alerts28d3 checks · Risk SAFE
  • Gen Agent Trust Hub28d

    This skill is a highly structured and security-conscious guide for managing Amazon OpenSearch Service and Serverless. It provides comprehensive instructions for migrations, provisioning, and analytics while strictly adhering to AWS security best practices, such as using SigV4 signing, IAM least-privilege, and AWS Secrets Manager for credential handling.

  • Socket28d

    No alerts

  • Snyk28d

    Risk: LOW · No issues

Signed by skilld at 04f39cf. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 2 months ago
metadata
{
  "version": "2"
}

README badge

README badge for aws/agent-toolkit-for-aws/amazon-opensearch-service