All skills
wshobson avatar

/embedding-strategies

@be57c0b
by Seth Hobsonwshobson/agents40k stars
4,281

Select and optimize embedding models for semantic search and RAG applications. Use when choosing embedding models, implementing chunking strategies, or optimizing embedding quality for specific domains.

Use this Skill: https://skilld.dev/gh/wshobson/agents/embedding-strategies

This session only. Nothing lands on disk.

SKILL.md

≈56 tokens always: the name and description. ≈649 when used: this file. ≈4.2k more on demand in 1 file.

Embedding Strategies

Guide to selecting and optimizing embedding models for vector search applications.

When to Use This Skill

  • Choosing embedding models for RAG
  • Optimizing chunking strategies
  • Fine-tuning embeddings for domains
  • Comparing embedding model performance
  • Reducing embedding dimensions
  • Handling multilingual content

Core Concepts

1. Embedding Model Comparison (2026)

Model Dimensions Max Tokens Best For
voyage-3-large 1024 32000 Claude apps (Anthropic recommended)
voyage-3 1024 32000 Claude apps, cost-effective
voyage-code-3 1024 32000 Code search
voyage-finance-2 1024 32000 Financial documents
voyage-law-2 1024 32000 Legal documents
text-embedding-3-large 3072 8191 OpenAI apps, high accuracy
text-embedding-3-small 1536 8191 OpenAI apps, cost-effective
bge-large-en-v1.5 1024 512 Open source, local deployment
all-MiniLM-L6-v2 384 256 Fast, lightweight
multilingual-e5-large 1024 512 Multi-language

2. Embedding Pipeline

Document → Chunking → Preprocessing → Embedding Model → Vector
                ↓
        [Overlap, Size]  [Clean, Normalize]  [API/Local]

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's

  • Match model to use case: Code vs prose vs multilingual
  • Chunk thoughtfully: Preserve semantic boundaries
  • Normalize embeddings: For cosine similarity search
  • Batch requests: More efficient than one-by-one
  • Cache embeddings: Avoid recomputing for static content
  • Use Voyage AI for Claude apps: Recommended by Anthropic

Don'ts

  • Don't ignore token limits: Truncation loses information
  • Don't mix embedding models: Incompatible vector spaces
  • Don't skip preprocessing: Garbage in, garbage out
  • Don't over-chunk: Lose important context
  • Don't forget metadata: Essential for filtering and debugging

Source: SKILL.md on GitHub

No alerts16d5 checks · Risk SAFE
  • Gen Agent Trust Hub16d

    This skill provides best practices, templates, and code implementations for working with text embedding models in semantic search and retrieval-augmented generation (RAG) applications. No security risks were identified.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW · No issues

  • Runlayer6mo

    1 file scanned · No issues

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at be57c0b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 4 months ago
  • embeddings
  • rag
  • semantic-search
  • vector-search
  • chunking
  • voyage-ai
  • openai
  • multilingual
  • fine-tuning

README badge

README badge for wshobson/agents/embedding-strategies

Guides selection and optimization of embedding models for vector search and RAG applications, with comparisons of Voyage AI, OpenAI, and open-source options across dimensions, token limits, and domain specialization. Covers chunking strategies, preprocessing, caching, and best practices for matching embeddings to use cases like code search, legal documents, and multilingual content.

Generated from the current SKILL.md.

Which embedding model should I use with Claude?
Voyage AI models (voyage-3-large, voyage-3, or domain-specific variants like voyage-code-3) are recommended by Anthropic for Claude applications. They support up to 32,000 tokens and are optimized for semantic search in RAG workflows.
Does this skill cover local embedding deployment?
Yes. The skill includes open-source models like bge-large-en-v1.5 and all-MiniLM-L6-v2 suitable for local deployment, though it does not provide deployment infrastructure — only model selection and optimization guidance.
Can I use this skill for multilingual content?
Yes. The skill covers multilingual embedding models like multilingual-e5-large and discusses handling multilingual content as a core use case.
What chunking strategies does this skill explain?
The skill outlines the embedding pipeline and chunking best practices (preserving semantic boundaries, appropriate chunk size and overlap), but refers to `references/details.md` for concrete templates and worked examples.
Can I reduce embedding dimensions to save storage?
Yes. The skill lists models with varying dimensions (from 384 to 3072) and covers dimension reduction as an optimization technique, though specific reduction methods are detailed in the referenced templates file.

Generated from the current SKILL.md. These answers refresh after source changes.