All skills
wshobson avatar

/similarity-search-patterns

@be57c0b
by Seth Hobsonwshobson/agents40k stars
4,281

Implement efficient similarity search with vector databases. Use when building semantic search, implementing nearest neighbor queries, or optimizing retrieval performance.

Use this Skill: https://skilld.dev/gh/wshobson/agents/similarity-search-patterns

This session only. Nothing lands on disk.

SKILL.md

β‰ˆ50 tokens always: the name and description. β‰ˆ637 when used: this file. β‰ˆ3.9k more on demand in 1 file.

Similarity Search Patterns

Patterns for implementing efficient similarity search in production systems.

When to Use This Skill

  • Building semantic search systems
  • Implementing RAG retrieval
  • Creating recommendation engines
  • Optimizing search latency
  • Scaling to millions of vectors
  • Combining semantic and keyword search

Core Concepts

1. Distance Metrics

| Metric | Formula | Best For | | ------------------ | ------------------ | --------------------- | --- | -------------- | | Cosine | 1 - (AΒ·B)/(β€–Aβ€–β€–Bβ€–) | Normalized embeddings | | Euclidean (L2) | √Σ(a-b)Β² | Raw embeddings | | Dot Product | AΒ·B | Magnitude matters | | Manhattan (L1) | Ξ£ | a-b | | Sparse vectors |

2. Index Types

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                 Index Types                      β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚    Flat     β”‚     HNSW      β”‚    IVF+PQ         β”‚
β”‚ (Exact)     β”‚ (Graph-based) β”‚ (Quantized)       β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ O(n) search β”‚ O(log n)      β”‚ O(√n)             β”‚
β”‚ 100% recall β”‚ ~95-99%       β”‚ ~90-95%           β”‚
β”‚ Small data  β”‚ Medium-Large  β”‚ Very Large        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's

  • Use appropriate index - HNSW for most cases
  • Tune parameters - ef_search, nprobe for recall/speed
  • Implement hybrid search - Combine with keyword search
  • Monitor recall - Measure search quality
  • Pre-filter when possible - Reduce search space

Don'ts

  • Don't skip evaluation - Measure before optimizing
  • Don't over-index - Start with flat, scale up
  • Don't ignore latency - P99 matters for UX
  • Don't forget costs - Vector storage adds up

Source: SKILL.md on GitHub

No alerts16d5 checks Β· Risk SAFE
  • Gen Agent Trust Hub16d

    The skill provides standard implementation templates and documentation for popular vector databases, including Pinecone, Qdrant, pgvector, and Weaviate. The code follows standard security practices for database interactions and uses well-known, legitimate libraries for search and machine learning tasks.

  • Socket16d

    No alerts

  • Snyk16d

    Risk: LOW Β· No issues

  • Runlayer6mo

    1/1 file flagged

  • ZeroLeaks5mo

    Score: 93/100 Β· 2 sections analyzed

Signed by skilld at be57c0b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 3 days ago.

Activeupdated 4 months ago
  • similarity-search
  • vector-databases
  • semantic-search
  • rag
  • nearest-neighbor
  • hnsw
  • embeddings
  • retrieval
  • recommendation-engines

README badge

README badge for wshobson/agents/similarity-search-patterns

Teaches patterns for efficient similarity search across vector databases, covering distance metrics, index types (flat, HNSW, IVF+PQ), and hybrid search strategies. Useful for building semantic search systems, RAG retrieval, and recommendation engines at scale.

Generated from the current SKILL.md.

What vector databases does this skill cover?
The skill teaches patterns applicable across vector databases (Faiss, Pinecone, Weaviate, etc.) rather than being tied to a specific one. It focuses on distance metrics, index types, and retrieval strategies that work across platforms.
Does this skill include code templates?
Yes. Detailed worked examples and concrete templates are in `references/details.md`, which you should read when implementing a specific pattern.
What distance metric should I use?
Cosine distance is the default for normalized embeddings. Euclidean (L2) works for raw embeddings, dot product when magnitude matters, and Manhattan (L1) for sparse vectors.
When should I use HNSW vs flat vs IVF+PQ indexes?
Use flat for small datasets (exact O(n) search), HNSW for medium to large (O(log n) with 95-99% recall), and IVF+PQ for very large scale (O(√n) with 90-95% recall and lower memory).

Generated from the current SKILL.md. These answers refresh after source changes.