All skills
redis avatar

/redis-semantic-cache

@66c796d official
by redisredis/agent-skills163 stars
30

Redis LangCache guidance for semantic caching of LLM responses on Redis Cloud — calling search/set via the SDK or REST API, tuning the similarity threshold, separating caches per task type, and filtering with custom attributes. Use when caching LLM completions or RAG answers to cut API cost and latency, building a cache-aside layer in front of OpenAI / Anthropic / etc., tuning hit rate vs precision, or splitting one app's LLM workloads into multiple LangCache caches.

Use this Skill: https://skilld.dev/gh/redis/agent-skills/redis-semantic-cache

This session only. Nothing lands on disk.

SKILL.md

≈124 tokens always: the name and description. ≈879 when used: this file. ≈965 more on demand in 2 files.

Redis Semantic Cache

Semantic caching for LLM responses with Redis Cloud's LangCache service. Stores prompts as embeddings; subsequent semantically-similar prompts return the cached response without re-calling the model.

LangCache is currently in preview on Redis Cloud. Features and behavior may change.

When to apply

  • Wrapping an LLM call (OpenAI, Anthropic, etc.) with a cache layer to cut cost and latency.
  • Caching RAG answers, classification outputs, or any deterministic LLM workload.
  • Tuning the precision/hit-rate trade-off for a semantic cache.
  • Splitting one application's LLM workloads across multiple cache instances.

1. The cache-aside flow

LangCache fits in front of any LLM call as a standard cache-aside pattern:

  1. Send the user's prompt to LangCache's search.
  2. Cache hit — return the stored response directly.
  3. Cache miss — call the LLM, then set the response so future similar prompts hit.
from langcache import LangCache
import os

lang_cache = LangCache(
    server_url=f"https://{os.getenv('HOST')}",
    cache_id=os.getenv("CACHE_ID"),
    api_key=os.getenv("API_KEY"),
)

result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.9)
if result:
    response = result[0]["response"]
else:
    response = llm.generate("What is Redis?")
    lang_cache.set(prompt="What is Redis?", response=response)

The same operations are available via REST (POST /v1/caches/{cacheId}/entries/search and POST /v1/caches/{cacheId}/entries) when an SDK isn't an option.

See references/langcache-usage.md for full SDK + REST samples and attribute-based storage.

2. Tune the similarity threshold

The threshold controls how close (in embedding cosine distance) a new prompt must be to a cached one to count as a hit. Higher = stricter match, fewer false positives. Lower = more hits, more risk of returning an off-topic answer.

Threshold Behavior Use when
0.95+ Near-exact match required Customer-facing answers where wrong responses are costly
0.9 Balanced default Most workloads — start here
0.8 Loose semantic match Internal tools, exploratory queries, FAQ deduplication
# Stricter — fewer false positives
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.95)

# Looser — higher hit rate
result = lang_cache.search(prompt="What is Redis?", similarity_threshold=0.8)

Adjust by watching the actual cache-hit rate and spot-checking that returned answers are still relevant.

See references/best-practices.md.

3. Separate caches per task type

Different LLM workloads should not share one cache — a "code question" prompt is semantically close to other code questions but has nothing to do with a password-reset support query, and crossing them returns garbage.

support_cache = LangCache(server_url=..., cache_id="support-cache-id", api_key=...)
code_cache    = LangCache(server_url=..., cache_id="code-cache-id",    api_key=...)

Create distinct cache IDs in Redis Cloud per task, and route each call to the right one. As a finer-grained alternative, store and search with custom attributes (e.g. {"category": "database"}) to keep tasks in the same cache but isolated by attribute filter — useful when the same prompt format spans subtopics.

References

Source: SKILL.md on GitHub

1 warning1mo3 checks · Risk SAFE
  • Gen Agent Trust Hub1mo

    This skill provides legitimate developer guidance and code examples for using Redis LangCache to implement semantic caching for LLM responses. It follows security best practices by utilizing environment variables for credentials and exclusively references official Redis resources.

  • Socket1mo

    No alerts

  • Snyk1mo

    Risk: MEDIUM · 2 issues

Signed by skilld at 66c796d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 days ago.

Activeupdated 4 months ago
metadata
{
  "author": "Redis, Inc.",
  "version": "0.1.0"
}
  • redis
  • semantic-caching
  • llm
  • embeddings
  • openai
  • anthropic
  • rag
  • cache-aside
  • langcache

README badge

README badge for redis/agent-skills/redis-semantic-cache

Implements semantic caching for LLM responses using Redis Cloud's LangCache service, storing prompts as embeddings to return cached answers for similar queries without re-calling the model. Use to reduce API costs and latency for OpenAI/Anthropic calls, RAG workloads, or classification tasks by tuning similarity thresholds and separating caches by task type.

Generated from the current SKILL.md.

What LLM providers does this work with?
LangCache works as a cache-aside layer in front of any LLM API — OpenAI, Anthropic, or others. You call the LLM only on cache misses.
Is LangCache production-ready?
LangCache is currently in preview on Redis Cloud. Features and behavior may change.
How do I choose a similarity threshold?
Start with 0.9 for most workloads. Use 0.95+ for customer-facing answers where wrong responses are costly, and 0.8 for internal tools where higher hit rate matters more than precision.
Can I cache different types of LLM tasks in one app?
Yes — create separate cache IDs per task type (e.g. support-cache-id, code-cache-id), or store custom attributes in the same cache and filter by them when searching.
Does this support REST API or only SDK?
Both. The SDK and REST API (`POST /v1/caches/{cacheId}/entries/search` and `POST /v1/caches/{cacheId}/entries`) offer the same operations.

Generated from the current SKILL.md. These answers refresh after source changes.