All skills

Safe Redis patterns for caching, sessions, rate limits, locks, messaging, key design, and connection setup in production apps.

  • 1 file
  • 14.1 KB
  • Updated 2 weeks ago
  • GitHub

Use this Skill: https://skilld.dev/gh/agenticluke/redis-production-patterns-plus/skill

This session only. Nothing lands on disk.

SKILL.md

≈33 tokens always: the name and description. ≈3.6k when used: this file.

Redis Patterns

Original work by ECC. Credit must stay with ECC.

Use this guide when you add Redis to a backend app.

Core Rules

  • Pick a data type that fits the job.
  • Reuse pooled connections.
  • Set a TTL on short-lived data.
  • Do not set a TTL on data that must stay.
  • Keep key names clear and stable.
  • Assume Redis can restart or lose data.
  • Use Lua when many steps must act as one step.
  • Track errors, memory use, and slow calls.
  • Do not put secrets or private data in key names.

One Redis command is atomic. A set of commands is not atomic unless you use Lua or a transaction.

Redis can save data with RDB files, AOF files, or both. Test restore steps before you trust Redis as a data store.

When to Use This Skill

Use these patterns for:

  • Caches
  • User sessions
  • Rate limits
  • Short-lived tokens
  • Counters and score boards
  • Locks on one Redis node
  • Pub/Sub messages
  • Redis Streams
  • Connection pools
  • Cluster or Sentinel setups

Do not use Redis as the only copy of key business data unless loss is safe and recovery is tested.

Data Types

Need Type Example key
Simple cache String product:123
User session Hash session:abc
Score board Sorted set scores:weekly
Unique values Set visitors:2026-09-17
Small ordered list List feed:user:456
Saved message log Stream events:orders
Counter String with INCR rate:user:123
Rough unique count HyperLogLog hll:pageviews

HyperLogLog gives a rough count. It does not store the values. A Bloom filter is a separate feature and may need a Redis module.

Cache-Aside

Read the cache first. On a miss, read the main data store and fill the cache.

import json
import random
import redis

r = redis.Redis(
    host="localhost",
    port=6379,
    decode_responses=True,
)

def get_product(product_id: int):
    key = f"product:{product_id}"
    cached = r.get(key)

    if cached is not None:
        return json.loads(cached)

    product = db.query_one(
        "SELECT * FROM products WHERE id = %s",
        (product_id,),
    )

    if product is None:
        # Save a short marker to stop many repeat misses.
        r.set(key, json.dumps({"missing": True}), ex=30)
        return None

    # Small random changes stop many keys from ending at once.
    ttl = 3600 + random.randint(0, 300)
    r.set(key, json.dumps(product), ex=ttl)
    return product

If you cache a missing marker, check for it before returning data:

value = json.loads(cached)
if value.get("missing") is True:
    return None
return value

A busy key can cause many app workers to load the same data after its TTL ends. Use a short lock, early refresh, or random TTL changes to limit this cache rush.

Safe Writes and Cache Removal

The main data store should stay the source of truth.

A safe default is:

  1. Write to the main data store.
  2. Remove the old cache value.
  3. Let the next read fill the cache.
def update_product(product_id: int, data: dict) -> None:
    db.execute(
        "UPDATE products SET name = %s WHERE id = %s",
        (data["name"], product_id),
    )
    r.delete(f"product:{product_id}")

Writing new cache data can race with another writer. Cache removal is often safer.

If the data store write and Redis call cannot be one atomic step, plan for a Redis error. A job queue or outbox can retry cache removal.

Tag-Based Cache Removal

Use a set to track related cache keys.

import json

def cache_product(product_id: int, category_id: int, data: dict) -> None:
    key = f"product:{product_id}"
    tag = f"tag:category:{category_id}"

    pipe = r.pipeline(transaction=True)
    pipe.set(key, json.dumps(data), ex=3600)
    pipe.sadd(tag, key)
    pipe.expire(tag, 3900)
    pipe.execute()

def invalidate_category(category_id: int) -> None:
    tag = f"tag:category:{category_id}"

    while True:
        keys = r.spop(tag, 100)
        if not keys:
            break
        r.delete(*keys)

    r.delete(tag)

Do not use KEYS in a live app. It can block Redis. Use SCAN when you must search.

Tag sets may hold keys that have already ended. Keep the tag TTL a little longer than the cache TTL.

Sessions

Use a random session ID. Keep it in a secure cookie. Do not put it in logs or URLs.

import secrets
import time

def create_session(user_id: int, ttl: int = 86400) -> str:
    session_id = secrets.token_urlsafe(32)
    key = f"session:{session_id}"

    pipe = r.pipeline(transaction=True)
    pipe.hset(
        key,
        mapping={
            "user_id": str(user_id),
            "created_at": str(int(time.time())),
        },
    )
    pipe.expire(key, ttl)
    pipe.execute()
    return session_id

def get_session(session_id: str) -> dict | None:
    data = r.hgetall(f"session:{session_id}")
    return data or None

def renew_session(session_id: str, ttl: int = 86400) -> bool:
    return bool(r.expire(f"session:{session_id}", ttl))

def delete_session(session_id: str) -> None:
    r.delete(f"session:{session_id}")

Set cookie flags such as Secure, HttpOnly, and a safe SameSite value.

Choose one session rule:

  • Fixed life: never renew the TTL.
  • Idle life: renew the TTL after valid use.

Rate Limits

Fixed Window

This Lua code adds the counter and TTL as one atomic step.

-- fixed_window.lua
local count = redis.call("INCR", KEYS[1])

if count == 1 then
    redis.call("EXPIRE", KEYS[1], ARGV[1])
end

return count
import time

fixed_window = r.register_script(
    open("fixed_window.lua", encoding="utf-8").read()
)

def is_rate_limited(
    user_id: int,
    limit: int = 100,
    window_seconds: int = 60,
) -> bool:
    window_id = int(time.time()) // window_seconds
    key = f"rate:user:{user_id}:{window_id}"
    count = fixed_window(keys=[key], args=[window_seconds])
    return int(count) > limit

A fixed window allows a burst near the edge of two windows.

Sliding Window

Use a sorted set when you need a smoother limit.

-- sliding_window.lua
local key = KEYS[1]
local seq_key = key .. ":seq"
local now_ms = tonumber(ARGV[1])
local window_ms = tonumber(ARGV[2])
local limit = tonumber(ARGV[3])
local ttl_seconds = math.ceil(window_ms / 1000)

redis.call("ZREMRANGEBYSCORE", key, 0, now_ms - window_ms)

if redis.call("ZCARD", key) >= limit then
    redis.call("EXPIRE", key, ttl_seconds)
    return 0
end

local seq = redis.call("INCR", seq_key)
local member = tostring(now_ms) .. ":" .. tostring(seq)

redis.call("ZADD", key, now_ms, member)
redis.call("EXPIRE", key, ttl_seconds)
redis.call("EXPIRE", seq_key, ttl_seconds)

return 1
sliding_window = r.register_script(
    open("sliding_window.lua", encoding="utf-8").read()
)

def allow_request(user_id: int) -> bool:
    now_ms = int(time.time() * 1000)
    result = sliding_window(
        keys=[f"rate:sliding:user:{user_id}"],
        args=[now_ms, 60_000, 100],
    )
    return bool(result)

Use trusted server time when clients may send false times. Limit by user, API key, or IP as needed. Take care with shared IP addresses.

Locks on One Redis Node

Use SET with NX and a TTL. Delete the lock only if its value still matches your token.

import secrets

RELEASE_LOCK = """
if redis.call("GET", KEYS[1]) == ARGV[1] then
    return redis.call("DEL", KEYS[1])
end
return 0
"""

def acquire_lock(resource: str, ttl_ms: int = 5000) -> str | None:
    token = secrets.token_hex(16)
    ok = r.set(f"lock:{resource}", token, nx=True, px=ttl_ms)
    return token if ok else None

def release_lock(resource: str, token: str) -> bool:
    result = r.eval(RELEASE_LOCK, 1, f"lock:{resource}", token)
    return bool(result)

Concrete use:

token = acquire_lock("order:payment:123", ttl_ms=30_000)

if token is None:
    raise RuntimeError("Payment is already being handled")

try:
    process_payment(order_id=123)
finally:
    release_lock("order:payment:123", token)

Keep these limits in mind:

  • The work may last longer than the lock TTL.
  • A paused worker may wake after its lock has ended.
  • Redis failover may lose a new lock.
  • A lock does not make unsafe work safe by itself.

Make the protected work safe to retry. For money, stock, or other key data, use a unique request ID or a rising fence number in the main data store.

Pub/Sub

Pub/Sub sends live messages only. It does not save them for later.

import json

def publish_event(channel: str, payload: dict) -> int:
    return r.publish(channel, json.dumps(payload))

def subscribe_events(channel: str) -> None:
    pubsub = r.pubsub(ignore_subscribe_messages=True)
    pubsub.subscribe(channel)

    try:
        for message in pubsub.listen():
            handle(json.loads(message["data"]))
    finally:
        pubsub.close()

A worker that is offline will miss messages. Use Streams when messages must be saved or tried again.

Redis Streams

Streams can keep messages and share work across a group.

from redis.exceptions import ResponseError

STREAM = "events:orders"
GROUP = "order-workers"

def ensure_group() -> None:
    try:
        r.xgroup_create(STREAM, GROUP, id="0", mkstream=True)
    except ResponseError as error:
        if "BUSYGROUP" not in str(error):
            raise

def emit(event: dict) -> str:
    return r.xadd(STREAM, event, maxlen=10_000, approximate=True)

def consume(consumer: str) -> None:
    while True:
        batches = r.xreadgroup(
            groupname=GROUP,
            consumername=consumer,
            streams={STREAM: ">"},
            count=10,
            block=2000,
        )

        for _, entries in batches or []:
            for message_id, data in entries:
                try:
                    process(data)
                except Exception:
                    # Leave it pending so it can be tried again.
                    continue

                r.xack(STREAM, GROUP, message_id)

Streams use at-least-once delivery. A message may be handled more than once. Make handlers safe to retry.

Also plan for:

  • Pending messages from dead workers
  • Retry limits
  • A dead-letter stream
  • Stream size limits
  • Alerts for old pending messages

Use XAUTOCLAIM to move old pending work to a live worker.

Key Design

Use one clear form:

app:kind:id
myapp:session:abc123
myapp:product:789
myapp:rate:user:123

Rules:

  • Use short, clear names.
  • Keep the same field order.
  • Do not include passwords, tokens, email addresses, or other private data.
  • Add a version when the stored shape may change, such as myapp:v2:product:789.
  • Set a size limit for lists, sets, hashes, and streams.
  • Avoid one huge key.

In Redis Cluster, a multi-key command must use one hash slot. Use a shared hash tag when needed:

order:{123}:data
order:{123}:items
order:{123}:lock

Do not add hash tags to all keys. That can send too much work to one node.

TTL Guide

Data Common starting TTL
User session 24 hours
API cache 5 to 15 minutes
Rate window Same as the window
Short token 5 to 10 minutes
Score board cache 1 to 24 hours
Reference cache 1 hour to 1 week

These are starting points, not fixed rules.

Set TTLs on cache data and other short-lived keys. A key without a TTL can stay until it is deleted or removed by the memory policy.

Check TTLs in tests:

ttl = r.ttl("product:123")
assert ttl > 0

Redis uses these TTL results:

  • -1: the key has no TTL.
  • -2: the key does not exist.

Connection Pool

Create one client per app process. Do not open a new client for each request.

from redis import ConnectionPool, Redis

pool = ConnectionPool(
    host="localhost",
    port=6379,
    db=0,
    max_connections=20,
    decode_responses=True,
    socket_connect_timeout=2,
    socket_timeout=2,
    health_check_interval=30,
)

r = Redis(connection_pool=pool)
r.ping()

Set pool size from real app load. A pool that is too small makes requests wait. A pool that is too large may flood Redis.

Use TLS and login rules when Redis is not on a trusted local network.

Cluster

from redis.cluster import RedisCluster

r = RedisCluster(
    host="redis-1",
    port=6379,
    decode_responses=True,
    socket_connect_timeout=2,
    socket_timeout=2,
)

Test failover and moved-key replies. Avoid old options that skip cluster safety checks unless you know why they are needed.

Sentinel

from redis.sentinel import Sentinel

sentinel = Sentinel(
    [
        ("sentinel-1", 26379),
        ("sentinel-2", 26379),
        ("sentinel-3", 26379),
    ],
    socket_timeout=0.5,
)

master = sentinel.master_for(
    "mymaster",
    decode_responses=True,
    socket_timeout=2,
)

replica = sentinel.slave_for(
    "mymaster",
    decode_responses=True,
    socket_timeout=2,
)

Read from the master when fresh data is required. A replica may be behind.

Memory Removal Rules

Rule What happens Good fit
noeviction New writes fail when memory is full Data that must not be removed
allkeys-lru Removes keys used least often General cache
allkeys-lfu Removes keys used least often over time Busy shared cache
volatile-lru Removes old TTL keys first Mixed data with careful TTL use
volatile-ttl Removes TTL keys that end soon Short-lived cache data

Do not mix critical data and throw-away cache data in one Redis store unless the memory rule is safe for both.

Production Checklist

Before release:

  • Set a memory limit and removal rule.
  • Set timeouts on clients.
  • Reuse a connection pool.
  • Use login rules and TLS when needed.
  • Block public network access.
  • Set TTLs on short-lived keys.
  • Set size limits on growing data types.
  • Avoid KEYS and other long blocking calls.
  • Handle Redis being slow or down.
  • Add retries only for safe calls.
  • Make message work safe to run twice.
  • Test restart, failover, backup, and restore.
  • Watch memory, errors, slow calls, hit rate, and pending stream work.
  • Keep the main data store as the source of truth when data loss is not safe.

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 1b89586. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 weeks ago.

Activeupdated 2 weeks ago
origin
ECC

README badge

README badge for agenticluke/redis-production-patterns-plus