All skills
vectorize-io avatar

/hindsight-docs

@5bfef3c
by vectorize-iovectorize-io/hindsight44k stars
5,845

Complete Hindsight documentation for AI agents. Use this to learn about Hindsight architecture, APIs, configuration, and best practices.

Use this Skill: https://skilld.dev/gh/vectorize-io/hindsight/hindsight-docs

This session only. Nothing lands on disk.

referencesdeveloperextensions.md

≈4.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Extensions

Extensions allow you to customize and extend Hindsight behavior without modifying core code. They enable multi-tenancy, custom authentication, additional HTTP endpoints, and operation hooks.


Available Extensions

TenantExtension

Handles multi-tenancy and API key authentication. Validates incoming requests and determines which PostgreSQL schema to use for database operations, enabling tenant isolation at the database level.

Built-in: ApiKeyTenantExtension

A simple implementation that validates API keys against an environment variable and uses the public schema for all authenticated requests.

HINDSIGHT_API_TENANT_EXTENSION=hindsight_api.extensions.builtin.tenant:ApiKeyTenantExtension
HINDSIGHT_API_TENANT_API_KEY=your-secret-key

No longer built in: SupabaseTenantExtension

Validates Supabase JWTs and gives each authenticated user their own PostgreSQL schema. It now lives in the extensions registry, which documents its configuration and ships a Dockerfile that builds an image with it.

:::warning Breaking change in 0.9.3 Up to 0.9.2 this extension was built in, at hindsight_api.extensions.builtin.supabase_tenant. That path no longer exists, so an install still pointing at it fails at startup with ModuleNotFoundError. Add the extension to your image and set HINDSIGHT_API_TENANT_EXTENSION=hindsight_ext_supabase_tenant:SupabaseTenantExtension. All HINDSIGHT_API_TENANT_* settings and the schema naming are unchanged. :::

External: StaticKeysTenantExtension

A fully self-hosted multi-user mode: users and their API keys are declared in environment variables (no external identity provider, no users table). Each user maps to their own PostgreSQL schema ({prefix}_{user_id}), provisioned lazily on first access, giving database-level memory isolation between users. Multiple API keys may map to the same user and schema.

User IDs are case-insensitive: they are lowercased (and dashes normalized to underscores) before building the schema name, so Rafael, rafael and RAFAEL all resolve to the same tenant schema.

It lives in the extensions registry, which documents its configuration and ships a Dockerfile that builds an image with it.

For other multi-tenant setups with separate schemas per tenant (e.g., custom JWT-based auth), implement a custom TenantExtension.


HttpExtension

Adds custom HTTP endpoints under the /ext/ path prefix. Useful for adding domain-specific APIs that integrate with Hindsight's memory engine.

Provides two router methods:

  • get_router(memory) — returns a FastAPI router mounted at /ext/
  • get_root_router(memory) — returns a FastAPI router mounted at the application root (for well-known endpoints or other paths that must be at specific locations). Returns None by default.

No built-in implementation - implement your own to add custom endpoints.

HINDSIGHT_API_HTTP_EXTENSION=mypackage.ext:MyHttpExtension

OperationValidatorExtension

Hooks into retain/recall/reflect operations for validation and monitoring. Use cases include:

  • Rate limiting and quota enforcement
  • Permission checks and content filtering
  • Audit logging and usage tracking
  • Custom metrics collection

No built-in implementation - implement your own based on your requirements.

self.context is the process-wide extension context — use it for process-global handles such as get_memory_engine(). It holds no per-request state: take the tenant and bank from each hook's own argument (ctx.bank_id and ctx.request_context, the latter holding the identity resolved by the tenant extension).

HINDSIGHT_API_OPERATION_VALIDATOR_EXTENSION=mypackage.validators:MyValidator

MCPExtension

Registers additional MCP (Model Context Protocol) tools on the Hindsight MCP server. Enables external packages to add custom tools without modifying core code.

No built-in implementation - implement your own to add custom MCP tools.

HINDSIGHT_API_MCP_EXTENSION=mypackage.mcp:MyMCPExtension

Writing Custom Extensions

Extension Basics

Extensions are Python classes loaded via environment variables:

HINDSIGHT_API_<TYPE>_EXTENSION=mypackage.module:MyExtensionClass

Configuration is passed via prefixed environment variables:

HINDSIGHT_API_<TYPE>_SOME_CONFIG=value
# Extension receives: {"some_config": "value"}

All extensions support lifecycle hooks:

  • on_startup() - Called when the application starts
  • on_shutdown() - Called when the application shuts down

Extensions have access to an ExtensionContext that provides:

  • run_migration(schema) - Run database migrations for a schema
  • get_memory_engine() - Get the MemoryEngine interface

Shipping your own database migrations

An extension that keeps state of its own returns the directory holding its Alembic revision files, and they are applied in the same migration run as Hindsight's own — same ordering guarantees, same alembic_version table:

class MyExtension(TenantExtension):
    def alembic_version_locations(self) -> list[str]:
        return [str(Path(__file__).parent / "alembic" / "versions")]

The directory holds revision files only — there is no env.py; Hindsight's own configures the schema and the connection. The tree is independent of Hindsight's: give its first revision down_revision = None and a branch_labels naming your extension, so an operator can address it (alembic upgrade <label>@head).

Independent branches have no ordering between them. A revision that needs a Hindsight table to exist first must say so explicitly, which orders it without making Hindsight's revision its parent:

depends_on = ("a1b2c3d4e5f6",)   # a Hindsight revision id

A directory that does not exist is skipped with a warning rather than failing the migration — a misconfigured extension must never leave a database unmigratable.

Example: Custom TenantExtension with JWT

import jwt
from hindsight_api.extensions import TenantExtension, TenantContext, AuthenticationError

class JwtTenantExtension(TenantExtension):
    def __init__(self, config: dict[str, str]):
        super().__init__(config)
        self.jwt_secret = config.get("jwt_secret")
        if not self.jwt_secret:
            raise ValueError("HINDSIGHT_API_TENANT_JWT_SECRET is required")

    async def authenticate(self, context: RequestContext) -> TenantContext:
        token = context.api_key
        if not token:
            # Optional headers dict is forwarded in HTTP/MCP error responses
            raise AuthenticationError("Bearer token required")

        try:
            payload = jwt.decode(token, self.jwt_secret, algorithms=["HS256"])
            tenant_id = payload.get("tenant_id")
            if not tenant_id:
                raise AuthenticationError("Missing tenant_id in token")
            return TenantContext(schema_name=f"tenant_{tenant_id}")
        except jwt.InvalidTokenError as e:
            raise AuthenticationError(str(e))

AuthenticationError accepts an optional headers dict that is forwarded in both HTTP and MCP error responses. This is useful for returning custom headers like WWW-Authenticate:

raise AuthenticationError(
    "Authorization required",
    headers={"WWW-Authenticate": 'Bearer realm="example"'},
)

Reading additional request headers

RequestContext carries the Authorization header as api_key. To authenticate on a different header — for instance when a gateway terminates auth with one shared identity and forwards the per-caller identity separately — name the headers you want forwarded:

HINDSIGHT_API_EXTENSION_PASSTHROUGH_HEADERS=x-user-assertion

They are available as context.extra_headers, keyed by lower-cased name:

async def authenticate(self, context: RequestContext) -> TenantContext:
    assertion = context.extra_headers.get("x-user-assertion")
    if not assertion:
        raise AuthenticationError("x-user-assertion header required")

    user_id = verify_assertion(assertion)  # your verification
    return TenantContext(schema_name=f"tenant_{user_id}")

This works on both the HTTP and MCP transports, and the same RequestContext is passed to OperationValidatorExtension hooks, so a validator can enforce rules against the identity resolved here.

Only headers you list are forwarded, and only when present on the request. The variable is unset by default, so extensions see no header data unless you opt in.

A header sent more than once is not forwarded at all, and a warning is logged. There is no safe way to choose between the copies — a proxy may append its trusted value either before or after a client-supplied one — so an extension reading it sees nothing and fails the request, rather than silently accepting a value that may be spoofed. Make sure your proxy replaces the identity header it injects instead of appending to it.

:::caution Deferred operations extra_headers describes the request being served. Operations that run later — a queued retain, a scheduled consolidation, a mental-model refresh — are executed by a background worker with no request behind them, so their RequestContext carries no headers. Authorize on the header at request time; do not rely on it inside work that continues after the response. :::

Example: Custom HttpExtension

from fastapi import APIRouter
from hindsight_api.extensions import HttpExtension

class MyHttpExtension(HttpExtension):
    def get_router(self, memory: MemoryEngine) -> APIRouter:
        router = APIRouter()

        @router.get("/hello")
        async def hello():
            return {"message": "Hello from extension!"}

        @router.post("/custom/{bank_id}/action")
        async def custom_action(bank_id: str):
            # Access memory engine for database operations
            pool = await memory._get_pool()
            # ... custom logic
            return {"status": "ok"}

        return router

    def get_root_router(self, memory: MemoryEngine) -> APIRouter | None:
        """Optional: mount routes at the application root (not under /ext/)."""
        router = APIRouter()

        @router.get("/.well-known/my-metadata")
        async def metadata():
            return {"version": "1.0"}

        return router

Routes from get_router are available at /ext/hello, /ext/custom/{bank_id}/action, etc. Routes from get_root_router are mounted at the app root (e.g., /.well-known/my-metadata).

Example: Custom OperationValidatorExtension

from hindsight_api.extensions import (
    OperationValidatorExtension,
    ValidationResult,
    PrecheckContext,
    RetainContext,
    RecallContext,
    ReflectContext,
    RetainResult,
)

class MyValidator(OperationValidatorExtension):
    # Pre-body validation (optional)
    async def precheck(self, ctx: PrecheckContext) -> ValidationResult:
        if ctx.content_length is not None and ctx.content_length > 10_000_000:
            return ValidationResult.reject("Payload is too large")
        return ValidationResult.accept()

    # Pre-operation validation (required)
    async def validate_retain(self, ctx: RetainContext) -> ValidationResult:
        # Implement your validation logic
        return ValidationResult.accept()
        # Or reject: return ValidationResult.reject("Reason")

    async def validate_recall(self, ctx: RecallContext) -> ValidationResult:
        return ValidationResult.accept()

    async def validate_reflect(self, ctx: ReflectContext) -> ValidationResult:
        return ValidationResult.accept()

    # Post-operation hooks (optional)
    async def on_retain_complete(self, result: RetainResult) -> None:
        # Log usage, update metrics, send notifications, etc.
        pass

precheck runs before the request body is read or deserialized. Its PrecheckContext.content_length is the parsed Content-Length header as an integer, or None when the header is missing or cannot be parsed (for example, chunked transfer encoding). Use it for cheap size-aware quota or cost guards; the full validate_* hooks still run after parsing and should enforce precise per-operation limits.

Memory curation hooks

Curating a memory (PATCH /v1/default/banks/{bank_id}/memories/{memory_id}: edit, invalidate, or revert) has its own pair of hooks, in addition to the validate_bank_write access check that runs first:

  • validate_memory_update(ctx: MemoryUpdateContext) runs before any work. The context carries the requested text (when editing it), the requested state, and edits_fields. Reject here to refuse the curation; the returned status_code is passed through to the HTTP response.
  • on_memory_update_complete(result: MemoryUpdateResult) runs once the change has committed. result.action is edit, invalidate, revert, or reason, and result.reembedded_tokens is the size of the text the engine embedded again (0 for a plain invalidation or a reason-only update). An edit or revert re-embeds the memory and re-consolidates the bank, so this is the figure to meter if curation should cost the same as ingesting that text.
Confining a caller to a tag scope

When several people share one bank and tags separate what each may see, resolve_tag_scope pins every operation of a caller to a tag filter. Return the tag groups the caller's reachable data must satisfy (they are AND-ed), or None for no restriction:

from hindsight_api.engine.search.tags import TagGroupLeaf
from hindsight_api.extensions import OperationValidatorExtension, TagScopeContext


class TeamScopes(OperationValidatorExtension):
    async def resolve_tag_scope(self, ctx: TagScopeContext):
        user = ctx.request_context.api_key_id  # or whatever your tenant extension resolved
        # Dan reads his own memories plus the shared team rules.
        return [TagGroupLeaf(tags=[f"user:{user}", "kind:rule"], match="any_strict")]

The engine combines the scope with whatever filter the caller sends, so a caller can narrow its view but never widen it:

  • Filtered reads — recall, reflect (and every tool it runs), the memory, document and mental-model lists, observation scopes, tag lists, the memory graph, the timeseries, entities and the entity graph, and the knowledge-base tree, search and export — only see rows inside the scope. Entity counts are recounted over the visible memories.
  • Reads and writes by id — a memory, a document and its chunks, a mental model or knowledge page — answer 404 when the item is outside the scope. A retain that names an existing document outside the scope is refused (403), since replacing or appending to it would touch someone else's text.
  • Source text follows the document. A memory can be visible through its own tags (say, a kind:rule fact extracted from someone's private note) while the note is not. Its chunk and document text are only returned when the document's tags pass the filter. This applies to every tag-filtered recall and reflect, with or without an extension.
  • Mental models and knowledge pages created or updated by a scoped caller must carry tags inside the scope (403 otherwise — the caller could not see what it made), and record the scope in their trigger (scope_tag_groups), which every refresh AND-s with the model's own filter, so they can never be built from memories their creator could not read. Background refreshes run unscoped and rely on that recorded scope. A recorded scope only ever narrows the model and is not editable through the API: to drop it, recreate the model.

Use a _strict match mode: the non-strict ones also admit untagged rows, and an untagged mental model is built from the whole bank.

Writing is a separate permission. A caller can read a shared scope without being allowed to change it. resolve_write_tag_scope returns the tags a caller may write, as shell-style patterns (user:dan, project:*), or None for no restriction (the default):

class TeamScopes(OperationValidatorExtension):
    async def resolve_write_tag_scope(self, ctx: TagScopeContext):
        user = ctx.request_context.api_key_id
        # Everyone writes their own memories; only Kate writes the team rules.
        return [f"user:{user}", "kind:rule"] if user == "kate" else [f"user:{user}"]

Every tag on anything the caller writes or changes must match one of the patterns, otherwise the write is refused with 403 (an item the caller cannot even read still answers 404). An untagged item belongs to everyone, so a restricted writer cannot produce or change one. That covers:

  • retain (text and files): each item's tags and the batch's document_tags, plus every tag the retain strategy's entity labels with tag: true could add (key:value for each allowed value, key:* for an open vocabulary), and explicit observation_scopes ("shared" writes untagged observations). They are checked before extraction, so a refused retain costs no LLM call;
  • memories and documents: editing, invalidating or clearing the observations of a memory; updating (including the new tags), reprocessing or deleting a document;
  • mental models and knowledge pages: creating one (its tags), updating, refreshing, clearing or deleting one, and renaming, moving or deleting a knowledge node (every page under it, and under the folder it moves into);
  • directives: creating, updating or deleting one. An untagged directive steers everyone's reflect, so a restricted writer cannot create one.

Directives are read like reflect reads them: untagged directives apply to everyone, tagged ones only inside the caller's read scope.

Whole-bank operations reach every memory regardless of tags, so a caller with a read or write scope cannot run them at all (403): export, clone and import, clearing the bank's memories or observations, deleting the bank, changing or resetting its config, mission or disposition, and running (or retrying failed) consolidation on request. The consolidation the engine queues after a scoped caller's own writes still runs. Operation status never returns the raw task payload to a scoped caller. Reads that are not tag-scoped (stats, webhooks) still go through validate_bank_read / validate_bank_write, which can only allow or deny them.

Deferring an operation

In addition to accept and reject, a validate_* hook can ask the worker to requeue the operation for a future time by raising DeferOperation. Use this for backpressure (rate-limited upstream, quota window not yet open, dependency warming up) — unlike a retry, it does not increment retry_count or write error_message. The worker sets next_retry_at to your exec_date and the task is invisible to claim queries until that time.

from datetime import datetime, timedelta, timezone

from hindsight_api.extensions import (
    DeferOperation,
    OperationValidatorExtension,
    RetainContext,
    ValidationResult,
)


class QuotaAwareValidator(OperationValidatorExtension):
    async def validate_retain(self, ctx: RetainContext) -> ValidationResult:
        if not await self._quota_available(ctx.bank_id):
            raise DeferOperation(
                exec_date=datetime.now(timezone.utc) + timedelta(minutes=5),
                reason="bank quota window exhausted",
            )
        return ValidationResult.accept()

DeferOperation is worker-only: do not raise it from validate_recall or validate_reflect in synchronous HTTP request paths — there is no queue to defer to and it will surface as a 500.

Example: Custom MCPExtension

from mcp.server.fastmcp import FastMCP
from hindsight_api.extensions import MCPExtension
from hindsight_api.engine import MemoryEngine

class MyMCPExtension(MCPExtension):
    async def register_tools(self, mcp: FastMCP, memory: MemoryEngine) -> None:
        @mcp.tool()
        async def custom_search(query: str) -> str:
            """Custom MCP tool for specialized search."""
            # Access memory engine for operations
            pool = await memory._get_pool()
            # ... custom logic
            return f"Results for: {query}"

Deploying Custom Extensions

With Docker

Extensions are not bundled in the image. Build one on top of Hindsight that installs your extension's dependencies and copies it in. Install into the image's virtualenv explicitly — it was created by uv sync and ships no pip of its own, so a bare pip install lands where the server can't see it:

FROM ghcr.io/vectorize-io/hindsight:latest

RUN uv pip install --python /app/api/.venv/bin/python --no-cache \
      'PyJWT[crypto]>=2.12.0' 'httpx>=0.27.0'

COPY my_extension /app/extensions/my_extension
ENV PYTHONPATH=/app/extensions

# Fail the build, not the first request, if it isn't importable.
RUN /app/api/.venv/bin/python -c "import my_extension"

Then point the service at that image and pass the extension's variables as environment. Give the API and worker containers the same extension configuration — the worker uses the tenant extension to enumerate schemas for background consolidation.

See the extensions registry README for the full recipe.

Bare Metal

Install your extension package in the same Python environment as Hindsight:

# Install Hindsight
pip install hindsight-api

# Install your extension package
pip install ./my-extensions
# or
pip install my-extensions-package

# Configure
export HINDSIGHT_API_TENANT_EXTENSION=my_extensions.auth:JwtTenantExtension
export HINDSIGHT_API_TENANT_JWT_SECRET=your-secret

# Run
hindsight-api

Contributing Extensions

Custom extensions that solve common use cases are welcome contributions to the Hindsight project. If you've built an extension for:

  • Authentication providers (OAuth, SAML, API gateways)
  • Rate limiting or quota management
  • Audit logging integrations
  • Metrics exporters (Datadog, New Relic, etc.)
  • Custom HTTP endpoints for specific platforms

Add it to the extensions registry — either as a directory under hindsight-extensions/, or as a registry entry linking to your own repository. That README covers the layout, the development workflow, and Docker packaging.

Extensions live outside the server so that installing Hindsight does not pull in a vendor's client library, and so changing an extension does not require a Hindsight release. Only extensions that add no dependencies and are useful to any deployment (ApiKeyTenantExtension, MemoryDefenseRegexExtension) stay in hindsight_api.extensions.builtin.

Source: SKILL.md on GitHub

1 alerttoday5 checks · Risk SAFE
  • Gen Agent Trust Hubtoday

    The skill is a comprehensive documentation set for the Hindsight memory system, providing architecture overviews, API references, and integration guides for multiple AI agent frameworks. No security risks were identified in the documentation or provided examples.

  • Sockettoday

    No alerts

  • Snyktoday

    Risk: LOW · No issues

  • Runlayer6mo

    30/42 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 5bfef3c. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 15 hours ago.

Activeupdated 2 months ago

README badge

README badge for vectorize-io/hindsight/hindsight-docs