All skills
vectorize-io avatar

/hindsight-docs

@5bfef3c
by vectorize-iovectorize-io/hindsight44k stars
5,845

Complete Hindsight documentation for AI agents. Use this to learn about Hindsight architecture, APIs, configuration, and best practices.

Use this Skill: https://skilld.dev/gh/vectorize-io/hindsight/hindsight-docs

This session only. Nothing lands on disk.

referencesdeveloperretain.md

≈4.7k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Retain: How Hindsight Stores Memories

When you call retain(), Hindsight transforms conversations and documents into structured, searchable memories that preserve meaning and context.

What Retain Does

Figure: What Retain Does. An animated diagram on the docs site; its narration, step by step:

  • new facts
    1. You send content, a context that says who is speaking, and when it was said.
    2. The original is kept as a document and split into chunks, so the exact passage can be handed back later.
    3. An LLM reads each chunk and pulls out facts with what, when, where, who and why. It keeps the reason, not just the event. Each fact is embedded right away.
    4. Entity resolution decides who each name is. It compares each name with the entities the bank already has. Neither is known yet, so both become new entities.
    5. Entities are stored once, and every fact that mentions them points to them.
    6. The facts are world facts: Bob is talking about Alice, not about the agent. Each keeps two times: when it happened and when Hindsight learned it.
    7. Finally the facts are linked: through shared entities, closeness in time, similar meaning (from the embeddings), and cause and effect.
    8. retain() is done. Observations are built from these facts later, in the background.
  • same person, new name
    1. Two months later, Bob mentions “Alice C.”
    2. Same path: stored as a new document, split into chunks…
    3. …and read by the LLM. The fact names “Alice C.”, a name the bank has never seen.
    4. The name is close to Alice Chen, and the same fact names the Zurich office she is already linked to. Together that is enough: same person.
    5. No new entity is created: this mention of “Alice C.” counts as Alice Chen.
    6. The fact keeps its own wording, but it points to Alice Chen, so it joins everything else about her. Ask about Alice later and you get all of it.
    7. Its entity link ties it to the three facts from January.
    8. Done: one more fact about the same Alice.

Rich Fact Extraction

Hindsight doesn't just store what was said — it captures why, how, and what it means.

What Gets Captured

When you retain "Alice joined Google last spring and was thrilled about the research opportunities", Hindsight extracts:

The core facts:

  • Alice joined Google
  • This happened last spring

The emotions and meaning:

  • She was thrilled
  • It represented an important opportunity

The reasoning:

  • She chose it for the research opportunities

This rich extraction means you can later ask "Why did Alice join Google?" and get a meaningful answer, not just "she joined Google."

Preserving Context

Traditional systems fragment information:

  • "Bob suggested Summer Vibes"
  • "Alice wanted something unique"
  • "They chose Beach Beats"

Hindsight preserves the full narrative:

  • "Alice and Bob discussed naming their summer party playlist. Bob suggested 'Summer Vibes' because it's catchy, but Alice wanted something unique. They ultimately decided on 'Beach Beats' for its playful tone."

This means search results include the full context, not disconnected fragments.


Two Types of Facts

Every fact is classified by whose perspective it captures — the agent that owns the bank, or the outside world:

Type What it captures Example
experience The bank's own agent acting, observing, or interacting — its first-person history "I recommended Python to Alice"
world Facts about other people, places, things, and events "Alice works at Google"

The split is decided by who is speaking, not by grammar. A first-person statement is an experience only when the speaker is the bank's agent. The same words said by someone else are a world fact about that person:

  • Agent's own log — "I patched the auth bug" → experience (the agent did it).
  • A user talking to the agent — "I bought a Tesla" → world (a fact about the user, not the agent).

Describe the speaker in each item's context to steer this correctly. When retaining transcripts or third-party content, a context like "Customer Maria is speaking" ensures her first-person statements are stored as world facts about Maria rather than mistaken for the agent's own experiences. For the agent's own logs, a context like "The assistant is speaking" attributes its first-person statements to the agent as experience facts.

Note: Observations are consolidated automatically in the background after retain() operations complete. This consolidation process synthesizes patterns from new facts into the bank's knowledge base.


Entity Recognition

Hindsight automatically identifies and tracks entities — the people, organizations, and concepts that matter.

What Gets Recognized

  • People: "Alice", "Dr. Smith", "Bob Chen"
  • Organizations: "Google", "MIT", "OpenAI"
  • Places: "Paris", "Central Park", "California"
  • Products & Concepts: "Python", "TensorFlow", "machine learning"

Entity Resolution

The same entity mentioned different ways gets unified through fuzzy name matching, reinforced by co-occurrence and temporal proximity:

  • "Alice" + "Alice Chen" + "Alice C." → one person

Because resolution keys off name similarity, close variants merge automatically. Names that do not resemble each other (a nickname and an unrelated formal name, for example) are not unified on the name alone, though shared co-occurring entities can still link them.

Why it matters: You can ask "What do I know about Alice?" and get everything, even if she was mentioned as "Alice Chen" in some conversations.

Resolution is a judgement call, so it can go the other way too: on a bank with a lot of history, a short new name that resembles an existing entity — and that turns up alongside entities the existing one is already associated with — can be absorbed into it instead of becoming its own entity. If you see facts about a new person attached to an unrelated entity, How entity resolution decides explains what is being compared and which setting makes matching stricter.

Context-Aware Disambiguation

If "Alice" appears with "Google" and "Stanford" multiple times, a new "Alice" mentioning those is likely the same person. Hindsight uses co-occurrence patterns to disambiguate common names.

Entity Labels

You can define a controlled vocabulary of key:value classification labels (e.g. pedagogy:scaffolding, engagement:active) that are extracted at retain time and stored as entities. Because labels become entities, they automatically link related memories in the knowledge graph and improve both semantic and keyword retrieval. Labels can optionally also write to the memory unit's tags, enabling standard tag-based filtering during recall and reflect.

Unlike regular entities, label entities never merge by name similarity — distinct label values must stay distinct, so they resolve by exact match only and are excluded from fuzzy name matching altogether.

See entity_labels in the bank config for full configuration details.


Building Connections

Memories aren't isolated — Hindsight creates a knowledge graph with four types of connections:

Entity Connections

All facts mentioning the same entity are linked together.

Enables: "Tell me everything about Alice" → retrieves all Alice-related facts

Time-Based Connections

Facts close in time are connected, with stronger links for closer dates.

Enables: "What else happened around then?" → finds contextually related events

Meaning-Based Connections

Semantically similar facts are linked, even if they use different words.

Enables: "Tell me about similar topics" → finds thematically related information

Causal Connections

Cause-effect relationships are explicitly tracked.

Enables: "Why did this happen?" → trace reasoning chains Example: "Alice felt burned out" ← caused by ← "She worked 80-hour weeks"


Understanding Time

Hindsight tracks two temporal dimensions:

When It Happened

For events (meetings, trips, milestones), Hindsight records when they occurred.

  • "Alice got married in June 2024" → occurred in June 2024

For general facts (preferences, characteristics), there's no specific occurrence time.

  • "Alice prefers Python" → ongoing preference

When You Learned It

Hindsight also tracks when you told it each fact.

Why both?

Imagine in January 2025, someone tells you "Alice got married in June 2024":

  • Historical queries work: "What did Alice do in 2024?" → finds the marriage
  • Recency ranking works: Recent mentions get priority in search
  • Temporal reasoning works: "What happened before her marriage?" → finds earlier events

Without this distinction, old information would either be unsearchable by date or treated as irrelevant.


Tagging Memories

Tags enable visibility scoping—useful when one memory bank serves multiple users but each should only see relevant memories.

  • Item tags: Tag individual memories with specific scopes
  • Document tags: Apply tags to all items in a batch
  • Tag filtering: Filter during recall/reflect by tags

See Retain API for code examples and Recall API for filtering options.


Images and Files in Your Content

A lot of what a document actually says lives in its pictures: the screenshot of the button an instruction refers to, the diagram holding an escalation path, the chart that is the data. content accepts an ordered list of blocks so those sit where they belong, instead of being summarised into a caption beforehand:

{
  "content": [
    {"type": "text",  "text": "To reset the VPN, click the button shown:"},
    {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": "..."}},
    {"type": "text",  "text": "...then reconnect."}
  ]
}

A plain string still works exactly as before — nothing about text-only retain changes.

The point is position. Extraction sends the text and the pictures together, in order, so the model reads the screenshot beside the sentence that introduces it. "Click the button shown below" is meaningless on its own; next to the picture it becomes a fact that names the button.

What comes back

Facts read naturally — [image: image/png] marks where an attachment was, never a content hash — and each memory carries the attachments it was drawn from, with a URL you can fetch:

{
  "text": "The escalation path for a stuck sync begins by contacting Tier 3 Platform.",
  "attachments": [
    {"id": "c414cd0e204d", "kind": "image", "media_type": "image/png",
     "url": "/v1/default/banks/my-bank/attachments/c414cd0e204d"}
  ]
}

Facts that came from the prose have no attachments, even when the same document is full of pictures. So an attachment shown next to a memory means the model looked at it to produce that memory — it is evidence, not decoration.

An observation carries the attachments of the facts it was consolidated from, so a screenshot still reaches you when recall or reflect answers from the observation rather than the raw fact. Reflect's based_on (with include.facts) returns each cited memory with its attachments, plus the document_id, chunk_id, tags and metadata it was stored with.

What to expect from charts and tables

An attachment carrying structured data is transcribed rather than summarised: each row, bar or labelled value becomes its own fact, carrying its label, its figure, and how it is drawn. A chart therefore produces many more memories than a screenshot does — that is deliberate, since a summary of a twenty-bar chart keeps three values and silently drops seventeen, and later questions ask about the ones it dropped.

Very dense pages are still sampled rather than exhausted: a single extraction pass over an infographic listing a hundred entries will capture a large share of them, not all of them.

Requirements

  • A model that can read images. By default that is the bank's retain model, but it does not have to be: HINDSIGHT_API_VLM_MODEL names a vision slot used only for the chunks that actually carry an attachment, leaving every text-only chunk on the retain LLM. So a bank whose documents are mostly prose can keep a cheap text model and pay for vision only where there is something to look at.
  • If that model cannot read images — or if Hindsight cannot tell, which is the case for gateway backends serving mixed catalogues — the retain is refused with 422 rather than dropping the attachment silently. See HINDSIGHT_API_LLM_VISION.
  • Batch retain (HINDSIGHT_API_RETAIN_BATCH_ENABLED) cannot carry attachments, and also refuses with 422.

This is different from POST /files/retain, which converts a whole file to markdown as its own document. That is still the right tool for a scanned report you want parsed — but it separates the file from the prose that referred to it, which is exactly what inline attachments avoid.

For storage, size limits and accepted media types, see Inline attachments in retain.


What You Get

After retain() completes:

  • Structured facts that preserve meaning, emotions, and reasoning
  • Unified entities that resolve different name variations
  • Knowledge graph with entity, temporal, semantic, and causal links
  • Temporal grounding for both historical and recency-based queries
  • Optional tags for filtering during recall

All stored in your isolated memory bank, ready for recall() and reflect().


Steering Extraction with a Mission

By default, retain() extracts all significant facts from the content. You can narrow this focus with a retain mission (retain_mission) — a plain-language description of what this bank should pay attention to.

e.g. Always include technical decisions, API design choices, and architectural trade-offs.
     Ignore meeting logistics, greetings, and social exchanges.

The mission is injected into the extraction prompt alongside the built-in rules — it steers the LLM without replacing the extraction logic. It works with the LLM-based extraction modes (concise, verbose, verbatim, custom) and is ignored in chunks mode.

For finer control, you can also change the extraction mode:

Mode When to use
concise (default) General-purpose — selective, fast
verbose When you need richer facts with full context and relationships
custom When you want to write your own extraction rules entirely
verbatim When you need the original chunk text preserved, with LLM-extracted metadata such as entities and dates
chunks When you want to store chunks as-is with no LLM call or extracted metadata

Set retain_mission and retain_extraction_mode via the bank config API, or with the HINDSIGHT_API_RETAIN_MISSION and HINDSIGHT_API_RETAIN_EXTRACTION_MODE environment variables.

When a mission excludes everything in a document

A mission narrows what becomes a memory — and content that produces no facts produces no memories at all. The document itself is still stored, but recall and reflect search memories, so a document with zero memories cannot be found by either. Tightening a mission therefore trades away retrieval of the raw source, not just fact creation.

This is a normal outcome, not an error: the retain succeeds and the operation is reported as completed. Two signals tell you it happened:

Where What to look for
retain.completed webhook data.memory_unit_count: 0
Metrics hindsight.retain.documents.total{outcome="no_facts"}

You can also audit after the fact: GET /documents returns memory_unit_count per document, so filtering for 0 lists everything currently unreachable.

Extraction is not fully deterministic — a borderline document can yield facts on one run and none on the next. Treat a zero as "this document needs another pass" rather than as a permanent verdict.

To recover a document, widen the mission and reprocess it — the stored text is re-extracted, and no re-upload is needed:

POST /v1/default/banks/{bank_id}/documents/{document_id}/reprocess

Observation Consolidation

After retain() completes, Hindsight automatically triggers observation consolidation in the background. This process:

  1. Analyzes new facts against existing observations
  2. Creates new observations when patterns emerge
  3. Refines existing observations with new evidence
  4. Tracks which facts support each observation

This happens asynchronously — your retain() call returns immediately while consolidation runs in the background.

See Observations for details on how consolidation works.


Memory Defense and Source Provenance

receipt_uri (optional)

Type: string.

Optional pointer into an external receipt or co-signature system. Stored as-is and surfaced in security_events.receipt_uri for any Memory Defense decision on this item.

422 — Memory Defense violation

When Memory Defense is enabled on the target bank and every item in the batch is blocked by policy, the request returns 422 with a violation list:

{
  "detail": {
    "violations": [
      { "index": 0, "detector": "prompt_injection", "severity": "high", "message": "..." }
    ]
  }
}

Partial-block batches return 200 with the un-blocked items processed; blocked items are silently dropped from the result with their decisions recorded in security_events.

See Memory Defense for the full guide.


Next Steps

  • Observations — How knowledge is consolidated after retain
  • Recall — How multi-strategy search retrieves relevant memories
  • Reflect — How the agentic loop uses observations
  • Retain API — Code examples and parameters

Source: SKILL.md on GitHub

1 alerttoday5 checks · Risk SAFE
  • Gen Agent Trust Hubtoday

    The skill is a comprehensive documentation set for the Hindsight memory system, providing architecture overviews, API references, and integration guides for multiple AI agent frameworks. No security risks were identified in the documentation or provided examples.

  • Sockettoday

    No alerts

  • Snyktoday

    Risk: LOW · No issues

  • Runlayer6mo

    30/42 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 5bfef3c. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 16 hours ago.

Activeupdated 2 months ago

README badge

README badge for vectorize-io/hindsight/hindsight-docs