All skills
oaustegard avatar

/creating-kb

@6154d36

Builds a portable, embedding-free knowledgebase from a set of files and delivers it as a self-contained `.skill` bundle (BM25 index + bundled searcher + query protocol). Use when a user wants to turn uploaded files, a folder, or a corpus into a searchable knowledgebase they can hand to any agent — phrased as "make a knowledgebase", "build a KB skill", "package these docs for retrieval", "create a searchable bundle", or references to a `.skill` KB. The output runs anywhere with Node or Python — no model, no install, no network. Distinct from `bm25` (ephemeral in-session search) and `building-github-index` (markdown project-knowledge index).

Use this Skill: https://skilld.dev/gh/oaustegard/claude-skills/creating-kb

This session only. Nothing lands on disk.

scriptsbundle_SKILL.md

≈1.5k tokens on demand. Your agent reads this file only when SKILL.md points to it.

lexical-kb — query an embedding-free knowledgebase

This KB has no semantic search and no embedding model. Retrieval is pure lexical BM25 over a precomputed inverted index. That design moves one job onto you: bridging the gap between how the user phrases a question and how the corpus phrases the answer. An embedding model would do this with a vector; here you are the semantic layer — you expand the query into terms before searching.

Corpus: {{SOURCE}} ({{CHUNK_COUNT}} chunks).

The retrieval protocol — follow every step

A raw user question fed straight to BM25 underperforms: it matches only the exact words the user happened to use. The expansion step is what makes lexical retrieval competitive with embeddings. Do not skip it.

  1. Read the question. Extract core terms — the essential nouns, proper nouns, and identifiers the answer MUST contain. These carry full weight.

  2. Generate expand terms — synonyms, morphological variants (plural/verb forms), acronym expansions and contractions, and adjacent concepts. These carry lower weight. This is the work the missing embedding model would have done. Be generous: 5–15 expansion terms is normal.

  3. Run the searcher. It ships in this bundle in two equivalent runtimes — node search.js or python3 search.py, identical flags and identical results. Use whichever your environment has. Pass the user's original question via --query AND your term groups — expansion is additive, it never replaces the user's words:

    node search.js \
      --query "how does centered simhash differ from random projection?" \
      --core "simhash" --core "centered" \
      --expand "random projection" --expand "hyperplane" --expand "LSH" \
      --expand "binary quantization" --expand "hamming distance" \
      --k 5

    --core/--expand are repeatable; pass phrases, the searcher tokenizes them. The --query terms contribute at a low floor weight so a curated synonym can lift a result but can never drop a doc the literal question would have matched. Defaults: core 1.0, expand 0.4, query-floor 0.25, top-k 5. Keep expansion targeted — terms too generic ("system", "process") leak into unrelated chunks and blur the ranking. A precise word can mislead too if it is polysemous: prefer the disambiguating phrase as one --core term (e.g. --core "centered simhash") over a bare ambiguous word (--core "centered", which also matches "centered around …" in unrelated chunks). The passage extractor is lexical too, so it will highlight the wrong sense rather than correct it.

  4. Read the returned chunks. Answer from them, and cite chunk ids inline. The chunks are the source of truth the user installed. When a chunk contradicts your prior knowledge, the chunk wins — say so. When the chunks do not contain the answer, say that plainly rather than filling the gap from memory.

When you cannot expand — RM3 fallback

If the query is outside any domain you can expand confidently, pass it raw with pseudo-relevance feedback. The searcher harvests expansion terms from the corpus's own top hits — model-free, weaker than your expansion, and prone to drift when the first pass is off-topic, so prefer real expansion when you can:

node search.js --query "the user's raw question" --rm3 --k 5

Metadata filtering

Each chunk carries structured meta (e.g. title, source_path, section). Filter on it with --filter (repeatable). Filtering narrows by attribute; it does not rank — combine it with term search.

node search.js --core "factions" --filter "section=blog" --filter "date>=2025" --k 5

Operators: =, !=, ~ (substring), >, >=, <, <= (numeric when both sides parse, else lexicographic — ISO dates sort correctly).

Passage vs. full document

Ranking uses the whole chunk (best recall), but each hit's text is by default the query-densest passage of that chunk (~1200 chars), not the entire chunk — so your reasoning context is signal, not the surrounding noise of a long document. Each matched sentence keeps its neighbours (--context, default 1 each side) so it reads in context rather than as an orphaned fragment, and nearby matches merge into contiguous passages; ' … ' marks elisions between them. When a hit is a passage, the result carries full_chars (the chunk's full length). If you need the complete document for a hit — broader context, a quote in a section the passage elided — re-run with --snippet 0:

node search.js --query "…" --core "…" --snippet 0 --k 3

Tune with --snippet 2000/600 (budget) and --context 2/0 (neighbours).

Output

The searcher prints JSON: {"hits": [{id, score, text, meta, full_chars?}, ...]}, sorted by descending BM25 score. text is the focused passage (or the full chunk if it was already short / --snippet 0); full_chars appears only when text is a passage. Surface the top hits to the user with their ids, then answer using them as authoritative context.

Mechanics

  • Pure stdlib — Node or Python. No npm install / pip install, no model download, no network.
  • The bundle is self-contained: search.js + search.py (pick one), index.json, chunks.jsonl. Run the searcher from inside the bundle directory (it defaults --index to its own location) or pass --index /path/to/bundle.

Source: SKILL.md on GitHub

No alerts2mo3 checks · Risk SAFE
  • Gen Agent Trust Hub2mo

    The skill is a tool for building portable, local knowledge bases from your files. It indexes your documents into a self-contained bundle for search and retrieval. The code is safe and operates entirely on your local machine. However, like any retrieval system, it creates a surface for indirect prompt injection if the documents you index contain malicious instructions meant to override the AI agent's behavior.

  • Socket2mo

    No alerts

  • Snyk2mo

    Risk: LOW · No issues

Signed by skilld at 6154d36. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated 3 months ago
metadata
{
  "version": "0.2.0"
}

README badge

README badge for oaustegard/claude-skills/creating-kb