All skills
oaustegard avatar

/hallucinating-labels

@7977b11

Assign items to a CLOSED label vocabulary that is too large to put in a prompt — product taxonomies, category hierarchies, tag vocabularies, routing tables, ICD/SIC-style code lists. A cheap model writes the label it thinks the vocabulary would use, and an embedder snaps that writing onto the nearest legal value, so the schema is never transmitted and the output is always in-vocabulary. Use for "classify these into our taxonomy", "tag these against the existing tag list", "map these queries to categories", "the enum is too big to send", or a Literal/enum that hits a provider cap. NOT for a vocabulary that fits in a prompt — structured output measured 0.701 acc@1 there against this pattern's 0.564. NOT for open-ended labelling with no fixed vocabulary, and not for ranked retrieval over documents (bm25).

  • 5 files
  • 25 KB
  • Updated last month
  • GitHub

Use this Skill: https://skilld.dev/gh/oaustegard/claude-skills/hallucinating-labels

This session only. Nothing lands on disk.

CHANGELOG.md

≈470 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Changelog — hallucinating-labels

0.1.0 — 2026-08-31

Initial. Implements the hallucinate-and-snap pattern from Doug Turnbull's "Don't classify. Hallucinate!" (softwaredoug.com, 2026-08-10), with three things measurement added that the post does not carry:

  • The boundary. Structured output over the full label set scored 0.701 acc@1 on WANDS against this pattern's 0.564. The post reports the pattern working and being cheaper, not the arm it loses to. The skill leads with it.
  • The register correction. The post's "novel, never-seen-before" prompt is safe only with a model too weak to obey it. A Haiku 4.5 subagent obeyed and scored 0.100 acc@1 against a 0.500 no-model control; re-anchored on register it scored 0.525/0.750. The register wording also beat novelty on Gemini across all 468 queries (0.564 vs 0.489), so it is strictly better.
  • The long-item case. On a 1,273-tag memory corpus of 1,500-character documents, the written labels and the direct embedding are complementary: 0.508 and 0.416 alone, 0.672 interleaved. --union exists for that. Long items also amplify the register error — under the novelty prompt the same corpus reads 0.208, half the control, which is how a prompt bug can look like a boundary on the pattern itself.

scripts/snap.py — build/snap CLI, tfidf and minilm backends, --union, --min-score. Arms and artifacts in oaustegard/experiments/hypothetical-classification.

Also carries the no-API story: gte-small int8 (33 MB) snaps the raw query at 0.455 acc@1 against MiniLM-L6's 0.417, and neither Pleias Monad (57M) nor Baguettotron (321M) earns a place in the pipeline — 0.425/0.400 as writers and 0.325/0.350 as rerankers, against a 0.500 no-model control. In a browser, ship the encoder alone.

[0.1.0] - 2026-09-01

Other

  • Add hallucinating-labels skill (#782)

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 7977b11. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last month
metadata
{
  "version": "0.1.0"
}

README badge

README badge for oaustegard/claude-skills/hallucinating-labels