All skills
anthropics avatar

/data-context-extractor

@7c35640

Generate or improve a company-specific data analysis skill by extracting tribal knowledge from analysts. BOOTSTRAP MODE - Triggers: "Create a data context skill", "Set up data analysis for our warehouse", "Help me create a skill for our database", "Generate a data skill for [company]" → Discovers schemas, asks key questions, generates initial skill with reference files ITERATION MODE - Triggers: "Add context about [domain]", "The skill needs more info about [topic]", "Update the data skill with [metrics/tables/terminology]", "Improve the [domain] reference" → Loads existing skill, asks targeted questions, appends/updates reference files Use when data analysts want Claude to understand their company's specific data warehouse, terminology, metrics definitions, and common query patterns.

Use this Skill: https://skilld.dev/gh/anthropics/knowledge-work-plugins/data-context-extractor

This session only. Nothing lands on disk.

referencesdomain-template.md

≈873 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Domain Reference File Template

Use this template when creating reference files for specific data domains (e.g., revenue, users, marketing).


# [DOMAIN_NAME] Tables

This document contains [domain]-related tables, metrics, and query patterns.

---

## Quick Reference

### Business Context

[2-3 sentences explaining what this domain covers and key concepts]

### Entity Clarification

**"[AMBIGUOUS_TERM]" can mean:**
- **[MEANING_1]**: [DEFINITION] ([TABLE]: [ID_FIELD])
- **[MEANING_2]**: [DEFINITION] ([TABLE]: [ID_FIELD])

Always clarify which one before querying.

### Standard Filters

For [domain] queries, always:
```sql
WHERE [STANDARD_FILTER_1]
  AND [STANDARD_FILTER_2]

Key Tables

[TABLE_1_NAME]

Location: [project.dataset.table] or [schema.table] Description: [What this table contains, when to use it] Primary Key: [COLUMN(S)] Update Frequency: [Daily/Hourly/Real-time] ([LAG] lag) Partitioned By: [PARTITION_COLUMN] (if applicable)

Column Type Description Notes
[column_1] [TYPE] [DESCRIPTION] [GOTCHA_OR_CONTEXT]
[column_2] [TYPE] [DESCRIPTION]
[column_3] [TYPE] [DESCRIPTION] Nullable

Relationships:

  • Joins to [OTHER_TABLE] on [JOIN_KEY]
  • Parent of [CHILD_TABLE] via [FOREIGN_KEY]

Nested/Struct Fields (if applicable):

  • [struct_name].[field_1]: [DESCRIPTION]
  • [struct_name].[field_2]: [DESCRIPTION]

[TABLE_2_NAME]

[REPEAT FORMAT]


Key Metrics

Metric Definition Table Formula Notes
[METRIC_1] [DEFINITION] [TABLE] [FORMULA] [CAVEATS]
[METRIC_2] [DEFINITION] [TABLE] [FORMULA]

Sample Queries

[QUERY_PURPOSE_1]

-- [Brief description of what this query does]
SELECT
    [columns]
FROM [table]
WHERE [standard_filters]
GROUP BY [grouping]
ORDER BY [ordering]

[QUERY_PURPOSE_2]

[ANOTHER_COMMON_QUERY]

[QUERY_PURPOSE_3]: [More Complex Pattern]

WITH [cte_name] AS (
    [CTE_LOGIC]
)
SELECT
    [final_columns]
FROM [cte_name]
[joins_and_filters]

Common Gotchas

  1. [GOTCHA_1]: [EXPLANATION]

    • Wrong: [INCORRECT_APPROACH]
    • Right: [CORRECT_APPROACH]
  2. [GOTCHA_2]: [EXPLANATION]


Related Dashboards (if applicable)

Dashboard Link Use For
[DASHBOARD_1] [URL] [DESCRIPTION]
[DASHBOARD_2] [URL] [DESCRIPTION]

---

## Tips for Creating Domain Files

1. **Start with the most-queried tables** - Don't try to document everything
2. **Include column-level detail only for important columns** - Skip obvious ones like `created_at`
3. **Real query examples > abstract descriptions** - Show don't tell
4. **Document the gotchas prominently** - These save the most time
5. **Keep sample queries runnable** - Use real table/column names
6. **Note nested/struct fields explicitly** - These trip people up

## Suggested Domain Files

Common domains to document (create separate files for each):

- `revenue.md` - Billing, subscriptions, ARR, transactions
- `users.md` - Accounts, authentication, user attributes
- `product.md` - Feature usage, events, sessions
- `growth.md` - DAU/WAU/MAU, retention, activation
- `sales.md` - CRM, pipeline, opportunities
- `marketing.md` - Campaigns, attribution, leads
- `support.md` - Tickets, CSAT, response times

Source: SKILL.md on GitHub

1 warning17d5 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    This skill acts as a developer tool to help analysts document data warehouse knowledge and generate specialized data analysis skills. It includes a Python script for packaging generated files. The primary security consideration is the processing of external data sources like database schemas to generate instructions, which presents a potential risk of indirect prompt injection.

  • Socket17d

    No alerts

  • Snyk17d

    Risk: LOW · No issues

  • Runlayer7mo

    6/6 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 7c35640. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub last week.

Activeupdated 8 months ago
  • Documentation
  • data-warehouse
  • knowledge-extraction
  • bigquery
  • snowflake
  • postgres
  • schema-discovery
  • metrics
  • sql

README badge

README badge for anthropics/knowledge-work-plugins/data-context-extractor

Extracts company-specific data warehouse schemas, terminology, and metric definitions from analysts, then generates a tailored data analysis skill with reference documentation. Used in bootstrap mode to create new skills from scratch or iteration mode to add domain-specific context to existing skills.

Generated from the current SKILL.md.

Does this skill connect directly to my data warehouse?
Yes, it uses `~~data warehouse` tools to query and explore your schema. It supports BigQuery, Snowflake, PostgreSQL/Redshift, and Databricks.
Can I use this to update an existing data skill?
Yes. Iteration Mode lets you load an existing skill and add new domains, metrics, or reference files without starting from scratch.
What format does the generated skill use?
It creates a directory with SKILL.md and a `references/` folder containing markdown files for entities, metrics, tables, and optionally a dashboards catalog.
Does this require Claude to know SQL?
No. The skill asks conversational questions to analysts and generates SQL queries itself during schema discovery.
What if my warehouse uses a non-standard SQL dialect?
The skill includes SQL dialect section documentation. Bootstrap Mode identifies your warehouse type and generates queries in the correct dialect.

Generated from the current SKILL.md. These answers refresh after source changes.