Official Skill Guide Reference
Source: "The Complete Guide to Building Skills for Claude" (Anthropic, 2025)
The official spec reference that Sigil consults during the CRAFT / VERIFY phases.
1. Progressive Disclosure: The 3-Level Architecture
Skills adopt three levels of context management:
| Level | What | When loaded | Token impact |
|---|---|---|---|
| 1st — YAML frontmatter | name + description | Always (system prompt) | Minimal |
| 2nd — SKILL.md body | Full instructions | When Claude judges the skill relevant | Medium |
| 3rd — Linked files | reference/, scripts/, assets/ |
On-demand, as Claude navigates | Variable |
Design principle: Keep SKILL.md focused on core instructions, and move detailed documentation to reference/, linking to it. SKILL.md is recommended to stay under 5,000 words.
2. YAML Frontmatter Full Specification
Required Fields
---
name: skill-name-in-kebab-case
description: What it does and when to use it. Include specific trigger phrases.
---name Rules
- kebab-case only (
my-cool-skill) - No spaces, no capitals, no underscores
- Must match the folder name
- Must not include
"claude"/"anthropic"(reserved)
description Rules
- Required structure:
[What it does]+[When to use it]+[Key capabilities] - Limit: 1024 characters (official)
- XML tags (
<>) are forbidden (security constraint: frontmatter is expanded into the system prompt) - Include concrete trigger phrases
- Mention relevant file types if applicable
Note: The ecosystem-internal
skill-templates.mdrestrictsdescriptionto "one Japanese sentence," but the official spec's limit is 1024 characters. The ecosystem restriction is an additional guardrail for context efficiency and routing decisions, and does not exceed the official spec.
Optional Fields
| Field | Constraint | Purpose |
|---|---|---|
license |
e.g. MIT, Apache-2.0 |
Open-source distribution |
allowed-tools |
e.g. "Bash(python:*) WebFetch" |
Tool access restriction |
compatibility |
1-500 characters | Environment requirements |
metadata |
Any YAML key-value | author, version, mcp-server, category, tags etc. |
Security Restrictions
- XML angle brackets (
<>) — must not appear in frontmatter "claude"/"anthropic"must not be used inname(reserved prefix)- Code execution in YAML is not allowed (safe YAML parsing)
3. Description Writing Rules
Good Examples
# Specific and actionable
description: Analyzes Figma design files and generates developer handoff documentation. Use when user uploads .fig files, asks for "design specs", "component documentation", or "design-to-code handoff".
# Includes trigger phrases
description: Manages Linear project workflows including sprint planning, task creation, and status tracking. Use when user mentions "sprint", "Linear tasks", "project planning", or asks to "create tickets".
# Clear value proposition
description: End-to-end customer onboarding workflow for PayFlow. Handles account creation, payment setup, and subscription management. Use when user says "onboard new customer", "set up subscription", or "create PayFlow account".Bad Examples
# Too vague — won't trigger reliably
description: Helps with projects.
# Missing triggers — Claude can't judge when to load
description: Creates sophisticated multi-page documentation systems.
# Too technical, no user triggers
description: Implements the Project entity model with hierarchical relationships.Debugging Approach
Ask Claude: "When would you use the [skill name] skill?" — Claude will quote the description back. Adjust based on what's missing.
4. Instruction Structure Best Practices
Recommended Structure
---
name: your-skill
description: [...]
---
# Your Skill Name
## Instructions
### Step 1: [First Major Step]
Clear explanation of what happens.
```bash
python scripts/fetch_data.py --project-id PROJECT_ID
Expected output: [describe what success looks like]Examples
Example 1: [common scenario] User says: "Set up a new marketing campaign" Actions:
- Fetch existing campaigns via MCP
- Create new campaign with provided parameters Result: Campaign created with confirmation link
Troubleshooting
Error: [Common error message] Cause: [Why it happens] Solution: [How to fix]
### Best Practices
| Practice | Detail |
|----------|--------|
| **Be Specific** | `Run python scripts/validate.py --input {filename}` > `Validate the data` |
| **Progressive Disclosure** | Core instructions in SKILL.md, detailed docs in `reference/` |
| **Error Handling** | Include MCP connection failure recovery steps |
| **Reference Bundled Resources** | `Before writing queries, consult reference/api-patterns.md` |
| **Critical Instructions First** | Put key rules at the top; use `## Important` / `## Critical` headers |
| **Concise Over Verbose** | Bullet points and numbered lists over prose |
### Advanced: Deterministic Validation
> For critical validations, consider bundling a script that performs checks programmatically rather than relying on language instructions. Code is deterministic; language interpretation isn't.
---
## 5. File Structure Requirements
your-skill-name/ ├── SKILL.md # Required — main skill file ├── scripts/ # Optional — executable code │ ├── process_data.py │ └── validate.sh ├── reference/ # Optional — documentation │ ├── api-guide.md │ └── examples/ └── assets/ # Optional — templates, fonts, icons └── report-template.md
**Critical Rules**:
- The filename must be exactly `SKILL.md` (case-sensitive). `SKILL.MD` and `skill.md` are not allowed.
- The folder name must be kebab-case: `notion-project-setup`
- Do not include a `README.md` inside the skill folder (documentation belongs in `SKILL.md` or `reference/`)
- Prepare a GitHub repo-level README separately (for distribution)
---
## 6. The MCP + Skills Relationship (Kitchen Analogy)
| Aspect | MCP (Connectivity) | Skills (Knowledge) |
|--------|--------------------|--------------------|
| Role | Professional kitchen (tools, ingredients, equipment) | Recipes (step-by-step instructions) |
| Function | Connects Claude to services (Notion, Asana, Linear, etc.) | Teaches Claude how to use services effectively |
| Provides | Real-time data access and tool invocation | Workflows and best practices |
| Answers | **What** Claude can do | **How** Claude should do it |
### Without Skills (MCP only)
- Users connect MCP but don't know what to do next
- Support tickets asking "how do I do X with your integration"
- Each conversation starts from scratch
- Inconsistent results
### With Skills + MCP
- Pre-built workflows activate automatically
- Consistent, reliable tool usage
- Best practices embedded in every interaction
- Lower learning curve
---
## 7. Composability & Portability
- **Composability**: Claude can load multiple skills simultaneously. Skills should work well alongside others, not assume exclusive capability.
- **Portability**: Skills work identically across Claude.ai, Claude Code, and API. Create once, works across all surfaces (provided environment supports dependencies).
---
## 8. Testing Methodology (3 Levels × 3 Areas)
### 3 Testing Levels
| Level | Method | Setup |
|-------|--------|-------|
| **Manual** | Run queries directly in Claude.ai | No setup, fast iteration |
| **Scripted** | Automate test cases in Claude Code | Repeatable validation |
| **Programmatic** | Run an evaluation suite via the Skills API | Systematic, defined test sets |
### 3 Testing Areas
#### Area 1: Triggering Tests
- ✅ Triggers on obvious tasks
- ✅ Triggers on paraphrased requests
- ❌ Doesn't trigger on unrelated topics
Should trigger:
- "Help me set up a new ProjectHub workspace"
- "I need to create a project in ProjectHub" Should NOT trigger:
- "What's the weather in San Francisco?"
- "Help me write Python code"
#### Area 2: Functional Tests
- Valid outputs generated
- API calls succeed
- Error handling works
- Edge cases covered
#### Area 3: Performance Comparison
Baseline (without skill) vs With skill comparison:
- Back-and-forth messages count
- Failed API calls
- Token consumption
---
## 9. Iteration Signals
### Undertriggering Signals
- Skill doesn't load when it should
- Users manually enabling it
- Support questions about when to use it
- **Solution**: Add more detail and nuance to description — include keywords, especially technical terms
### Overtriggering Signals
- Skill loads for irrelevant queries
- Users disabling it
- Confusion about purpose
- **Solution**: Add negative triggers, be more specific. Clarify scope.
```yaml
# Negative trigger example
description: Advanced data analysis for CSV files. Use for statistical modeling, regression, clustering. Do NOT use for simple data exploration (use data-viz skill instead).Execution Issues
- Inconsistent results
- API call failures
- User corrections needed
- Solution: Improve instructions, add error handling. Put critical instructions first. Use scripts for deterministic validation.
10. Official Quick Checklist
Before You Start
- 2-3 concrete use cases identified
- Tools identified (built-in or MCP)
- Reviewed guide and example skills
- Planned folder structure
During Development
- Folder named in kebab-case
-
SKILL.mdfile exists (exact spelling) - YAML frontmatter has
---delimiters -
namefield: kebab-case, no spaces, no capitals -
descriptionincludes WHAT and WHEN - No XML tags (
<>) anywhere - Instructions are clear and actionable
- Error handling included
- Examples provided
- References clearly linked
Before Upload
- Tested triggering on obvious tasks
- Tested triggering on paraphrased requests
- Verified doesn't trigger on unrelated topics
- Functional tests pass
- Tool integration works (if applicable)
After Upload
- Test in real conversations
- Monitor for under/over-triggering
- Collect user feedback
- Iterate on description and instructions
Core Contract — Rationale and Measured Rates
Canonical home for the numbers and citations behind the Core Contract in SKILL.md.
- Analyze project context (stack, conventions, existing skills) before any generation.
- Discover high-value skill opportunities ranked by Priority = Frequency x Complexity x Risk.
- Mirror the project's actual naming, imports, testing, and error handling conventions.
- Default to Micro Skills (
10-80lines,< 2,000tokens); promote to Full only when complexity requires it. Skills exceeding2,000tokens degrade activation reliability and consume disproportionate context window budget. Absolute cap per Anthropic best-practices is500SKILL.md lines — beyond this, split intoreference/*.mdloaded on-demand via Read tool (three-level progressive disclosure: frontmatter → body → linked files). - Write skill
descriptionas a trigger phrase (how the user would naturally ask), not a summary — properly optimized descriptions improve activation from~20%to50%, and adding usage examples raises it from72%to~90%. Use Anthropic's skill-creator train/test split method (60/40 on ~20 synthetic prompts) to validate description activation before install. Always write in third person ("Processes Excel files and generates reports"); first/second-person POV drifts from the system-prompt voice and degrades discovery. - Counter Claude's documented undertriggering tendency — make descriptions explicit about when to activate, not just what the skill does. Include concrete trigger contexts ("Use when the user mentions dashboards, metrics, data visualization, or internal reporting, even if they don't say 'dashboard'"); passive summaries (e.g. "helps with documents") lose measurable activation rate.
- Skill description budget has two distinct limits — distinguish them: (a) per-description hard cap is
1,024characters (agentskills.io spec — exceeding this risks parser rejection or truncation), (b) per-description quality target is< 250characters (signal density goal — shorter descriptions improve routing precision and increase coexisting skill capacity). The runtime aggregate budget defaults to~2%of the context window (fallback~16,000characters total across all loaded skill descriptions, overridable viaSLASH_COMMAND_TOOL_CHAR_BUDGET). Always validate against the hard cap; treat the target as a strong recommendation. - Validate skill
nameagainst agentskills.io spec: kebab-case only, max64characters, must not start/end with hyphen, no consecutive hyphens, must not contain"claude"or"anthropic"(reserved words). Prefer gerund form (verb +-ing, e.g.,processing-pdfs,analyzing-spreadsheets,managing-databases) — this signals activity/capability more clearly than noun-only names and improves discovery. Do not add namespace prefixes (myorg/skillname,myorg:skillname) — Claude Code silently fails to load such skills without error. - Emit an
agents/eval-set.jsontrigger dataset alongside each non-trivial skill:13+queries mixing positive + negative + edge cases, each tagged withshould_trigger: true|false. Run skill-creator 2.0 loop at--max-iterations 5 --holdout 0.4with3evaluations per query for stable trigger rate; pick the winning description by held-out test score, never train score, to avoid overfitting the trigger heuristic. - Validate every skill against the 12-point rubric; install only at
9+/12. Run3independent grading passes per evaluation and use majority vote to counter LLM grader non-determinism. - Sync-write to both
.claude/skills/and.agents/skills/. - Avoid duplicating ecosystem agent functionality.
- Set
disable-model-invocation: trueonly for skills that must be explicitly invoked by the user (e.g., destructive operations, one-off migrations). - Use ATTUNE data to improve future discovery and ranking; adopt evolutionary self-modification — compare child skill performance against parent baseline before archiving improvements (HyperAgents pattern).
- Apply
_common/OPUS_5_AUTHORING.mdfor portable authoring; resolve runtime facts only through_common/CLI_COMPATIBILITY.md.