Full Quality Criteria (49 Checks)
Score each AGENTS.md root file against this checklist. Standard: the file helps an agent execute correctly with minimal context.
Contents
- Scoring
- A. Commands and execution readiness (12 checks)
- B. Gotchas and repeated mistakes (10 checks)
- C. Conventions and decision boundaries (11 checks)
- D. Signal-to-noise and bloat control (9 checks)
- E. Currency and validation (7 checks)
- Grade mapping
- Automatic fails
Scoring
Yes= 1,No= 0,N/A= excluded from the denominator- Grade uses
earned / applicable - Target: >= 91% of applicable points (grade A)
A. Commands and execution readiness (12)
- Working
devcommand (or equivalent local run) - Working
testcommand - Working
buildcommand - Working
lintand/ortypecheckcommand - Deploy/release command when applicable
- Migration/seed/db command when applicable
- Commands are copy-paste ready (no placeholders)
- Commands match the package manager and scripts
- Required env bootstrap steps (incl. secondary runtimes like Python venvs)
- Where to run commands (root/workspace)
- A command for targeted test/debug iteration, quiet flags included so a full passing suite does not linger in later turns
- No duplicate or conflicting command variants
B. Gotchas and repeated mistakes (10)
- At least one high-frequency failure mode
- Gotchas are project-specific, not generic
- Gotchas include corrective action (what to do instead)
- Gotchas include trigger context (when the rule applies)
- Captures at least one issue discovered from PR/review feedback
- Ordering/dependency gotchas where order matters
- Data/env gotchas where setup mistakes cause failures
- Avoids vague advice like "be careful" or "follow patterns"
- Separates universal rules from edge-case rules
- Removes gotchas that no longer happen
C. Conventions and decision boundaries (11)
- States conventions that materially change implementation choices
- States naming/path conventions when CI/tooling depends on them
- States test strategy conventions (unit/e2e boundaries) when relevant
- Every rule the whole repo must obey is inline in root; nothing must-obey lives only behind an
@import, a.claude/rules/file, or a.cursor/rules/file that a single tool resolves - Marks scope boundaries: monorepo root vs workspace files
- Avoids restating what the agent or harness already does (tool-use conventions, read before edit, todo tracking, run tests after a change)
- Rules name the condition that triggers them, so precision lands on when a rule applies rather than on forbidding a whole class of action
- Emphasis markers (IMPORTANT, NEVER, YOU MUST) used sparingly on critical rules agents skip
- Guidance states the outcome wanted; blanket prohibitions appear only where the harmful-precision test clears them
- No rule contradicts a parent instruction file, an installed skill, or another section of the same file; precedence is stated where overlap is deliberate
- Conventions with an exemplar in the repo name that file path instead of describing the pattern in prose
D. Signal-to-noise and bloat control (9)
- Root file concise for repo complexity (60-150 lines for active app repos; 200 is Claude Code's stated ceiling, and Codex stops reading at 32 KiB across all instruction files)
- No full framework documentation pasted inline
- No copy-pasted full templates
- No exhaustive file tree or "every file" inventory
- No long architecture deep dives in root file
- Non-universal guidance lives where it loads on demand (nested file, path-scoped rule, skill, or plain link), not in root and not behind an
@importthat loads at launch anyway - No duplicate guidance across sections
- No content auto-memory owns (user preferences, personal feedback, evolving project status)
- Each section passes the litmus test: removing it would cause mistakes
E. Currency and validation (7)
- Referenced file paths exist
- Referenced tools/dependencies are still in use
- Commands have been run (or limitations documented when run isn't possible)
- Removed references to deleted folders/APIs
- Version-sensitive guidance is date/version scoped where needed
- Clear maintenance loop (how to keep the file current)
- Personal overrides stay private; any CLAUDE.local.md fallback blocker is identified and its loading mode verified
Grade mapping
Use earned / applicable percentage:
- A: >= 91%
- B: 76% to < 91%
- C: 59% to < 76%
- D: 39% to < 59%
- F: < 39%
Example: 36/40 = 90% -> Grade B.
Automatic fails
Mark grade F regardless of score if any hold:
- Commands are mostly broken/stale
- Instructions are primarily generic advice, or restatements of default agent behavior
- File is dominated by copied docs/templates rather than executable guidance
- The intended tool does not load the shared instructions: check Claude Code version, built-in mod, Project instructions mode, and leftover project Claude files; absence of a CLAUDE.md wrapper is not a failure