All skills
softaworks avatar

/session-handoff

@1c0662a
by softaworkssoftaworks/agent-toolkit2.5k stars
230

Creates comprehensive handoff documents for seamless AI agent session transfers. Triggered when: (1) user requests handoff/memory/context save, (2) context window approaches capacity, (3) major task milestone completed, (4) work session ending, (5) user says 'save state', 'create handoff', 'I need to pause', 'context is getting full', (6) resuming work with 'load handoff', 'resume from', 'continue where we left off'. Proactively suggests handoffs after substantial work (multiple file edits, complex debugging, architecture decisions). Solves long-running agent context exhaustion by enabling fresh agents to continue with zero ambiguity.

Use this Skill: https://skilld.dev/gh/softaworks/agent-toolkit/session-handoff

This session only. Nothing lands on disk.

evalsmodel-expectations.md

≈1.4k tokens on demand. Your agent reads this file only when SKILL.md points to it.

Model-Specific Expectations

This document describes expected behavior differences across Claude models when using the session-handoff skill.

Model Characteristics

Haiku (Fast, Lightweight)

  • Strengths: Quick responses, follows explicit instructions well
  • Limitations: May need more guidance, less proactive
  • Skill adjustments: May need more explicit prompts for complex scenarios

Sonnet (Balanced)

  • Strengths: Good balance of speed and capability, handles workflows well
  • Limitations: May occasionally miss subtle triggers
  • Skill adjustments: Should work well with default instructions

Opus (Most Capable)

  • Strengths: Excellent context understanding, proactive suggestions
  • Limitations: May over-elaborate when not needed
  • Skill adjustments: May add extra context/suggestions beyond requirements

Expected Behaviors by Scenario

Scenario 1: Basic Handoff Creation

Aspect Haiku Sonnet Opus
Trigger recognition Should trigger Should trigger Should trigger
Script execution Runs script Runs script Runs script
TODO completion May need prompting Fills reasonable defaults Rich, detailed content
Validation reminder May skip Usually includes Always includes

Haiku-specific guidance:

  • May need explicit "now fill in the TODO sections"
  • Keep prompts simple and direct

Opus-specific notes:

  • May proactively suggest additional sections
  • May add extra context without being asked

Scenario 2: Handoff Chaining

Aspect Haiku Sonnet Opus
Finds previous handoffs With explicit prompt Usually automatic Always automatic
Uses --continues-from May need reminder Usually correct Always correct
Context from previous Basic reference Good summary Detailed synthesis

Haiku-specific guidance:

  • Explicitly mention "link to the previous handoff"
  • May need to specify exact filename

Scenario 3: Resume from Handoff

Aspect Haiku Sonnet Opus
Lists handoffs With prompt Automatic Automatic
Staleness check May skip Usually runs Always runs
Context absorption Basic Good Excellent
Next steps focus May need guidance Usually clear Proactive planning

Haiku-specific guidance:

  • Explicitly ask "check the staleness first"
  • May need "what are the next steps from the handoff?"

Scenario 4: Proactive Handoff Suggestion

Aspect Haiku Sonnet Opus
Recognizes substantial work Unlikely without prompt Sometimes Usually
Suggests handoff Rarely proactive Sometimes proactive Often proactive
Timing of suggestion N/A After 5+ major items After 3-5 items

Notes:

  • Haiku will rarely proactively suggest handoffs
  • Sonnet may suggest after explicit substantial work description
  • Opus most likely to suggest unprompted

Scenario 5: Validation Flow

Aspect Haiku Sonnet Opus
Runs validation script With explicit request Usually automatic Always automatic
Interprets score Basic Good Detailed
Actionable feedback May need prompting Usually provides Detailed plan

Scenario 6: Staleness Check

Aspect Haiku Sonnet Opus
Runs staleness script With explicit request Usually Always
Interprets results Basic Good Detailed analysis
Recommendations Repeats script output Contextualizes Strategic advice

Scenario 7: Secret Detection

Aspect Haiku Sonnet Opus
Detects secrets Via script Via script Via script + may notice more
Warning clarity Basic Clear Detailed security advice
Remediation guidance Script output Clear steps Comprehensive plan

Tuning Recommendations

For Haiku Optimization

If Haiku struggles:

  1. Add more explicit trigger phrases to description
  2. Include step-by-step numbered instructions
  3. Add explicit checkpoints ("After creating, run validation")
  4. Reduce ambiguity in instructions

For Sonnet Optimization

If Sonnet misses triggers:

  1. Ensure key terms are in description
  2. Add example trigger phrases
  3. Make workflow decision points clearer

For Opus Optimization

If Opus over-elaborates:

  1. Add "keep responses concise" guidance
  2. Specify when NOT to add extra content
  3. Define clear scope boundaries

Pass/Fail Criteria by Model

Minimum Pass Thresholds

Model Min Score Notes
Haiku 49/70 (70%) Allow some missed proactive triggers
Sonnet 56/70 (80%) Should handle most scenarios well
Opus 63/70 (90%) Should excel at all scenarios

Critical Failures (Any Model)

These should always work regardless of model:

  • Basic handoff creation with explicit request
  • Script execution when instructed
  • Secret detection warning
  • File creation in correct location

Testing Protocol

  1. Run setup script:

    python evals/setup_test_env.py
    cd /tmp/handoff-eval-project
  2. Test each scenario in new conversation

    • Start fresh conversation for each scenario
    • Use exact trigger phrases from test-scenarios.md
    • Record scores using results template
  3. Compare across models

    • Note significant behavior differences
    • Identify skill improvements needed
    • Update SKILL.md if Haiku needs more guidance
  4. Document findings

    • Use results template for each model
    • Note specific failure modes
    • Recommend skill adjustments

Source: SKILL.md on GitHub

1 warning17d5 checks · Risk SAFE
  • Gen Agent Trust Hub17d

    The Session Handoff skill manages markdown documents to preserve context between AI agent sessions. It uses Python scripts to collect project metadata via Git and validate document quality. Analysis identified that the skill executes local shell commands (Git) and ingests external project metadata, such as commit messages, which presents a surface for indirect prompt injection.

  • Socket17d

    No alerts

  • Snyk17d

    Risk: LOW · No issues

  • Runlayer7mo

    4/12 files flagged

  • ZeroLeaks5mo

    Score: 93/100 · 2 sections analyzed

Signed by skilld at 1c0662a. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 2 months ago.

Dormantupdated 9 months ago
  • session-handoff
  • context-management
  • agent-continuity
  • memory-preservation
  • long-running-tasks
  • workflow-persistence
  • handoff-documentation
  • state-capture

README badge

README badge for softaworks/agent-toolkit/session-handoff

Generates handoff documents that preserve project state, context, and pending work so fresh AI agents can resume without ambiguity. Addresses long-running sessions where context windows fill up by capturing current state, decisions made, critical files, and immediate next steps in a chainable format.

Generated from the current SKILL.md.

When should I create a handoff?
Create a handoff when the user requests to save state, pause work, or says context is getting full. Also proactively suggest one after substantial work like 5+ file edits, complex debugging, or major architecture decisions.
Can I chain handoffs together for long-running projects?
Yes. Use the `--continues-from` flag when creating a new handoff to link it to a previous one, creating a context chain that new agents can follow.
What does the validation script check for?
The validator checks for remaining TODO placeholders, required sections, potential secrets (API keys, passwords, tokens), file existence, and generates a quality score. Do not finalize a handoff with detected secrets or a score below 70.
How do I know if a handoff is still usable when resuming?
Run the staleness checker, which assesses based on time elapsed, git commits, file changes, branch divergence, and missing files. It returns FRESH, SLIGHTLY_STALE, STALE, or VERY_STALE status.
Where are handoffs stored?
Handoffs are stored in `.claude/handoffs/` with timestamped filenames like `2024-01-15-143022-implementing-auth.md`.

Generated from the current SKILL.md. These answers refresh after source changes.