All skills
aktsmm avatar

/ai-cli-benchmark

@2ce5c0b
by yamapanaktsmm/agent-skills26 stars
4

Design, run, and review fair performance, token, and cost comparisons across GitHub Copilot CLI, Codex CLI, and direct model APIs. Use for AI CLI benchmark, Copilot vs Codex, prompt-cache diagnosis, cache read/write analysis, fresh-process latency measurement, token-cost comparison, or when a benchmark must separate provider caching from local CLI caches.

  • 5 files
  • 11 KB
  • License
  • Updated last week
  • GitHub

Use this Skill: https://skilld.dev/gh/aktsmm/agent-skills/ai-cli-benchmark

This session only. Nothing lands on disk.

SKILL.md

≈94 tokens always: the name and description. ≈958 when used: this file. ≈178 more on demand in 1 file.

AI CLI Benchmark

When to Use

  • Copilot CLI、Codex CLI、直接 API の速度・token・費用を比較する。
  • prompt cache の read / write を区別し、cache-offを主張できるか確認する。
  • fresh process、temporary home、instruction/tool isolationを揃えた再現可能なbenchmarkを作る。

過去sessionの事後分析だけならanalyze-copilot-sessionsを使う。このskillは比較条件の設計、実行、証拠化を担当する。

Cache Boundaries

  • Copilot CLI / Codex CLIには、provider側model prompt cacheを無効化する公開optionがない。fresh process、一時home、ephemeral、nonceでもcache offと呼ばない。
  • COPILOT_CACHE_HOMEはMarketplace、自動更新packageなどのローカル一時データ用。COPILOT_MCP_TOOL_CACHEはローカルMCP serverのtool-list snapshot用。model prompt cacheではない。
  • OpenAI prompt cacheは手動消去できない。GPT-5.6以降をResponses APIから直接呼ぶ場合だけ、explicit-only modeでbreakpointを置かず、そのrequestのcache writeを避けられる。CLIへ推測適用しない。
  • nonceはuser prompt全体の一致を避ける補助であり、その前のhidden system / agent prefixのcache readを防がない。

Workflow

  1. Freeze the task
  • dataset、expected labels、prompt、schema、input order、model、reasoningを固定し、hashを取る。
  • accuracyと性能を同時に比較する場合、同じ合格条件を全providerへ適用する。
  1. Define measurement boundaries
  • one request、serial total、fresh-process end to endを区別する。
  • API直呼びとCLI起動込み時間をmodel speedとして並べない。
  1. Isolate each CLI
  • repo外のowned temp directory、fresh process、一時COPILOT_HOME / auth-only CODEX_HOMEを使う。
  • custom instructions、tools、MCP、session reuseを無効にし、secretを子processとlogへ渡さない。
  1. Observe cache; do not assume control
  • 必要ならprompt先頭へ一意nonceを入れ、nonce本体ではなくhashだけ保存する。
  • zero-cache-read gateは診断用。失敗はprovider cacheの観測であり、benchmark失敗として隠さない。
  1. Record raw categories
  • normal input、cache read、cache write、output、reasoningを別fieldで保存する。
  • provider schemaで包含・排他関係を確認するまでtoken区分を足さない。
  1. Repeat before claiming performance
  • 単発差をcacheの因果効果にしない。median / p95 / rangeとaccuracyを報告する。
  1. Validate the artifact
python scripts/validate_benchmark.py result.json `
  --require-cache-metadata `
  --report validation.json

Reporting Rules

  • cache writeはcache hitではない。readと分け、該当modelの料金区分で費用を計算する。
  • uncached runではなく、観測どおりcache read 0 runまたはcache-observed runと書く。
  • CLI version、model、reasoning、prompt hash、nonce hash、fresh-process条件、測定境界、usage schemaを記載する。
  • cache read 0でも将来のwriteやhidden prefixの制御を保証しない。

Done Criteria

  • dataset/prompt/schema/model/reasoningが固定され、hashがある。
  • 全providerでfresh processと一時homeを使う。
  • cache read/writeを別fieldで保存する。
  • 測定境界とaccuracyを一緒に報告する。
  • 複数runなしの速度差へ因果を割り当てない。
  • benchmark JSON validatorがPASSする。

References

Source: SKILL.md on GitHub

No third-party reports yet.

Signed by skilld at 2ce5c0b. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 15 hours ago.

Activeupdated last week
argument-hint
比較対象、dataset/prompt、model、測定回数、cache観測条件、出力先
user-invocable
true
metadata
{
  "author": "yamapan (https://github.com/aktsmm)"
}

README badge

README badge for aktsmm/agent-skills/ai-cli-benchmark