All skills
cursor avatar

/poteto-mode

@12d587d
by cursorcursor/plugins9.1k stars
857

poteto's agent style for concise, detailed responses, deliberate subagents, unslopped prose, simple code, and verified work. Use for poteto, /poteto-mode, or requests to work in this style.

Use this Skill: https://skilld.dev/gh/cursor/plugins/poteto-mode

This session only. Nothing lands on disk.

playbookseval.md

≈631 tokens on demand. Your agent reads this file only when SKILL.md points to it.

Eval

You own the experiment design. Plan, blind, run, synthesize.

Non-negotiables for blinding:

  • No eval, test, judge, experiment, rubric, score, compare, benchmark, candidate, or arena in any directory, file, or prompt the candidate sees.
  • The candidate prompt looks like an organic user request. State the goal, not the meta.
  • No chain-eliciting cues. Don't ask the candidate to list which skills, principles, or files they applied. Ask for design notes generally and grade chain-following from code shape, not self-report.
  • Sanitize directory and slug names. Use project-shaped names a user might pick.
  • Don't tell the candidate other candidates exist.
  • The judge can know it's judging but sees outputs by sanitized label only, never by model name.
  • Comparing two variants: one judge scores both sets in a single pass on one scale, blind to which set each came from.

Steps:

  1. Frame. State what variant is under test and what behavior counts as success. Write the rubric (3-6 concrete criteria) for the judge only. Hold it back from candidates.
  2. Set up sanitized environments. Per-candidate working dir with the variant in place. Plant any context an organic task would have: a project skeleton, the skills the candidate would naturally read.
  3. Author one organic prompt. What a user would type. No leakage of what's being measured.
  4. Spawn N parallel candidates on different models per the arena skill's Phase B. Each works in its own sanitized dir. Same prompt to each.
  5. Spawn one blinded judge on a different model family per the arena skill's Phase C. Judge sees outputs by sanitized label and the rubric, never a model name.
  6. Verify the chain from transcripts, not self-report. Read each candidate's local transcript under the active workspace's agent-transcripts/ directory (the system prompt names this path). Do not glob across ~/.cursor/projects/*/. That crosses workspace boundaries and reads private chats from unrelated projects. Look at which files each candidate actually opened. Grade chain-following from the files it really read plus the shape of the code, never from the candidate's own claims.
  7. Read every candidate output yourself end to end. Compare to the judge's verdict. Disagreement means a model is biased or the rubric is ambiguous. Synthesize.

Reply: variant under test, rubric, per-candidate notes, judge's verdict, your synthesis, and a recommendation for whether to promote the variant.

Source: SKILL.md on GitHub

2 warnings7d3 checks · Risk SAFE
  • Gen Agent Trust Hub7d

    The skill provides a comprehensive framework for agents to handle complex software development tasks, including PR management and project orchestration. It uses local scripts to automate dependency installation and interface with the GitHub and Graphite CLIs. The primary security considerations are its high degree of autonomy and the processing of external data such as PR comments and session transcripts, which present a surface for indirect prompt injection.

  • Socket7d

    1 alert: gptSecurity

  • Snyk7d

    Risk: MEDIUM · 1 issue

Signed by skilld at 12d587d. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub yesterday.

Activeupdated last week
disable-model-invocation
true
mode
true
icon
crown
Other metadata
color
yellow
reminder
New task? Playbook match or rigor needed -> apply /poteto-mode. Casual turn or user opts out -> don't.

README badge

README badge for cursor/plugins/poteto-mode