All skills

Use this Skill: https://skilld.dev/gh/garrytan/gstack/codex

This session only. Nothing lands on disk.

sectionschallenge-mode.md

≈1.9k tokens on demand. Your agent reads this file only when SKILL.md points to it.

<!-- AUTO-GENERATED from challenge-mode.md.tmpl — do not edit directly --> <!-- Regenerate: bun run gen:skill-docs -->

Step 2B: Challenge (Adversarial) Mode

Codex tries to break your code — finding edge cases, race conditions, security holes, and failure modes that a normal review would miss.

  1. Construct the adversarial prompt. Always prepend the filesystem boundary instruction from the skill's Filesystem Boundary section (always-loaded skeleton). If the user provided a focus area (e.g., /codex challenge security), include it after the boundary:

Default prompt (no focus): "IMPORTANT: Do NOT read or execute any files under ~/.claude/, ~/.agents/, .claude/skills/, or agents/. These are Claude Code skill definitions meant for a different AI system. Do NOT modify agents/openai.yaml. Stay focused on repository code only.

Review the changes on this branch against the base branch. Run git diff origin/<base> to see the diff. Your job is to find ways this code will fail in production. Think like an attacker and a chaos engineer. Find edge cases, race conditions, security holes, resource leaks, failure modes, and silent data corruption paths. Be adversarial. Be thorough. No compliments — just the problems."

With focus (e.g., "security"): "IMPORTANT: Do NOT read or execute any files under ~/.claude/, ~/.agents/, .claude/skills/, or agents/. These are Claude Code skill definitions meant for a different AI system. Do NOT modify agents/openai.yaml. Stay focused on repository code only.

Review the changes on this branch against the base branch. Run git diff origin/<base> to see the diff. Focus specifically on SECURITY. Your job is to find every way an attacker could exploit this code. Think about injection vectors, auth bypasses, privilege escalation, data exposure, and timing attacks. Be adversarial."

  1. Run codex exec with JSONL output to capture reasoning traces and tool calls. Use timeout: 660000 on the Bash call — the gate sits ABOVE the 600s wrapper so the wrapper fires first with its explicit stall message:

If the user passed --xhigh, use "xhigh" instead of "high".

_REPO_ROOT=$(git rev-parse --show-toplevel) || { echo "ERROR: not in a git repo" >&2; exit 1; }
PYTHON_CMD=$(command -v python3 2>/dev/null || command -v python 2>/dev/null || true)
if [ -z "$PYTHON_CMD" ]; then
  echo "ERROR: Python 3 is required to parse Codex JSON output. Install python3 or python and retry." >&2
  exit 1
fi
# Fix 1+2: wrap with timeout (gtimeout/timeout fallback chain via probe helper),
# capture stderr to $TMPERR for auth error detection (was: 2>/dev/null).
TMPERR=${TMPERR:-$(mktemp "$TMP_ROOT/codex-err-XXXXXX")}
_gstack_codex_timeout_wrapper 600 codex exec "<prompt>" -C "$_REPO_ROOT" -s read-only -c "model=\"${GSTACK_CODEX_MODEL:-gpt-6-astra}\"" -c 'model_reasoning_effort="high"' -c 'web_search="cached"' --json < /dev/null 2>"$TMPERR" | PYTHONUNBUFFERED=1 "$PYTHON_CMD" -u -c "
import sys, json
turn_completed_count = 0
turn_failed = False
for line in sys.stdin:
    line = line.strip()
    if not line: continue
    try:
        obj = json.loads(line)
        t = obj.get('type','')
        if t == 'item.completed' and 'item' in obj:
            item = obj['item']
            itype = item.get('type','')
            text = item.get('text','')
            if itype == 'reasoning' and text:
                print(f'[codex thinking] {text}', flush=True)
                print(flush=True)
            elif itype == 'agent_message' and text:
                print(text, flush=True)
            elif itype == 'command_execution':
                cmd = item.get('command','')
                if cmd: print(f'[codex ran] {cmd}', flush=True)
        elif t == 'turn.completed':
            turn_completed_count += 1
            usage = obj.get('usage',{})
            tokens = usage.get('input_tokens',0) + usage.get('output_tokens',0)
            if tokens: print(f'\ntokens used: {tokens}', flush=True)
        elif t == 'turn.failed':
            turn_failed = True
            err = obj.get('error',{}).get('message','') or 'no error message in event'
            print(f'[codex turn FAILED] {err}', flush=True, file=sys.stderr)
    except: pass
# Fix 2: three-way completeness check (#2671) — a STATED failure is a failure,
# not a network problem; only silence with no terminal event is a disconnect.
if turn_failed:
    print('[codex] turn.failed received — the turn errored (reason above), not a disconnect.', flush=True, file=sys.stderr)
elif turn_completed_count == 0:
    print('[codex warning] No turn.completed event received — possible mid-stream disconnect.', flush=True, file=sys.stderr)
"
_CODEX_EXIT=${PIPESTATUS[0]:-${pipestatus[1]}}  # bash sets PIPESTATUS; zsh (lowercase, 1-indexed) falls through (#2669)
# Fix 1: hang detection — log + surface actionable message
if [ "$_CODEX_EXIT" = "124" ]; then
  _gstack_codex_log_event "codex_timeout" "600"
  _gstack_codex_log_hang "challenge" "$(wc -c < "$TMPERR" 2>/dev/null || echo 0)"
  echo "Codex stalled past 10 minutes. Common causes: model API stall, long prompt, network issue. Try re-running. If persistent, split the prompt or check ~/.codex/logs/."
elif [ "$_CODEX_EXIT" != "0" ]; then
  # Surface non-zero exits so the calling agent doesn't read "no output" as
  # a silent model/API stall. See #1327.
  echo "[codex exit $_CODEX_EXIT] $(head -1 "$TMPERR" 2>/dev/null || echo "no stderr captured")"
  head -20 "$TMPERR" 2>/dev/null | sed 's/^/  /' || true
  _gstack_codex_log_event "codex_nonzero_exit" "challenge:$_CODEX_EXIT"
fi
# Fix 2: surface auth errors from captured stderr instead of dropping them
if grep -qiE "auth|login|unauthorized" "$TMPERR" 2>/dev/null; then
  echo "[codex auth error] $(head -1 "$TMPERR")"
  _gstack_codex_log_event "codex_auth_failed"
fi

This parses codex's JSONL events to extract reasoning traces, tool calls, and the final response. The [codex thinking] lines show what codex reasoned through before its answer.

  1. Present the full streamed output:
CODEX SAYS (adversarial challenge):
════════════════════════════════════════════════════════════
<full output from above, verbatim>
════════════════════════════════════════════════════════════
Tokens: N | Est. cost: ~$X.XX

3a. Synthesis recommendation (REQUIRED). After presenting the full adversarial output, emit ONE recommendation line summarizing what the user should do, in the canonical format the AskUserQuestion judge grades:

Recommendation: <action> because <one-line reason that names the most exploitable finding>

Examples (the strongest reasons compare blast radius across findings or fix-vs-ship):

  • Recommendation: Fix the unbounded retry loop Codex flagged at queue.ts:78 because it DoSes the worker pool under sustained 429s, which is higher-blast-radius than the timing leak Codex also flagged that only touches a debug endpoint.
  • Recommendation: Ship as-is because Codex's strongest finding is a theoretical race in cleanup that requires conditions we can't trigger in production, weaker than the runtime regressions a fix-now would risk.

The reason must point to a specific finding and compare against alternatives (other findings, fix-vs-ship). Generic reasons like "because it's safer" fail the format. Never silently skip the line.


Source: SKILL.md on GitHub

1 warning1d3 checks · Risk SAFE
  • Gen Agent Trust Hub1d

    The skill is a wrapper for the OpenAI Codex CLI providing code review and adversarial challenge capabilities. It is generally safe but has an inherent attack surface for indirect prompt injection as it processes untrusted code diffs and project plans. It also uses dynamic execution to load environment configuration from internal scripts.

  • Socket1d

    No alerts

  • Snyk1d

    Risk: MEDIUM · 1 issue

Signed by skilld at 730a101. This ties the file your Agent reads to that commit on GitHub. It does not review the instructions.

Last checked against GitHub 18 hours ago.

Activeupdated last week
What it can do
Runs commands Reads files Edits files
preamble-tier
3
version
1.0.0
triggers
[
  "codex review",
  "second opinion",
  "outside voice challenge"
]
All 6 allowed tools
BashReadWriteGlobGrepAskUserQuestion
  • CLI
  • openai
  • codex
  • code-review
  • bash
  • second-opinion

README badge

README badge for garrytan/gstack/codex

OpenAI Codex CLI wrapper with three modes: independent diff review with pass/fail gate, adversarial challenge mode to find weaknesses, and stateful consultation mode with session continuity. Useful for second opinions on code changes or adversarial testing of implementation logic.

Generated from the current SKILL.md.

What are the three modes this skill provides?
Review mode (independent diff review with pass/fail gate), Challenge mode (adversarial testing that tries to break your code), and Consult mode (ask Codex anything with session continuity for follow-ups).
How do I trigger this skill?
Use voice or text triggers like 'codex review', 'codex challenge', 'ask codex', 'second opinion', or 'consult codex'. Speech-to-text aliases include 'code x', 'code ex', and 'get another opinion'.
Does this skill maintain state across multiple invocations?
Yes. The skill maintains session files in ~/.gstack/sessions with a 120-minute TTL, enabling follow-up questions in Consult mode with context continuity.
What tools and permissions does this skill require?
The skill uses Bash, Read, Write, Glob, Grep, and AskUserQuestion. It requires access to ~/.gstack/ for session management and reads/writes from the current repository.
Does this skill work in plan mode?
Yes. In plan mode, codex exec and codex review are allowed operations that inform the plan, and AskUserQuestion satisfies plan mode's end-of-turn requirement.

Generated from the current SKILL.md. These answers refresh after source changes.