Multi-Engine Cognitive Walkthrough
Shared engine selection, capability/authorization gates, dispatch, capture, attribution and degraded-mode policy: _common/MULTI_ENGINE_RECIPE.md and _common/CLI_COMPATIBILITY.md. This reference defines only the domain payload and integration rules.
Default flow for /echo multi. Run subagents in parallel — one per AVAILABLE engine — to perform cognitive walkthroughs of the same UI flow across the same persona set, then integrate the results across three axes — per-persona concurrence, cross-persona universality, and engine-specific blind-spot fills.
Flow
SCOPE → CAST → PREFLIGHT → FAN-OUT (parallel subagents) → NORMALIZE → CLUSTER → SCORE (per-persona + cross-persona + per-engine) → GROUND → SYNTHESIZE → DELIVERThe flow mirrors Echo[demand]'s tri-engine demand generation but the unit of work is a step-level walkthrough trace, not a feature demand.
1. SCOPE
Define the walkthrough target once. All selected subagents share:
- UI flow under evaluation — name, entry condition, success condition (and failure conditions if any).
- Step list — ordered atomic steps the persona is expected to traverse (e.g.,
S1 land → S2 sign up CTA → S3 form → S4 verify email → S5 dashboard). Each step is the unit subagents must score. - Artifact references — screenshot paths, copy excerpts, route names, or live URL. Subagents that can read images should be told so explicitly; subagents that cannot should be given verbal descriptions of the screen plus key copy.
- Mode —
walkthrough(default),confusion(focus on cognitive load),emotion(focus on emotion scoring),dark-pattern(focus on bias/deception),a11y(accessibility persona only). - Constraints — task limit (1-4 tasks per persona per session) and walkthrough depth.
2. CAST — persona selection (shared across engines)
Select at least 3 personas spanning at least 2 axes of the Persona Diversity Matrix. All three engines walk the same persona set through the same UI flow — divergence comes from how each engine channels the persona, not from different persona pools.
- Prefer Cast registry at
.agents/personas/registry.yamlif available. - When Cast is absent, generate proto-personas internally under AI persona guardrails (
_common/AI_PERSONA_RISKS.md) and cap their confidence at 0.50. - For each persona include in the FAN-OUT prompt:
persona_id, archetype, environmental context (device, connectivity, attention level, time pressure), known mental model gaps, prior tool exposure, and last frustration. This is the samePERSONA_CHANNELblock Echo[demand] uses.
3. PREFLIGHT — engine availability detection (Echo main context, never delegated)
Identical to _common/MULTI_ENGINE_RECIPE.md §2. Probe codex, agy, claude from Echo main context with the canonical probe; subagent PATH is narrower and would produce false-negative verdicts.
4. FAN-OUT — parallel subagents
Dispatch one independent task per selected, authorized engine using the shared CLI adapter so they run concurrently. Each subagent receives the same persona set + same step list + same artifacts and is asked to walk every persona through every step independently. The matrix unit is one (persona, step) cell per engine — so 3 personas × 5 steps × 3 engines = 45 cells per session.
| Subagent | Engine | Baseline command |
|---|---|---|
walkthrough-codex |
Codex CLI | Authorized invocation via _common/CLI_COMPATIBILITY.md |
walkthrough-agy |
Antigravity CLI | Authorized headless/native dispatch → _common/CLI_COMPATIBILITY.md §9; validate outputs under _common/MULTI_ENGINE_RECIPE.md §3.5 |
walkthrough-claude |
Claude Code CLI (subagent) | Agent tool with subagent_type: general-purpose |
Loose prompt rule: pass only Role + Personas + UI flow + Step list + Artifacts + Output schema. Do NOT pass Echo's Nielsen heuristics, NASA-TLX rubric, dark-pattern taxonomy, or Peak-End synthesis rules — those frames apply at SYNTHESIZE. Each engine should channel the persona through the steps in its own way.
Required JSON output schema:
{
"engine": "codex|agy|claude",
"walkthroughs": [
{
"persona_id": "Persona name or archetype (must match input)",
"environmental_context": "Device, connectivity, attention, time pressure",
"steps": [
{
"step_id": "S1 | S2 | ...",
"step_label": "Short description of what the persona is trying to do",
"expected_behavior": "What a successful persona would do",
"predicted_behavior": "What the channeled persona actually does (in first-person voice if natural)",
"friction_points": [
{
"friction_id": "engine-local id, e.g. codex-F1",
"friction_class": "copy-ambiguity | trust-signal-missing | navigation-loss | cognitive-overload | form-friction | feedback-absence | dark-pattern | a11y-block | mental-model-gap | other",
"description": "What the persona experiences — first-person where natural",
"severity": "0 (none) | 1 (cosmetic) | 2 (minor) | 3 (major) | 4 (blocker)",
"evidence_locator": "Copy text / element / route / screenshot region the friction attaches to"
}
],
"cognitive_load": "1 (very low) | 2 | 3 | 4 (very high) — SEQ-style single-item rating",
"emotional_score": -3,
"emotional_dimension": "Valence (-3 to +3) — required; Arousal / Dominance optional",
"confusion_moments": ["Concrete moment of confusion in persona voice"],
"mental_model_gap": "Optional: what the persona expected vs what happened",
"suggested_fix": "Optional: one-line fix from persona perspective (NOT implementation jargon)"
}
],
"journey_summary": {
"overall_emotion": "-3 to +3 average / Peak-End estimate",
"completion_likelihood": "0.0-1.0",
"task_success": "true | false | partial"
}
}
],
"engine_notes": "Optional: which persona this engine felt strongest channeling, and any persona × step it explicitly skipped or struggled with"
}If an engine returns free-form Markdown, ask its subagent to re-emit as JSON before integrating.
5. NORMALIZE
Parse the usable JSON outputs into a unified walkthrough cell list. Tag each cell with (engine, persona_id, step_id). Preserve per-engine wording — divergent persona voice is signal.
6. CLUSTER — dedup within the same (persona, step)
Group friction points that describe the same UX issue inside the same (persona_id, step_id) cell. Two friction points match when all three hold:
- same
persona_id— never cluster across personas (see cross-persona signal below) - same
step_id— never cluster across steps; the same friction class at a different step is a separate finding - same
friction_classAND overlappingevidence_locator— e.g., two engines both flagcopy-ambiguityon the same CTA copy onS2for the same persona
Critical rule: Do NOT cluster across personas. Friction X for persona-A on S2 and friction X for persona-B on S2 are separate — the fact that two personas independently surface the same friction is the cross-persona signal that SYNTHESIZE highlights, not a deduplication target.
Record the set of engines that produced each cluster (within the same persona × step).
7. SCORE — per-persona concurrence + cross-persona universal + per-engine consistency (Pattern H)
Echo scores along three axes, all of which can appear on the final report.
Axis 1 — Per-cluster concurrence (within one persona × step)
Confidence axis from _common/MULTI_ENGINE_RECIPE.md Pattern H:
| Engines in cluster | Confidence label | Interpretation |
|---|---|---|
| 3 / 3 | CONFIRMED |
All engines channeled this persona to the same friction at the same step — high-confidence UX problem. |
| 2 / 3 | LIKELY |
Two engines concur; one missed it — note the silent engine. |
| 1 / 3 | CANDIDATE |
Only one engine surfaced this. Must pass GROUND (step 8) to ship. |
Axis 2 — Cross-persona universality (across persona clusters at the same step)
After per-cluster scoring, identify friction classes that surface across multiple personas on the same step:
CROSS-PERSONA-UNIVERSAL— friction class appears in ≥2 personas at the same step, AND wasCONFIRMEDorLIKELYin each persona's cluster. Strongest signal. This is friction every kind of user feels at the same point in the flow — usually the most important finding in the report.CROSS-PERSONA-SEGMENT— friction class appears in 2+ personas at the same step but only asCANDIDATEin each. Surface as "segment-level hypothesis; needs grounding."PERSONA-SPECIFIC— friction surfaced for exactly one persona. Persona-specific insight; do not generalize. (Often valuable: e.g., a11y persona surfacing a screen-reader issue.)
Axis 3 — Per-engine consistency (across personas, per engine)
For each engine, check whether its emotional scoring and severity is internally consistent across personas. Strong engine inconsistency (e.g., one engine scores S2 as -3 for the beginner but +1 for the senior while the other two engines score both negative) is a persona-channeling weakness signal for that engine — flag it in engine_notes synthesis. Do not drop the data; surface it as a calibration note.
Perspective axis (Pattern H)
Tag the consolidated finding on each step:
CONVERGENT— all engines reach the same friction verdict at this step (regardless of count of frictions). The story of the step is settled.DIVERGENT-N— engines split on the step's verdict (e.g., one engine says "fine," two say "broken," or three engines each surface a different friction).N= number of distinct dissenting positions. Preserve the split as a feature, not a bug — it's information about edge-cases or engine bias.
Critical Echo-specific rule (does NOT exist in Judge): CANDIDATE / DIVERGENT-N findings are NOT auto-low-value. A single engine surfacing the "silent majority" friction — the one users complain about least loudly because they've normalized it — is exactly the finding the other two engines might smooth over. Preserve all CANDIDATE findings through GROUND.
8. GROUND — verify CANDIDATE / DIVERGENT findings (Echo main context, never delegated)
For every CANDIDATE cluster, the Echo main context must:
- Artifact existence check — does the cited
evidence_locator(copy / element / route / screenshot region) actually exist in the supplied artifacts? If hallucinated, markREJECTED-HALLUCINATION. - Persona-voice authenticity check — does the
descriptionandconfusion_momentssound like the named persona? Or does it slip into developer/PM language? If inauthentic, markREJECTED-VOICE-MISMATCH. - Severity sanity check — does the cited severity match the friction class? Cosmetic items scored 4 (blocker) or a11y blockers scored 1 (cosmetic) are calibration errors — adjust or mark
NEEDS-INFO. - Already-mitigated check — is the cited friction already addressed by copy / UI / a11y affordances the engine missed? If yes,
REJECTED-ALREADY-MITIGATED. - Real-data calibration — if Voice / Trace / Field / session-replay data exists in the project (e.g.,
.agents/voice.md,.agents/trace.md), cross-check. Apply confidence tag:[validated]— synthetic friction matches real user signal[supported]— partial evidence (one source agrees)[hypothesis]— no real-data conflict, no real-data support[synthetic-only]— no real-data sources available
- Mark each as
VERIFIED-DIVERGENT(keep with confidence tag),REJECTED-{reason}(drop), orNEEDS-INFO(escalate).
For CONFIRMED and LIKELY clusters, verify artifact existence, persona-voice authenticity and real-data calibration. Engine agreement does not replace evidence; apply any additional grounding check required by the claim.
9. SYNTHESIZE — persona × engine matrix + priority-ranked friction list
Echo's multi-engine output presents both the persona × engine matrix (raw view) and a priority-ranked friction list (action view). This is the opposite of Judge's "by severity" structure — Echo's value lives in the matrix where Pattern H's both axes are visible.
Synthesis steps:
Cross-persona universal frictions (top section) — list every
CROSS-PERSONA-UNIVERSALfinding, ordered by severity × #personas × #engines. These are the strongest signals; they should drive Palette / Experiment handoff first.Per-step matrix view — for each step in the flow, present a compact matrix:
Step S2: Sign up CTA | codex | agy | claude Beginner persona | -3 | -3 | -2 CONFIRMED CONVERGENT Senior persona | -1 | 0 | -2 LIKELY DIVERGENT-2 Mobile user persona | -3 | -3 | -3 CONFIRMED CONVERGENT ────────────────────────────────────────────── Cross-persona verdict: CROSS-PERSONA-UNIVERSAL copy-ambiguity, trust-signal-missingShow emotional score in each cell; annotate the verdict row with cross-persona tag and dominant friction classes.
Per-persona journey summary — for each persona, list the friction points ranked by
CONFIRMED → LIKELY → VERIFIED-DIVERGENT, with the step and engine attribution attached. Include each persona's overall emotion (Peak-End computed across the run) and completion likelihood.Divergent-voice insights — call out every
VERIFIED-DIVERGENTfinding in its own section with engine attribution and the rationale for keeping it. This is where breakthrough findings live.Engine-channeling notes — short section noting which engine channeled which persona most distinctively, and any per-engine consistency weakness flagged at SCORE step. Useful for next-session persona prep and for spotting when one engine systematically under- or over-emotes.
Recommended fixes — aggregate
suggested_fixstrings from all clusters, dedupe, and rank by the friction's severity × cross-persona count. These feed the A/B hypothesis section.A/B test hypotheses — generate hypotheses from
CROSS-PERSONA-UNIVERSALand high-severityCONFIRMEDfindings (matches Echo's standardwalkthroughoutput).Rejection ledger (condensed) — count by category —
REJECTED-HALLUCINATION / VOICE-MISMATCH / ALREADY-MITIGATED / NEEDS-INFO. Preserves SNR transparency without re-introducing noise.
Engine-attribution rule
Every friction that ships must carry both an engine-concurrence tag AND a perspective tag (Pattern H two-axis):
[codex+agy+claude] [CONVERGENT] [validated]— strongest signal: 3/3 concurrence + all engines converged + real-data validation[codex+agy] [CONVERGENT] [supported]— 2/3 + converged + partial real-data[codex+agy+claude] [DIVERGENT-2] [hypothesis]— 3 engines all surfaced friction at this step but reached 2 distinct verdicts; surface the split[claude-verified] [DIVERGENT-3] [synthetic-only]— 1/3 verified divergent voice, the three engines disagreed across the step, no real-data calibration available
Cross-persona-universal findings additionally carry a [CROSS-PERSONA-UNIVERSAL] or [CROSS-PERSONA-SEGMENT] tag.
10. DELIVER — output structure
Output follows the existing Echo walkthrough report template (echo/reference/output-templates.md) with these tri-engine additions:
- Header summary table gains engine-status line and dual concurrence stats:
CONFIRMED: N / LIKELY: N / VERIFIED-DIVERGENT: NANDCROSS-PERSONA-UNIVERSAL: N / SEGMENT: N / PERSONA-SPECIFIC: N. - Cross-persona universal section is mandatory in multi mode (single-engine mode treats cross-persona analysis as optional).
- Per-step matrix view is mandatory; the compact emotion × persona × engine grid is the signature multi-mode deliverable.
- Dark pattern auto-promotion — a dark-pattern friction flagged by ≥2 engines retains the
CONFIRMEDwalkthrough-priority tag and incrementsdark_pattern_auto_promoted. This is the local risk-asymmetry rule, not a finding of legal violation or real-user validation. Verify the artifact and escalate regulatory applicability to Canon; even a single engine may flag a risk for verification. - AI persona bias disclosure — every multi-engine report must include the synthetic-only / hypothesis / supported / validated calibration distribution, even if no real-data sources existed. This makes the bias surface visible to the team.
Do not include rejected friction in the main list. Do not surface engine-raw output. Synthetic-true tagging applies to every finding unless calibration upgraded it to [validated].
Parallel Subagent Invocation
Use the canonical spawn/capture template in _common/CLI_COMPATIBILITY.md with the JSON schema in this reference. Spawn once per selected available engine, not a fixed three. Add these domain fields; the main context owns normalization, grounding and synthesis.
Role: For each persona below, walk through every step of the UI flow independently. You are one of the selected engines channeling the same personas through the same flow; do not smooth out persona quirks to match expected best practices. Report what the channeled persona actually experiences at each step, even when the friction is small or the experience is positive.
Constraints:
- Walk EVERY persona through EVERY step — do not skip cells; if a persona would abandon mid-flow, still record the abandonment step and mark
task_success: false - Speak in persona voice in
predicted_behavior,description, andconfusion_moments— first person where natural - Assign emotional_score at every step (-3 to +3); use cognitive_load 1-4 every step
- Do not invent UI elements, copy, or routes not present in the supplied artifacts — if the artifact is insufficient for a step, mark friction_class
otherwith description "evidence insufficient" rather than fabricating - Report friction even when severity = 1 (cosmetic) — surface, do not pre-filter
- Note in engine_notes which persona you felt strongest channeling and any cell you genuinely struggled with
Degraded Modes
Use _common/MULTI_ENGINE_RECIPE.md § Engine Availability Modes and its actual-engine denominator. A healthy Claude+Codex pair is the normal dual-engine baseline, not a 2/3 degraded result.
With one usable engine, mark findings synthetic-only and do not claim cross-engine universality. With zero, use walkthrough. Persona coverage, unavailable UI states and missing real-user calibration remain explicit limitations.
Cross-References
_common/MULTI_ENGINE_RECIPE.md— Pattern H protocol (this skill's pattern type); shared SCOPE/PREFLIGHT/FAN-OUT/NORMALIZE/CLUSTER mechanics_common/SUBAGENT.md §MULTI_ENGINE— base engine dispatch and loose-prompt rules_common/AI_PERSONA_RISKS.md— persona-bias guardrails (still apply in multi mode, per engine)echo/reference/tri-engine-demand.md— sibling persona × engine matrix flow (demand-generation domain); cross-persona signal logic mirrored herejudge/reference/tri-engine-review.md— canonical Pattern C tri-engine flow; PREFLIGHT/FAN-OUT sharedecho/reference/ux-frameworks.md— emotion model, cognitive load index, Peak-End rule applied at SYNTHESIZEecho/reference/output-templates.md— base walkthrough report template extended by multi-engine additionsecho/reference/cognitive-persona-model.md— CPM persona dimensions fed into the FAN-OUT prompt