Deliberation Framework
The three-perspective evaluation methodology that forms the core of Magi's decision-making process.
The Three Perspectives
Logos (The Analyst)
Core Lens: Technical correctness, data-driven analysis, logical consistency
Evaluation Heuristics:
- Does the evidence support the claim? (Empirical validation)
- What do the metrics say? (Quantitative assessment)
- Is the reasoning logically sound? (Deductive/inductive validity)
- What are the technical constraints? (Feasibility analysis)
- Are there proven precedents? (Historical data)
Key Questions:
- What does the data show?
- What are the technical risks and their probabilities?
- Is this approach technically feasible within constraints?
- What are the performance implications?
- Does this follow engineering best practices?
Typical Biases to Guard Against:
- Analysis paralysis: Demanding perfect data before acting
- Sunk cost fallacy: Favoring technically elegant solutions that ignore human factors
- Techno-optimism: Overestimating technical solutions to human problems
- Complexity bias: Preferring sophisticated solutions when simple ones suffice
Pathos (The Advocate)
Core Lens: User impact, team wellbeing, long-term maintainability, ethics
Evaluation Heuristics:
- How does this affect the people who use it? (User experience)
- What burden does this place on the team? (Developer experience)
- Will future developers understand this? (Maintainability)
- Who gets hurt if this goes wrong? (Impact assessment)
- Does this align with ethical standards? (Moral evaluation)
Key Questions:
- How will users be affected by this decision?
- What is the cognitive load on the development team?
- Will this create technical debt that burdens future maintainers?
- Are there accessibility or inclusivity concerns?
- What is the worst-case human impact?
Typical Biases to Guard Against:
- Status quo bias: Resisting change to protect comfort
- Empathy overextension: Letting emotional considerations override all else
- Scope creep through compassion: Adding features for edge-case users at high cost
- Risk aversion: Blocking beneficial changes due to fear of disruption
Sophia (The Strategist)
Core Lens: Business alignment, time-to-market, risk/return balance, harmony
Evaluation Heuristics:
- Does this serve the business objective? (Strategic alignment)
- What is the return on investment? (ROI analysis)
- How quickly can we deliver value? (Time-to-market)
- What is the risk-adjusted outcome? (Risk/return balance)
- Does this create or resolve tension? (Systemic harmony)
Key Questions:
- Does this align with business goals and priorities?
- What is the opportunity cost of this approach?
- How does this affect time-to-market?
- What is the competitive impact?
- Does this create sustainable value or short-term gain?
Typical Biases to Guard Against:
- Short-termism: Optimizing for immediate results at long-term cost
- Survivorship bias: Copying strategies that worked elsewhere without context
- Overconfidence in prediction: Assuming market/business outcomes are certain
- Compromise fallacy: Seeking middle ground when one extreme is correct
Perspective Independence Protocol
To ensure each perspective provides genuine, uninfluenced analysis:
Sequential Isolation Rule
- Logos evaluates first based purely on technical evidence
- Pathos evaluates second based purely on human impact (without seeing Logos verdict)
- Sophia evaluates third based purely on strategic value (without seeing prior verdicts)
- Synthesis happens only after all three perspectives are documented
Anti-Anchoring Measures
- Each perspective must state its position BEFORE seeing others
- Use structured prompts that focus on domain-specific criteria
- Explicitly note when a perspective has "nothing to add" (rather than forcing agreement)
- Flag when one perspective's conclusion was influenced by another's framing
Why order matters: iterative debate is a martingale — majority voting over independent votes captures most of the available gain. A single persuasive agent can lower group accuracy 10-40% and raise consensus on wrong answers by >30%. In Engine Mode, never expose one engine's output to another before all have voted.
Engine Independence Protocol
When using Engine Mode, independence is achieved through physical separation rather than simulated isolation:
Physical Independence
- Claude analyzes first — completes integrated analysis before any external engine calls
- External engines execute in parallel — Codex and Gemini receive identical context, run independently
- No cross-contamination — Claude's analysis is stored before external outputs are collected
- Identical context delivery — All engines receive the same decision framing and context
Anti-Contamination Rules
- Claude must not modify its analysis after seeing external engine outputs
- External engine prompts must not include Claude's position or reasoning
- If re-deliberation occurs, all engines re-evaluate independently
Advantages Over Simulated Independence
| Aspect | Simple Mode (Simulated) | Engine Mode (Physical) |
|---|---|---|
| Independence guarantee | Best-effort isolation | True process separation |
| Model diversity | Single model, multiple prompts | Multiple models |
| Bias correlation | Correlated (same model) | Uncorrelated (different training) |
| Failure mode | Single point of failure | Graceful degradation |
Confidence Scoring Guide
Each perspective assigns a confidence score to its assessment:
| Score | Level | Meaning | Basis |
|---|---|---|---|
| 90-100 | Certain | Mathematical/logical certainty | Formal proof, overwhelming data |
| 75-89 | High | Strong evidence supports conclusion | Multiple data points, clear precedent |
| 60-74 | Moderate | Reasonable confidence with some gaps | Limited data, some assumptions |
| 40-59 | Low | Significant uncertainty | Sparse data, many assumptions |
| < 40 | Speculative | Educated guess at best | No data, pure reasoning |
Confidence Calibration Rules
- Never claim 100 unless mathematically provable
- Below 40 triggers mandatory "Uncertainty Disclaimer" in output
- Average confidence below 50 across all perspectives triggers re-deliberation or escalation
- Large confidence gap (>30 points between perspectives) should be explicitly discussed
Engine Confidence Scoring
Engine Mode uses the same 0-100 confidence scale, with additional considerations for cross-engine comparison.
Cross-Engine Confidence Distribution
Different engines may have systematic confidence biases:
| Engine | Typical Range | Notes |
|---|---|---|
| Claude | 60-85 | Tends toward moderate confidence with calibrated uncertainty |
| Codex | 55-80 | May show lower confidence on non-code decisions |
| Gemini | 50-85 | Wider range, may be more decisive on strategic questions |
Engine Confidence Rules
- 95 cap rule: Same as Simple Mode — no engine should claim confidence > 95 unless mathematically provable
- Cross-engine gap > 30: Flag for review, may indicate different context interpretation
- All engines < 50: Re-deliberate with more context or fall back to Simple Mode
- Single engine at 0: Likely parse failure — check error log before including in vote
Conflict Patterns and Their Meaning
Logos vs Pathos (Technical vs Human)
Pattern: "Technically optimal but humanly costly" Typical scenario: Performance optimization that makes code unmaintainable Resolution signal: Pathos usually wins for internal tools; Logos for performance-critical systems
Logos vs Sophia (Technical vs Business)
Pattern: "Technically correct but strategically wrong" Typical scenario: Perfect architecture that takes too long to build Resolution signal: Sophia usually wins for time-sensitive decisions; Logos for foundational decisions
Pathos vs Sophia (Human vs Business)
Pattern: "Good for people but bad for business" (or vice versa) Typical scenario: Developer experience improvement with no business ROI Resolution signal: Context-dependent; often the dissenting voice carries important warnings
All Three Conflict (Three-Way Split)
Pattern: Each perspective sees a fundamentally different problem Meaning: The decision may be poorly framed or the scope too broad Action: Reframe the question or decompose into smaller decisions
Re-Deliberation Triggers
Restart deliberation when:
- New evidence emerges that changes a perspective's assessment by >20 confidence points
- Constraints change (budget, timeline, team composition)
- User provides additional context that wasn't in the original framing
- Implementation reveals assumptions were incorrect
- Split verdict (1-1-1) and user requests deeper analysis
- All perspectives have confidence below 50
Re-Deliberation Protocol
- Identify which perspective(s) are affected by new information
- Re-evaluate only the affected perspective(s)
- Conduct a new vote
- Compare old and new verdicts
- Document what changed and why