Evidence
Moved verbatim from the README during the docs unification so the register keeps one canonical home. Summarized in status.md.
📖 Evidence & Limitations
✅ Evidence-Based Design
This workflow is grounded in empirical evidence from the 2025-2026 AI agent research boom. Every architectural decision - from parallel subagent orchestration to cross-session learning - is backed by peer-reviewed papers, open-source tools, and industry benchmarks.
| Practice | Source | Evidence | Where We Implement |
|---|---|---|---|
| Structured parallel execution | CAID (Geng & Neubig, CMU, 2026) | +25.6pp PaperBench / +14.7pp Commit0 (absolute, vs same-framework single-agent; latest revision — v1 reported 26.7/14.3) via central manager + isolated git worktree per engineer + test-gated integration. Benchmarks are from-scratch construction, not brownfield edits | Plan Critique borrows the isolation pattern only (fresh context, independent files, single consolidator) — not worktrees or test-gated merges. Code execution stays sequential by default; the experimental structured-parallel path is specified in rfc-parallel-scope-execution.md |
| Cross-session learning | Cat (Liu et al., Beihang, 2025); Memory Transfer (Kim et al., KAIST, 2026) | Context as callable tool; +3.7% via abstract memory pools | Session knowledge from past cycles read during workflow setup |
| Output validation guards | Stage-Gate Agentic (PDMA, 2026); Phaselock (2026) | AI agents with gates reduce execution failures; 80 enforceable rules | Shape Up output guard + Tech Planning validation guard |
| Context isolation | Clean Context Pattern (Agent Factory, 2026); GAM (Zhejiang U., 2026) | Fresh context per agent outperforms shared pipelines; write isolation prevents contamination | subagents.md - context:"fresh" per subagent; disk-based artifacts |
| Visual review gate | Plannotator (backnotprop, 2025); Placement Theory (Tian Pan, 2026) | Browser-based plan annotation with structured feedback loop | Plannotator gate active when Review Mode > Auto; skipped in Auto |
| Intra-step recovery | Try-Heal-Retry (Nweke, 2026); PALADIN (Chaudhary et al., 2025) | 89.68% recovery rate via annotated failure trajectories | subagents.md - Retry 1× + skip with logged error per subagent |
| Unstructured cooperation penalty | CooperBench (Khatua et al., 2026) | 600+ tasks, 12 libraries, 4 languages: peer agents score on average 30% lower together than solo; monotonic decline 68.6% → 46.5% → 30.0% as teams grow 2 → 3 → 4 | Plan Critique uses fresh-context subagents with zero inter-agent communication and independent file outputs; no peer-to-peer coordination exists anywhere in the workflow |
| Self-organizing team penalty | Multi-Agent Teams Hold Experts Back (Pappu et al., 2026) | Self-organizing teams (no fixed roles) trail their own best member by up to 41.1% on ML benchmarks; failure mode is integrative compromise (averaging expert and non-expert views), worsening with team size | No self-organizing teams: every scope runs under a manager-owned plan with acceptance contracts and a single consolidator — never peer negotiation |
| Git-primitive coordination at scale | Building a C compiler with a team of parallel Claudes (Anthropic, Feb 2026) | Demonstration, not controlled evidence (no single-agent baseline reported): ~16 agents, ~2,000 sessions, ~$20k building a Rust-based C compiler (builds Linux 6.9); coordination through task lock-files written into git, merge conflicts frequent | scripts/stelow lock borrows the lock-file coordination idea; worktree isolation deliberately not adopted — merge burden judged disproportionate to 2–3-scope risk (see scope-execution-strategy.md) |
| Communication topology limits | clawRxiv 2604.00736 (2026) | Overhead grows quadratically: C(n)=0.023n²+0.04n; 50% at n=7; agents inflate 34% when aware of peers | Max 4-5 parallel subagents (n≤5 optimal zone); no message passing between agents — each writes independent file |
| Partitioning decides parallelism | Co-Coder (Yang et al., 2026) | 28 real-world projects on DevEval + CodeProjectEval: 56.8% → 68.1% (+11.3pp) with 2.10× speedup and −28% cost on DevEval; gains largest on densest cross-file dependencies. Naive file-parallel: no latency gain (806s vs 800s), +44% cost, negligible pass-rate gain; Claude Code Agent Teams lowest pass rate (54.1%) | Code execution defaults to sequential as a brownfield risk posture — not a Co-Coder deduction (the paper favors partitioning coupled work, not avoiding it). Parallel scope dispatch is opt-in with post-hoc git diff --name-only overlap audit; target_files intersection is coarser than cohesion partitioning — see the experimental upgrade path in rfc-parallel-scope-execution.md |
| Metric-driven optimization | ReflexGrad (Kadu et al., 2025); ReliabilityBench (Gupta et al., 2026) | +40pp lift via dual-process routing; standardized reliability measurement | optimization scopes routed to optimization goals (subagent + acceptance) |
| Acceptance-based execution | Pattern inspired by Try-Heal-Retry (Nweke, 2026) and PALADIN (Chaudhary et al., 2025) | Self-correction in same context outperforms fresh re-delegation | Scope executor delegates with acceptance contract - child self-corrects (harness-dependent) before parent evaluates |
| Audit gap-to-scope loop | Pattern inspired by Agentic Debugging (Zhang et al., 2025) | Multi-agent feedback loops improve fix rate | Audit classifies gaps → ESCALATED become new scopes → /sw-next enforces loop back to Execution |
Structured parallelism, not more agents. The evidence distinguishes how agents run in parallel, not whether: manager-led execution with isolated workspaces and test-gated integration beats the same single agent by large margins on long-horizon greenfield construction (CAID 2026: +25.6pp PaperBench, +14.7pp Commit0; Co-Coder 2026: +11.3pp on DevEval with 2.10× speedup) — with gains largest where cross-file dependencies are densest, provided work is partitioned by cohesion rather than by file. Unstructured cooperation degrades instead: peer agents score on average 30% lower together than solo (CooperBench 2026), and self-organizing teams trail their own expert by up to 41.1% (Hold Experts Back 2026). Stelow's sequential default for code execution is therefore a scope-size and brownfield posture — 1–5 scopes have a short critical path where worktree/merge overhead eats the gain, and legacy coupling makes cohesion partitioning unreliable — not a deduction from CAID/Co-Coder, which point the other way. Parallel scope dispatch stays opt-in with post-hoc overlap audit; the experimental structured-parallel path (Complete appetite, cymbal-verified transitive disjointness, parent-owned test-gated merge, measured sidecar) is specified in rfc-parallel-scope-execution.md.
⚠️ Known Limitations & Radical Transparency
Even with these guardrails, the AI agent still exhibits predictable failure modes. This workflow is a tool for amplifying human judgment, not a substitute for it.
How to read this table: Each row is honest about what the workflow can and cannot do. Every mitigation has a corresponding "not solved" assessment. Read both before deciding whether this workflow helps your context.
| # | Limitation | Impact | What the workflow tries to do | Why it's not solved |
|---|---|---|---|---|
| 1 | Context rot - compliance with own rules drops from ~73% (turn 5) to ~33% (turn 16) in long sessions | Gamage 2026, 4,416 trials, 12 models/8 providers. Replicated by Liu et al. 2023 "Lost in the Middle". | Subagents use context: "fresh". Ordered-execution-goal creates isolated scope execution. Execution stage has explicit "Context Rot Check" re-reading plan from disk. | Reduced but not solved. The orchestrator itself can forget its own rules in long sessions spanning multiple stages. The core transformer limitation (U-shaped attention curve) remains intrinsic. |
| 2 | Confabulated research references - Agents cite nonexistent papers or books (~11-57% hallucination rate across models) | arXiv 2604.03173 - 10 models/3 databases/69K citation instances | Claim verification via Lessons Learned cross-referencing during setup. | Caught by structure, not guaranteed. Multi-model consensus (≥3 LLMs citing same work) yields 95.6% accuracy, but the workflow doesn't enforce this. |
| 3 | Silent wrong answers - Cross-task state leakage produces plausible but incorrect outputs | UCC (arXiv 2604.01350), 2026 | Write isolation per subagent; clean context pattern | Mitigated by isolation, not by detection. No mechanism to detect when contamination happens despite isolation. |
| 4 | Overconfidence in estimates - AI systematically underestimates implementation complexity | Agentic Overconfidence (ICLR 2026) - all tested agents exhibit agentic overconfidence | Run knobs are declared by human as constraints, not estimated by the LLM. The LLM only checks appetite_fit (fits/cuts_needed/reshape). No estimation step. | Addressed by design - knobs are constraints, not estimates. The human sets the bounds before shaping. The LLM checks fit, not effort. But the human still needs to set knobs honestly. |
| 5 | Approval gate fatigue - Users can desensitize to visual gates and approve without scrutiny | Tian Pan Apr 2026 - HITL queues have dynamics | Plannotator requires active annotations (deletions, comments, labels). Auto/Product Spec Gate review modes skip gates entirely when appropriate. | Delayed, not prevented. Review Mode selection helps reduce unnecessary gates, but if the human always picks Complete+Product Spec + Interface + Scopes, fatigue still sets in. |
| 6 | 80% Problem - AI ships the happy path (CRUD, main flow) but omits error handling, observability, security, retry, rollback, edge cases | Osmani Jan 2026 (coined the term); GitClear 2025 | Tech Planning requires NFRs per scope. Acceptance contracts can include NFR criteria (if the plan specifies them). Audit classifies omissions as gaps - ESCALATED ones become new scopes. | Partially mitigated, not solved. NFRs must be in the plan to appear in the contract. Audit classification depends on the LLM - misclassification means gaps slip through. Same model evaluates both stages. |
| 7 | Model dependency - Claude Opus, Gemini Flash, GPT-4o produce significantly different quality | Veracode 2025 - 45% of AI-generated code contains flaws across 100+ models; Anthropic Jan 2026 - RCT: AI-assisted devs score 17% lower on comprehension tests | Every artifact tracks generated_by: {model_name} in frontmatter. Gate stage shows provenance before Plannotator review. | Transparency, not mitigation. Knowing the model helps calibrate expectations, but it doesn't fix the quality gap. The comprehension penalty (Anthropic 2026) affects users regardless. |
| 8 | Constraint decay - AI progressively violates its own self-imposed rules over time | arXiv 2026 (Constraint Decay) - structural constraints drift in backend code generation; HORIZON - agents break on long-horizon tasks | Context rot rules explicitly warn about this. "No patching in degraded context" rule blocks the most common decay pattern. | Same root cause as context rot. The warning helps, but stopping a session mid-flow is disruptive and users rarely do it. |
| 9 | Code hallucination - AI invents APIs, functions, or contracts that don't exist (~20% of failures) | CloudAPIBench - 20.41% of failures are hallucinated APIs; Code LLM failures | Verification stage runs the test suite, which catches some hallucinated APIs. | Caught by tests, not by the workflow. If tests don't exist (or are also hallucinated), neither Verification nor Critique detects it. |
| 10 | Shallow review trap - same LLM that wrote the code also reviews it | Ox Security 2025 - 300+ repos, 10 anti-patterns, AI code in production with critical flaws | Verification uses context: "fresh" subagent reviewers - same model but fresh session context. | Automatic via context: "fresh" - fresh context restores full rule awareness lost to context rot (~33% rule adherence at turn 16 vs ~73% at turn 5). True cross-model independence offers marginal additional benefit. |
| 11 | Expertise cliff - AI fails in mature codebases with implicit conventions, undocumented architecture | Tian Pan Mai 2026; METR 2025 RCT - experienced devs 19% slower with AI | Domain libraries and structured specs help surface some conventions. Execution Critique checks for broken refs and anti-patterns. | Not addressed. This workflow was designed for greenfield or well-documented features. If your codebase has 10 years of undocumented architecture decisions, the AI will violate them. |
| 12 | Plan staleness - plans generated against one snapshot; by execution time, target has changed | Superpowers Issue #989 - parallel sessions cause spec/plan staleness | Git diff check before scope execution detects if target files changed since plan creation. | Staleness detected but not auto-resolved. Only detects file-level changes, not semantic staleness. LLM decides whether staleness matters - no forced re-plan. |
| 13 | Pipeline memory loss - no cross-session memory of own failure patterns | Flamehaven 2026 - cross-session memory, MICA governance schema | Execution Critique saves lessons from each cycle. Setup stage automatically reads past lessons with forced reflection. | Captured and injected, but not verified. Same model that made mistakes reads the lessons. Context rot can still cause mid-session forgetting. Cannot auto-verify lesson adherence. |
| 14 | Code complexity growth - AI-generated code increases complexity over time | Cursor Study (MSR 2026) - static analysis warnings +30%, code complexity +41% after month 2 | Execution Critique includes anti-pattern detection (god functions >100 lines, global mutable state). Optional Code Quality Gate with static analysis. | Caught too late. Complexity analysis happens after code is written. No mechanism to prevent complexity during generation - only flag it after. |
| 15 | Activity ≠ productivity - more PRs, more commits does not mean more value delivered | METR 2025 RCT - 19% slower for experienced devs; Faros AI 2025 - 9% more tasks, 0% DORA improvement | Run knobs anchor scope size to human attention budget. OUT/IN scoping keeps proposals focused. Execution Critique includes "close without follow-up" as valid outcome. | Honest assessment: Knobs mitigate scope bloat, but require humans to set them honestly. appetite_fit is validated by the Plan Critique stage's fresh-context feasibility reviewer (reusing existing 5-reviewer infrastructure). The knob system is new - its real-world effectiveness is not yet measured. |
| 16 | Coordination overhead — adding agents to shared-state coding tasks degrades quality | CooperBench 2026 — peer agents score on average 30% lower together than solo (600+ tasks); clawRxiv 2604.00736 — broadcast/P2P overhead C(n)=0.023n²+0.04n (R²=0.98), 50% at n=7 | Unstructured and self-organizing cooperation are banned outright (manager-owned plan + acceptance contracts everywhere; CooperBench, Hold Experts Back). Research/review parallelism uses fresh context, zero inter-agent communication, independent file outputs. Code execution defaults to sequential; parallel scope dispatch is opt-in via post-hoc overlap detection (git diff --name-only per scope) + opt-in file-reservation locks (see file-locking.md). Full pipeline in scope-execution-strategy.md; experimental structured-parallel path in rfc-parallel-scope-execution.md. | Partially addressed. Sequential default + audit + cooperation bans cover the unstructured failure modes the papers demonstrate. What is not settled: whether structured parallel scope execution (the CAID/Co-Coder pattern) transfers to brownfield stelow scopes — the measurement mirror ships since v0.65.0, but no comparative data has been collected yet. The RFC defines the experiment that would settle it; until then the default stays sequential. |
What this means for you
- Every artifact is a draft. Treat spec-product.md, spec-tech.md, critique reports, and interface proposals as first drafts that need human eyes.
- Results vary by model and codebase. A small model generating a plan for a mature codebase is a recipe for failure - regardless of how structured the workflow is.
- Human review is required. The workflow catches structural gaps (missing scopes, contradictory requirements, some untested edge cases). It does NOT catch logic errors in individual lines, security flaws in business logic, or nuanced architectural trade-offs - those need you.
We don't claim to solve product planning. We claim to structure the thinking so you catch more before you code. The rest is still up to you.
Research sourced May 2026. All references are hyperlinked for verification.