Evidence

Moved verbatim from the README during the docs unification so the register keeps one canonical home. Summarized in status.md.

📖 Evidence & Limitations

✅ Evidence-Based Design

This workflow is grounded in empirical evidence from the 2025-2026 AI agent research boom. Every architectural decision - from parallel subagent orchestration to cross-session learning - is backed by peer-reviewed papers, open-source tools, and industry benchmarks.

PracticeSourceEvidenceWhere We Implement
Structured parallel executionCAID (Geng & Neubig, CMU, 2026)+25.6pp PaperBench / +14.7pp Commit0 (absolute, vs same-framework single-agent; latest revision — v1 reported 26.7/14.3) via central manager + isolated git worktree per engineer + test-gated integration. Benchmarks are from-scratch construction, not brownfield editsPlan Critique borrows the isolation pattern only (fresh context, independent files, single consolidator) — not worktrees or test-gated merges. Code execution stays sequential by default; the experimental structured-parallel path is specified in rfc-parallel-scope-execution.md
Cross-session learningCat (Liu et al., Beihang, 2025); Memory Transfer (Kim et al., KAIST, 2026)Context as callable tool; +3.7% via abstract memory poolsSession knowledge from past cycles read during workflow setup
Output validation guardsStage-Gate Agentic (PDMA, 2026); Phaselock (2026)AI agents with gates reduce execution failures; 80 enforceable rulesShape Up output guard + Tech Planning validation guard
Context isolationClean Context Pattern (Agent Factory, 2026); GAM (Zhejiang U., 2026)Fresh context per agent outperforms shared pipelines; write isolation prevents contaminationsubagents.md - context:"fresh" per subagent; disk-based artifacts
Visual review gatePlannotator (backnotprop, 2025); Placement Theory (Tian Pan, 2026)Browser-based plan annotation with structured feedback loopPlannotator gate active when Review Mode > Auto; skipped in Auto
Intra-step recoveryTry-Heal-Retry (Nweke, 2026); PALADIN (Chaudhary et al., 2025)89.68% recovery rate via annotated failure trajectoriessubagents.md - Retry 1× + skip with logged error per subagent
Unstructured cooperation penaltyCooperBench (Khatua et al., 2026)600+ tasks, 12 libraries, 4 languages: peer agents score on average 30% lower together than solo; monotonic decline 68.6% → 46.5% → 30.0% as teams grow 2 → 3 → 4Plan Critique uses fresh-context subagents with zero inter-agent communication and independent file outputs; no peer-to-peer coordination exists anywhere in the workflow
Self-organizing team penaltyMulti-Agent Teams Hold Experts Back (Pappu et al., 2026)Self-organizing teams (no fixed roles) trail their own best member by up to 41.1% on ML benchmarks; failure mode is integrative compromise (averaging expert and non-expert views), worsening with team sizeNo self-organizing teams: every scope runs under a manager-owned plan with acceptance contracts and a single consolidator — never peer negotiation
Git-primitive coordination at scaleBuilding a C compiler with a team of parallel Claudes (Anthropic, Feb 2026)Demonstration, not controlled evidence (no single-agent baseline reported): ~16 agents, ~2,000 sessions, ~$20k building a Rust-based C compiler (builds Linux 6.9); coordination through task lock-files written into git, merge conflicts frequentscripts/stelow lock borrows the lock-file coordination idea; worktree isolation deliberately not adopted — merge burden judged disproportionate to 2–3-scope risk (see scope-execution-strategy.md)
Communication topology limitsclawRxiv 2604.00736 (2026)Overhead grows quadratically: C(n)=0.023n²+0.04n; 50% at n=7; agents inflate 34% when aware of peersMax 4-5 parallel subagents (n≤5 optimal zone); no message passing between agents — each writes independent file
Partitioning decides parallelismCo-Coder (Yang et al., 2026)28 real-world projects on DevEval + CodeProjectEval: 56.8% → 68.1% (+11.3pp) with 2.10× speedup and −28% cost on DevEval; gains largest on densest cross-file dependencies. Naive file-parallel: no latency gain (806s vs 800s), +44% cost, negligible pass-rate gain; Claude Code Agent Teams lowest pass rate (54.1%)Code execution defaults to sequential as a brownfield risk posture — not a Co-Coder deduction (the paper favors partitioning coupled work, not avoiding it). Parallel scope dispatch is opt-in with post-hoc git diff --name-only overlap audit; target_files intersection is coarser than cohesion partitioning — see the experimental upgrade path in rfc-parallel-scope-execution.md
Metric-driven optimizationReflexGrad (Kadu et al., 2025); ReliabilityBench (Gupta et al., 2026)+40pp lift via dual-process routing; standardized reliability measurementoptimization scopes routed to optimization goals (subagent + acceptance)
Acceptance-based executionPattern inspired by Try-Heal-Retry (Nweke, 2026) and PALADIN (Chaudhary et al., 2025)Self-correction in same context outperforms fresh re-delegationScope executor delegates with acceptance contract - child self-corrects (harness-dependent) before parent evaluates
Audit gap-to-scope loopPattern inspired by Agentic Debugging (Zhang et al., 2025)Multi-agent feedback loops improve fix rateAudit classifies gaps → ESCALATED become new scopes → /sw-next enforces loop back to Execution
Structured parallelism, not more agents. The evidence distinguishes how agents run in parallel, not whether: manager-led execution with isolated workspaces and test-gated integration beats the same single agent by large margins on long-horizon greenfield construction (CAID 2026: +25.6pp PaperBench, +14.7pp Commit0; Co-Coder 2026: +11.3pp on DevEval with 2.10× speedup) — with gains largest where cross-file dependencies are densest, provided work is partitioned by cohesion rather than by file. Unstructured cooperation degrades instead: peer agents score on average 30% lower together than solo (CooperBench 2026), and self-organizing teams trail their own expert by up to 41.1% (Hold Experts Back 2026). Stelow's sequential default for code execution is therefore a scope-size and brownfield posture — 1–5 scopes have a short critical path where worktree/merge overhead eats the gain, and legacy coupling makes cohesion partitioning unreliable — not a deduction from CAID/Co-Coder, which point the other way. Parallel scope dispatch stays opt-in with post-hoc overlap audit; the experimental structured-parallel path (Complete appetite, cymbal-verified transitive disjointness, parent-owned test-gated merge, measured sidecar) is specified in rfc-parallel-scope-execution.md.

⚠️ Known Limitations & Radical Transparency

Even with these guardrails, the AI agent still exhibits predictable failure modes. This workflow is a tool for amplifying human judgment, not a substitute for it.

How to read this table: Each row is honest about what the workflow can and cannot do. Every mitigation has a corresponding "not solved" assessment. Read both before deciding whether this workflow helps your context.

#LimitationImpactWhat the workflow tries to doWhy it's not solved
1Context rot - compliance with own rules drops from ~73% (turn 5) to ~33% (turn 16) in long sessionsGamage 2026, 4,416 trials, 12 models/8 providers. Replicated by Liu et al. 2023 "Lost in the Middle".Subagents use context: "fresh". Ordered-execution-goal creates isolated scope execution. Execution stage has explicit "Context Rot Check" re-reading plan from disk.Reduced but not solved. The orchestrator itself can forget its own rules in long sessions spanning multiple stages. The core transformer limitation (U-shaped attention curve) remains intrinsic.
2Confabulated research references - Agents cite nonexistent papers or books (~11-57% hallucination rate across models)arXiv 2604.03173 - 10 models/3 databases/69K citation instancesClaim verification via Lessons Learned cross-referencing during setup.Caught by structure, not guaranteed. Multi-model consensus (≥3 LLMs citing same work) yields 95.6% accuracy, but the workflow doesn't enforce this.
3Silent wrong answers - Cross-task state leakage produces plausible but incorrect outputsUCC (arXiv 2604.01350), 2026Write isolation per subagent; clean context patternMitigated by isolation, not by detection. No mechanism to detect when contamination happens despite isolation.
4Overconfidence in estimates - AI systematically underestimates implementation complexityAgentic Overconfidence (ICLR 2026) - all tested agents exhibit agentic overconfidenceRun knobs are declared by human as constraints, not estimated by the LLM. The LLM only checks appetite_fit (fits/cuts_needed/reshape). No estimation step.Addressed by design - knobs are constraints, not estimates. The human sets the bounds before shaping. The LLM checks fit, not effort. But the human still needs to set knobs honestly.
5Approval gate fatigue - Users can desensitize to visual gates and approve without scrutinyTian Pan Apr 2026 - HITL queues have dynamicsPlannotator requires active annotations (deletions, comments, labels). Auto/Product Spec Gate review modes skip gates entirely when appropriate.Delayed, not prevented. Review Mode selection helps reduce unnecessary gates, but if the human always picks Complete+Product Spec + Interface + Scopes, fatigue still sets in.
680% Problem - AI ships the happy path (CRUD, main flow) but omits error handling, observability, security, retry, rollback, edge casesOsmani Jan 2026 (coined the term); GitClear 2025Tech Planning requires NFRs per scope. Acceptance contracts can include NFR criteria (if the plan specifies them). Audit classifies omissions as gaps - ESCALATED ones become new scopes.Partially mitigated, not solved. NFRs must be in the plan to appear in the contract. Audit classification depends on the LLM - misclassification means gaps slip through. Same model evaluates both stages.
7Model dependency - Claude Opus, Gemini Flash, GPT-4o produce significantly different qualityVeracode 2025 - 45% of AI-generated code contains flaws across 100+ models; Anthropic Jan 2026 - RCT: AI-assisted devs score 17% lower on comprehension testsEvery artifact tracks generated_by: {model_name} in frontmatter. Gate stage shows provenance before Plannotator review.Transparency, not mitigation. Knowing the model helps calibrate expectations, but it doesn't fix the quality gap. The comprehension penalty (Anthropic 2026) affects users regardless.
8Constraint decay - AI progressively violates its own self-imposed rules over timearXiv 2026 (Constraint Decay) - structural constraints drift in backend code generation; HORIZON - agents break on long-horizon tasksContext rot rules explicitly warn about this. "No patching in degraded context" rule blocks the most common decay pattern.Same root cause as context rot. The warning helps, but stopping a session mid-flow is disruptive and users rarely do it.
9Code hallucination - AI invents APIs, functions, or contracts that don't exist (~20% of failures)CloudAPIBench - 20.41% of failures are hallucinated APIs; Code LLM failuresVerification stage runs the test suite, which catches some hallucinated APIs.Caught by tests, not by the workflow. If tests don't exist (or are also hallucinated), neither Verification nor Critique detects it.
10Shallow review trap - same LLM that wrote the code also reviews itOx Security 2025 - 300+ repos, 10 anti-patterns, AI code in production with critical flawsVerification uses context: "fresh" subagent reviewers - same model but fresh session context.Automatic via context: "fresh" - fresh context restores full rule awareness lost to context rot (~33% rule adherence at turn 16 vs ~73% at turn 5). True cross-model independence offers marginal additional benefit.
11Expertise cliff - AI fails in mature codebases with implicit conventions, undocumented architectureTian Pan Mai 2026; METR 2025 RCT - experienced devs 19% slower with AIDomain libraries and structured specs help surface some conventions. Execution Critique checks for broken refs and anti-patterns.Not addressed. This workflow was designed for greenfield or well-documented features. If your codebase has 10 years of undocumented architecture decisions, the AI will violate them.
12Plan staleness - plans generated against one snapshot; by execution time, target has changedSuperpowers Issue #989 - parallel sessions cause spec/plan stalenessGit diff check before scope execution detects if target files changed since plan creation.Staleness detected but not auto-resolved. Only detects file-level changes, not semantic staleness. LLM decides whether staleness matters - no forced re-plan.
13Pipeline memory loss - no cross-session memory of own failure patternsFlamehaven 2026 - cross-session memory, MICA governance schemaExecution Critique saves lessons from each cycle. Setup stage automatically reads past lessons with forced reflection.Captured and injected, but not verified. Same model that made mistakes reads the lessons. Context rot can still cause mid-session forgetting. Cannot auto-verify lesson adherence.
14Code complexity growth - AI-generated code increases complexity over timeCursor Study (MSR 2026) - static analysis warnings +30%, code complexity +41% after month 2Execution Critique includes anti-pattern detection (god functions >100 lines, global mutable state). Optional Code Quality Gate with static analysis.Caught too late. Complexity analysis happens after code is written. No mechanism to prevent complexity during generation - only flag it after.
15Activity ≠ productivity - more PRs, more commits does not mean more value deliveredMETR 2025 RCT - 19% slower for experienced devs; Faros AI 2025 - 9% more tasks, 0% DORA improvementRun knobs anchor scope size to human attention budget. OUT/IN scoping keeps proposals focused. Execution Critique includes "close without follow-up" as valid outcome.Honest assessment: Knobs mitigate scope bloat, but require humans to set them honestly. appetite_fit is validated by the Plan Critique stage's fresh-context feasibility reviewer (reusing existing 5-reviewer infrastructure). The knob system is new - its real-world effectiveness is not yet measured.
16Coordination overhead — adding agents to shared-state coding tasks degrades qualityCooperBench 2026 — peer agents score on average 30% lower together than solo (600+ tasks); clawRxiv 2604.00736 — broadcast/P2P overhead C(n)=0.023n²+0.04n (R²=0.98), 50% at n=7Unstructured and self-organizing cooperation are banned outright (manager-owned plan + acceptance contracts everywhere; CooperBench, Hold Experts Back). Research/review parallelism uses fresh context, zero inter-agent communication, independent file outputs. Code execution defaults to sequential; parallel scope dispatch is opt-in via post-hoc overlap detection (git diff --name-only per scope) + opt-in file-reservation locks (see file-locking.md). Full pipeline in scope-execution-strategy.md; experimental structured-parallel path in rfc-parallel-scope-execution.md.Partially addressed. Sequential default + audit + cooperation bans cover the unstructured failure modes the papers demonstrate. What is not settled: whether structured parallel scope execution (the CAID/Co-Coder pattern) transfers to brownfield stelow scopes — the measurement mirror ships since v0.65.0, but no comparative data has been collected yet. The RFC defines the experiment that would settle it; until then the default stays sequential.

What this means for you

We don't claim to solve product planning. We claim to structure the thinking so you catch more before you code. The rest is still up to you.

Research sourced May 2026. All references are hyperlinked for verification.