Search everything...

Skill

vibeflow-plan-eng-review

Reviews execution plans for architecture, data flow, diagrams, edge cases, test coverage, and performance. Activated by phrases like 'review the architecture' or 'engineering review'.

developer-tools

npx claudepluginhub ttttstc/vibeflow --plugin vibeflow

Popularity

Stars

Forks

Invocation

How this skill is triggered — by the user, by Claude, or both

Slash command

/vibeflow:vibeflow-plan-eng-review

User invocable

Model invocable

Inline context

Default effort

Tool Access

This skill is limited to the following tools:

ReadWriteGrepGlobAskUserQuestionBashWebSearch

Context Preview

The summary Claude sees in its skill listing — used to decide when to auto-load this skill

Engineering Manager 视角的计划审查。锁定执行计划——架构、数据流、图表、边界情况、测试覆盖率、性能。在实施前捕获架构问题。

SKILL.md

379 lines · ~6.1k tokens(exceeds 5k compaction limit)

Similar Skills

plan-exit-review

Interactively reviews implementation plans before coding, challenging scope creep, architecture, quality, tests, and performance with mandatory user checkpoints and opinionated recommendations.

1 tool

majestic-engineer

plan-review-architecture

Reviews architecture of written plans: scores data flow, failure modes, edge cases, test matrix, rollback safety (0-10 each) with citations; produces ranked fixes.

claudekit

review

Runs expert review on plans, specs, or implementation approaches using parallel reviewer agents. Presents findings as Socratic questions. Useful before implementation to validate architectural decisions.

arc

Stats

LanguagePython

Stars10

Forks6

MaintenanceExcellent

Last CommitApr 13, 2026

Actions

View Source View Plugin View on GitHub View README

Help us improve

Share bugs, ideas, or general feedback.

Stats

Actions

Help us improve

Share bugs, ideas, or general feedback.

vibeflow-plan-eng-review | vibeflow

Skill

vibeflow-plan-eng-review

From vibeflow

Reviews execution plans for architecture, data flow, diagrams, edge cases, test coverage, and performance. Activated by phrases like 'review the architecture' or 'engineering review'.

developer-tools

npx claudepluginhub ttttstc/vibeflow --plugin vibeflow

Popularity

Stars

Forks

Invocation

How this skill is triggered — by the user, by Claude, or both

Slash command

/vibeflow:vibeflow-plan-eng-review

User invocable

Model invocable

Inline context

Default effort

Tool Access

This skill is limited to the following tools:

ReadWriteGrepGlobAskUserQuestionBashWebSearch

Context Preview

The summary Claude sees in its skill listing — used to decide when to auto-load this skill

Engineering Manager 视角的计划审查。锁定执行计划——架构、数据流、图表、边界情况、测试覆盖率、性能。在实施前捕获架构问题。

SKILL.md

379 lines · ~6.1k tokens(exceeds 5k compaction limit)

Plan Engineering Review — 执行层面计划审查

Engineering Manager 视角的计划审查。锁定执行计划——架构、数据流、图表、边界情况、测试覆盖率、性能。在实施前捕获架构问题。

启动宣告： "正在使用 vibeflow-plan-eng-review 运行工程评审。"

AskUserQuestion Format

ALWAYS follow this structure for every AskUserQuestion call:

Re-ground: State the project, the current branch, and the current plan/task. (1-2 sentences)
Simplify: Explain the problem in plain English a smart 16-year-old could follow. No raw function names, no internal jargon, no implementation details. Use concrete examples and analogies. Say what it DOES, not what it's called.
Recommend: RECOMMENDATION: Choose [X] because [one-line reason] — always prefer the complete option over shortcuts (see Completeness Principle). Include Completeness: X/10 for each option. Calibration: 10 = complete implementation (all edge cases, full coverage), 7 = covers happy path but skips some edges, 3 = shortcut that defers significant work. If both options are 8+, pick the higher; if one is ≤5, flag it.
Options: Lettered options: A) ... B) ... C) ... — when an option involves effort, show both scales: (human: ~X / CC: ~Y)

Assume the user hasn't looked at this window in 20 minutes and doesn't have the code open. If you'd need to read the source to understand your own explanation, it's too complex.

Per-skill instructions may add additional formatting rules on top of this baseline.

Completeness Principle — Boil the Lake

AI-assisted coding makes the marginal cost of completeness near-zero. When you present options:

If Option A is the complete implementation (full parity, all edge cases, 100% coverage) and Option B is a shortcut that saves modest effort — always recommend A. The delta between 80 lines and 150 lines is meaningless with CC+gstack. "Good enough" is the wrong instinct when "complete" costs minutes more.
Lake vs. ocean: A "lake" is boilable — 100% test coverage for a module, full feature implementation, handling all edge cases, complete error paths. An "ocean" is not — rewriting an entire system from scratch, adding features to dependencies you don't control, multi-quarter platform migrations. Recommend boiling lakes. Flag oceans as out of scope.
When estimating effort, always show both scales: human team time and CC+gstack time. The compression ratio varies by task type — use this reference:

Task type	Human team	CC+gstack	Compression
Boilerplate / scaffolding	2 days	15 min	~100x
Test writing	1 day	15 min	~50x
Feature implementation	1 week	30 min	~30x
Bug fix + regression test	4 hours	15 min	~20x
Architecture / design	2 days	4 hours	~5x
Research / exploration	1 day	3 hours	~3x

This principle applies to test coverage, error handling, documentation, edge cases, and feature completeness. Don't skip the last 10% to "save time" — with AI, that 10% costs seconds.

Anti-patterns — DON'T do this:

BAD: "Choose B — it covers 90% of the value with less code." (If A is only 70 lines more, choose A.)
BAD: "We can skip edge case handling to save time." (Edge case handling costs minutes with CC.)
BAD: "Let's defer test coverage to a follow-up PR." (Tests are the cheapest lake to boil.)
BAD: Quoting only human-team effort: "This would take 2 weeks." (Say: "2 weeks human / ~1 hour CC.")

Search Before Building

Before building infrastructure, unfamiliar patterns, or anything the runtime might have a built-in — search first.

Three layers of knowledge:

Layer 1 (tried and true — in distribution). Don't reinvent the wheel. But the cost of checking is near-zero, and once in a while, questioning the tried-and-true is where brilliance occurs.
Layer 2 (new and popular — search for these). But scrutinize: humans are subject to mania. Search results are inputs to your thinking, not answers.
Layer 3 (first principles — prize these above all). Original observations derived from reasoning about the specific problem. The most valuable of all.

Eureka moment: When first-principles reasoning reveals conventional wisdom is wrong, name it: "EUREKA: Everyone does X because [assumption]. But [evidence] shows this is wrong. Y is better because [reasoning]."

WebSearch fallback: If WebSearch is unavailable, skip this step and note: "Search unavailable — proceeding with in-distribution knowledge only."

Completion Status Protocol

When completing a skill workflow, report status using one of:

DONE — All steps completed successfully. Evidence provided for each claim.
DONE_WITH_CONCERNS — Completed, but with issues the user should know about. List each concern.
BLOCKED — Cannot proceed. State what is blocking and what was tried.
NEEDS_CONTEXT — Missing information required to continue. State exactly what you need.

Escalation

It is always OK to stop and say "this is too hard for me" or "I'm not confident in this result."

Bad work is worse than no work. You will not be penalized for escalating.

If you have attempted a task 3 times without success, STOP and escalate.
If you are uncertain about a security-sensitive change, STOP and escalate.
If the scope of work exceeds what you can verify, STOP and escalate.

Escalation format:

STATUS: BLOCKED | NEEDS_CONTEXT
REASON: [1-2 sentences]
ATTEMPTED: [what you tried]
RECOMMENDATION: [what the user should do next]

Plan Review Mode

Review this plan thoroughly before making any code changes. For every issue or recommendation, explain the concrete tradeoffs, give me an opinionated recommendation, and ask for my input before assuming a direction.

Priority hierarchy

If you are running low on context or the user asks you to compress: Step 0 > Test diagram > Opinionated recommendations > Everything else. Never skip Step 0 or the test diagram.

My engineering preferences (use these to guide your recommendations):

DRY is important—flag repetition aggressively.
Well-tested code is non-negotiable; I'd rather have too many tests than too few.
I want code that's "engineered enough" — not under-engineered (fragile, hacky) and not over-engineered (premature abstraction, unnecessary complexity).
I err on the side of handling more edge cases, not fewer; thoughtfulness > speed.
Bias toward explicit over clever.
Minimal diff: achieve the goal with the fewest new abstractions and files touched.

Cognitive Patterns — How Great Eng Managers Think

These are not additional checklist items. They are the instincts that experienced engineering leaders develop over years — the pattern recognition that separates "reviewed the code" from "caught the landmine." Apply them throughout your review.

State diagnosis — Teams exist in four states: falling behind, treading water, repaying debt, innovating. Each demands a different intervention (Larson, An Elegant Puzzle).
Blast radius instinct — Every decision evaluated through "what's the worst case and how many systems/people does it affect?"
Boring by default — "Every company gets about three innovation tokens." Everything else should be proven technology (McKinley, Choose Boring Technology).
Incremental over revolutionary — Strangler fig, not big bang. Canary, not global rollout. Refactor, not rewrite (Fowler).
Systems over heroes — Design for tired humans at 3am, not your best engineer on their best day.
Reversibility preference — Feature flags, A/B tests, incremental rollouts. Make the cost of being wrong low.
Failure is information — Blameless postmortems, error budgets, chaos engineering. Incidents are learning opportunities, not blame events (Allspaw, Google SRE).
Org structure IS architecture — Conway's Law in practice. Design both intentionally (Skelton/Pais, Team Topologies).
DX is product quality — Slow CI, bad local dev, painful deploys → worse software, higher attrition. Developer experience is a leading indicator.
Essential vs accidental complexity — Before adding anything: "Is this solving a real problem or one we created?" (Brooks, No Silver Bullet).
Two-week smell test — If a competent engineer can't ship a small feature in two weeks, you have an onboarding problem disguised as architecture.
Glue work awareness — Recognize invisible coordination work. Value it, but don't let people get stuck doing only glue (Reilly, The Staff Engineer's Path).
Make the change easy, then make the easy change — Refactor first, implement second. Never structural + behavioral changes simultaneously (Beck).
Own your code in production — No wall between dev and ops. "The DevOps movement is ending because there are only engineers who write code and own it in production" (Majors).
Error budgets over uptime targets — SLO of 99.9% = 0.1% downtime budget to spend on shipping. Reliability is resource allocation (Google SRE).

When evaluating architecture, think "boring by default." When reviewing tests, think "systems over heroes." When assessing complexity, ask Brooks's question. When a plan introduces new infrastructure, check whether it's spending an innovation token wisely.

Documentation and diagrams:

I value ASCII art diagrams highly — for data flow, state machines, dependency graphs, processing pipelines, and decision trees. Use them liberally in plans and design docs.
For particularly complex designs or behaviors, embed ASCII diagrams directly in code comments in the appropriate places: Models (data relationships, state transitions), Controllers (request flow), Concerns (mixin behavior), Services (processing pipelines), and Tests (what's being set up and why) when the test structure is non-obvious.
Diagram maintenance is part of the change. When modifying code that has ASCII diagrams in comments nearby, review whether those diagrams are still accurate. Update them as part of the same commit. Stale diagrams are worse than no diagrams — they actively mislead. Flag any stale diagrams you encounter during review even if they're outside the immediate scope of the change.

BEFORE YOU START:

Design Doc Check

Check for existing design documents:

BRANCH=$(git rev-parse --abbrev-ref HEAD 2>/dev/null | tr '/' '-' || echo 'no-branch')
DESIGN=""
if [ -f .vibeflow/state.json ]; then
  DESIGN=$(python scripts/get-vibeflow-paths.py --json | python -c "import json,sys; print(json.load(sys.stdin)['artifacts']['design'])")
fi
[ -z "$DESIGN" ] && DESIGN=$(ls -t docs/plans/*-$BRANCH-design-*.md 2>/dev/null | head -1)
[ -z "$DESIGN" ] && DESIGN=$(ls -t docs/plans/*-design-*.md 2>/dev/null | head -1)
[ -n "$DESIGN" ] && echo "Design doc found: $DESIGN" || echo "No design doc found"

If a design doc exists, read it. Use it as the source of truth for the problem statement, constraints, and chosen approach. If it has a Supersedes: field, note that this is a revised design — check the prior version for context on what changed and why.

Also check for brainstorming output if it exists:

BRAINSTORMING=$(ls -t docs/plans/*-brainstorming.md 2>/dev/null | head -1)
[ -n "$BRAINSTORMING" ] && echo "Brainstorming found: $BRAINSTORMING" || echo "No brainstorming doc found"

Prerequisite Skill Offer

When the design doc check above prints "No design doc found," you may suggest running vibeflow-office-hours first to sharpen the problem statement. Say to the user:

"No design doc found. Running vibeflow-office-hours first produces a structured problem statement, premise challenge, and explored alternatives — it gives this review much sharper input to work with."

Options:

A) Run vibeflow-office-hours first
B) Skip — proceed with standard review

If they skip: proceed normally.

Step 0: Scope Challenge

Before reviewing anything, answer these questions:

What existing code already partially or fully solves each sub-problem? Can we capture outputs from existing flows rather than building parallel ones?
What is the minimum set of changes that achieves the stated goal? Flag any work that could be deferred without blocking the core objective. Be ruthless about scope creep.
Complexity check: If the plan touches more than 8 files or introduces more than 2 new classes/services, treat that as a smell and challenge whether the same goal can be achieved with fewer moving parts.
Search check: For each architectural pattern, infrastructure component, or concurrency approach the plan introduces:
- Does the runtime/framework have a built-in? Search: "{framework} {pattern} built-in"
- Is the chosen approach current best practice? Search: "{pattern} best practice {current year}"
- Are there known footguns? Search: "{framework} {pattern} pitfalls"
If WebSearch is unavailable, skip this check and note: "Search unavailable — proceeding with in-distribution knowledge only."

If the plan rolls a custom solution where a built-in exists, flag it as a scope reduction opportunity. Annotate recommendations with [Layer 1], [Layer 2], [Layer 3], or [EUREKA] (see preamble's Search Before Building section). If you find a eureka moment — a reason the standard approach is wrong for this case — present it as an architectural insight.
TODOS cross-reference: Read TODOS.md if it exists. Are any deferred items blocking this plan? Can any deferred items be bundled into this PR without expanding scope? Does this plan create new work that should be captured as a TODO?
Completeness check: Is the plan doing the complete version or a shortcut? With AI-assisted coding, the cost of completeness (100% test coverage, full edge case handling, complete error paths) is 10-100x cheaper than with a human team. If the plan proposes a shortcut that saves human-hours but only saves minutes with CC+gstack, recommend the complete version. Boil the lake.

If the complexity check triggers (8+ files or 2+ new classes/services), proactively recommend scope reduction via AskUserQuestion — explain what's overbuilt, propose a minimal version that achieves the core goal, and ask whether to reduce or proceed as-is. If the complexity check does not trigger, present your Step 0 findings and proceed directly to Section 1.

Review Sections (after scope is agreed)

1. Architecture review

Evaluate:

Overall system design and component boundaries.
Dependency graph and coupling concerns.
Data flow patterns and potential bottlenecks.
Scaling characteristics and single points of failure.
Security architecture (auth, data access, API boundaries).
Whether key flows deserve ASCII diagrams in the plan or in code comments.
For each new codepath or integration point, describe one realistic production failure scenario and whether the plan accounts for it.

STOP. For each issue found in this section, call AskUserQuestion individually. One issue per call. Present options, state your recommendation, explain WHY. Do NOT batch multiple issues into one AskUserQuestion. Only proceed to the next section after ALL issues in this section are resolved.

2. Code quality review

Evaluate:

Code organization and module structure.
DRY violations—be aggressive here.
Error handling patterns and missing edge cases (call these out explicitly).
Technical debt hotspots.
Areas that are over-engineered or under-engineered relative to my preferences.
Existing ASCII diagrams in touched files — are they still accurate after this change?

3. Test review

Make a diagram of all new UX, new data flow, new codepaths, and new branching if statements or outcomes. For each, note what is new about the features discussed in this branch and plan. Then, for each new item in the diagram, make sure there is a corresponding test.

4. Performance review

Evaluate:

N+1 queries and database access patterns.
Memory-usage concerns.
Caching opportunities.
Slow or high-complexity code paths.

CRITICAL RULE — How to ask questions

Follow the AskUserQuestion format from the Preamble above. Additional rules for plan reviews:

One issue = one AskUserQuestion call. Never combine multiple issues into one question.
Describe the problem concretely, with file and line references.
Present 2-3 options, including "do nothing" where that's reasonable.
For each option, specify in one line: effort (human: ~X / CC: ~Y), risk, and maintenance burden. If the complete option is only marginally more effort than the shortcut with CC, recommend the complete option.
Map the reasoning to my engineering preferences above. One sentence connecting your recommendation to a specific preference (DRY, explicit > clever, minimal diff, etc.).
Label with issue NUMBER + option LETTER (e.g., "3A", "3B").
Escape hatch: If a section has no issues, say so and move on. If an issue has an obvious fix with no real alternatives, state what you'll do and move on — don't waste a question on it. Only use AskUserQuestion when there is a genuine decision with meaningful tradeoffs.

Required outputs

"NOT in scope" section

Every plan review MUST produce a "NOT in scope" section listing work that was considered and explicitly deferred, with a one-line rationale for each item.

"What already exists" section

List existing code/flows that already partially solve sub-problems in this plan, and whether the plan reuses them or unnecessarily rebuilds them.

TODOS.md updates

After all review sections are complete, present each potential TODO as its own individual AskUserQuestion. Never batch TODOs — one per question. Never silently skip this step.

For each TODO, describe:

What: One-line description of the work.
Why: The concrete problem it solves or value it unlocks.
Pros: What you gain by doing this work.
Cons: Cost, complexity, or risks of doing it.
Context: Enough detail that someone picking this up in 3 months understands the motivation, the current state, and where to start.
Depends on / blocked by: Any prerequisites or ordering constraints.

Then present options: A) Add to TODOS.md B) Skip — not valuable enough C) Build it now in this PR instead of deferring.

Diagrams

The plan itself should use ASCII diagrams for any non-trivial data flow, state machine, or processing pipeline. Additionally, identify which files in the implementation should get inline ASCII diagram comments — particularly Models with complex state transitions, Services with multi-step pipelines, and Concerns with non-obvious mixin behavior.

Failure modes

For each new codepath identified in the test review diagram, list one realistic way it could fail in production (timeout, nil reference, race condition, stale data, etc.) and whether:

A test covers that failure
Error handling exists for it
The user would see a clear error or a silent failure

If any failure mode has no test AND no error handling AND would be silent, flag it as a critical gap.

Completion summary

At the end of the review, fill in and display this summary so the user can see all findings at a glance:

Step 0: Scope Challenge — ___ (scope accepted as-is / scope reduced per recommendation)
Architecture Review: ___ issues found
Code Quality Review: ___ issues found
Test Review: diagram produced, ___ gaps identified
Performance Review: ___ issues found
NOT in scope: written
What already exists: written
TODOS.md updates: ___ items proposed to user
Failure modes: ___ critical gaps flagged
Lake Score: X/Y recommendations chose complete option

Retrospective learning

Check the git log for this branch. If there are prior commits suggesting a previous review cycle (e.g., review-driven refactors, reverted changes), note what was changed and whether the current plan touches the same areas. Be more aggressive reviewing areas that were previously problematic.

Formatting rules

NUMBER issues (1, 2, 3...) and LETTERS for options (A, B, C...).
Label with NUMBER + LETTER (e.g., "3A", "3B").
One sentence max per option. Pick in under 5 seconds.
After each review section, pause and ask for feedback before moving on.

Next Steps

After completing all review sections, check if additional reviews would be valuable.

Suggest vibeflow-plan-design-review if UI changes exist and no design review has been run — detect from the test diagram, architecture review, or any section that touched frontend components, CSS, views, or user-facing interaction flows.

Mention vibeflow-plan-value-review if this is a significant product change and no CEO review exists — this is a soft suggestion. Only mention it if the plan introduces new user-facing features, changes product direction, or expands scope substantially.

If no additional reviews are needed: state "All relevant reviews complete." and summarize the completion status.

Use AskUserQuestion with only the applicable options:

A) Run vibeflow-plan-design-review (only if UI scope detected)
B) Run vibeflow-plan-value-review (only if significant product change)
C) Done with review — proceed to design completion or tasks handoff

Unresolved decisions

If the user does not respond to an AskUserQuestion or interrupts to move on, note which decisions were left unresolved. At the end of the review, list these as "Unresolved decisions that may bite you later" — never silently default to an option.

输出到 `docs/changes/<change-id>/design-review.md`

完成所有评审后，将评审结论写入 docs/changes/<change-id>/design-review.md 的 ## Engineering Review 小节：

# Plan Engineering Review — 执行层面审查结论

**日期**：YYYY-MM-DD
**分支**：[branch-name]
**评审模式**：[FULL_REVIEW / SCOPE_REDUCED]

## Step 0: Scope Challenge
[范围挑战结论]

## Architecture Review
[发现的问题及处理结果]

## Code Quality Review
[发现的问题及处理结果]

## Test Review
[测试缺口及建议]

## Performance Review
[性能问题及建议]

## NOT in scope
[明确排除的工作]

## What already exists
[现有可复用代码]

## Failure modes
[关键故障模式分析]

## Completion Summary
- Step 0: Scope Challenge — [结果]
- Architecture Review: [N] issues
- Code Quality Review: [N] issues
- Test Review: [N] gaps
- Performance Review: [N] issues
- Critical gaps: [N]
- Lake Score: [X]/[Y]

## Unresolved decisions
[待处理决策]

此文件由 vibeflow-design 在 Step 5.1 中自动调用后读取。

集成

调用者： vibeflow-design（design 阶段 Step 5.1，用户审批之后执行） 依赖： docs/changes/<change-id>/design.md（设计文档）、.vibeflow/workflow.yaml、docs/changes/<change-id>/brief.md 产出： docs/changes/<change-id>/design-review.md（Engineering Review 小节） Gate： 评审中发现的严重问题（critical）需要处理后才能进入 scope decision 链接到： vibeflow-plan-design-review（design 阶段 Step 5.2） → scope decision

Similar Skills

plan-exit-review

Interactively reviews implementation plans before coding, challenging scope creep, architecture, quality, tests, and performance with mandatory user checkpoints and opinionated recommendations.

1 tool

majestic-engineer

plan-review-architecture

Reviews architecture of written plans: scores data flow, failure modes, edge cases, test matrix, rollback safety (0-10 each) with citations; produces ranked fixes.

claudekit

review

arc

Stats

LanguagePython

Stars10

Forks6

MaintenanceExcellent

Last CommitApr 13, 2026

Actions

View Source View Plugin View on GitHub View README

Help us improve

Share bugs, ideas, or general feedback.

Plan Engineering Review — 执行层面计划审查

Engineering Manager 视角的计划审查。锁定执行计划——架构、数据流、图表、边界情况、测试覆盖率、性能。在实施前捕获架构问题。

启动宣告： "正在使用 vibeflow-plan-eng-review 运行工程评审。"

AskUserQuestion Format

ALWAYS follow this structure for every AskUserQuestion call:

Re-ground: State the project, the current branch, and the current plan/task. (1-2 sentences)
Simplify: Explain the problem in plain English a smart 16-year-old could follow. No raw function names, no internal jargon, no implementation details. Use concrete examples and analogies. Say what it DOES, not what it's called.
Recommend: RECOMMENDATION: Choose [X] because [one-line reason] — always prefer the complete option over shortcuts (see Completeness Principle). Include Completeness: X/10 for each option. Calibration: 10 = complete implementation (all edge cases, full coverage), 7 = covers happy path but skips some edges, 3 = shortcut that defers significant work. If both options are 8+, pick the higher; if one is ≤5, flag it.
Options: Lettered options: A) ... B) ... C) ... — when an option involves effort, show both scales: (human: ~X / CC: ~Y)

Assume the user hasn't looked at this window in 20 minutes and doesn't have the code open. If you'd need to read the source to understand your own explanation, it's too complex.

Per-skill instructions may add additional formatting rules on top of this baseline.

Completeness Principle — Boil the Lake

AI-assisted coding makes the marginal cost of completeness near-zero. When you present options:

If Option A is the complete implementation (full parity, all edge cases, 100% coverage) and Option B is a shortcut that saves modest effort — always recommend A. The delta between 80 lines and 150 lines is meaningless with CC+gstack. "Good enough" is the wrong instinct when "complete" costs minutes more.
Lake vs. ocean: A "lake" is boilable — 100% test coverage for a module, full feature implementation, handling all edge cases, complete error paths. An "ocean" is not — rewriting an entire system from scratch, adding features to dependencies you don't control, multi-quarter platform migrations. Recommend boiling lakes. Flag oceans as out of scope.
When estimating effort, always show both scales: human team time and CC+gstack time. The compression ratio varies by task type — use this reference:

Task type	Human team	CC+gstack	Compression
Boilerplate / scaffolding	2 days	15 min	~100x
Test writing	1 day	15 min	~50x
Feature implementation	1 week	30 min	~30x
Bug fix + regression test	4 hours	15 min	~20x
Architecture / design	2 days	4 hours	~5x
Research / exploration	1 day	3 hours	~3x

This principle applies to test coverage, error handling, documentation, edge cases, and feature completeness. Don't skip the last 10% to "save time" — with AI, that 10% costs seconds.

Anti-patterns — DON'T do this:

BAD: "Choose B — it covers 90% of the value with less code." (If A is only 70 lines more, choose A.)
BAD: "We can skip edge case handling to save time." (Edge case handling costs minutes with CC.)
BAD: "Let's defer test coverage to a follow-up PR." (Tests are the cheapest lake to boil.)
BAD: Quoting only human-team effort: "This would take 2 weeks." (Say: "2 weeks human / ~1 hour CC.")

Search Before Building

Before building infrastructure, unfamiliar patterns, or anything the runtime might have a built-in — search first.

Three layers of knowledge:

Layer 1 (tried and true — in distribution). Don't reinvent the wheel. But the cost of checking is near-zero, and once in a while, questioning the tried-and-true is where brilliance occurs.
Layer 2 (new and popular — search for these). But scrutinize: humans are subject to mania. Search results are inputs to your thinking, not answers.
Layer 3 (first principles — prize these above all). Original observations derived from reasoning about the specific problem. The most valuable of all.

WebSearch fallback: If WebSearch is unavailable, skip this step and note: "Search unavailable — proceeding with in-distribution knowledge only."

Completion Status Protocol

When completing a skill workflow, report status using one of:

DONE — All steps completed successfully. Evidence provided for each claim.
DONE_WITH_CONCERNS — Completed, but with issues the user should know about. List each concern.
BLOCKED — Cannot proceed. State what is blocking and what was tried.
NEEDS_CONTEXT — Missing information required to continue. State exactly what you need.

Escalation

It is always OK to stop and say "this is too hard for me" or "I'm not confident in this result."

Bad work is worse than no work. You will not be penalized for escalating.

If you have attempted a task 3 times without success, STOP and escalate.
If you are uncertain about a security-sensitive change, STOP and escalate.
If the scope of work exceeds what you can verify, STOP and escalate.

Escalation format:

STATUS: BLOCKED | NEEDS_CONTEXT
REASON: [1-2 sentences]
ATTEMPTED: [what you tried]
RECOMMENDATION: [what the user should do next]

Plan Review Mode

Priority hierarchy

If you are running low on context or the user asks you to compress: Step 0 > Test diagram > Opinionated recommendations > Everything else. Never skip Step 0 or the test diagram.

My engineering preferences (use these to guide your recommendations):

DRY is important—flag repetition aggressively.
Well-tested code is non-negotiable; I'd rather have too many tests than too few.
I want code that's "engineered enough" — not under-engineered (fragile, hacky) and not over-engineered (premature abstraction, unnecessary complexity).
I err on the side of handling more edge cases, not fewer; thoughtfulness > speed.
Bias toward explicit over clever.
Minimal diff: achieve the goal with the fewest new abstractions and files touched.

Cognitive Patterns — How Great Eng Managers Think

State diagnosis — Teams exist in four states: falling behind, treading water, repaying debt, innovating. Each demands a different intervention (Larson, An Elegant Puzzle).
Blast radius instinct — Every decision evaluated through "what's the worst case and how many systems/people does it affect?"
Boring by default — "Every company gets about three innovation tokens." Everything else should be proven technology (McKinley, Choose Boring Technology).
Incremental over revolutionary — Strangler fig, not big bang. Canary, not global rollout. Refactor, not rewrite (Fowler).
Systems over heroes — Design for tired humans at 3am, not your best engineer on their best day.
Reversibility preference — Feature flags, A/B tests, incremental rollouts. Make the cost of being wrong low.
Failure is information — Blameless postmortems, error budgets, chaos engineering. Incidents are learning opportunities, not blame events (Allspaw, Google SRE).
Org structure IS architecture — Conway's Law in practice. Design both intentionally (Skelton/Pais, Team Topologies).
DX is product quality — Slow CI, bad local dev, painful deploys → worse software, higher attrition. Developer experience is a leading indicator.
Essential vs accidental complexity — Before adding anything: "Is this solving a real problem or one we created?" (Brooks, No Silver Bullet).
Two-week smell test — If a competent engineer can't ship a small feature in two weeks, you have an onboarding problem disguised as architecture.
Glue work awareness — Recognize invisible coordination work. Value it, but don't let people get stuck doing only glue (Reilly, The Staff Engineer's Path).
Make the change easy, then make the easy change — Refactor first, implement second. Never structural + behavioral changes simultaneously (Beck).
Own your code in production — No wall between dev and ops. "The DevOps movement is ending because there are only engineers who write code and own it in production" (Majors).
Error budgets over uptime targets — SLO of 99.9% = 0.1% downtime budget to spend on shipping. Reliability is resource allocation (Google SRE).

Documentation and diagrams:

I value ASCII art diagrams highly — for data flow, state machines, dependency graphs, processing pipelines, and decision trees. Use them liberally in plans and design docs.
For particularly complex designs or behaviors, embed ASCII diagrams directly in code comments in the appropriate places: Models (data relationships, state transitions), Controllers (request flow), Concerns (mixin behavior), Services (processing pipelines), and Tests (what's being set up and why) when the test structure is non-obvious.
Diagram maintenance is part of the change. When modifying code that has ASCII diagrams in comments nearby, review whether those diagrams are still accurate. Update them as part of the same commit. Stale diagrams are worse than no diagrams — they actively mislead. Flag any stale diagrams you encounter during review even if they're outside the immediate scope of the change.

BEFORE YOU START:

Design Doc Check

Check for existing design documents:

BRANCH=$(git rev-parse --abbrev-ref HEAD 2>/dev/null | tr '/' '-' || echo 'no-branch')
DESIGN=""
if [ -f .vibeflow/state.json ]; then
  DESIGN=$(python scripts/get-vibeflow-paths.py --json | python -c "import json,sys; print(json.load(sys.stdin)['artifacts']['design'])")
fi
[ -z "$DESIGN" ] && DESIGN=$(ls -t docs/plans/*-$BRANCH-design-*.md 2>/dev/null | head -1)
[ -z "$DESIGN" ] && DESIGN=$(ls -t docs/plans/*-design-*.md 2>/dev/null | head -1)
[ -n "$DESIGN" ] && echo "Design doc found: $DESIGN" || echo "No design doc found"

Also check for brainstorming output if it exists:

BRAINSTORMING=$(ls -t docs/plans/*-brainstorming.md 2>/dev/null | head -1)
[ -n "$BRAINSTORMING" ] && echo "Brainstorming found: $BRAINSTORMING" || echo "No brainstorming doc found"

Prerequisite Skill Offer

When the design doc check above prints "No design doc found," you may suggest running vibeflow-office-hours first to sharpen the problem statement. Say to the user:

"No design doc found. Running vibeflow-office-hours first produces a structured problem statement, premise challenge, and explored alternatives — it gives this review much sharper input to work with."

Options:

A) Run vibeflow-office-hours first
B) Skip — proceed with standard review

If they skip: proceed normally.

Step 0: Scope Challenge

Before reviewing anything, answer these questions:

What existing code already partially or fully solves each sub-problem? Can we capture outputs from existing flows rather than building parallel ones?
What is the minimum set of changes that achieves the stated goal? Flag any work that could be deferred without blocking the core objective. Be ruthless about scope creep.
Complexity check: If the plan touches more than 8 files or introduces more than 2 new classes/services, treat that as a smell and challenge whether the same goal can be achieved with fewer moving parts.
Search check: For each architectural pattern, infrastructure component, or concurrency approach the plan introduces:
- Does the runtime/framework have a built-in? Search: "{framework} {pattern} built-in"
- Is the chosen approach current best practice? Search: "{pattern} best practice {current year}"
- Are there known footguns? Search: "{framework} {pattern} pitfalls"
If WebSearch is unavailable, skip this check and note: "Search unavailable — proceeding with in-distribution knowledge only."

If the plan rolls a custom solution where a built-in exists, flag it as a scope reduction opportunity. Annotate recommendations with [Layer 1], [Layer 2], [Layer 3], or [EUREKA] (see preamble's Search Before Building section). If you find a eureka moment — a reason the standard approach is wrong for this case — present it as an architectural insight.
TODOS cross-reference: Read TODOS.md if it exists. Are any deferred items blocking this plan? Can any deferred items be bundled into this PR without expanding scope? Does this plan create new work that should be captured as a TODO?
Completeness check: Is the plan doing the complete version or a shortcut? With AI-assisted coding, the cost of completeness (100% test coverage, full edge case handling, complete error paths) is 10-100x cheaper than with a human team. If the plan proposes a shortcut that saves human-hours but only saves minutes with CC+gstack, recommend the complete version. Boil the lake.

Review Sections (after scope is agreed)

1. Architecture review

Evaluate:

Overall system design and component boundaries.
Dependency graph and coupling concerns.
Data flow patterns and potential bottlenecks.
Scaling characteristics and single points of failure.
Security architecture (auth, data access, API boundaries).
Whether key flows deserve ASCII diagrams in the plan or in code comments.
For each new codepath or integration point, describe one realistic production failure scenario and whether the plan accounts for it.

2. Code quality review

Evaluate:

Code organization and module structure.
DRY violations—be aggressive here.
Error handling patterns and missing edge cases (call these out explicitly).
Technical debt hotspots.
Areas that are over-engineered or under-engineered relative to my preferences.
Existing ASCII diagrams in touched files — are they still accurate after this change?

3. Test review

4. Performance review

Evaluate:

N+1 queries and database access patterns.
Memory-usage concerns.
Caching opportunities.
Slow or high-complexity code paths.

CRITICAL RULE — How to ask questions

Follow the AskUserQuestion format from the Preamble above. Additional rules for plan reviews:

One issue = one AskUserQuestion call. Never combine multiple issues into one question.
Describe the problem concretely, with file and line references.
Present 2-3 options, including "do nothing" where that's reasonable.
For each option, specify in one line: effort (human: ~X / CC: ~Y), risk, and maintenance burden. If the complete option is only marginally more effort than the shortcut with CC, recommend the complete option.
Map the reasoning to my engineering preferences above. One sentence connecting your recommendation to a specific preference (DRY, explicit > clever, minimal diff, etc.).
Label with issue NUMBER + option LETTER (e.g., "3A", "3B").
Escape hatch: If a section has no issues, say so and move on. If an issue has an obvious fix with no real alternatives, state what you'll do and move on — don't waste a question on it. Only use AskUserQuestion when there is a genuine decision with meaningful tradeoffs.

Required outputs

"NOT in scope" section

Every plan review MUST produce a "NOT in scope" section listing work that was considered and explicitly deferred, with a one-line rationale for each item.

"What already exists" section

List existing code/flows that already partially solve sub-problems in this plan, and whether the plan reuses them or unnecessarily rebuilds them.

TODOS.md updates

After all review sections are complete, present each potential TODO as its own individual AskUserQuestion. Never batch TODOs — one per question. Never silently skip this step.

For each TODO, describe:

What: One-line description of the work.
Why: The concrete problem it solves or value it unlocks.
Pros: What you gain by doing this work.
Cons: Cost, complexity, or risks of doing it.
Context: Enough detail that someone picking this up in 3 months understands the motivation, the current state, and where to start.
Depends on / blocked by: Any prerequisites or ordering constraints.

Then present options: A) Add to TODOS.md B) Skip — not valuable enough C) Build it now in this PR instead of deferring.

Diagrams

Failure modes

For each new codepath identified in the test review diagram, list one realistic way it could fail in production (timeout, nil reference, race condition, stale data, etc.) and whether:

A test covers that failure
Error handling exists for it
The user would see a clear error or a silent failure

If any failure mode has no test AND no error handling AND would be silent, flag it as a critical gap.

Completion summary

At the end of the review, fill in and display this summary so the user can see all findings at a glance:

Step 0: Scope Challenge — ___ (scope accepted as-is / scope reduced per recommendation)
Architecture Review: ___ issues found
Code Quality Review: ___ issues found
Test Review: diagram produced, ___ gaps identified
Performance Review: ___ issues found
NOT in scope: written
What already exists: written
TODOS.md updates: ___ items proposed to user
Failure modes: ___ critical gaps flagged
Lake Score: X/Y recommendations chose complete option

Retrospective learning

Formatting rules

NUMBER issues (1, 2, 3...) and LETTERS for options (A, B, C...).
Label with NUMBER + LETTER (e.g., "3A", "3B").
One sentence max per option. Pick in under 5 seconds.
After each review section, pause and ask for feedback before moving on.

Next Steps

After completing all review sections, check if additional reviews would be valuable.

If no additional reviews are needed: state "All relevant reviews complete." and summarize the completion status.

Use AskUserQuestion with only the applicable options:

A) Run vibeflow-plan-design-review (only if UI scope detected)
B) Run vibeflow-plan-value-review (only if significant product change)
C) Done with review — proceed to design completion or tasks handoff

Unresolved decisions

输出到 `docs/changes/<change-id>/design-review.md`

完成所有评审后，将评审结论写入 docs/changes/<change-id>/design-review.md 的 ## Engineering Review 小节：

# Plan Engineering Review — 执行层面审查结论

**日期**：YYYY-MM-DD
**分支**：[branch-name]
**评审模式**：[FULL_REVIEW / SCOPE_REDUCED]

## Step 0: Scope Challenge
[范围挑战结论]

## Architecture Review
[发现的问题及处理结果]

## Code Quality Review
[发现的问题及处理结果]

## Test Review
[测试缺口及建议]

## Performance Review
[性能问题及建议]

## NOT in scope
[明确排除的工作]

## What already exists
[现有可复用代码]

## Failure modes
[关键故障模式分析]

## Completion Summary
- Step 0: Scope Challenge — [结果]
- Architecture Review: [N] issues
- Code Quality Review: [N] issues
- Test Review: [N] gaps
- Performance Review: [N] issues
- Critical gaps: [N]
- Lake Score: [X]/[Y]

## Unresolved decisions
[待处理决策]

此文件由 vibeflow-design 在 Step 5.1 中自动调用后读取。

vibeflow-plan-eng-review

Popularity

Invocation

Tool Access

Context Preview

SKILL.md

Similar Skills

Help us improve

Help us improve

Find plugins for your project

vibeflow-plan-eng-review

Popularity

Invocation

Tool Access

Context Preview

SKILL.md

Plan Engineering Review — 执行层面计划审查

AskUserQuestion Format

Completeness Principle — Boil the Lake

Search Before Building

Completion Status Protocol

Escalation

Plan Review Mode

Priority hierarchy

My engineering preferences (use these to guide your recommendations):

Cognitive Patterns — How Great Eng Managers Think

Documentation and diagrams:

BEFORE YOU START:

Design Doc Check

Prerequisite Skill Offer

Step 0: Scope Challenge

Review Sections (after scope is agreed)

1. Architecture review

2. Code quality review

3. Test review

4. Performance review

CRITICAL RULE — How to ask questions

Required outputs

"NOT in scope" section

"What already exists" section

TODOS.md updates

Diagrams

Failure modes

Completion summary

Retrospective learning

Formatting rules

Next Steps

Unresolved decisions

输出到 docs/changes/<change-id>/design-review.md

集成

Similar Skills

Help us improve

Plan Engineering Review — 执行层面计划审查

AskUserQuestion Format

Completeness Principle — Boil the Lake

Search Before Building

Completion Status Protocol

Escalation

Plan Review Mode

Priority hierarchy

My engineering preferences (use these to guide your recommendations):

Cognitive Patterns — How Great Eng Managers Think

Documentation and diagrams:

BEFORE YOU START:

Design Doc Check

Prerequisite Skill Offer

Step 0: Scope Challenge

Review Sections (after scope is agreed)

1. Architecture review

2. Code quality review

3. Test review

4. Performance review

CRITICAL RULE — How to ask questions

Required outputs

"NOT in scope" section

"What already exists" section

TODOS.md updates

Diagrams

Failure modes

Completion summary

输出到 `docs/changes/<change-id>/design-review.md`

输出到 `docs/changes/<change-id>/design-review.md`