structured-agentic-workflow

agent
Guvenlik Denetimi
Basarisiz
Health Uyari
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Basarisiz
  • rm -rf — Recursive force deletion command in pull-superpowers.sh
  • network request — Outbound network request in pull-superpowers.sh
Permissions Gecti
  • Permissions — No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

Set of skills and workflows to ensure humans and agents stay focused while vibe-coding.

README.md

The Structured Agentic Development Workflow

Agent skills for intent-audited development: you decide what to build, a frontier model
plans and supervises, any model builds, and nothing is done until the work is checked against what
you meant.

Every significant change follows one cycle:
Brainstorm → Plan → Build → 3rd-Person Review → Verify.
Phases are never skipped, plans are files rather than conversations, and nothing is "done"
without evidence.

Intent-audited development

Spec-driven development builds to a spec. Intent-driven development writes the intent down.
Intent-audited development makes your intent the baseline and checks the work against it
before anything ships: not "does the code match the spec?" but "is this what you meant, and can
the agent prove it?"

  • Your intent is the baseline. What you chose, what you ruled out and what the user will see
    are recorded before anything is planned. Once the plan is approved, only you can change them.
  • A frontier model supervises. It designs, writes a plan that leaves nothing to guess, and
    reviews through seven lenses, from impact to business sense. Another vendor's model can review
    too, from a ready-made brief.
  • Any model or tool builds. Claude, Codex, Cursor, Gemini, across vendors, at a fraction of the
    cost. A gap in the plan makes it stop and report, never improvise.
  • Drift blocks "done". Code is audited against the plan, and the plan against your intent.
    A difference nobody agreed to stops the completion claim.

Highlights

  • Redundant full-suite runs retired. A four-rung test ladder, scoped per gate, instead of the same suite six or seven times per feature.
  • A rule change is a table, not a paragraph. Every case the rule can meet, today's outcome beside each proposal's; the rows that differ are the change, and they become the plan's test cases.
  • One set of review lenses, from first idea to finished code. Impact, removal, logic, behaviour, business sense, proof and coherence, one question each, on the approaches, the plan and the built code.
  • A review that owns the code. /3p-review loops to zero findings at every severity, minors included.
  • No mock on the value path. A path with no test across the real seam cannot pass review.
  • Nothing finishes smaller than it started. Every removal is audited with its replacement named.
  • Evidence before claims. No "done" without a run you made yourself this session.
  • Works with any agent. Claude Code, Cursor, Gemini CLI, Copilot, and 40+ others.

Each of these is argued out in what makes this different, further down.


Setup

1. Install graphify (recommended)

Powers codebase search in /brainstorm and /write-plan. The workflow runs without it —
both skills fall back to grep after saying so once — but this is where a lot of the quality
comes from.

You should follow the setup instructions as per the graphify repo, but for a quick reference we are including them here:

uv tool install graphifyy   # the official PyPI package, per graphify's own README
graphify install            # registers the /graphify skill with your agent

The double-y is graphify's own naming, not a typo: its README says "The PyPI package is
graphifyy (double-y). Other graphify* packages on PyPI are not affiliated."
Check there
before installing if in doubt.

If the graphify command isn't found afterwards, run uv tool update-shell. pipx install graphifyy and pip install graphifyy also work.

Add graphify-out/ to your project's .gitignore; it is a build artifact.

2. Install the workflow skills

# Works with Claude Code, Cursor, Gemini CLI, Copilot, and 40+ other agents
npx skills add nikhilw/structured-agentic-workflow

# superpowers supplies TDD and systematic debugging — strongly recommended
npx skills add obra/superpowers -s test-driven-development -s systematic-debugging

Prefer one command that pulls everything? Clone the repo and run ./install.sh
(.\install.ps1 on Windows). Per-agent targets, manual steps, and the full skill inventory
are in installation.md.

Upgrading from an earlier install? Re-run it. Two skills are new (verify-completion,
existing-mechanisms), and superpowers' verification-before-completion is no longer part of
this workflow: /verify-completion replaces it. ./install.sh removes the link the old install
created, leaving any copy you installed another way alone.

3. Point your project's agent config at the workflow

Add the workflow skills to CLAUDE.md / AGENTS.md (or .cursorrules, GEMINI.md,
.github/copilot-instructions.md) so the agent starts every conversation knowing the
workflow exists and can suggest phase transitions itself:

## Workflow Skills

ALWAYS follow the Structured Agentic Development Workflow. These skills are installed
globally and define the development lifecycle:

- `agentic-workflow` — orchestrates the full lifecycle; suggests phase transitions automatically
- `/brainstorm` — explore the problem space before planning (no code, no plans)
- `/decision-summary` : what has been decided so far, in plain words; ask any time, changes nothing
- `/write-plan` — write phased plans to `docs/plans/new/` (agent-decoupled)
- `/build-phase` — execute one plan phase: test-first → implement → scoped tests → self-review
- `/build-model` — dedicated build-model session: build → 3p-review → handoff-summary → stop
- `/3p-review` — independent code review; the reviewer owns the code
- `/handoff-summary` — emit the fixed-format Build Handoff Summary
- `/verify-completion` — the final gate: fresh suite, requirements tick-off, plan-drift audit
- `test-scope` : how wide each test run must be, what a run may execute, and when a run can be cited instead of re-run
- `existing-mechanisms` : the eight questions about what the codebase already does
- `review-lenses` : the perspectives every review looks through, which gate runs which, and the outside-review brief
- `/triage` — recommend the next task, minimizing context thrash

Startup default: load `agentic-workflow` at startup.

A fuller template — standing quality bar, architecture facts, hard rules, and what to keep
out of the file — is in agent-config.md.


Use

/brainstorm add offline support to the sync layer   # explore, challenge, decide
/write-plan offline-sync                            # phased plan → docs/plans/new/
                                                    # you review it, then: mv to docs/plans/
/build-phase docs/plans/offline-sync.md Phase 1     # TDD → test → self-review, per phase
/3p-review                                          # holistic review, loops until clean
/verify-completion                                  # fresh evidence + drift audit, then archive

agentic-workflow drives these transitions for you — you rarely type the middle three. Ask
/triage when you are not sure what to pick up next.

To hand the build to a cheaper model (Sonnet is the usual choice, or another tool
entirely): approve the plan, mv it to docs/plans/, then in that session run
/build-model docs/plans/offline-sync.md. It reviews the plan first and halts if it finds a
defect, then builds every phase, reviews its own work, emits a handoff summary, and stops.
Bring that summary back to your main model, which re-reviews with fresh eyes and verifies.
Tier guidance is in multi-model.md.

The skills

Skill Does
agentic-workflow Orchestrates the lifecycle and drives phase transitions
/brainstorm Explores approaches, challenges the design, writes a decision document
/decision-summary Summarises what has been decided, in plain words: what you will get and how it works. A capability, not a phase
/write-plan Writes a phased, fully-decided plan to docs/plans/new/
/build-phase Executes one phase: test-first → implement → scoped tests → self-review
/build-model Entry point for a dedicated build model: build → review → handoff → stop
/3p-review Independent review that owns the code; loops until clean
/handoff-summary Emits the fixed-format Build Handoff Summary
/verify-completion The final gate: fresh suite, requirements tick-off, decision-to-code drift audit
test-scope The shared test-run ladder, the one boundary on what a run may execute, and the citable-run rule
existing-mechanisms The eight questions about what already exists, shared by brainstorm, plan, build and review
review-lenses The review perspectives (impact, removal, logic, behaviour, business sense, proof, coherence), which gate runs each, and the brief for an outside reviewer
/triage Recommends the next task, minimizing context thrash
/github-backlog Maintains features and bugs as GitHub issues
/workflow-config Sets TDD/BDD, output brevity, and backlog source

Plus test-driven-development and systematic-debugging from
superpowers. /verify-completion replaces superpowers'
verification-before-completion, which this workflow no longer installs: it keeps that skill's Iron
Law and adds the requirements tick-off and the drift audit.


Documentation

Document What's in it
workflow.md The full lifecycle, phase by phase, with the complete diagram
multi-model.md The planning/build/review model split and the two contracts
agent-config.md What to put in CLAUDE.md / AGENTS.md
installation.md Per-agent targets, script options, manual install, skill inventory
configuration.md /workflow-config — TDD/BDD, caveman brevity, GitHub issues
practices.md Task selection, refactoring monoliths, the "no surprises" rule
philosophy.md Why the workflow is shaped this way
comparison.md Side by side with Spec Kit, Kiro, BMAD, Superpowers, gstack, Compound Engineering and Traycer
extras/driving-cursor-as-build-model.md Worked recipe: running Cursor's CLI agent headless as the build model
extras/driving-agy-as-build-model.md The same for Google's Antigravity CLI (agy): what differs from Cursor
extras/driving-agy-as-plan-and-code-reviewer.md agy as the outside reviewer: one conversation, a plan review then a code review
extras/agent-cli-setup.md One-time setup for cursor-agent and agy: skills, one allowlist in two formats, two probes

What makes this different

There are other agent-skill libraries — obra/superpowers
is the best known, and this workflow composes with it rather than competing. Eight things set
this one apart (side by side with the alternatives):

1 · The build model doesn't have to be the planning model.
Because the plan is a file that resolves every decision, you can plan with Opus and build
with Sonnet — or hand the plan to Cursor, Gemini Flash, Copilot, or a local model entirely. Frontier reasoning is the scarcest resource in agentic development, and most of
the tokens a coding agent burns are not reasoning at all — they are reading files, writing
boilerplate, and re-running tests. This workflow is built so you stop paying frontier prices
for typing. → multi-model.md

Concretely: plan with Opus, then let Cursor's agent build the plan headless in the
background. The whole build comes back to your expensive model as a few hundred bytes —
a result line, not a transcript. Here's the exact
recipe.

2 · A review gate that takes ownership.
/3p-review switches persona to an independent Senior Architect who owns the code on
sign-off
, loops until zero findings of any severity (minors get fixed, not waved), and
cannot pass without a no-mock test across the real integration seam. Past a volume threshold
it refuses to fix things itself and emits a Rework Brief instead — a reviewer who rewrites
half the feature has become its author.

3 · An enforced lifecycle, not a toolbox.
One orchestrator drives the whole cycle with an explicit phase-transition table and a plan
lifecycle (new/ → plans/ → done/). The agent owns forward motion instead of stalling
between phases, and each phase has an exit gate that user approval cannot retroactively
satisfy.

4 · The plan carries the foresight.
A small build model fills every silence with the happy path. So /write-plan forces the
expensive model to write the foresight down: failure modes, lifetimes, error codes and
their owners, concurrency and aliasing, named seam tests per value path, and exact
command-plus-expected-output test criteria. Every name in the plan must be verified to exist,
or marked new. Then nine review passes before it is saved, one question each (see 7), and the
first starts from the codebase rather than the document: a caller that breaks the build is never in
the plan, so checking the plan's own names will not find it.

5 · Index-first codebase search, with the questions that go with it.
/brainstorm and /write-plan build and query a graphify
knowledge graph of the repo before proposing anything, and count what a change reaches with the
type checker. The graph links symbols by name, so it finds where to look but misses a call made
through an instance and merges two symbols that share a name; your language's type checker or language server (tsc,
pyright, gopls, rust-analyzer and the like) resolves the types and gives the real count. Graph for where to look, types for the count, grep for names hidden
in strings. The most expensive mistake in a
brainstorm is reimplementing something that already exists under a name nobody grepped for, so
existing-mechanisms makes the search into eight required answers: every caller, every related
flow, what already does this job, whether you are extending or replacing it, what becomes dead code
if you do, and whether you are quietly adding a second pathway beside the first.

It also holds the impact trace, which both phases run and which has three axes rather than the
one everybody runs. Structural is the call graph. Functional is the end-to-end flows, because a
change can leave every call site compiling and still break the product. Consolidation asks what the
codebase looks like afterwards: what is left unused, what got abandoned without anyone deciding to,
whether this unifies two pathways or adds a third, and what should be extracted and reused. The
structural axis feels like a finished answer, which is exactly why it is usually the only one run.

6 · Drift is measured, not hoped for.
Plans change during a build, and a feature that ends up somewhere other than where it was aimed is
usually the sum of a dozen reasonable corrections. So corrections are made in one place and written
down: the build model halts and reports rather than working around a gap, the planning model amends
the plan and appends to its Amendment Log, and /verify-completion reads the decision document
against the plan and the plan against the code before anything is called done. Undocumented drift
blocks the completion claim.

7 · Every review looks through the same lenses, and one at a time.
A review that carries several questions at once answers the easiest and reports clean on the rest.
So review-lenses gives each question its own pass: impact (what else this reaches),
removal (is every deletion justified and its job still done), logic (step it through real
cases, crashes and concurrent runs), behaviour (what the user actually sees, on every surface),
business sense (would a typical user who never heard the reasoning find it odd), proof (would
the tests fail if it were built wrong) and coherence (does the document still agree with
itself). /brainstorm looks through them at every approach, where a lens can still rule one out;
/write-plan runs them as its passes; /3p-review runs them on the built code. What the user will
see is written into the decision document and checked at every later gate, and each gate can hand a
filled-in brief to another model, because the reader who never heard the reasoning is the one who
spots what makes no business sense.

8 · A rule change is a table, not a paragraph.
Some changes add a thing; others change when something happens: a trigger, a filter, a default, a
threshold. Prose describes the cases someone thought of, and the case nobody thought of is the one
that ships. So /brainstorm enumerates the cases from the state the rule reads, puts today's outcome
beside each approach's, and labels every row that differs as the fix or as collateral nobody asked
for. Two proposals that sound equally targeted turn out to move three rows and nine, and the
differing rows become the plan's test cases.

flowchart LR
    Start([Task]) --> P

    subgraph P ["1 · Planning model — frontier, high reasoning"]
        direction TB
        P1["/brainstorm"] --> P2["/write-plan"]
    end

    P -->|"plan file<br/>the forward contract"| B

    subgraph B ["2 · Build model — cheap and fast, or another tool"]
        direction TB
        B1["/build-model"] --> B2["/build-phase × N"]
        B2 --> B3["/3p-review"]
        B3 --> B4["/handoff-summary"]
    end

    B -->|"handoff summary<br/>the return contract"| R

    subgraph R ["3 · Review model — fresh eyes, planning model by default"]
        direction TB
        R1["/3p-review"] --> R2["/verify-completion<br/>suite · requirements · drift"]
    end

    R --> Done([Feature complete])

Both build lanes are optional: the same model can carry the whole cycle. The split is there
when you want it.


Worth pairing with

None of these are part of the workflow and none are required. They solve problems this
workflow runs into once the tasks get long or the agents get plural.

A shared memory between agents. Agents in separate sessions cannot see each other's
context, so the same file gets re-read and the same decision gets re-made in parallel. A
shared memory store, agentic-tools or anything equivalent, gives them one place to write
findings and read someone else's. The plan file already does this for the build lane; a
memory store extends it to work that is not phase-shaped.

Another tool as the build model. Anything that can read a plan file and edit a repo can
take the build lane, which is the point of making the plan a file rather than a conversation.
Driving Cursor's agent headless in the background is written up end to end, including the
launch prompt, the failure modes and the lines that prevent them, in
driving-cursor-as-build-model.md, and the same for
Google's Antigravity CLI in driving-agy-as-build-model.md.

Long-lived subagents for long-running work. A one-off subagent starts cold every time: it
re-establishes context, re-explores, reports, and throws all of it away. Across a long task
you pay that setup on every call. A subagent that holds its conversation instead answers the
second question knowing what it learned on the first, which is the difference between
delegating a question and delegating a job.


License

MIT. Superpowers skills are vendored under their own MIT license — see
vendor/superpowers/LICENSE.

Yorumlar (0)

Sonuc bulunamadi