groundwork
Health Warn
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 5 GitHub stars
Code Pass
- Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
A tool-agnostic workflow for building software with AI agents: spec → plan → test-first build loop → human-approved ship.
Groundwork
Spec, proof, approval. A workflow your AI agent can't skip.
Your agent writes the code. Groundwork is the project layer: it owns the spec, the cards, the
evidence and the approvals, so any session can pick up where the last one stopped.

cd your-project
npx groundwork-ai init
Needs Node.js 22+ and a git repo. Then open your AI tool and type /gw.
New to Groundwork? Read the 5-minute quickstart.
The problem
AI agents are fast and confident. Real projects need more: the chat fills up and the plan goes
with it, "done" is whatever the model says, and the rules you agreed yesterday are gone by
tomorrow. A week in, nobody can say what shipped, what's next, or why a choice was made.
Before and after
Without Groundwork
You: Add voting to my book club app.
Agent: Done! A clean voting page with results and a summary.
You: ... (Is it tested? What was I doing yesterday? Which parts moved?)
With Groundwork
You: /gw I want a page where my book club votes on next month's book
Agent: A few specific questions: who can vote, one vote or rank several, hide results?
You: Anyone with the link. One vote. Hide results until it closes.
Agent: Spec saved. Card 1.1: tests written, code passing, review passed.
What changed: the home page lists this month's books.
How to check: run npm run dev. Caveats: no styling yet.
You: /gw-approve
Agent: Committed [1.1] Book list page. Next up: card 1.2 Voting.
The project survives the session, so the next session starts where this one stopped.
How it works
Three steps, one card at a time. Your main session is the runner: it hands each card to three
roles and stops at your door.
- Spec and plan. The agent asks specific questions, suggests the smallest useful version
first, then writes phases and small cards. Stack choices are yours: options with trade-offs,
and your pick is recorded. - Build one card. The tester writes failing tests, the implementer makes them pass, and a
reviewer with fresh context re-runs everything and checks every criterion. The proof is saved
next to the card, with steps you can follow to check it yourself. - You approve. Nothing is committed until you do, card by card or once per phase. When the
same mistake happens twice it becomes a rule; if it still happens, a guard blocks it.
flowchart TD
idea([Your idea]) --> spec[Spec<br/>a few specific questions]
spec --> plan[Plan<br/>phases and small cards]
plan --> card
subgraph card [For each card]
direction LR
tester[Tester<br/>writes failing tests] --> impl[Implementer<br/>makes them pass]
impl --> review[Reviewer<br/>checks the evidence]
review -. sent back .-> impl
end
card --> approve{You approve?}
approve -- yes --> commit([Committed, on to the next card])
approve -. no, with a reason .-> card
classDef you fill:#fef3c7,stroke:#d97706,color:#451a03
classDef agent fill:#e0e7ff,stroke:#6366f1,color:#1e1b4b
classDef done fill:#dcfce7,stroke:#16a34a,color:#052e16
class idea,approve you
class spec,plan,tester,impl,review agent
class commit done
style card fill:none,stroke:#94a3b8,stroke-dasharray:4 3
Read the deep version: the build loop and evidence and approval.
What's new in 0.7.0
- A lighter runner. Each role gets a short hand-off, checks run once instead of between roles, and
/gw-approvesuggests a fresh session when a phase closes. - Roles keep their own records. The tester and implementer write their own History and
call:lines on the card, and anything missing from the files counts as not agreed. doctorreports startup cost. It shows what/gwand each role read before any code, flags long cards, and finds evidence saved as UTF-16.- OpenCode sessions start briefed. On OpenCode 2.x, every new session gets the handoff lines Claude Code shows at start.
- Safer defaults. New projects start with the
no-ai-trailersguard on, and evidence images are marked binary so git can't corrupt screenshots. gw-ui-specnames what tests look for (an id, label or role), so testers don't guess selectors.
Full notes: the v0.7.0 release.
What it costs
Every session loads about 480 tokens of Groundwork files, and npx groundwork-ai doctor shows
your project's number. A card runs three fresh role sessions of about 9k tokens each, at $0.04 to
$0.13 per card on the calculator build. On a small app in one sitting, a plain chat is probably
cheaper; once a project outlives one context window the flat cost per card wins, and resuming the
finished calculator took 6.8x fewer tokens at $0.009, against $0.023 without Groundwork.
What it costs, and when it pays off.
Proof
Every number here comes from one recorded build. The calculator walkthrough has the sessions, transcripts and metrics, and the proof page has the charts.
Who it's for
- Experienced developers. A real review step before anything merges, evidence saved in the repo, and guards for rules you are done repeating.
- Early career, no senior around. Work one small card at a time, with a reviewer that catches what you would miss and proof you can point to.
- Vibecoders. Answer a few questions, approve in plain English, and ship; small changes take the quick path and features get structure.
Honest comparison
| Groundwork | A raw agent session | |
|---|---|---|
| Project state | Files: spec, cards, handoff, decisions | The chat scrollback |
| "Done" | Evidence attached; approval refuses without it | Whatever the model summarizes |
| Review | A separate reviewer with fresh context, can't edit code | The context that wrote it |
| Learning | Repeated mistakes become rules, then guards | Starts fresh every session |
| Resume | Any session, tool or model continues from the handoff | Re-explain everything |
When not to use it: one-off scripts and throwaway experiments (ask your agent). Fully
autonomous overnight runs: Groundwork stops for you by design. Teams and pull-request flows: not
built yet. There are no benchmarks or evals; the claim is the mechanism, not a score.
Quickstart
The 5-minute quickstart goes from install to your first approval. These pages cover the rest:

- Your first 10 minutes walks through a full session.
- Existing projects covers installing into a repo that already has code.
- Command reference lists every
/gw-*command; the CLI reference coversgroundwork-ai. - The build loop and evidence and approval explain the mechanics.
- Adapters for Claude Code, OpenCode and other tools.
- What it costs explains where the tokens go; the FAQ answers the rest.
- Upgrading.
npx groundwork-ai upgraderefreshes Groundwork's files and keeps your spec, cards and lessons. How updating works.
FAQ
- Does it work with an existing project? Yes. Setup maps the codebase, records what is already in use, and saves a test baseline so old failures don't block new work. Existing projects.
- Can I use any AI tool? Any tool that reads files. Claude Code and OpenCode get native commands, subagents and guards; others follow the markdown.
- What does it cost in context? About 480 tokens are always loaded; everything else loads when a step needs it, and
npx groundwork-ai doctorshows your project's number.
More questions: the FAQ.
Contributing
Issues and pull requests are welcome. The docs live in docs/. Run npm test before opening a PR;
CI runs Windows and Linux on Node 22 and 24.
License
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found