done-is-a-claim
Health Gecti
- License — License: Apache-2.0
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Community trust — 28 GitHub stars
Code Uyari
- fs module — File system access in tools/render_stamp.py
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
18 incident-backed rules and 6 skills for coding agents. Evidence templates, runnable demos, and drop-in AGENTS.md / CLAUDE.md.
Done is a claim
English · Türkçe
Your agent says “done”. What would prove it?
18 incident-backed rules and 6 focused skills for coding agents. Give the agent a
clear acceptance target, ask for the actual result, and keep the evidence attached
to the work it tested. Plain Markdown. Use the pieces you need.
Get started · Choose a skill · Try the demo · Read the incidents
Watch a false green
A command fails. The process reading its output succeeds. Report the wrong exit
code, and a failed check becomes a green report.
Runnable source. This is a synthetic demonstration,
not a production log or a benchmark.
Run it yourself from the cloned repository's root with Python 3.10+:
python examples/false_green.py
The demo succeeds when it exposes the mismatch. The verifier inside it still
fails. That distinction is the point.
Get started
git clone https://github.com/Grit-77/done-is-a-claim.git
cd done-is-a-claim
| You use | Add to your project |
|---|---|
Codex or an agent that reads AGENTS.md |
Copy AGENTS.md to the project root. |
| Claude Code | Copy AGENTS.md and CLAUDE.md, which imports it with @AGENTS.md. |
| Existing instructions | Merge the relevant rules into your file. Keep your project-specific commands and constraints. |
You can start with that one rules file. The optional skills add depth when a
particular problem comes up. Install a skill on macOS, Linux or Windows →
Try this first request in your own project:
Read the project instructions. Before making changes, identify the acceptance
command for this task and the user-visible result it checks. When reporting the
result, include the tested revision, actual output and anything you did not check.
This checks whether the agent understood the request. It does not enforce compliance.
What changes in the report
| A claim | Evidence that can support it |
|---|---|
| “The tests passed.” | The command, collected cases, actual result, its own exit status and the tested tree. |
| “That failure was already on main.” | Comparable branch and clean-base runs, with matching failure signatures. |
| “The worker finished.” | The worker's actual diff, checked inputs and acceptance result. Delivery is recorded separately. |
| “The screenshot was saved.” | An image that decodes and has been visually inspected. |
| “Ready to publish.” | The reviewed artifact still matches the artifact being published. |
Use the completion receipt,
acceptance brief and handoff
as small, reusable formats. A missing check belongs in the report.
Practical Bash and PowerShell evidence recipes →
Choose a skill
| When this happens | Use this skill | It helps you decide |
|---|---|---|
| The task's “pass” condition is vague | acceptance-design | Does this check cross the boundary the user actually cares about? |
| A number looks convincing | reading-measurements | What does this output establish, and what does it leave unknown? |
| A test fails on your branch | whose-red | Is there comparable evidence for a regression, an existing failure or an unresolved cause? |
| A subagent reports success | collecting-worker-results | Is there relevant work, was that work tested, and was it delivered? |
| Work changed after a check | evidence-freshness | Does the receipt still describe the inputs you are about to act on? |
| A README or launch post makes a claim | public-claims | Can the reader trace it, and are the verification limits stated? |
Each skill is a standalone SKILL.md. No Grit service, account or CLI is needed.
The field rules
The complete wording lives in AGENTS.md. Each original rule has a
failure story in INCIDENTS.md.
| Moment | Rules |
|---|---|
| Before you start | 01 Define acceptance before the work. 02 Check the paths. |
| Before you say “done” | 03 Re-run on the current work. 04 Check the wider suite. 05 Capture the command's own exit code. 06 No tests is no pass. 07 Unknown is not pass. 08 Read every required check. |
| When you report | 09 Read the artifact. 10 Open what was written. 11 Look at it yourself. 12 Verify citations. 13 State the base and re-check it before publication. |
| When you test and fix | 14 Test behavior, not agreement with a constant. 15 Fix the defect class. 16 Test the detector on a known-good case. 17 Register cleanup before the work. |
| Always preserve the work | 18 Never use a destructive command to answer a question. |
Where this came from
These rules grew out of running Claude Code and Codex on Grit's own repository.
The original internal records report 3,489 task acceptance re-runs between
14–29 September 2026, with 2,282 passing the first independent re-run. A
separate internal count reports 736 tasks whose own acceptance passed while
the full suite broke on main.
These are author-reported historical observations, not a public benchmark. The
raw internal logs are not included. Some unsuccessful re-runs were environment
failures; these figures do not establish an agent-error rate or the effectiveness
of this toolkit. Claims and limits →
The additional workflows distill Grit's operational runbooks. They are identified
as guidance, not newly measured incidents. We also studied how related projects
organize installation, skills and verification. Sources and design decisions →
Try to fool it
An agent can repeat a rule and still make the wrong decision. The
scenario pack puts the instructions under pressure: a passing
pipe, mismatched controls, a worker with no relevant changes, stale evidence and
other traps. Run the prompts in fresh sessions and keep the responses.
These are manual behavioral evaluations, not published success-rate claims.
For the repository itself:
python tools/run_tests.py
python tools/check_repository.py
The test runner rejects empty or entirely skipped suites. The checks exercise the
tools and validate local document targets, skill metadata
and imports. CI runs on Windows and Linux. Passing
these checks does not prove that an agent follows the instructions.
Rules are not a gate
This repository supplies instructions, examples and evaluation material. It does
not intercept an agent's tools, block a merge or enforce a deployment policy.
Use your project's actual test and release gates for enforcement. Grit's separate
RADAR project explores that operational layer.
Bring the failure that taught you
A useful contribution starts with what the agent claimed, what was true, and the
evidence that revealed the difference. Propose a rule
or read the contribution guide.
Apache-2.0 · Made by Grit, Ankara.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi