done-is-a-claim

agent
Security Audit
Warn
Health Pass
  • License — License: Apache-2.0
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Community trust — 28 GitHub stars
Code Warn
  • fs module — File system access in tools/render_stamp.py
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

18 incident-backed rules and 6 skills for coding agents. Evidence templates, runnable demos, and drop-in AGENTS.md / CLAUDE.md.

README.md

Done is a claim

English · Türkçe

Done is a claim. A red SHOW THE RECEIPT stamp lands on the ivory cover.

Replay stamp · Static cover

Your agent says “done”. What would prove it?

18 incident-backed rules and 6 focused skills for coding agents. Give the agent a
clear acceptance target, ask for the actual result, and keep the evidence attached
to the work it tested. Plain Markdown. Use the pieces you need.

Get started · Choose a skill · Try the demo · Read the incidents

Watch a false green

A command fails. The process reading its output succeeds. Report the wrong exit
code, and a failed check becomes a green report.

Runnable source. This is a synthetic demonstration,
not a production log or a benchmark.

Run it yourself from the cloned repository's root with Python 3.10+:

python examples/false_green.py

The demo succeeds when it exposes the mismatch. The verifier inside it still
fails.
That distinction is the point.

Get started

git clone https://github.com/Grit-77/done-is-a-claim.git
cd done-is-a-claim
You use Add to your project
Codex or an agent that reads AGENTS.md Copy AGENTS.md to the project root.
Claude Code Copy AGENTS.md and CLAUDE.md, which imports it with @AGENTS.md.
Existing instructions Merge the relevant rules into your file. Keep your project-specific commands and constraints.

You can start with that one rules file. The optional skills add depth when a
particular problem comes up. Install a skill on macOS, Linux or Windows →

Try this first request in your own project:

Read the project instructions. Before making changes, identify the acceptance
command for this task and the user-visible result it checks. When reporting the
result, include the tested revision, actual output and anything you did not check.

This checks whether the agent understood the request. It does not enforce compliance.

What changes in the report

A claim Evidence that can support it
“The tests passed.” The command, collected cases, actual result, its own exit status and the tested tree.
“That failure was already on main.” Comparable branch and clean-base runs, with matching failure signatures.
“The worker finished.” The worker's actual diff, checked inputs and acceptance result. Delivery is recorded separately.
“The screenshot was saved.” An image that decodes and has been visually inspected.
“Ready to publish.” The reviewed artifact still matches the artifact being published.

Use the completion receipt,
acceptance brief and handoff
as small, reusable formats. A missing check belongs in the report.

Practical Bash and PowerShell evidence recipes →

Choose a skill

When this happens Use this skill It helps you decide
The task's “pass” condition is vague acceptance-design Does this check cross the boundary the user actually cares about?
A number looks convincing reading-measurements What does this output establish, and what does it leave unknown?
A test fails on your branch whose-red Is there comparable evidence for a regression, an existing failure or an unresolved cause?
A subagent reports success collecting-worker-results Is there relevant work, was that work tested, and was it delivered?
Work changed after a check evidence-freshness Does the receipt still describe the inputs you are about to act on?
A README or launch post makes a claim public-claims Can the reader trace it, and are the verification limits stated?

Each skill is a standalone SKILL.md. No Grit service, account or CLI is needed.

The field rules

The complete wording lives in AGENTS.md. Each original rule has a
failure story in INCIDENTS.md.

Moment Rules
Before you start 01 Define acceptance before the work. 02 Check the paths.
Before you say “done” 03 Re-run on the current work. 04 Check the wider suite. 05 Capture the command's own exit code. 06 No tests is no pass. 07 Unknown is not pass. 08 Read every required check.
When you report 09 Read the artifact. 10 Open what was written. 11 Look at it yourself. 12 Verify citations. 13 State the base and re-check it before publication.
When you test and fix 14 Test behavior, not agreement with a constant. 15 Fix the defect class. 16 Test the detector on a known-good case. 17 Register cleanup before the work.
Always preserve the work 18 Never use a destructive command to answer a question.

Where this came from

These rules grew out of running Claude Code and Codex on Grit's own repository.
The original internal records report 3,489 task acceptance re-runs between
14–29 September 2026, with 2,282 passing the first independent re-run. A
separate internal count reports 736 tasks whose own acceptance passed while
the full suite broke on main.

These are author-reported historical observations, not a public benchmark. The
raw internal logs are not included. Some unsuccessful re-runs were environment
failures; these figures do not establish an agent-error rate or the effectiveness
of this toolkit. Claims and limits →

The additional workflows distill Grit's operational runbooks. They are identified
as guidance, not newly measured incidents. We also studied how related projects
organize installation, skills and verification. Sources and design decisions →

Try to fool it

An agent can repeat a rule and still make the wrong decision. The
scenario pack puts the instructions under pressure: a passing
pipe, mismatched controls, a worker with no relevant changes, stale evidence and
other traps. Run the prompts in fresh sessions and keep the responses.

These are manual behavioral evaluations, not published success-rate claims.

For the repository itself:

python tools/run_tests.py
python tools/check_repository.py

The test runner rejects empty or entirely skipped suites. The checks exercise the
tools and validate local document targets, skill metadata
and imports. CI runs on Windows and Linux. Passing
these checks does not prove that an agent follows the instructions.

Rules are not a gate

This repository supplies instructions, examples and evaluation material. It does
not intercept an agent's tools, block a merge or enforce a deployment policy.
Use your project's actual test and release gates for enforcement. Grit's separate
RADAR project explores that operational layer.

Bring the failure that taught you

A useful contribution starts with what the agent claimed, what was true, and the
evidence that revealed the difference. Propose a rule
or read the contribution guide.

Apache-2.0 · Made by Grit, Ankara.

Reviews (0)

No results found