config-drift-checker
Health Pass
- License — License: NOASSERTION
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Community trust — 30 GitHub stars
Code Warn
- fs module — File system access in .github/workflows/bump.yml
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
CI for your coding-agent setup (CLAUDE.md, skills, hooks): pinned baselines, a canary on every Claude Code release, model and version matrix, context-cost drift, self-diagnosing reports, repair. Claude Code, Codex, Gemini.
config-drift-checker
Your agent conventions are code. This is their CI. The rules your team taught its coding
agent (CLAUDE.md or AGENTS.md, skills, hooks) decide how your software gets written now, and
everything underneath them moves without asking: Claude Code alone ships about 25 releases a
month, and the model behind an alias changes server-side with no changelog
(it already has, silently, for weeks).
This turns those rules into eval cases, runs them on every PR and every release against a pinned
baseline, and tells you the moment something stops working: when, why, and what moved.
One suite, three agents. First-class on Claude Code (skills, hooks, release canaries, the
whole drift machinery); the same cases also run through OpenAI's Codex and Google's Gemini CLIs,
with live passing runs on all three, experimental labels on the newer two until full published
comparisons.
Proof it works: we broke our own setup, and the repair skill fixed it.
A skill's trigger description rewritten the way a careless PR would: the suite fell 1.00 → 0.56,
the tripwire case read 0.00, and the report named the cause itself: the skill was discovered but
never invoked, so fix the trigger wording, not the packaging. Then the repair skill restored the
behaviour on its own and proved it with a green re-run, for $0.28. Both artifacts are unedited.
See it live
| What you're looking at | |
|---|---|
| The sabotage report | what a real break looks like: a deliberately broken skill trigger, the tripwire at 0.00, the report naming the cause itself |
| The repair that fixed it | the repair skill's own PR-ready summary from fixing that break live: what drifted, the smallest edit, the green re-run as evidence, $0.28 spent |
| The drift observatory | this plugin's own suite re-run on every Claude Code release, with a live stability streak, and a subscribable feed of per-release verdicts: the behavioural changelog nobody publishes |
| The demo repo | a small Spring Boot API whose whole setup (cases, config, workflow) was written by /config-drift-checker:setup unattended, kept exactly as generated |
| The demo's drift index | the same observatory for that demo repo, built by its own CI |
| What a setup is worth | the first community suite (Spring Boot conventions) with a published with/without measurement: the guard hook is worth +0.75, the run itself |
| The Claude Code release report | every public claude plugin eval suite we can find, loaded on each new Claude Code release for $0: which still load, which broke on this release, which never loaded, with the fix for each |
| The site | one page with all of the above |
Quick start
One command in the repo whose setup you want protected:
claude plugin marketplace add jameskomo/config-drift-checker && claude plugin install config-drift-checker@jameskomo && claude "/config-drift-checker:setup"
Five minutes: it finds your CLAUDE.md, skills and hooks, writes starter eval cases from them,
smoke-runs them, and writes .cdc.yml plus the GitHub workflow. Prefer never opening Claude at
all? node <plugin-root>/tools/cdc-bootstrap.mjs . runs the same setup headlessly, and--no-agent scaffolds everything for $0 (blank starter case included) so you fill in the prompts
yourself. Add one secret
(CLAUDE_CODE_OAUTH_TOKEN from claude setup-token to run on a Pro/Max subscription at no extra cost, within its usage limits, orANTHROPIC_API_KEY), push, done.
Try the two free checks on any plugin, nothing to install. Each is a single file with no
dependencies (read it first, it's short):
curl -fsSLO https://raw.githubusercontent.com/jameskomo/config-drift-checker/v1/config-drift-checker/tools/skill-lint.mjs && node skill-lint.mjs . # your SKILL.md files
curl -fsSLO https://raw.githubusercontent.com/jameskomo/config-drift-checker/v1/config-drift-checker/tools/suite-doctor.mjs && node suite-doctor.mjs . # your eval suite vs the installed Claude Code; add --fix
Neither starts a model run. What they check.
Already have a suite in the claude plugin eval format? One step:
- uses: jameskomo/config-drift-checker/action@v1
with: { plugin-dir: . }
(@v0 keeps working; both moving tags point at the same latest release.)
What you get
Every row is shipped and tested; where a public receipt exists, it's linked.
Detect: know the moment behaviour moves
| Pinned baseline + canary | the baseline never moves under you; the canary tests each new Claude Code release and alias model before your team meets it |
| Noise bands with guards | each case's allowed wobble is learned from its own history; a real break can't hide in the band (no recovering run, or a persisting drop, stays red) |
| Refusal labels | a model guardrail change is labelled a refusal, never blamed on your setup |
| Discovered vs invoked | a red skill case says which repair it needs: fix the trigger wording, or fix the packaging. Watch it self-diagnose a real break |
| Efficiency drift | slower, pricier, longer gets flagged even when every case still passes |
| MCP plugins without the real service | mocks in the official evals/mocks/ format answer for your MCP servers under both runners, real servers stay down by default, and target: mock_calls grades what the agent sent |
| Coverage | which of your rules have no test: a percentage, a badge, and a coverage-min gate |
| Skill linter | skill-lint checks every SKILL.md before any model run: frontmatter that strict parsers reject, missing or vague trigger descriptions, overlapping skills without negative scope, broken file references |
| Format drift | suite-doctor checks your eval cases against the runner of the Claude Code you're about to test, for free (no model runs), names every case the new version rejects, and --fix migrates the known changes. Runs as a preflight on every release, so a schema change shows up as "this case no longer loads", not as mysterious regressions |
| Context-cost drift | context-cost measures the tokens your skills, agents, commands and CLAUDE.md add to every session, with Claude Code's own plugin details, per release. "Your setup got 18% more expensive on 2.1.295" shows up as a line, not as a surprise bill |
| Real usage vs evals | usage-check reads your local session history and sets it against your eval cases: skills that are tested but never used, used but never tested, or dead weight that costs context every session. Counts only, no prompt text ever leaves the transcripts |
Diagnose: red comes with answers, not homework
| Reports that show their work | every report lists the whole suite including skipped cases, every discovered skill and whether it fired, and exactly which checks ran beyond a bare claude plugin eval |
| Model and version matrix | drift-matrix runs your suite across models and Claude Code versions and draws one page: which model and release your setup survives, with a plain verdict per model ("haiku: safe on 2.1.290 to 2.1.295; fails one case on 2.1.288"). Budget-capped, --dry-run for $0 |
drift-bisect |
a case passed weeks ago and fails today: binary-search the Claude Code releases in between, log2(N) runs, get the culprit version and the bug-report sentence |
trace-keeper |
preserves the official runner's transcripts, which it otherwise deletes on exit |
| What a setup is worth | the same tasks with and without your setup, published: the guard hook measures +0.75 |
Repair: and prove the fix
| Autonomous repair | on red, the smallest setup edit that restores the behaviour, verified by re-running the failing cases. A real repair, $0.28, first attempt green |
| Bump and pin PRs | two proven-green canaries open the PR that moves your pins, evidence attached, never auto-merged |
Operate: it runs itself, and reports to you
| The observatory | stat tiles, a stability streak, and a verdict timeline per Claude Code release. Ours, live |
| The drift wire | a subscribable Atom feed plus verdicts.json: one behavioural verdict per release, the changelog nobody else publishes |
| Setup health | every run's report gets a setup-health panel (skills linted, eval suite loaded by the real runner, each finding with its fix), and the observatory adds a format drift per Claude Code release strip: the releases where your test suite itself stopped being valid, separate from behaviour drift |
| The status badge | embeddable like a coverage badge: green "cc2.1.269 · 3 releases clean", red naming the version the day something breaks |
| Posts and digest, pre-written | drift-digest turns every new verdict into ready-to-paste posts per platform (X, LinkedIn, Reddit when something actually broke) plus a weekly digest; a workflow can post to Bluesky/Mastodon automatically |
| Hard budget caps | per run and per month, enforced from a ledger; on a Claude Pro/Max subscription token, $0 API, counted against the plan's usage limits |
| One-command onboarding | cdc-bootstrap runs the whole setup headlessly with an auth preflight, or scaffolds everything for $0 with --no-agent |
| Fleet + org rollout | one dashboard and pin policy across every repo, and a reusable org workflow that installs the check with a three-line caller. No hosted server, ever |
| Public release report | on every Claude Code release, a free job loads every public eval suite we can find on the new version and the one before. Live page |
| Community suites | maintained setups with published worth numbers (Spring Boot first); a hosted-tier waitlist decides what we run for you |
Across agents: one suite, three CLIs
| Claude Code | first-class: skills, hooks, plugins, release canaries, the whole drift machinery |
| Codex (experimental) | --agent codex with AGENTS.md bridging; live-calibrated on a ChatGPT plan |
| Gemini (experimental) | --agent gemini with GEMINI.md bridging; live run scored 1.00 on the free Google tier |
| skill-creator suites | evals-convert import turns Anthropic's skill-creator evals.json into plugin-eval cases (and export goes back), so one set of evals runs under both |
How this relates to claude plugin eval
Claude Code ships an eval runner, and it's good: it runs your cases, grades them, generates
starter cases with init, and writes a report. We build on it, not beside it: cases are in that
exact format, the Action prefers the official runner (bundled fallback for older versions), and aclaude plugin eval . --json out.json result feeds our diff, report and drift index directly.
claude plugin eval (built in) |
config-drift-checker | |
|---|---|---|
| Run cases, grade, report on one run | yes, it's the runner we build on | uses it |
| Generate starter cases | init |
/config-drift-checker:setup, same format |
| A stored baseline to diff against | no, you compare runs by eye | pinned baseline, promoted deliberately |
| History across releases | no, the docs advise pinning your model | every run kept; a drift index over every version |
| Flake vs break | no, a noisy case just fails sometimes | per-case noise bands with anti-masking guards |
| Watching Claude Code and model releases | no | release watch + canary, throttled by your budget |
| Pin bump PRs | no | two green canaries open a PR with evidence |
| Red check, PR comment, Slack | exit code | all three |
| Spend control | a per-run ceiling flag | per-run and per-month caps, enforced from a ledger |
| Coverage of your rules | no | percent, badge, coverage-min gate |
| Repair proposal on red | no | a PR with the smallest fix, proven live |
| Why a skill case failed | a score | discovered vs invoked: trigger wording or packaging |
| Transcripts | deleted when the command exits | trace-keeper copies them next to the JSON |
| Which model and version is safe | one --model per run |
drift-matrix: a grid of models and releases on one page |
| Context cost | plugin details shows today's number |
context-cost keeps it per release and names what moved |
| Does real use match the suite | no (/skill-doctor shows usage, not evals) |
usage-check crosses the two |
skill-creator evals.json |
a separate format it doesn't read | evals-convert imports and exports |
| Is my suite still valid on the new release | load errors when you run it | suite-doctor preflight for $0, plus a public report across every suite we can find |
One sentence: their command answers "does my plugin work right now on my machine"; this answers
"did anything stop working since the baseline, across every release, without me watching".
Trust
Stable since v1.0: the v0/v1 moving tag never breaks your workflow; breaking changes mean
a new major with an upgrade note in CHANGELOG.md.
Zero npm dependencies (Node builtins only), 246 tests that run against a fake claude with no API
key, every third-party action pinned to a verified commit SHA, CodeQL on every push. Runs on your
runner with your key; nothing is sent to us, because there is no us to send it to. Details in
docs/security.md.
What's here
config-drift-checker/ the plugin: skills (setup · run · write-case · repair) + the tools
tools/ shim runner · diff · classify · report · dashboard · coverage · watch · gate · promote · trace-keeper · fleet · bootstrap · drift-bisect · drift-digest
test/ node --test suite, fake claude, npm test
action/ composite GitHub Action: gate → run → diff → store → PR → repair → alert
examples/komo-stack/ a full example suite with .cdc.yml and baseline results
docs/ user guide · architecture · eval format & runner · runbook · security
Hosted tier (waitlist)
Self-hosting is free forever, that never changes. A managed tier (we run the canaries, fleet
dashboards and alerts; you get the PRs and the pages) gets built when enough teams want it:
join the waitlist,
two questions, no commitment. Discussions are open
for everything else.
Documentation
Start with the user guide (.cdc.yml reference here); full index in docs/.
Licence
FSL-1.1-Apache-2.0: free to use, modify and self-host; not to be offered as a competing
commercial service; each release becomes Apache-2.0 two years after publication.
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found