config-drift-checker

agent
Guvenlik Denetimi
Uyari
Health Gecti
  • License — License: NOASSERTION
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Community trust — 30 GitHub stars
Code Uyari
  • fs module — File system access in .github/workflows/bump.yml
Permissions Gecti
  • Permissions — No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

CI for your coding-agent setup (CLAUDE.md, skills, hooks): pinned baselines, a canary on every Claude Code release, model and version matrix, context-cost drift, self-diagnosing reports, repair. Claude Code, Codex, Gemini.

README.md

config-drift-checker

Your agent conventions are code. This is their CI. The rules your team taught its coding
agent (CLAUDE.md or AGENTS.md, skills, hooks) decide how your software gets written now, and
everything underneath them moves without asking: Claude Code alone ships about 25 releases a
month, and the model behind an alias changes server-side with no changelog
(it already has, silently, for weeks).
This turns those rules into eval cases, runs them on every PR and every release against a pinned
baseline, and tells you the moment something stops working: when, why, and what moved.

One suite, three agents. First-class on Claude Code (skills, hooks, release canaries, the
whole drift machinery); the same cases also run through OpenAI's Codex and Google's Gemini CLIs,
with live passing runs on all three, experimental labels on the newer two until full published
comparisons.

tests
release
agent setup

Proof it works: we broke our own setup, and the repair skill fixed it.
A skill's trigger description rewritten the way a careless PR would: the suite fell 1.00 → 0.56,
the tripwire case read 0.00, and the report named the cause itself: the skill was discovered but
never invoked
, so fix the trigger wording, not the packaging. Then the repair skill restored the
behaviour on its own and proved it with a green re-run, for $0.28. Both artifacts are unedited.

The whole story in 20 seconds: a careless PR breaks a skill, the suite goes red, the report names the cause, the repair skill fixes it

See it live

What you're looking at
The sabotage report what a real break looks like: a deliberately broken skill trigger, the tripwire at 0.00, the report naming the cause itself
The repair that fixed it the repair skill's own PR-ready summary from fixing that break live: what drifted, the smallest edit, the green re-run as evidence, $0.28 spent
The drift observatory this plugin's own suite re-run on every Claude Code release, with a live stability streak, and a subscribable feed of per-release verdicts: the behavioural changelog nobody publishes
The demo repo a small Spring Boot API whose whole setup (cases, config, workflow) was written by /config-drift-checker:setup unattended, kept exactly as generated
The demo's drift index the same observatory for that demo repo, built by its own CI
What a setup is worth the first community suite (Spring Boot conventions) with a published with/without measurement: the guard hook is worth +0.75, the run itself
The Claude Code release report every public claude plugin eval suite we can find, loaded on each new Claude Code release for $0: which still load, which broke on this release, which never loaded, with the fix for each
The site one page with all of the above

Quick start

One command in the repo whose setup you want protected:

claude plugin marketplace add jameskomo/config-drift-checker && claude plugin install config-drift-checker@jameskomo && claude "/config-drift-checker:setup"

Five minutes: it finds your CLAUDE.md, skills and hooks, writes starter eval cases from them,
smoke-runs them, and writes .cdc.yml plus the GitHub workflow. Prefer never opening Claude at
all? node <plugin-root>/tools/cdc-bootstrap.mjs . runs the same setup headlessly, and
--no-agent scaffolds everything for $0 (blank starter case included) so you fill in the prompts
yourself. Add one secret
(CLAUDE_CODE_OAUTH_TOKEN from claude setup-token to run on a Pro/Max subscription at no extra cost, within its usage limits, or
ANTHROPIC_API_KEY), push, done.

Try the two free checks on any plugin, nothing to install. Each is a single file with no
dependencies (read it first, it's short):

curl -fsSLO https://raw.githubusercontent.com/jameskomo/config-drift-checker/v1/config-drift-checker/tools/skill-lint.mjs && node skill-lint.mjs .        # your SKILL.md files
curl -fsSLO https://raw.githubusercontent.com/jameskomo/config-drift-checker/v1/config-drift-checker/tools/suite-doctor.mjs && node suite-doctor.mjs .    # your eval suite vs the installed Claude Code; add --fix

Neither starts a model run. What they check.

Already have a suite in the claude plugin eval format? One step:

- uses: jameskomo/config-drift-checker/action@v1
  with: { plugin-dir: . }

(@v0 keeps working; both moving tags point at the same latest release.)

What you get

Every row is shipped and tested; where a public receipt exists, it's linked.

Detect: know the moment behaviour moves

Pinned baseline + canary the baseline never moves under you; the canary tests each new Claude Code release and alias model before your team meets it
Noise bands with guards each case's allowed wobble is learned from its own history; a real break can't hide in the band (no recovering run, or a persisting drop, stays red)
Refusal labels a model guardrail change is labelled a refusal, never blamed on your setup
Discovered vs invoked a red skill case says which repair it needs: fix the trigger wording, or fix the packaging. Watch it self-diagnose a real break
Efficiency drift slower, pricier, longer gets flagged even when every case still passes
MCP plugins without the real service mocks in the official evals/mocks/ format answer for your MCP servers under both runners, real servers stay down by default, and target: mock_calls grades what the agent sent
Coverage which of your rules have no test: a percentage, a badge, and a coverage-min gate
Skill linter skill-lint checks every SKILL.md before any model run: frontmatter that strict parsers reject, missing or vague trigger descriptions, overlapping skills without negative scope, broken file references
Format drift suite-doctor checks your eval cases against the runner of the Claude Code you're about to test, for free (no model runs), names every case the new version rejects, and --fix migrates the known changes. Runs as a preflight on every release, so a schema change shows up as "this case no longer loads", not as mysterious regressions
Context-cost drift context-cost measures the tokens your skills, agents, commands and CLAUDE.md add to every session, with Claude Code's own plugin details, per release. "Your setup got 18% more expensive on 2.1.295" shows up as a line, not as a surprise bill
Real usage vs evals usage-check reads your local session history and sets it against your eval cases: skills that are tested but never used, used but never tested, or dead weight that costs context every session. Counts only, no prompt text ever leaves the transcripts

Diagnose: red comes with answers, not homework

Reports that show their work every report lists the whole suite including skipped cases, every discovered skill and whether it fired, and exactly which checks ran beyond a bare claude plugin eval
Model and version matrix drift-matrix runs your suite across models and Claude Code versions and draws one page: which model and release your setup survives, with a plain verdict per model ("haiku: safe on 2.1.290 to 2.1.295; fails one case on 2.1.288"). Budget-capped, --dry-run for $0
drift-bisect a case passed weeks ago and fails today: binary-search the Claude Code releases in between, log2(N) runs, get the culprit version and the bug-report sentence
trace-keeper preserves the official runner's transcripts, which it otherwise deletes on exit
What a setup is worth the same tasks with and without your setup, published: the guard hook measures +0.75

Repair: and prove the fix

Autonomous repair on red, the smallest setup edit that restores the behaviour, verified by re-running the failing cases. A real repair, $0.28, first attempt green
Bump and pin PRs two proven-green canaries open the PR that moves your pins, evidence attached, never auto-merged

Operate: it runs itself, and reports to you

The observatory stat tiles, a stability streak, and a verdict timeline per Claude Code release. Ours, live
The drift wire a subscribable Atom feed plus verdicts.json: one behavioural verdict per release, the changelog nobody else publishes
Setup health every run's report gets a setup-health panel (skills linted, eval suite loaded by the real runner, each finding with its fix), and the observatory adds a format drift per Claude Code release strip: the releases where your test suite itself stopped being valid, separate from behaviour drift
The status badge embeddable like a coverage badge: green "cc2.1.269 · 3 releases clean", red naming the version the day something breaks
Posts and digest, pre-written drift-digest turns every new verdict into ready-to-paste posts per platform (X, LinkedIn, Reddit when something actually broke) plus a weekly digest; a workflow can post to Bluesky/Mastodon automatically
Hard budget caps per run and per month, enforced from a ledger; on a Claude Pro/Max subscription token, $0 API, counted against the plan's usage limits
One-command onboarding cdc-bootstrap runs the whole setup headlessly with an auth preflight, or scaffolds everything for $0 with --no-agent
Fleet + org rollout one dashboard and pin policy across every repo, and a reusable org workflow that installs the check with a three-line caller. No hosted server, ever
Public release report on every Claude Code release, a free job loads every public eval suite we can find on the new version and the one before. Live page
Community suites maintained setups with published worth numbers (Spring Boot first); a hosted-tier waitlist decides what we run for you

Across agents: one suite, three CLIs

Claude Code first-class: skills, hooks, plugins, release canaries, the whole drift machinery
Codex (experimental) --agent codex with AGENTS.md bridging; live-calibrated on a ChatGPT plan
Gemini (experimental) --agent gemini with GEMINI.md bridging; live run scored 1.00 on the free Google tier
skill-creator suites evals-convert import turns Anthropic's skill-creator evals.json into plugin-eval cases (and export goes back), so one set of evals runs under both

How this relates to claude plugin eval

Claude Code ships an eval runner, and it's good: it runs your cases, grades them, generates
starter cases with init, and writes a report. We build on it, not beside it: cases are in that
exact format, the Action prefers the official runner (bundled fallback for older versions), and a
claude plugin eval . --json out.json result feeds our diff, report and drift index directly.

claude plugin eval (built in) config-drift-checker
Run cases, grade, report on one run yes, it's the runner we build on uses it
Generate starter cases init /config-drift-checker:setup, same format
A stored baseline to diff against no, you compare runs by eye pinned baseline, promoted deliberately
History across releases no, the docs advise pinning your model every run kept; a drift index over every version
Flake vs break no, a noisy case just fails sometimes per-case noise bands with anti-masking guards
Watching Claude Code and model releases no release watch + canary, throttled by your budget
Pin bump PRs no two green canaries open a PR with evidence
Red check, PR comment, Slack exit code all three
Spend control a per-run ceiling flag per-run and per-month caps, enforced from a ledger
Coverage of your rules no percent, badge, coverage-min gate
Repair proposal on red no a PR with the smallest fix, proven live
Why a skill case failed a score discovered vs invoked: trigger wording or packaging
Transcripts deleted when the command exits trace-keeper copies them next to the JSON
Which model and version is safe one --model per run drift-matrix: a grid of models and releases on one page
Context cost plugin details shows today's number context-cost keeps it per release and names what moved
Does real use match the suite no (/skill-doctor shows usage, not evals) usage-check crosses the two
skill-creator evals.json a separate format it doesn't read evals-convert imports and exports
Is my suite still valid on the new release load errors when you run it suite-doctor preflight for $0, plus a public report across every suite we can find

One sentence: their command answers "does my plugin work right now on my machine"; this answers
"did anything stop working since the baseline, across every release, without me watching".

Trust

Stable since v1.0: the v0/v1 moving tag never breaks your workflow; breaking changes mean
a new major with an upgrade note in CHANGELOG.md.
Zero npm dependencies (Node builtins only), 246 tests that run against a fake claude with no API
key, every third-party action pinned to a verified commit SHA, CodeQL on every push. Runs on your
runner with your key; nothing is sent to us, because there is no us to send it to. Details in
docs/security.md.

What's here

config-drift-checker/   the plugin: skills (setup · run · write-case · repair) + the tools
  tools/                shim runner · diff · classify · report · dashboard · coverage · watch · gate · promote · trace-keeper · fleet · bootstrap · drift-bisect · drift-digest
  test/                 node --test suite, fake claude, npm test
action/                 composite GitHub Action: gate → run → diff → store → PR → repair → alert
examples/komo-stack/    a full example suite with .cdc.yml and baseline results
docs/                   user guide · architecture · eval format & runner · runbook · security

Hosted tier (waitlist)

Self-hosting is free forever, that never changes. A managed tier (we run the canaries, fleet
dashboards and alerts; you get the PRs and the pages) gets built when enough teams want it:
join the waitlist,
two questions, no commitment. Discussions are open
for everything else.

Documentation

Start with the user guide (.cdc.yml reference here); full index in docs/.

Licence

FSL-1.1-Apache-2.0: free to use, modify and self-host; not to be offered as a competing
commercial service; each release becomes Apache-2.0 two years after publication.

Yorumlar (0)

Sonuc bulunamadi