claude-harness

agent
Security Audit
Fail
Health Warn
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 7 GitHub stars
Code Fail
  • child_process — Shell command execution capability in hooks/adapters/agy-adapter.cjs
  • spawnSync — Synchronous process spawning in hooks/adapters/agy-adapter.cjs
  • fs module — File system access in hooks/adapters/agy-adapter.cjs
  • exec() — Shell command execution in hooks/agent-model-guard.cjs
  • os.homedir — User home directory access in hooks/agent-model-guard.cjs
  • process.env — Environment variable access in hooks/agent-model-guard.cjs
  • fs module — File system access in hooks/agent-model-guard.cjs
  • os.homedir — User home directory access in hooks/batch-guard.cjs
  • process.env — Environment variable access in hooks/batch-guard.cjs
  • fs module — File system access in hooks/batch-guard.cjs
  • exec() — Shell command execution in hooks/capability-graph-guard.cjs
  • fs.rmSync — Destructive file system operation in hooks/capability-graph-guard.cjs
  • os.homedir — User home directory access in hooks/capability-graph-guard.cjs
  • process.env — Environment variable access in hooks/capability-graph-guard.cjs
  • fs module — File system access in hooks/capability-graph-guard.cjs
  • child_process — Shell command execution capability in hooks/creds-resolve.cjs
  • spawnSync — Synchronous process spawning in hooks/creds-resolve.cjs
  • fs module — File system access in hooks/creds-resolve.cjs
  • exec() — Shell command execution in hooks/fable-delegate-guard.cjs
  • fs.rmSync — Destructive file system operation in hooks/fable-delegate-guard.cjs
  • os.homedir — User home directory access in hooks/fable-delegate-guard.cjs
  • process.env — Environment variable access in hooks/fable-delegate-guard.cjs
  • fs module — File system access in hooks/fable-delegate-guard.cjs
  • os.homedir — User home directory access in hooks/git-tree-guard.cjs
  • fs module — File system access in hooks/git-tree-guard.cjs
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

A daily-driver Claude Code harness, mirrored public: incident-born guard hooks, enforced model routing, zero-dependency tools, and a nightly self-reflection ladder.

README.md

claude-harness

The Claude Code setup I run every day, mirrored public. Incident-born guard
hooks, a model-routing policy that is enforced rather than suggested, a dozen
small zero-dependency tools, and a nightly self-reflection loop that promotes
observations into rules only when the evidence earns it.

License: MIT
Platform
Runtime
Last sync
Sponsor

This is not a starter kit designed in an afternoon. It grew rule by rule out of
real incidents over six months of daily use, and most of it exists because
something broke first. The private repo that backs it also holds the agent's
memory; that stays on the machine. Everything here was swept file by file
before publishing (see Security).


Contents


Why it exists

Three examples of what "incident-born" means:

  • hooks/process-kill-guard.cjs exists because a subagent cleaned up its
    test window with Stop-Process -Name notepad and killed a real Notepad
    session with about 40 tabs of unsaved work. Name-based process kills are now
    blocked at the tool layer. PID-based kills still work.
  • hooks/agent-model-guard.cjs exists because one workflow spawned 110
    agents on the most expensive model through inherited defaults and burned a
    full five-hour usage window. Every agent spawn now requires an explicit
    model, and the expensive tier is capped per session.
  • git-hooks/pre-commit runs hooks/secret-guard.cjs over every staged
    file in every repo on the machine, whatever the language, because keys were
    once found sitting in plaintext on disk.
  • The four-round fix session of 2026-09-03 is written up in
    docs/postmortem-2026-09-03-declick-launch.md:
    what a file-scoped fix workflow, a capped advisor, an over-tight edit budget
    and a missing git guard cost on one launch day, and what each became.
  • hooks/git-tree-guard.cjs exists because a reviewer in a 17-agent fix
    workflow ran git stash to watch a test fail without its fix, the pop
    conflicted on a sibling's edit, and a 25-file fix pass sat silently reverted
    under six agents still working. Stash, path checkouts, restore, hard reset
    and clean are now denied in shell calls; baselines come from copies
    (git show HEAD:<path> into a scratch dir). workflows/fix-findings.js
    says the same thing in every prompt it sends.

The design stance behind all of it: a rule written in prose is a hope. A rule
that matters gets a hook, the hook gets a probe that makes it fail on purpose,
and the probe runs on a schedule. A guard that has never been observed failing
has been installed, not verified.

How it fits together

flowchart LR
    subgraph session["Claude Code session"]
        S[SessionStart] --> P[UserPromptSubmit]
        P --> T{Tool call}
        T -->|PreToolUse| G[Guards]
        G -->|allow| X[Tool runs]
        G -->|deny + reason| P
        X -->|PostToolUse| A[Advisory hooks]
        A --> P
        P -->|Stop| L[Telemetry]
    end
    C[CLAUDE.md<br/>global agreement] -.loaded.-> S
    SOUL[SOUL.md<br/>identity] -.read first.-> S
    L --> M[Nightly meditation]
    M -->|observation → fact → rule → trait| C
    M --> SOUL
    X -->|git commit| H[pre-commit chain<br/>secret scan, doc gates, lint]

Two layers do the work. Hooks sit on Claude Code's lifecycle events and
either deny a call with a reason the model can act on, or inject context the
model would otherwise forget. The ladder runs once a night, reads the
telemetry and the day's work, and decides whether anything earned promotion
into a standing rule. Everything else in the repo is tooling in service of
those two.

Guards

Every guard is a single file, no dependencies, wired in settings.json. Each
carries an override marker so a deliberate exception is one comment away and
gets logged, rather than a reason to switch the guard off.

Hook Event What it does Override
secret-guard.cjs PreToolUse, pre-commit Scans tool inputs and every staged file for key shapes, key=value secrets and env files. Placeholder-aware. none
output-secret-watch.cjs MessageDisplay The only output-side guard: watches what the model prints for the same shapes. none
process-kill-guard.cjs PreToolUse Denies name-based process termination (Stop-Process -Name, taskkill /IM, pkill, killall) and dynamic invocations carrying -Name. PID forms pass. KILL_BY_NAME_OK
agent-model-guard.cjs PreToolUse Every Agent, Task or Workflow spawn must name its model. Caps top-tier spawns per session. In Workflow scripts the top tier may only appear as a top-level synthesizer after the fan-out. AGENT_GUARD_FABLE_CAP
capability-graph-guard.cjs PreToolUse, SubagentStart/Stop Models delegate downward only (Fable → Opus → Sonnet → Haiku). Peers are not edges. The advisor agent is always placed one rung above its caller. CAPABILITY_GRAPH_GUARD=off
fable-delegate-guard.cjs PreToolUse, SessionStart, UserPromptSubmit When the main loop runs on the top model, budgets its direct edits and denies shell code-writing, so implementation goes to cheaper subagents. Decisions, review and synthesis stay. # FABLE_OK: <why>
batch-guard.cjs PreToolUse Denies the fourth consecutive single-statement shell call or single Read/Glob/Grep. Profiling showed one call per turn was the largest single cost. # SEQ: <dependency>
slow-command-guard.cjs PreToolUse Denies backgrounded finite test runs and recursive grep/find rooted at a projects dir, home or a drive. Both hang sessions. BG_TEST_OK, SLOW_OK
git-tree-guard.cjs PreToolUse Denies git stash, checkout/restore of paths, reset --hard, clean in shell calls: a working tree that several agents edit is read-only to git. Reads, branches and commits pass. Prose mentioning the words is not a hit. # GIT_TREE_OK: <why>
dev-server-guard.cjs PreToolUse, PostToolUse Denies dev servers piped through head/tail or backgrounded without an explicit opt-in; reminds that stopping a wrapper on Windows leaves the children alive. DEV_SERVER_BG_OK
scope-lock.cjs PreToolUse, UserPromptSubmit Confines Edit/Write to a directory for the session. Arm with scope-lock <dir> as a prompt. scope-unlock
repeat-tool-guard.cjs PostToolUse Counts identical consecutive calls and escalates a reminder at 3, 5 and 8. Advisory, never blocks. REPEAT_GUARD_OFF=1
no-auto-compact.cjs PreCompact Turns "never auto-compact, ask at 80%" from prose into a hook. none
context-nudge.py UserPromptSubmit One nudge per high-context crossing to consider /compact or /clear. none
opus-handoff-inject.cjs SessionStart, UserPromptSubmit Detects an Opus session and injects the lower-cost operating notes once. none
creds-resolve.cjs SessionStart Fills .env from a local vault when .env.example exists, so the agent never asks for a key it already has. none
session-count.py SessionStart Warns when several sessions share one rate limit. none
guard-canary.ps1 SessionStart, ~20h Makes each guard fail on purpose and confirms it blocks. none
correction-tracker.ps1 UserPromptSubmit Buckets user corrections so a repeated one surfaces as a rule candidate instead of waiting for a human to notice. none
skill-telemetry.py Stop One JSONL record per turn: skills, agents, MCP servers, tools, tokens. The ladder's data layer. none
lsp-reaper.ps1 scheduled Kills orphaned TypeScript language servers (once found 663 of them holding 13.6 GB). none

hooks/tests/ holds the probes. hooks/adapters/ wires the same files into
Codex and Antigravity so there is one guard suite, not three copies
(docs/harness-parity.md). Mechanism and incident
history for each guard: docs/harness-guards.md.

Model routing and delegation

The global agreement routes work by cost and the hooks enforce it.

flowchart TD
    F[Fable<br/>decisions, review, synthesis] --> O[Opus<br/>planning, orchestration, hard debugging]
    O --> S[Sonnet<br/>implementation, exploration]
    S --> H[Haiku<br/>lookups, mechanical edits]
    S -. advisor .-> O
    O -. advisor .-> F
  • Downward only. A spawn that crosses a missing edge is denied. A fork
    inherits its caller's model, so it counts as a peer edge from any subagent.
  • Upward is consultation, not delegation. A worker that hits an
    architecture choice, a security boundary, or a second failed fix spawns
    agents/advisor.md. The guard ignores any model it asks for and places it
    one rung up. Guidance comes back; ownership stays with the caller.
  • The economics are measured, not assumed. A subagent costs roughly 60k
    input tokens before its first tool call, then 2-4k per call. Under ten calls
    or eighty edited lines, doing it inline is cheaper on any model. The numbers
    and the method are in
    tools/tokflow/AUDIT-2026-09-02.md.
  • Lean agent types by default. haiku-scout, sonnet-implementer and
    opus-owner carry restricted tool sets and cost about a third of a
    general-purpose spawn.

Tools

Zero-dependency, one directory each, each with its own README.

Tool What it does
gates Mechanical checks over the harness's own docs, hooks and skills: link rot, hook wiring, declared-vs-actual counts. Runs --staged in pre-commit.
prove Automates "a check never observed failing has been run, not verified": breaks the watched thing, confirms red, restores, confirms green.
spend Token and dollar ledger from local transcripts, per session and per day.
tokflow Transcript miner behind the token audit: where the fixed cost per turn actually goes.
recall One search across every institutional-memory store on the machine.
skillfind Finds any skill on the machine, including the ones no session can see.
fleet Live board of running Claude Code sessions.
gitradar Status board of every git repo on the machine: dirty trees, unpushed commits, stale branches.
cronwatch Health board for Windows Task Scheduler jobs: last run, last result, next due.
envdoctor Read-only checkup of secrets wiring. Reports names and locations, never values.
procledger Every process an agent starts gets a PID entry; cleanup is by PID, never by name.
errorlog Two daily error logs: one harvests itself from transcripts, one you type into.
deskclaw Read-only eye on the Windows desktop for native apps and dialogs; a hand only when a human arms it. Redacts before anything reaches a transcript.
ears Hears any audio or video file and returns a transcript.
mouth Minimal Windows text-to-speech so a long job can say it finished.
harness-sync Generates AGENTS.md and GEMINI.md from CLAUDE.md and reports parity across the three harnesses.

Scheduled jobs

Job Cadence What it does
meditation nightly The reflection loop. See the ladder.
fleet-briefing daily, 7am Collects overnight deploys, CI, revenue and error-tracker issues into one HTML board and emails it.
errorlog daily, before meditation Harvests errors out of the day's transcripts so the meditation session can read them.
harness-audit weekly A headless agent reads the harness itself in an isolated worktree and reports drift: registered hooks that do not dispatch, docs that lie, counts that are wrong.
deploy-sentinel every 30 min Polls deploy states and CI, opens an incident exactly once per new failure.
costclaw-watchdog nightly Diffs cumulative spend against last night, alerts on the incident shapes that once cost a real bill, renders a 30-night trend.
launch-board on demand Local web console: per-project launch readiness with re-check buttons.
harness-health.ps1 on demand Read-only check of hooks, MCP servers and plugins after any settings change.

Subagents

Agent Model Role
haiku-scout Haiku Mechanical lookups: file searches, symbol hunting, inventory tables, git history.
sonnet-implementer Sonnet Feature slices and refactors within a defined scope. Given files, acceptance criteria and a verify command.
opus-owner Opus A large or risky task the main loop has scoped. May delegate downward.
advisor one rung above the caller One focused decision. Read-only, guidance only, never capped: a blocked consultation becomes a guess, and a guess costs more than the advice.
security-reviewer Opus Read-only review of anything touching auth, billing, secrets, webhooks or database access. Findings only, never edits.

The meditation ladder

The part people ask about most. A scheduled session runs every morning,
reflects on recent work, and appends dated observations. Ideas climb a ladder:

observation  →  fact (memory)  →  rule (CLAUDE.md)  →  trait (SOUL.md)

Each rung has explicit graduation gates. A rule needs three or more signals
across two or more distinct sessions, with signals older than thirty days
counting half. Every promotion cites the dated evidence that earned it. The
ladder runs both ways: one contradiction is recorded, two demote. Failure
lessons are written as evidence ("when X broke, Y fixed it"), not commands, so
a hostile input cannot become a standing rule in one session.

The design goal is that refusing a promotion is the normal outcome. A gate that
has never once refused anything is not a gate. The gates, the demotion path
and the write rails for SOUL.md are in
meditations/MEDITATIONS.md.

Layout

Path What it is
CLAUDE.md The global working agreement, loaded into every session. Generated from a single source so Claude Code, Codex and Antigravity read the same text.
SOUL.md Who the agent is. Read before CLAUDE.md. Template here, see below.
RTK.md Notes for rtk, a Rust CLI proxy that compresses shell output 60 to 90 percent via a hook.
settings.json Hook wiring, permissions, env. Secrets live in a separate untracked file it points at.
hooks/ The guards above, their probes under tests/, and the Codex and Antigravity adapters.
git-hooks/ The global pre-commit chain (core.hooksPath): secret scan, staged doc gates, Python lint and dead-code gate.
tools/ The tools above.
scripts/ The scheduled jobs above.
agents/ The subagent definitions above.
docs/ Reference docs CLAUDE.md points at, plus reddit-claude-setup-share.md, a guided tour written to be pasted into Claude Code and adapted to your project.
meditations/ The nightly loop and the promotion ladder. Templates here, see below.

Stealing pieces

Do not clone this expecting a turnkey install. Paths are Windows and specific
to one machine. The useful move is taking one piece at a time.

One guard. Copy the file, then register it. A PreToolUse hook that prints a
deny reason to stderr and exits 2 blocks the call and hands the model the
reason:

{
  "hooks": {
    "PreToolUse": [
      {
        "matcher": "Bash|PowerShell",
        "hooks": [
          { "type": "command", "command": "node \"C:/path/to/hooks/process-kill-guard.cjs\"" }
        ]
      }
    ]
  }
}

The pre-commit chain. Point git at the directory once and every repo on
the machine gets the secret scan:

git config --global core.hooksPath /path/to/git-hooks

The ladder. meditations/MEDITATIONS.md is self-contained. Start with an
empty CANDIDATES.md, run the nightly prompt in scripts/meditation/, and
refuse the first few promotions on purpose to see the gates work.

docs/reddit-claude-setup-share.md walks
the whole setup with a "steal this pattern" line per item.

What is templated, what is not here

SOUL.md and everything under meditations/ are structural templates in this
mirror. The real files accumulate personal and business context that stays
private. The mechanism, the gates and the write rails are here unchanged.

Not in the mirror:

  • projects/, the agent's memory store.
  • skills/, personal skill definitions, some of them voice and identity
    material.
  • The profile block of the private CLAUDE.md (memory wiring, project map, a
    remote devbox) and the autoMode trust-boundary block of settings.json.
    Both map my machines and business context.
  • .secrets.env and anything else untracked.
  • The scheduled jobs that drive product repos and the autonomous company that
    runs on top of this harness. Different repos, private.

Security

The private repo never contained credentials. Before each sync this mirror is
swept file by file for key shapes, bearer tokens, credentialed URLs,
key=value secrets, emails and phone numbers, and the sweep prints the file
count beside its verdict so a clean result on zero files cannot pass as clean.
The last sync scanned 156 files; the only hits were fake keys inside
tools/deskclaw/tests/, which exist to prove the redaction works.

If you find something that should not be here, see SECURITY.md.

Contributing, license, support

Issues and pull requests are welcome, especially incident reports of the form
"this guard let X through" with a probe that reproduces it. See
CONTRIBUTING.md. Changes to the mirror are listed in
CHANGELOG.md.

MIT. If these tools save you time:

Sponsor on GitHub
Buy Me a Coffee

Reviews (0)

No results found