agent-flow
Health Uyari
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 6 GitHub stars
Code Uyari
- process.env — Environment variable access in bin/agent-flow.js
- process.env — Environment variable access in extensions/cross-harness.ts
- process.env — Environment variable access in extensions/guard.ts
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
A self-bootstrapping, self-healing agent engineering skill for any repository.
Agent Flow
Your agents trust the context you give them. Agent Flow makes sure it's still true.
Context-drift detection, guardrails & an Implement → Review → QA pipeline for AI coding agents — Claude Code, Codex CLI, Gemini CLI, Cursor, Copilot, Windsurf and any tool that reads AGENTS.md.
Agent Flow demo video: an AI coding agent runs the Implement → Review → QA pipeline, catching context drift and risky diffs.
Agent Flow is a zero-runtime-dependency CLI plus skills layer for AI coding agents — Claude Code, Codex CLI, Gemini CLI, Cursor, GitHub Copilot, Windsurf, and anything that reads AGENTS.md — that keeps context files honest, gates risky diffs, and runs Implement → Review → QA as separate, enforceable processes.
Coding agents fail quietly. The model is usually fine. What goes wrong is everything around it:
- an
AGENTS.mdthat still names a file renamed last quarter; - a "reviewer" that is just the implementer in a different hat;
- a "read-only" rule that is only a sentence in a prompt;
- a new payment dependency nobody looked at.
Agent Flow is a small, auditable layer that catches these:
| What it does | How it's enforced | |
|---|---|---|
| 🩺 Drift detection | Checks every path your context files mention (manifest and the backticked/paths in the prose) against the real filesystem. Case-exact, so it works on Windows and macOS too. |
agent-flow doctor in CI and in the pre-commit hook |
| 🛡️ Guardrails | Reviewer and QA can't write. Implementers stay in their worktree (file tools enforced; shell best-effort). No agent session — pipeline role or not — skips hooks, force-pushes, or pushes to main, and nobody writes protected paths. |
Claude Code subagent tool restrictions, Codex read-only sandbox, or (on Pi) the tool_call hook — plus the pre-commit hook everywhere. See per-harness table. |
| ⚖️ Mechanical risk | Classifies the actual diff against your protected paths and risk boundaries → reviewer tier, draft PR, human gate. | risk_classify / agent-flow classify |
| 🔁 Bounded review loop | Implement → Review → QA as separate processes. Round 3 auto-escalates to Needs Me with a decision brief. | State machine rejects illegal transitions and rounds that go backwards |
| 🔎 Risk audit | Flags new dependencies, auth/payment code, destructive data ops, outbound calls, and secrets (values never printed) against a baseline. | agent-flow audit-risk --fail-on-new |
Zero runtime dependencies. No network calls. No telemetry. Every check is plain TypeScript you can read in an afternoon.
30 seconds
Someone renamed src/users/service.ts. Nobody told AGENTS.md. No setup, no config — run it in any repo:
$ npx @drix10/agent-flow doctor
✓ context files found: AGENTS.md
✗ referenced paths exist — 1
AGENTS.md:3 src/users/service.ts → did you mean src/users/user-service.ts?
✓ no unfilled placeholders
Then an agent adds stripe and edits a payments file on a branch:
$ npx @drix10/agent-flow audit-risk --fail-on-new
✗ 2 new risk surface(s):
payment package.json stripe
dependency package.json stripe
$ npx @drix10/agent-flow classify
risk: critical reviewer: high-reasoning human approval: required 4 files vs main
- critical: src/payments/charge.ts is under protected path src/payments/
- medium: dependency manifest changed (package.json) — risk review required
✗ protected paths touched: src/payments/charge.ts
That is real output (trimmed). doctor reads every AGENTS.md, CLAUDE.md, GEMINI.md, .cursorrules, Copilot and Windsurf rules file it finds; agent-flow init adds a manifest for staleness tracking. Each check exits non-zero, so CI goes red before an agent builds on a false premise.
Install
Claude Code, Codex CLI, Gemini CLI, Cursor, Copilot, Windsurf — the primary targets — get the skills, a read-only reviewer subagent (where the harness has one), and the zero-dependency CLI:
npm install -D @drix10/agent-flow
npx @drix10/agent-flow install --harness claude # or codex | gemini | cursor | copilot | windsurf | agents
# — or just the skills, via the skills CLI:
npx skills add Drix10/agent-flow
Any other tool that reads AGENTS.md — Aider, Zed, Warp, JetBrains Junie, RooCode, Amp, opencode, goose, and more — already gets drift detection and risk classification from the CLI with no install step at all; --harness agents just adds a conventional .agents/skills/ folder on top.
install works on Windows, macOS and Linux, is idempotent, and never overwrites a skill you've edited (--force to override). Every command it needs — doctor, audit-risk, classify, state, worktree — ships in the CLI, so nothing here depends on a harness-specific extension API.
Pi additionally gets tool-level enforcement, because Pi exposes a tool_call hook the others don't:
pi install npm:@drix10/agent-flow
Everyone should add the pre-commit gate and the CI checks. In a non-Node repo (Python, Go, Rust…), install the CLI once with npm i -g @drix10/agent-flow; the hook finds it there.
npx @drix10/agent-flow hook install # protected paths, secrets, broken context refs
# .github/workflows/agent-context.yml
name: agent-context
on: [push, pull_request]
jobs:
context:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with: { node-version: 22 }
- run: npm ci
- run: npx @drix10/agent-flow doctor
# needs a committed .risk-baseline.json: review `audit-risk` once, then
# `npx @drix10/agent-flow baseline accept --all --yes` and commit it
- run: npx @drix10/agent-flow audit-risk --fail-on-new
Use
On Claude Code, Codex, Gemini CLI, Cursor and Copilot, these run as skills the agent invokes by name (or you ask for by task — "bootstrap this repo", "implement issue 42") plus the CLI commands the skill shells out to. Pi additionally wires them up as slash commands.
| Skill / CLI | Pi shortcut | What happens |
|---|---|---|
bootstrap skill + npx @drix10/agent-flow scan |
/bootstrap |
Scans read-only and proposes AGENTS.md + CONTEXT_MANIFEST.json. Every claim carries [HIGH CONFIDENCE], [INFERRED] or [NEEDS VERIFICATION]. Asks you for protected paths and risk boundaries. Writes each file only after you approve it. |
invoking-agents skill |
/implement 42 |
Runs Implementer → classify → Reviewer → QA → PR, each as a separate process. See the pipeline. |
npx @drix10/agent-flow doctor |
/doctor |
Drift report. Changes nothing. |
gardener skill |
/repair-docs |
Re-reads the code, fixes the prose, then refreshes the manifest. |
npx @drix10/agent-flow audit-risk |
/audit-risk |
New risk surfaces since the baseline. |
gardener skill |
/sync-context |
Updates DOCS_INDEX.md with stale, missing and archived docs. |
| all of the above, in order | /garden |
Full maintenance pass. |
The pipeline state lives in AGENT_STATE.md, sorted into 🔴 Needs Me, 🔵 Working and 🟢 Completed.
The pipeline
issue ─▶ worktree ─▶ Implementer ─▶ classify ─▶ Reviewer ─▶ QA ─▶ PR
.worktrees/ fast model real diff high-reasoning fast draft if critical
issue-N confined to vs manifest no write, re-runs
its worktree no shell flakes once
▲ │
└──── ≤ 2 rounds ──────────┘──▶ round 3 = 🔴 Needs Me
- Separate processes. Each role is its own process with its own
AGENT_FLOW_ROLE—claude -pon Claude Code,codex execon Codex CLI,pi -pon Pi. The Reviewer launches read-only: thereviewersubagent on Claude Code,codex exec --sandbox read-onlyon Codex,--tools read,grep,find,lson Pi. See skills/invoking-agents/SKILL.md for the exact command on each. - Artifacts only. Roles hand off
issue.md,diff.patch,classification.json,review-rN.jsonandqa-rN.json. They never hand off reasoning. - Issue text is untrusted. It is wrapped in
<untrusted_issue>. Every skill tells the model to treat it as requirements, never instructions. The guard blocks role changes and nested agent launches from inside a role. - Critical changes get a draft PR and a human. Low-risk changes can auto-merge only if you opt in (
pipeline.auto_merge_low_risk).
What is enforced, per harness
Being straight about this is the point of the project.
| Claude Code | Codex CLI | Gemini CLI | Cursor / Copilot / Windsurf | Pi | |
|---|---|---|---|---|---|
| Reviewer can't write | ✅ subagent tools: Read, Grep, Glob |
✅ codex exec --sandbox read-only |
⚠️ tool list (verify) | ❌ instruction only | ✅ --tools + guard |
| Protected paths | ✅ hook | ✅ hook | ✅ hook | ✅ hook | ✅ guard + hook |
| Round cap / transitions | ✅ npx @drix10/agent-flow state |
✅ CLI | ✅ CLI | ✅ CLI | ✅ state_update |
| Drift + risk checks | ✅ CLI | ✅ CLI | ✅ CLI | ✅ CLI | ✅ tools + CLI |
| Shell writes by read-only roles | n/a (no shell) | sandbox | — | — | ⚠️ best-effort pattern block |
The ✅ cells in "Protected paths" come from the pre-commit hook. It runs at commit time, not at edit time, and a human can bypass it with --no-verify. The guard stops agents from doing that. Full details, plus every other AGENTS.md-reading tool the CLI already works with unmodified (Aider, Zed, Warp, JetBrains Junie, RooCode, and more): docs/HARNESS-MATRIX.md.
What allowed-tools in a SKILL.md does: nothing, as far as enforcement goes. We tested it (FM-16): a Reviewer skill without write in allowed-tools still wrote a file when asked. That result is why the guard exists.
Context that stays true
Agent Flow writes AGENTS.md, the file Pi, Codex, Cursor, Copilot and most agents already load. Claude Code picks it up through a one-line CLAUDE.md (@AGENTS.md). One source of truth for every harness.
AGENTS.md always loaded · purpose, commands, paved paths, local traps (<150 lines)
<module>/AGENTS.md loaded in that module · only what differs from the root
DOCS_INDEX.md which design docs exist, which are stale
CONTEXT_MANIFEST.json every referenced path + last_verified · protected_paths · risk_boundaries
Drift gets caught three ways:
- Mechanically, by
doctorin CI and the pre-commit hook. - In review, where the Reviewer flags claims the diff made false. These are
[CONTEXT_STALE]flags, and they don't block the PR. - By the Gardener, which fixes the prose before it refreshes a timestamp. A timestamp refreshed without re-reading the code turns stale context into "fresh" context, and that is the failure this project exists to prevent.
Fix mistakes where they'll stay fixed
When an agent makes the same mistake twice, fix it at the highest level you can:
- Architecture: make the mistake impossible.
- Static analysis: a lint rule or a type.
- Hooks and guard: add it to
protected_paths, or add a pre-commit check. - Skills and context: a local trap in the module's
AGENTS.md. - Style guide: the weakest fix, because it depends on someone remembering it.
If the Implementer keeps editing your migrations, a better prompt won't stop it. Add **/migrations/** to protected_paths, and the guard and the hook will.
Honest limits
- Shell analysis is best-effort, on the harnesses that only offer pattern-matched shell checks. Pattern matching can't catch every way to write a file through an interpreter. For a hard guarantee, launch read-only roles without a shell tool at all: Claude Code's subagent
tools:list, Codex's--sandbox read-only, Pi's--tools read,grep,find,ls, or any of them inside a container. - The risk audit is heuristic. It gives you leads to review, not verdicts. It is tuned to rarely false-alarm, so it will miss things.
- The Pi guard and tools only exist in Pi. Other harnesses get the same logic through the CLI and hook, which work at commit time rather than per tool call. Real cross-harness enforcement at tool-call time needs an MCP server (ROADMAP).
- One repo at a time (FM-13). There is no model-provider fallback yet (FM-15).
All known failure modes, including the ones we found in our own code, are listed in FAILURE_MODES.md. If you find one that isn't there, open an issue. Each confirmed one becomes a test.
Security
- Zero runtime dependencies. Only
node:built-ins, so there is less supply chain to trust. - No network. ctxlint runs only if it's already installed locally. v1.0.x used
npx, which could download code; that is fixed. - Every git call uses an argv array. No shell strings.
- Every write stays inside the repo. No
.., no absolute paths, no escaping through symlinks. - Human confirmation is real. Pi shows a dialog that only a human can click. Headless writes need
AGENT_FLOW_HEADLESS_WRITES=1, set by whoever launches the process. - Guard blocks and state transitions are logged to
.agent-flow/audit.jsonl. Agents can't edit that file directly.
See SECURITY.md to audit these claims yourself.
Contributing
npm install && npm test runs 65+ tests against real git repos in temp directories, on Linux, macOS and Windows in CI, covering every install --harness target. Read CONTRIBUTING.md, then pick a failure mode.
License
MIT
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi