longgraph-skill

agent
Security Audit
Pass
Health Pass
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Community trust — 53 GitHub stars
Code Pass
  • Code scan — Scanned 3 files during light audit, no dangerous patterns found
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Long-horizon agent skill for Claude Code / Cursor / Codex / Grok — multi-task ledger loop (related or not), host-portable (re-send the prompt to continue), clean-context supervisor, verified gates. Markdown library (loop-graph · quest), not a framework.

README.md

longgraph

Long-horizon agent skill for Claude Code, Cursor, Codex & Grok.

Stop agent drift with a durable ledger, a clean-context supervisor, and verified gates.
Queue many long tasks in one loop — even unrelated ones — and keep going after a host switch
by re-sending the same prompt against the files.

Design once → compile a durable loop-graph → verify all the way to done.

GitHub stars
License: MIT
PRs welcome
Hosts: Claude Code · Grok · Cursor · Codex
Type: Claude Code skill · prompt library

English · 简体中文

Executor and clean-context supervisor loops running side by side

longgraph (longgraph-skill) is a curated Claude Code skill / agent skill and
cross-host prompt library for long-running / long-horizon agent work —
multi-hour coding, multi-milestone migrations, a queue of long tasks in one
loop
(they need not be related), and anything that outlives one context window.
It is graph engineering for agents: specialized roles (executor · supervisor ·
scout) connected through durable, inspectable files — not another orchestration
runtime. Because the scoreboard lives on disk, you can change hosts mid-run:
open the same workspace, re-send the frozen node prompt, and continue.

One durable graph, portable across hosts. For a simple self-contained goal,
use the host's normal task or goal directly; longgraph starts where durable graph
structure adds value.

Renamed from octopus. Same library; primary brand is now longgraph /
longgraph-skill / /longgraph / .longgraph/. Older posts that say
octopus-skill, /octopus, or .octopus/ still work via GitHub repo redirect,
install legacy symlinks, and in-flight run-dir alias (see Install
and Migration).

Evidence

These are not one-shot demos. longgraph is a Markdown skill / prompt library
(not an orchestration runtime). The table mixes checkable public Git, a
function-only redacted multi-day pattern, and synthetic pedagogy.

Case What a reader can verify Kind
Self-iteration of this skill 87 public commits across ~14 calendar days (2026-07-19 → 2026-08-02), 74 files, method rules written back into the library (no wake edge, gate-wait backlog, blocked≠parked, bounded live edges, authoring≠runtime) Public Git facts — fixed anchor 6efcb7f
Multi-day control-plane pattern Multi-day wall-clock, tens of rounds, many directives: durable ledger, clean-context supervisor overturns self-reported evidence, non-skippable gates, blocked-work lane, owner A/B/C — functions only, no private payload Redacted real-run pattern
migrate-blob-storage Multi-milestone ledger: pilot → cohort, forced convergence, supervisor overturns self-reported evidence, non-skippable gate + blocked-work lane Synthetic pedagogy (fictional app)
add-tests-to-cli Smallest full run: three rounds, register-then-defer, clean-context supervisor intent Synthetic pedagogy (fictional CLI)

How to read the clock. The self-iteration window’s ~14 days / ~340 hours is
project wall-clock (first public commit → frozen anchor), not continuous model
execution and not a claim of unattended production autonomy. Re-check Git with the
commands in the self-iteration case.
The redacted multi-day card uses coarse buckets only and is not private-Git
re-checkable — see its evidence boundary.

Publication rules for future cases:
public / private boundary.

When to use this

Reach for longgraph when you need any of:

  • A long-horizon agent that keeps working after context compaction / session resets
  • A durable task ledger (single scoreboard) instead of chat-memory progress
  • Several long tasks in one loop — a continuous queue, even when items are unrelated
  • Host-portable continuity — switch Claude Code ↔ Cursor ↔ Codex ↔ Grok mid-run by re-sending the prompt against the same files
  • An independent clean-context supervisor — not the same agent grading itself
  • Verified done: acceptance gates re-run against real output, not self-reported “done”
  • Multi-milestone work with non-skippable gates and explicit owner red lines
  • A Markdown skill / prompt library that works across Claude Code · Cursor · Codex · Grok

When not to use this

  • One-shot edits, small PR-sized tasks, or anything that fits a single clean session
  • You want a runtime framework (LangGraph, CrewAI, AutoGen, custom agent server)
  • You only need a single short prompt with no ledger, gates, or independent review

How it compares

Approach Runtime / server? Independent verifier Durable scoreboard Multi-task queue + mid-run host switch
LangGraph / CrewAI / AutoGen Yes You build it Usually yes Framework-bound; often one deployment stack
One mega-prompt / single skill No No (self-check) Weak (chat memory) Weak — progress dies with the session
longgraph (this repo) No — Markdown only Yes (supervisor node) Yes (ledger.md) Yes — files are the run; re-send the prompt

Also called / related searches: long-running agent skill, prevent agent drift,
multi-task agent loop, switch AI coding host mid-task, Claude Code multi-agent
supervisor
, agent ledger, loop skill, graph engineering for
agents
, clean-context review.

Why longgraph

Long-running agents tend to drift in predictable ways: scope expands, “done”
becomes self-reported, tests stop proving the real path, and early decisions
disappear from context. longgraph moves the safeguards outside the model’s memory:

  • Verified, not merely written — acceptance gates are rerun against real output.
  • Durable state — the ledger survives context loss and remains the single scoreboard.
  • Many long tasks, one loop — the ledger is a continuous queue; items can be
    independent (migrations, test debt, docs, gates) without forcing one mega-goal.
  • Host-portable — progress is files under .longgraph/<date-slug>/, not chat
    history. Point another host at the same workspace, re-send the compiled node
    prompt, and pick up the next open ledger item.
  • Clean-context review — an independent supervisor can catch drift the executor cannot see.
  • Forced convergence — growth is periodically stopped, measured, and simplified.
  • Low-friction owner decisions — genuine owner-only calls arrive as a short
    recommended A/B/C choice, not a technical homework assignment.

It is Markdown, not an orchestration framework: no application runtime, server,
or vendor lock-in. Install it as a Claude Code plugin or symlink the skills
into Cursor / Codex / Grok.

Multi-task loops & switching hosts

One loop is a queue, not a single story. Each round still does one ledger item
end-to-end (implement → verify → record), but the ledger can hold many long items
at once — related milestones or unrelated backlog (the gate-wait backlog pattern
is the extreme case: useful work with no dependency on the item under audit). You
do not need a new graph every time the next long task is about something else.

The host is swappable; the files are not. A compiled loop-graph run freezes
prompts and state under .longgraph/<date-slug>/. To continue elsewhere:

  1. Use a workspace that can see those files (and the project).
  2. Re-send the same frozen executor (and, if used, supervisor) prompt on the new host.
  3. The node reads ledger.md / directives.md and continues from the next open item.

You are not exporting chat transcripts. Invocation syntax still follows each host’s
dialect (per-host references) — only the progress is portable.

Is longgraph the right tool?

Your task shape Choose What you get
One self-contained goal that fits a normal task/session Use the host's ordinary task or goal directly No longgraph wrapper or extra prompt layer
Many rounds, durable state, non-skippable gates, owner boundaries, host switching, or independent verification longgraph / loop-graph An executor loop plus a clean-context supervisor, coordinated through durable files

Rule of thumb: if you do not need the graph, do not use longgraph.

Quick start

Claude Code

Install the plugin from the marketplace:

/plugin marketplace add levi-qiao/longgraph-skill
/plugin install longgraph@longgraph-skill

Codex or Cursor

Install the library and symlink it into supported hosts:

curl -fsSL https://raw.githubusercontent.com/levi-qiao/longgraph-skill/main/install.sh | sh

To install from a local clone, run ./install.sh from the repository root.

Design a run

Invoke /longgraph. It detects Codex or Claude Code, inspects the workspace, and asks
only for unresolved owner decisions before compiling the run. Choose direct creation
to have it start both same-host runtime nodes, or prompts-only for manual/cross-host
launch. You can also invoke loop-graph directly.

Authoring and runtime stay separate: the author skill compiles the work but never
executes it. Generated nodes follow their frozen run contract under
.longgraph/<date-slug>/.

Renamed from octopus

Was (legacy / still accepted) Now (primary)
Product octopus, repo octopus-skill longgraph, repo longgraph-skill
Slash /octopus /longgraph (install still symlinks /octopus → same tree)
Plugin octopus@octopus-skill longgraph@longgraph-skill
Run dir .octopus/<date-slug>/ .longgraph/<date-slug>/ (continue in-flight .octopus/ runs in place)
Contract octopus.loop-graph.* longgraph.loop-graph.* (old headers on existing files still mean the same family)

GitHub renames redirect old clone/curl URLs (…/octopus-skill/……/longgraph-skill/…).
Re-run install.sh or reinstall the plugin once so primary names win on disk.

How the graph works

Role Responsibility Durable edge
Executor Works one ledger item, verifies it in the same round, then records the result Reads and writes ledger.md
Supervisor Re-verifies from its own separate context, checkpoints passing work, and corrects drift Reads the ledger; writes only the directives edge (live queue + cold archive)
Scout (optional) Researches a bounded question away from the critical path Writes a findings file read only on reference

The load-bearing rule is one node = one prompt + one single-writer edge.
The ledger has exactly one writer. The supervisor never shares the executor’s
context, never edits its scoreboard, and steers only through the one-way
directives edge.

For the rationale behind every constraint, read
the methodology. For the node and edge model, see
the loop-graph model.

Host compatibility

Host loop-graph execution
Codex ✅ detects the host and directly creates both runtime nodes
Claude Code ✅ detects the host and directly creates two background runtime sessions when capability checks pass
Grok prompts-only execution target
Cursor prompts-only execution target
shell / cron prompts-only execution target

Authoritative syntax, pacing, context carry, and hooks live in separate
per-host references, so authoring loads only the selected host. Mid-run host switches reuse the same
durable run directory; only how you start each tick changes.

Repository map

Path Purpose
Root SKILL.md /longgraph entrypoint; checks fit and delegates authoring to loop-graph
Loop-graph author Generates executor, supervisor, ledger, and directive artifacts
lib/ Shared methodology
Host references One independently loaded owner for each host's runtime facts
Worked examples Public-Git self-iteration plus fictional ledgers showing gates in action
Public / private boundary What may enter the public tree vs stay project-local

Governance

longgraph applies its own anti-bloat rule to the library: no prompt enters
without a real run that proved its value.
Curated and opinionated beats
comprehensive.

Contributions are welcome. Start with the contribution guide.

Credits

The loop-graph skill grew from real runs and community input. A
public-Git self-iteration case
records how the method was hardened into this library. Special thanks to
@BrightProgrammer7 for the
migrate-blob-storage example and the discussions that sharpened milestone
gates and the node/edge vocabulary.

License

MIT © 2026 levi-qiao

Reviews (0)

No results found