mnemosyne

mcp
Guvenlik Denetimi
Uyari
Health Uyari
  • License — License: Apache-2.0
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Gecti
  • Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Gecti
  • Permissions — No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

Version control for AI agent memory. Commit, branch, merge, blame and time-travel what your agents know.

README.md

Mnemosyne

Version control for AI agent memory.

Licence: Apache 2.0
 CI
 crates.io
 PyPI
 PyPI downloads
 Docs
 Status: early development

A support agent is told a past answer was wrong. It runs bisect to find the commit where the belief entered, blame to trace it to a misread billing note, then merges in a parallel branch that had the right answer, resolving a real conflict. The history keeps every line of it.

An AI agent builds up memory as it works: facts it learns, decisions it makes. Frameworks store that as state it overwrites as it goes. Mnemosyne gives agent memory what Git gives code: commits, branches, merge, blame and bisect. A Rust core, a mnem CLI, and a Python SDK. Local, deterministic, no network, no model calls.

The problem

Your agent runs for an hour, makes forty tool calls, and updates its memory the whole way. Then it gets something wrong, and all you have is the memory as it stands right now. You cannot see when the bad fact got in, what the agent knew before it went off track, or why it believes what it believes. Version control solved exactly this for code.

What you get

Command What it does
commit · log record memory as an immutable, content-addressed snapshot with its provenance, and walk the history
show <commit> reconstruct the memory exactly as it stood at any past commit
branch · checkout · diff fork memory for the price of a pointer to explore a hypothesis in isolation
merge combine two lines of memory, surfacing conflicts as objects you inspect and resolve
blame resolve any belief to the commit (through merges) and the observation that introduced it
bisect binary-search a run for the first commit where a wrong belief appears
export · import dump the whole history as one portable JSON file, and rebuild a store from it

A Python agent loop: a Claude call, store.add with provenance, store.commit, and store.blame tracing a belief to its origin

Every correctness property above - exact reconstruction, precise bisect, accurate blame, no lost writes on merge - holds at 100% across an 80-seed, 720-run sweep, plus a separate 50,000-case fuzz test on the merge algorithm, gated in CI on every change. Full detail: the benchmark.

Quickstart

curl --proto '=https' --tlsv1.2 -LsSf https://github.com/Nabzx/mnemosyne/releases/latest/download/mnem-git-installer.sh | sh
# no Rust toolchain needed; PowerShell equivalent: mnem-git-installer.ps1 on the same release page
pip install mnem-agents     # the Python SDK; `import mnem`

Already have Rust? cargo install mnem-git builds the CLI from source instead.

Three commands, and an agent's memory is under version control:

mnem init ./agent-memory && cd ./agent-memory
mnem add customer-4821 "on the Enterprise plan" --source ticket-4821
mnem commit -m "open the case" --author agent

Now watch it catch a real mistake:

# an hour later, the agent misreads a billing note:
mnem add customer-4821 "downgraded to Pro last month" --source billing-note-8842 --step step-31
mnem commit -m "reconcile the plan tier" --author agent

# a wrong answer surfaces. find where it entered, and why:
mnem bisect --node customer-4821 --equals '"downgraded to Pro last month"'
mnem blame customer-4821

The Python SDK:

import mnem

store = mnem.init("./agent-memory")
store.add("customer-4821", "on the Enterprise plan",
          provenance=mnem.Provenance(source="ticket-4821"))
store.commit("learn the plan tier", author="support-agent")

b = store.blame("customer-4821")       # which commit set this, and why
print(b.commit[:8], b.provenance.source)

Shell completions:

mnem completions zsh > ~/.zfunc/_mnem     # or bash, fish, powershell, elvish

More worked examples, one per surface (SDK, Claude, LangGraph, MCP), live in
examples/ - the README GIF above is examples/claude_agent.py.

Coming from Git

Git mnem
git init mnem init
git add mnem add
git commit mnem commit
git log mnem log
git branch mnem branch
git checkout mnem checkout
git diff mnem diff
git merge mnem merge
git blame mnem blame, same idea
git bisect mnem bisect, same idea

What is missing on purpose, for now: push / pull / clone (the sync protocol between stores is Era 2), a staged hunk (add -p, staging is a whole node), and a text-merge conflict marker (a conflict is an object you resolve with --resolve <id>=ours|theirs|base|delete or --strategy, not an inline marker).

The GitHub for AI agents

Git made source code collaborative; GitHub made it social. As software becomes agents working in teams, they need the same stack underneath. Mnemosyne is building it in three eras:

  1. The substrate (now): single-agent versioned memory, local and deterministic.
  2. The collaboration layer: semantic merge that reasons about contradiction, a sync protocol, and a review step before a memory update lands in shared memory. Pull requests, for agent memory.
  3. The platform (1.0): the whole agent (prompt, tools, memory, policy, evals) as one versioned, signed, forkable artefact, with a registry.

The later two are the point. Everything today is 0.0.x groundwork (ADR-0010).

Status

Tag What landed
v0.0.2 the object store, commit · add · log, the Python binding
v0.0.3 branch · checkout · diff, time travel
v0.0.4 deterministic three-way merge with conflict objects
v0.0.5 blame · bisect, the provenance index
v0.0.6 an MCP server and a LangGraph adapter, so an agent uses mnem as its memory
v0.0.7 a benchmark, a docs site, the on-disk format frozen; published to crates.io and PyPI
v0.0.8 prebuilt mnem binaries (no Rust toolchain needed), export/import in the CLI and the Python SDK

Era 1, the substrate, is complete. Era 2, the collaboration layer, is next. See ROADMAP.md and CHANGELOG.md.

How it works

  • mnem-store (Rust): the object model, the content-addressed store, the commit graph. Never touches the network or a model; a cargo deny check enforces it.
  • mnem: the CLI (crate mnem-git, Rust) and the SDK (mnem-agents, Python) built on the core.
  • The store is a .mnem/ directory, format specified in docs/format/ and frozen for 0.0.x at format_version 1. Decisions live in docs/adr/.
  • Plug into an agent: an MCP server (mnem-mcp, works with Claude Desktop and Claude Code), a LangGraph BaseStore (mnem-langgraph), an OpenAI Agents SDK Session (mnem-openai-agents), or a CrewAI StorageBackend (mnem-crewai).

Prior work

A wave of 2026 research points at this idea (Git4Data, GitOfThoughts, StateFuse, MemTX, LatticeMind), each a paper or a prototype. One finding is worth stating plainly: versioned memory does not make an agent give better answers. What it gives you is history, audit, and safe merging. That is the whole pitch, and it is enough. The benchmark has the numbers, why I built this has the longer version, and why agent memory needs version control makes the general case.

Want the same story worked end to end, one command at a time, with real commit ids instead of a compressed summary? Debugging a poisoned agent with bisect narrates exactly what the GIF above is doing.

FAQ

Why not a vector store or RAG? A different axis. Retrieval ranks by similarity; mnem versions and audits exact state. They compose rather than compete: searching inside a mnem-versioned memory is a reasonable future direction, not something this replaces.

Why not mem0, Zep, or a memory-layer product? Those manage what an agent remembers: extraction, summarisation, retrieval. mnem manages the history of whatever memory representation you already have, closer to the substrate under those tools than a competitor to them.

Is mnem the only one doing git-shaped things for agent state? No, and it isn't the first. docs/comparisons.md checks Letta, Memoria and ByteRover CLI against their actual source, not their marketing pages, and says plainly where each one is ahead and where mnem's claim differs.

Why not LangGraph's own checkpointer? A checkpoint resumes a run. There is no blame, bisect, merge, or a long-lived branch model. mnem-langgraph's MnemosyneStore targets BaseStore (long-term memory), not BaseCheckpointSaver.

Why not just append to a JSONL file? That is Baseline B in the benchmark. It answers "what was the state at step t", but not "which observation set this" or "merge two agents' memories, surfacing the conflicts". The benchmark's audit-query table has the full comparison.

Does this make my agent smarter? No. See "Prior work" above: the pitch is history, audit, and safe merging, not accuracy.

Contributing & licence

Developed by a single maintainer; issues and feedback welcome, but not open to external pull requests at this stage (CONTRIBUTING.md). Apache 2.0 (LICENSE).

Yorumlar (0)

Sonuc bulunamadi