Veracium

mcp
Security Audit
Fail
Health Warn
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Fail
  • eval() — Dynamic code execution via eval() in bench/run_bench.py
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Provenance-aware memory for AI agents.

README.md

Veracium

tests
PyPI
Python
license

Veracium is a provenance-aware memory plug-in for agentic systems
durable, per-user memory that resists the injection and confabulation failures
that plague naive agent memory. Provenance means every fact tracks who said
it
: a claim from an email your agent merely read can never become a "fact" it
asserts. It remembers what the user said, past interactions, and what worked —
and it remembers where each of those came from.

Veracium is the production distillation of an evaluation-driven research project
(agent-memory): every design choice below traces to a measured finding, and the
research's synthetic-corpus harness is reused as the regression suite.

Research: the evaluation instrument behind those findings — a longitudinal
benchmark for agent memory — is described in Q. Spencer, "Ground Truth First:
A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover
in Memory-Architecture Rankings"
(arXiv:2607.21962, 2026).

Why it's shaped this way

  • Typed graph + dated episodes are the store of record. Entity facts live as
    relational edges (with unforgeable provenance); interaction history lives as
    dated episodes. A curated "wiki" view is compiled from them and cached — never
    the source of truth. (The layered design won on both short and 9-week horizons;
    flat stores each failed one regime.)
  • Supersession, never erasure. Functional facts (preference, employer,
    deadline) keep one current value with the prior value retained as history —
    "what did X used to be?" stays answerable. (The category commercial memory
    systems handle worst; Veracium's strongest.)
  • Representation is a security control. Third-party claims (received email,
    external docs) are quarantined structurally — stored as third_party_claim
    edges with the claimant as subject, never as user facts. Content-type quarantine
    catches obligation/debt/renewal claims regardless of how plausible they look.
    (Held against a full plausibility ladder incl. contact-impersonation.)
  • Bring your own model. Veracium never owns your API keys or model choice; it
    calls a Complete callable you supply. A reference Anthropic provider ships in
    the box.
  • Embedded by default. Zero external services: one SQLite file. Swap in
    Neo4j/Postgres later via the Store interface.

Install

pip install "veracium[anthropic]"   # core + the reference LLM provider

Extras: [mcp] adds the MCP server, [dev] adds pytest. The core alone depends
only on pydantic. To work from source instead:

git clone https://github.com/veracium-ai/Veracium.git && cd Veracium
pip install -e ".[anthropic,dev]"

Links: docs · veracium.ai · PyPI

Use (library)

from veracium import Memory, EvidenceAuthor
from veracium.llm.anthropic import AnthropicComplete

mem = Memory(llm=AnthropicComplete())   # or pass your own Complete callable

# Remember interactions. `author` is the trust-critical input.
mem.remember("alice", "USER: I'm vegetarian and have a dog named Ollie.")
mem.remember("alice", "From billing@scam: you owe $900.",
             author=EvidenceAuthor.THIRD_PARTY, event_type="email")

# Recall grounded, provenance-flagged context for a prompt.
ctx = mem.recall("alice", "suggest a lunch spot")
print(ctx.context)   # states the vegetarian constraint; the $900 "claim" is
                     # rendered under a never-assert flag, not as a fact.

No Anthropic API key? AnthropicComplete is just a convenience — Veracium calls any
Complete callable you supply. To run without SDK/key setup, wrap a client you
already have; examples/claude_cli_provider.py wraps the claude CLI as a
drop-in provider (from claude_cli_provider import ClaudeCLIComplete), and
examples/openai_provider.py wraps any OpenAI-compatible chat-completions API
(OpenAI itself, vLLM, Ollama's /v1 endpoint) via OpenAIComplete — point it
at a local server with OpenAIComplete(base_url=...) and override models with
whatever model name your server serves.

Use (MCP)

veracium-mcp exposes remember / recall / answer / maintain tools to any
MCP-compatible agent (Claude Desktop/Code, others) with no host-side Python. See
docs/mcp.md for the config JSON and tool reference.

Documentation

Hosted docs: veracium-ai.github.io/Veracium

  • examples/demo.ipynb — the scam-email injection demo,
    runnable end to end (open in Colab).
  • examples/langchain_memory.py — Veracium as
    the long-term memory layer of a LangChain chat app (session-keyed hybrid:
    LangChain buffers recent turns, Veracium holds durable facts with provenance
    and quarantine; your existing LangChain model powers both sides).
  • docs/concepts.md — the mental model: edges vs episodes
    vs the compiled wiki, provenance & authorship, quarantine, the abstention gate,
    lifecycle.
  • docs/recipes.md — short copy-paste examples, one per
    capability (quarantine, mixed provenance, budgeted recall, portability,
    feedback verbs, audit, local models).
  • docs/api.md — the public API: Memory, MemoryConfig,
    EvidenceAuthor, providing your own LLM callable or store.
  • docs/mcp.md — running and registering the MCP server.
  • docs/design-rationale.md — why there's no
    update()/delete(), no LLM-free extraction, no TTL purging — and what's
    genuinely on the roadmap.
  • docs/telemetry.md — the opt-in, anonymous, content-free usage statistics (off by default).
  • docs/diagnostics.md — opt-in error reporting: local-first error log, consented + redacted send.
  • ROADMAP.md · CHANGELOG.md

Status

The validated layered design is implemented, tested (44 offline tests, plus
opt-in live tiers: the acceptance eval and a real-corpus robustness harness),
and passes its own research-claim bar (5/5, 0 injection asserts). Roadmap
v0.1–v0.7 complete, plus opt-in telemetry, a self-check, consented error
reporting, and an operation audit log. See ROADMAP.md.

License

MIT

Reviews (0)

No results found