kb-arena

mcp
Security Audit
Warn
Health Warn
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 9 GitHub stars
Code Pass
  • Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Benchmark 19 retrieval architectures (vector, contextual, QnA, knowledge graph, LightRAG, hybrid, RAPTOR, PageIndex, BM25, metadata-filtered, temporal, rerank, HyDE, multi-query, late-interaction, SPLADE, agentic, plus two quantum-inspired rerankers) on your own docs. Citable evidence bundles, bootstrap CIs, significance tests.

README.md

KB Arena

Compare retrieval architectures on your own documentation and choose with evidence.

PyPI
Python
CI
License
DOI

KB Arena runs the same corpus and question set through 19 built-in retrieval strategies,
lexical, dense, graph, hybrid, hierarchical, reranked, and more. It records retrieval quality,
answer quality, latency, cost, and run artifacts, so you can decide which design fits your data
before you build on it.

Use it before you commit to a retrieval architecture, or as a regression lab after your corpus,
chunking, model, or index changes.

Historical KB Arena retrieval result

Explore a checked result

The packaged demo uses precomputed AWS Compute results. It needs no API key, Docker service, or
Neo4j instance.

pip install kb-arena
kb-arena demo

Open the URL printed by the command. The demo ships 8 benchmark result files for the
aws-compute corpus, one per strategy, so the benchmark table, the per-tier breakdown, the
source drill-down and the strategy comparison all have real numbers behind them.

The Retriever Lab and the spread across repeated runs are empty until you produce them. They
read run directories, and the package bundles none. Run kb-arena retriever-lab and
kb-arena benchmark --runs 3 against a corpus to fill them.

Point it at your own documents

Ingest your files, build the strategy your provider budget allows, and score retrieval for free
through BM25, the one built-in strategy that needs no key. Run it on your
documents
below has the setup for local and hosted providers, and
the own-corpus walkthrough
runs ingest through an exported evidence bundle with the real console output from each step.

Four ways to run a comparison

  • CLI runs kb-arena, with commands from ingest through evidence. See the command
    reference
    .
  • HTTP API starts with kb-arena serve, which runs the FastAPI server behind the dashboard.
    See the HTTP reference
    for every route and its auth gate.
  • MCP server installs with pip install 'kb-arena[mcp]', then runs with python3 -m kb_arena.mcp.server. It exposes corpus, strategy, benchmark, and evidence tools over stdio to
    an MCP client such as Claude Code or Codex. See
    the server module.
  • GitHub Action
    retrieval-regression-gate
    ingests a corpus, builds an index, and fails a pull request when a named metric drops past a
    threshold. See
    the example workflow.

What KB Arena helps decide

Question Comparison
Do exact terms matter more than semantic similarity? BM25 against dense retrieval
Does document context improve chunk retrieval? Naive against contextual vector
Do cross-document relationships justify a graph? Dense against graph and hybrid
Does hierarchy help on broad questions? Dense against RAPTOR and PageIndex
Is a reranker worth its latency and cost? Dense against reranked dense
Is a proposed method meaningfully different? Paired scores, confidence intervals, and rank overlap

KB Arena does not give a universal leaderboard. A result applies to the corpus, questions,
ground truth, configuration, and models named in that run.

Evidence included in this repository

The tracked Retriever Lab run 855aac4e is a historical, reproducible example:

  • Corpus: aws-compute, three documents and 1,549 words
  • Run date: 2026-04-26
  • Questions: 75 across five tiers
  • Cutoff: top 5 chunks
  • Chunk-level labels: 35 questions have labels, and 40 do not
  • Scope: eight strategies from the version available on the run date
Strategy Recall@5 MRR NDCG@5
Contextual Vector 0.355 0.433 0.388
Naive Vector 0.352 0.414 0.367
RAPTOR 0.352 0.414 0.367
BM25 0.275 0.352 0.278
Hybrid 0.080 0.093 0.086
PageIndex 0.061 0.111 0.076

Historical retrieval metrics from run 855aac4e

Source: tracked report and
run artifact.

These numbers show the report format and calculation path. The corpus is too small, and
its chunk labels are too incomplete, to support a general winner claim. Q&A Pairs and Knowledge
Graph from that run have zero chunk-level scores because their retrieved identifiers did not map
to the available labels. Those zeroes reflect gaps in the current evaluation, not evidence that the
methods cannot retrieve useful context.

A larger public corpus for method development

The repository includes a deterministic NIST SP 800-171 Revision 3 corpus built from the
official publication. It has 130 control documents and 80 questions across direct,
paraphrased, scenario, boundary, and multi-control categories. Each question maps to source control
sections, with 48 development, 12 validation, and 20 holdout items.

The question set is not a human-approved benchmark but a machine-generated draft. Do not publish a
strategy winner from it until a qualified reviewer checks the questions, answers, constraints, and
holdout isolation. See the corpus notes,
source manifest, and
evaluation method.

Run it on your documents

Local models with Ollama

Install and start Ollama, then pull both generation and embedding models:

ollama pull llama3.1:8b
ollama pull nomic-embed-text

export KB_ARENA_LLM_PROVIDER=ollama
export KB_ARENA_EMBEDDING_PROVIDER=ollama

Create a corpus, place files in raw/, and run the pipeline:

pip install 'kb-arena[all-formats]'
kb-arena init-corpus my-docs
cp -R /path/to/docs/. datasets/my-docs/raw/
kb-arena run --corpus my-docs --skip-graph

Create and ingest an example documentation corpus

run ingests a populated raw/ directory automatically. Remove --skip-graph after starting
Neo4j when you want graph and hybrid comparisons.

Hosted generation and embeddings

The generation and embedding providers are independent:

export KB_ARENA_LLM_PROVIDER=anthropic
export KB_ARENA_ANTHROPIC_API_KEY=...
export KB_ARENA_EMBEDDING_PROVIDER=openai
export KB_ARENA_OPENAI_API_KEY=...

kb-arena run --corpus my-docs

Supported embedding providers are OpenAI, Voyage, Cohere, Gemini, local BGE, and Ollama. See the
getting-started guide for formats, Neo4j setup, checkpoints, and provider
configuration.

Evaluation paths

Use the retrieval-only path when you need to isolate the index and ranking behavior:

kb-arena label-chunks --corpus my-docs
kb-arena retriever-lab --corpus my-docs --top-k 5

It reports Recall@k, Precision@k, Hit@k, MRR, NDCG@k, MAP, R-Precision, bpref,
bootstrap confidence intervals, and per-tier breakdowns.

Use the full benchmark for generated-answer scoring, source attribution, latency,
cost, and reliability:

kb-arena benchmark --corpus my-docs --top-k 5
kb-arena report --corpus my-docs --format markdown

Use optimization after you define a development split. Keep a separate holdout for the published
comparison:

kb-arena optimize \
  --corpus my-docs \
  --split development \
  --strategies bm25,naive_vector,contextual_vector,raptor \
  --top-ks 3,5,10 \
  --metric ndcg

After you choose the configuration, run the public comparison with
kb-arena benchmark --split holdout.

Optimization remains retrieval-only. QnA Pairs and RAPTOR reuse their prebuilt indexes and sweep
top-k only. They do not regenerate pairs or summaries during a search.

Read the evaluation method before interpreting small score differences or
synthetic question sets.

Strategy catalog

The catalog holds 19 strategies. The default all benchmark runs 9 of them and leaves
out 10: LightRAG, Metadata Filtered, Temporal, Rerank Vector, SQR, HyDE, Multi-Query, Late
Interaction, SPLADE and Agentic. Each one is out for one of two reasons the catalog records.

Four need an optional dependency group a plain install does not carry: Rerank Vector,
SQR, Late Interaction and SPLADE. Rerank Vector installs per backend, so read its entry
in the strategy catalog rather than one command.

Seven are marked experimental, which means they run but carry no general performance
claim: LightRAG, Metadata Filtered, Temporal, SQR, HyDE, Multi-Query and Agentic. SQR is
in both lists.

The API reports loaded and unavailable strategies at GET /strategies.

Strategy Architecture Default Notes
Naive Vector Dense Yes Chunk, embed, cosine retrieval
Contextual Vector Dense Yes Adds parent context before embedding
Q&A Pairs Generated index Yes Creates likely questions at index time
Knowledge Graph Graph Yes Retrieves through Neo4j entities and relationships
LightRAG Experimental No Local entity neighborhood plus a global community summary, needs Neo4j
Hybrid Hybrid Yes Routes and fuses vector and graph results with RRF
RAPTOR Hierarchical Yes Retrieves chunks and recursive summaries
PageIndex Hierarchical Yes Uses document structure and LLM tree traversal
BM25 Lexical Yes Keyless keyword baseline
Metadata Filtered Access-aware dense No Applies a tag, owner, classification, and doc ID filter inside retrieval
Temporal Version-aware dense No Prefers the newest document version and supports an as-of date
Rerank Vector Reranked dense No The BGE backend uses kb-arena[rerank].
QISS Experimental Yes Pure NumPy fidelity reranker over dense candidates
SQR Experimental No Qiskit Aer SWAP-test reranker, install kb-arena[quantum]
HyDE Experimental No Embeds an LLM-written hypothetical answer instead of the question
Multi-Query Experimental No Asks the LLM for several sub-queries and fuses their results with RRF
Late Interaction Token-level dense No ColBERT-style MaxSim reranker, install kb-arena[late-interaction]
SPLADE Learned sparse No Term-weight expansion over its own sparse index, install kb-arena[splade]
Agentic Experimental No Retrieve-judge-refine loop under a hard iteration and call budget

See strategy details and the
plugin guide.

Data and method limits

  • Auto-generated questions help expand coverage, but production queries and human review give
    stronger deployment evidence.
  • An LLM judge can introduce model bias. Use a different judge family, keep the prompts and model
    versions, and inspect disagreements.
  • Architecture-native indexes do not always return the same chunk identifiers. Validate qrel
    mappings before comparing retrieval metrics.
  • Cost and latency depend on provider, model, cache state, hardware, concurrency, and region.
  • Tune on development data. Publish results only from a sealed holdout.
  • Quantum strategies are experiments. The AWS sample does not show a Recall@5 gain over the dense
    baseline.

The method guide defines the evidence that belongs with a public result.

Project references

Checking a result

A benchmark number is only worth as much as what sits beside it. These three
commands are how KB Arena says what a number is, and what it is not.

# Repeat a benchmark, so a difference can be told from noise
kb-arena benchmark --corpus my-docs --runs 3 --seed 7

# Read the spread across those repeats
kb-arena variance --corpus my-docs

# Write the record that travels with a run
kb-arena evidence --corpus my-docs --run-id <id>

Spread across repeats

variance groups runs by experiment and by build. Two runs from different
commits measured different code, so it lists their values and reports no mean.
Two runs is a range a reader can misread as a bound, so it says when a row rests
on fewer than three.

evidence writes the command, the package version, the commit, the platform and
the seed beside the result. It also writes whether the run may be cited:

"citable": false,
"why_not_citable": "publishable is true only when every scored question is human-reviewed."

What a run says about itself

kb-arena evidence --check <path> reads a bundle back. It refuses one that calls
itself citable while its own review says otherwise, and one that is not citable
and does not say why.

The committed example lives at
results/run_422209dd.
It needs no API key to repeat.

Public datasets

kb-arena datasets                                  # what is available, and its terms
kb-arena datasets --name crag --destination ~/data/crag

An adapter records who made the data, which revision, under what licence, and
what KB Arena did to it before scoring. A moving revision such as latest is
refused, because a run against one cannot be repeated.

A dataset whose licence forbids redistribution is never bundled. CRAG is CC BY-NC
4.0, so its adapter ships nothing, fetches nothing for you, and refuses to write
inside the checkout. See dataset adapters.

The sealed holdout

kb-arena optimize refuses --split holdout without --confirm-holdout. Every
run that reads holdout questions appends to results/holdout_uses.jsonl, judged
by the questions it read rather than by the split it named. kb-arena holdout-uses prints that ledger.

Development

git clone https://github.com/xmpuspus/kb-arena
cd kb-arena
pip install -e '.[dev]'
ruff check .
ruff format --check .
pytest tests/ -q --ignore=tests/live

The frontend uses Next.js 16 and needs Node.js 20.9 or later:

cd web
npm ci
npm run lint
npm run build

Citation

The canonical citation metadata is in CITATION.cff. GitHub can export it through
the repository's Cite this repository action. The archived software record is available through
the DOI badge above.

License

MIT

Reviews (0)

No results found