infona-oss

mcp
Security Audit
Warn
Health Warn
  • License — License: Apache-2.0
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Warn
  • fs module — File system access in .github/workflows/catalog-freshness.yml
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Messy CSV, JSON, or text → one LLM schema pass → deterministic rows into Neo4j → ask in English (Cypher).

README.md

Infona

Messy CSV, JSON, or text → one LLM schema pass → deterministic rows into Neo4j → ask in English (Cypher).

Luna (or your configured model) sees the file once and names types, attributes, relationships. Every cell maps through that schema via insert_facts. Then ask compiles English to Cypher on the populated graph. Ingest does not merge intra-file fragments — er rebuild is the second pass. Ontology extend and field-conflict handling are part of the same loop.

infona.ai (waitlist / demo) · what's free · API · agents

docs Apache-2.0 Python 3.12+ Neo4j npm @infona-ai/mcp tests

English question compiled to Cypher on Neo4j: AstraZeneca runs FLAURA2, indication NSCLC.

Ask in English → Cypher on Neo4j. The graph is the trials.csv sample. Zero-key FLAURA2 is a cached-plan replay, not live inference.

This repo is the OSS runtime (infona_client, @infona-ai/cli, @infona-ai/mcp). Paid adapters, production Clerk, Explorer, billing/entitlement, and cloud infra live in the private parent, not here. Full table: docs/BOUNDARY.md.


10-minute quickstart

Need: Docker + Node 20+ (for the infona CLI). A stranger gets a real answer with no API key.

./scripts/oss_up.sh compose-ups Neo4j + API, waits until /health reports Neo4j up, writes ~/.infona/config.json, and loads the prebuilt trials graph. After that, bare infona works. --local / --no-login are first-run / one-off flags; they do not rewrite config. Re-connect without compose: infona init --local.

Zero-key (cached-plan replay)

The prebuilt path replays a cached Cypher plan. It is not live inference.
/ask stays always-LLM Cypher whenever a real model key (or INFONA_LLM_BASE_URL) is configured.

git clone https://github.com/infona-ai/infona-oss.git && cd infona-oss
cp .env.example .env          # leave OPENROUTER_API_KEY empty / as the placeholder
npm i -g @infona-ai/cli       # or use npx @infona-ai/cli in place of infona
./scripts/oss_up.sh           # Neo4j + API + ~/.infona/config.json + prebuilt trials graph
infona ask "Which Phase 3 NSCLC trials is AstraZeneca running?" --kg trials

That question should return FLAURA2, labelled as a cached-plan replay (not live inference).

Reload the snapshot later (still no key):

./scripts/load_prebuilt_trials.sh
infona ask "Which Phase 3 NSCLC trials is AstraZeneca running?" --kg trials

Advertised bound stays 10 minutes. Measured 1 min 42 s cold on macOS 26.5.1 + Colima (Ubuntu 24.04 VM, 4 CPU, 6 GB; warm daemon, empty project, docker compose build --no-cache) from git clone to that zero-key ask. Native Linux was not measured. First-time neo4j:5-community pull is extra; 10 minutes still covers it. Timing notes: docs/quickstart-timing.md.

Placeholder keys from .env.example (sk-or-...) count as no key.
INFONA_ASK_CACHED_PLAN=1 forces replay even with a key (tests). =0 disables it.

1. Messy suppliers — URI merge

examples/suppliers-messy.csv is synthetic (Acme / Globex / Initech, fake tax IDs). No real customer data.

Schema inference needs a key (paste OPENROUTER_API_KEY=sk-or-... into .env):

infona ingest examples/suppliers-messy.csv --kg suppliers
infona er rebuild --kg suppliers

Ingest writes every row as its own Supplier fragment. er rebuild re-blocks the already-ingested graph and collapses fragment URIs (6→3). Headquarters and credit_rating then land as graph state: Austin is the current HQ (San Francisco stays stored, closed); equal-trust credit_rating stays dual-current and flagged.

Rebuilding entity resolution for suppliers…
  Supplier         6 → 3  (−3 fragments across 2 clusters)

  merge  https://graph.infona.ai/entities/Supplier/ERP-1001
         losers:     https://graph.infona.ai/entities/Supplier/CRM-4402, https://graph.infona.ai/entities/Supplier/DIR-8891
         reason:     signal-richest
         score:      1.00
         provenance: erp @ 2026-03-01T12:00:00+00:00 (source_of_truth)

  merge  https://graph.infona.ai/entities/Supplier/ERP-2001
         losers:     https://graph.infona.ai/entities/Supplier/CRM-5503
         reason:     signal-richest
         score:      1.00
         provenance: erp @ 2026-03-01T12:00:00+00:00 (source_of_truth)

  conflict  headquarters
         entity:     https://graph.infona.ai/entities/Supplier/ERP-1001
         winner:     Austin  (erp, source_of_truth, 2026-03-01T12:00:00+00:00)
         loser:      San Francisco  (directory, supplementary, 2024-06-01T00:00:00+00:00)
         reason:     authority

  unresolved  credit_rating
         entity:     https://graph.infona.ai/entities/Supplier/ERP-1001
         crm: BBB @ 2026-03-01T12:00:00+00:00 (source_of_truth)
         erp: A @ 2026-03-01T12:00:00+00:00 (source_of_truth)
         flagged: equal-trust sources — not silently guessed

Done. 3 fragments absorbed.
  • merge — three Acme name variants (and two Globex) became one entity each. The surviving URI is the signal-richest fragment; its provenance is the source row that won (erp, timestamp, authority). That URI collapse is applied to the graph.
  • conflict / headquarters — the report names Austin as the authority-axis winner (ERP source_of_truth over a stale directory scrape). Austin is current; San Francisco stays stored and closed.
  • unresolved / credit_rating — ERP says A, CRM says BBB. Same authority, same timestamp. The report flags the pair instead of silently picking; both stay dual-current.

Fixture notes: examples/suppliers-messy.md.
Hermetic proof: tests/test_suppliers_messy_fixture.py.

2. ingest → ask — the payoff

With a key, live /ask is always-LLM Cypher. The cached plan is not consulted when a real key is present.

infona ingest examples/trials.csv --kg my-data
infona ask "Which Phase 3 NSCLC trials is AstraZeneca running?" --kg my-data

That question should return FLAURA2. examples/trials.csv is a 16-row oncology sample (8 sponsors, 11 drugs, 7 indications) — public program names, synthetic TRIAL-* IDs, no patient data.

infona ingest inferring a schema from trials.csv, then writing Trial, Sponsor, Drug, Indication nodes into the graph

infona ask compiling English to Cypher and lighting three sponsor paths — FLAURA2, MARIPOSA, CROWN — into NSCLC

The looping SVGs are generated from scripts/render_readme_demos.py. Local Neo4j notes: docs/neo4j-local.md. If something fails, the CLI should name the next command.

Python package (library, not the infona CLI — that is @infona-ai/cli). Same version as the npm packages:

pip install infona-client

Import path is infona_client. Graph IRIs live under https://graph.infona.ai/. Env prefix is INFONA_* only.


How to use

CLI, MCP, and HTTP share one canonical route per operation. Do not invent a bespoke path.

Command What it does
infona use <kg> Save the working graph. Later ingest / ask can drop --kg.
infona ingest <file> --kg <kg> Schema once, then deterministic rows. Does not merge intra-file fragments.
infona er rebuild --kg <kg> Second-pass URI collapse: apply winners, report leftovers. --kg is required (does not read infona use).
infona ask "…" --kg <kg> Always-LLM Cypher with a key; cached-plan replay with none.
infona ontology types List types / attributes.
infona ontology resolve "…" Evolve the ontology from English (--kg is optional context; not the use default).
infona export --kg <kg> JSON or CSV out. --kg is required.
infona use trials
infona ingest examples/trials.csv
infona ask "Which Phase 3 NSCLC trials is AstraZeneca running?"
infona ontology types
infona ontology resolve "add a runs relationship from Sponsor to Trial"
infona er rebuild --kg trials
infona export --kg trials -f json -o trials.json

Writes go through insert_facts + refresh_after_write. Entity URIs via graph.ontology_queries.entity_uri. Instance relationships use https://graph.infona.ai/onto/<leaf>.

More CLI: packages/cli/README.md. HTTP: docs/API.md.

Optional: 3rd-party REST / SQL extract (dlt)

pip install infona-client does not pull dlt. Install the extra on the backend only (pip install 'infona-client[dlt]'). Infona is the destination — there is no dlt warehouse sink.

There is no infona ingest --dlt. The same spec is POST /graphs/{tenant}/ingest/dlt (SDK Client.ingestDlt, MCP ingest_dlt). The CLI is not on that route.

# frozen body for POST /graphs/{tenant}/ingest/dlt
cat > spec.json <<'EOF'
{
  "source": {
    "kind": "rest_api",
    "base_url": "https://api.example.com",
    "auth": {"type": "bearer", "secret_ref": "env:EXAMPLE_TOKEN"},
    "resources": ["v1/contacts"]
  },
  "map": {"v1/contacts": {"type": "Contact", "id_field": "id"}},
  "kg": "crm"
}
EOF
curl -sS -X POST http://localhost:8000/graphs/default/ingest/dlt \
  -H 'Content-Type: application/json' \
  --data-binary @spec.json

SQL is the same shape with "kind": "sql" and "dsn": "env:EXAMPLE_DSN". Hosted Explorer Connect / Run is premium (ONTA-554) and hits this same route.


MCP (agents)

Same ask, same graph, same exact rows — as a tool result. Same HTTP routes the CLI hits.

{
  "mcpServers": {
    "infona": {
      "command": "npx",
      "args": ["-y", "-p", "@infona-ai/mcp", "infona-mcp"],
      "env": {
        "INFONA_API_URL": "http://localhost:8000",
        "INFONA_TENANT": "default"
      }
    }
  }
}

ask, search, agent, ingest_csv, ingest_dlt, er_rebuild, export_kg, ontology, jobs. packages/mcp/README.md.


Eval

Query accuracy is a live always-LLM Cypher pin; the small-n 8-question run is historical. Protocol, dated table, and repro: docs/EVAL.md. Eval is Python-only; there is no infona eval CLI.


What leaves your machine

Infona does not phone home unless you turn it on. Default off.

export INFONA_TELEMETRY=1          # opt in
export INFONA_TELEMETRY=0          # force off (wins over a previous yes)

The first-run CLI prompt (infona / infona init on a TTY) asks the same question and writes ~/.infona/telemetry.json. There is no opt-out default.

Only when enabled, one anonymous JSON object per job:

  • job type (ingest / ask / er rebuild / export)
  • a row-count bucket (not the exact count)
  • source type (csv / json / jsonl / text / http — never a filename)
  • error class (exception type or HTTP family — never the message)

A random install_id (UUID) identifies the install, not you.

Never leaves: your data, column names, file names, graph content, workspace / tenant ids, prompts, answers, Cypher, emails, API keys.

When enabled, the default collector is the public Infona-oss PostHog project (write-only project token). Override with INFONA_TELEMETRY_URL, set it to off, or use INFONA_TELEMETRY_SINK=stderr / file locally.

Full contract: docs/TELEMETRY.md.


What this is not

Infona is not a memory or context layer that stuffs retrieved chunks into a prompt window. This repo registers no default open-web page fetcher; you bring retrieval or you skip web fetch (docs/BOUNDARY.md). It is not RAG over a vector index, and it is not "chat with your CSV."

Product path is Neo4j GraphStore / Cypher. SPARQL is not the product query language. Neptune is not the product store.


What you get

Entity resolution infona er rebuild collapses fragment URIs (6→3 on the suppliers fixture). Winner URI, reason, score, provenance timestamp. Authority-axis winners become the current graph value (Austin HQ; SF stored/closed). Equal-trust credit_rating stays dual-current and flagged.
Provenance Source + timestamp + authority on the winning fact in the report. Answers carry per-fact citations (tests/test_answer_citations.py).
Schema from one pass Luna (or your configured model) sees the file once. Types, attributes, relationships. No per-row LLM.
Deterministic rows Every cell maps through that schema via insert_facts.
A real graph Neo4j. Sponsors, trials, drugs, indications are nodes.
Ask Always-LLM Cypher when a key is present. Cached-plan replay when it is not. Fail-closed when the plan is a silent wrong total.
CLI + MCP + HTTP Same canonical routes. infona, @infona-ai/mcp, POST /graphs/{tenant}/ask.
Export JSON or CSV back out. The graph is yours.
CSV / JSON / text
  → schema inference (1 LLM call; skipped for the prebuilt snapshot)
  → deterministic row mapping
  → Neo4j knowledge graph (GraphStore / Cypher)
  → er rebuild (URI collapse; field winners applied; equal-trust flagged)
  → ask (cached-plan replay with no key; always-LLM Cypher with a key)

Ask is always-LLM Cypher when a model is configured. Grounding, probes, and few-shots inform the model; they do not replace it.

export OPENROUTER_API_KEY=sk-or-...
export INFONA_QUERY_PROVIDER=openrouter
export INFONA_QUERY_MODEL=openai/gpt-oss-120b

What's free

  • OSS (this repo): ingest, ontology, ask, MCP / CLI / HTTP, export, free sources, BYOK registry, plugin seams.
  • Bring your own retrieval: OSS registers no open-web page fetcher. Enrichment that needs a URL fetch declines unless you register one — or you use hosted Infona.
  • Hosted-only: managed keys Infona bills, paid search/scrape ladders, curated Enhanced ontology, Explorer, billing, Clerk, cloud infra.

Full table: docs/BOUNDARY.md.

Product path: FastAPI + Neo4j GraphStore (Cypher). SPARQL / Neptune are not product backends.


For coding agents

Humans and coding agents follow the same contract. Read AGENTS.md before writing code. How to set up a clone and open a PR: CONTRIBUTING.md. First-time authors sign CLA.md on the PR (I have read the CLA Document and I hereby sign the CLA). Apache-2.0; public publication is a one-way door. Never commit secrets.

Do

  • Import infona_client. npm: @infona-ai/cli, @infona-ai/mcp. Env: INFONA_*. IRIs: https://graph.infona.ai/….
  • Writes through insert_facts + refresh_after_write. Mint entities with entity_uri only.
  • File budget ~500 lines; hard 550 for new infona_client / packages / tests files. Oversized files are pinned in tests/test_file_size_budget.py and must not grow.
  • Hermetic tests (MemoryGraphStore, mocks). No live Neo4j required. Optional: docs/neo4j-local.md.
  • Run scripts/check_boundary.sh and the tests that import what you touched.
pytest tests/test_file_size_budget.py -q
pytest tests/<touched>.py -q
npm test --workspace packages/cli   # if you touched TS

Do not

  • from infona.* / import infona.* (this package imports on its own).
  • Short-circuit production /ask with golden-string Cypher, or hardcode persona-CSV answers.
  • Treat SPARQL as the product query language, or Neptune as the product store. Residual NeptuneClient imports stay — do not delete them in drive-by cleanup.
  • Re-register a default StaticHttpFetcher (BYOR). Do not ship or imply a shared platform key (BYOK = the caller's env).

License and contributing

Apache 2.0 — LICENSE, NOTICE.

Shipped packages share one lockstep version: infona-client on PyPI and @infona-ai/cli / @infona-ai/mcp on npm. Release notes: CHANGELOG.md.

docs/API.md · docs/BOUNDARY.md · ROADMAP.md · SECURITY.md · CODE_OF_CONDUCT.md · CHANGELOG.md · CONTRIBUTING.md · CLA.md · AGENTS.md

Reviews (0)

No results found