minilake

mcp
Guvenlik Denetimi
Uyari
Health Uyari
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 6 GitHub stars
Code Gecti
  • Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Gecti
  • Permissions — No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

Minilake is a free, local Databricks API emulator — a single-developer tool for testing databricks-sdk/Terraform code against real SQL, real Delta Lake, and real Job execution, without paying for cloud compute.

README.md

minilake — Local Databricks API Emulator

MiniLake

Free, open-source local Databricks emulator for offline development and testing.

Real SQL & Spark execution · Unity Catalog hierarchy · Databricks SDK compatible · Terraform compatible · MIT licensed

GitHub release CI Release GHCR image License Python

Documentation · GitHub · Container Image (GHCR)


MiniLake is a free, local Databricks API emulator — a single-developer tool for testing
databricks-sdk/Terraform code against real SQL, real Delta Lake, and real Job execution,
without paying for cloud compute.

Quick Start

# Option 1: PyPI
pip install minilake
minilake --port 8000

# Option 2: GitHub Container Registry
docker run -p 8000:8000 ghcr.io/dmux/minilake:latest

# Option 3: Clone and build
git clone https://github.com/dmux/minilake && cd minilake
docker compose up -d

# Verify (any option)
curl http://localhost:8000/_minilake/health

No account, no API key, no sign-up — and the container image downloads nothing at runtime:
DuckDB's delta extension and the Delta / Unity Catalog Spark jars are baked in at build
time, so it works air-gapped (details).

Then point any Databricks client at it:

from databricks.sdk import WorkspaceClient

w = WorkspaceClient(host="http://localhost:8000", token="dev")
w.catalogs.create(name="vendas")

More in Getting Started.

Documentation

Getting Started Install, first catalog and query, internal endpoints
Configuration Every environment variable, persistence, HTTPS/TLS
Databricks SDK Unity Catalog, warehouses, SQL and jobs from Python
Terraform & Asset Bundles The provider, and bundle deploy / bundle run
Spark & Delta Lake Real Delta files, real Spark jobs, spark.table() by name
MCP Server 67 tools for LLM agents — examples, tool reference, troubleshooting
Testing & development Running the suite, adding an API group
Releases & CI/CD How a tag becomes a published image
Feature status Endpoint-by-endpoint status and design rationale

Supported Services

Service Status Notes
Unity Catalog (catalogs, schemas, tables, volumes) ✅ Real Each catalog = its own DuckDB database (ATTACH), native catalog.schema.table addressing
EXTERNAL Delta Tables ✅ Real Real Delta files; INSERT/UPDATE/DELETE via a generated Spark job, reads via delta_scan()
SQL Statement Execution ✅ Real Real DuckDB; JSON_ARRAY/ARROW_STREAM/CSV, INLINE/EXTERNAL_LINKS
SQL Warehouses ✅ Real Full CRUD + lifecycle
Jobs ✅ Real Sibling Docker container execution (Spark) or subprocess fallback; real DAG scheduling (depends_on/run_if); sql_task.file
Workspace ✅ Real File-backed notebook/script storage; raw-bytes workspace-files sync powers databricks bundle deploy / bundle run
DBFS & Files API ✅ Real File-backed storage, chunked upload
Secrets ✅ Real Real CRUD; values only resolvable inside job env vars, never via direct API (matches real Databricks)
Clusters ✅ Real state machine CRUD + timed lifecycle transitions; no real Spark compute (by design)
Permissions ✅ Real CRUD Single-user "allow-all" default (by design — see Gaps)
Identity (SCIM) ✅ Static Fake current-user endpoint
Persistence (MINILAKE_PERSIST=1) ✅ Real JSON snapshot on shutdown, restored on startup
Unity Catalog protocol for Spark ✅ Real spark.table("cat.sch.tbl") resolves against minilake — see Spark & Delta Lake
JupyterLab + PySpark + Delta (optional) ✅ Real docker compose --profile notebook up
MCP Server (optional, MINILAKE_MCP=1) ✅ Real 67 tools + resources + prompts at /mcp — see MCP Server
Secrets ACLs, Repos/Git, multi-language notebooks, DBT/pipeline tasks, Model Registry, Vector Search, Dashboards 🚫 Not implemented Returns 501 NOT_IMPLEMENTED

Known Gaps

These are deliberate, not oversights — minilake targets one developer running it
locally, not a shared or multi-tenant server:

  • No real authentication — any token is accepted; there's only ever one real user.
  • No access-control enforcement — the Permissions API is real CRUD but always allow-all,
    so a test that passes here says nothing about grants in a real workspace.
  • No real Spark compute for Clusters — state machine only; real compute happens through
    Jobs' sibling containers instead.
  • Single process, no HA — and DuckDB's single-writer model means concurrent load
    contends on locks.
  • Uneven test coveragejobs.py, sql_statements.py and unity_catalog.py are
    covered mostly on happy paths, not edge cases.
  • Secrets ACLs not implemented — scope/secret CRUD is real, ACL endpoints aren't.

Contributing

See CONTRIBUTING.md for the project structure, how to add a new API
group, and the PR checklist.

License

MIT — see LICENSE.

Yorumlar (0)

Sonuc bulunamadi