minilake
Health Warn
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Low visibility — Only 6 GitHub stars
Code Pass
- Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Pass
- Permissions — No dangerous permissions requested
No AI report is available for this listing yet.
Minilake is a free, local Databricks API emulator — a single-developer tool for testing databricks-sdk/Terraform code against real SQL, real Delta Lake, and real Job execution, without paying for cloud compute.
MiniLake
Free, open-source local Databricks emulator for offline development and testing.
Real SQL & Spark execution · Unity Catalog hierarchy · Databricks SDK compatible · Terraform compatible · MIT licensed
Documentation · GitHub · Container Image (GHCR)
MiniLake is a free, local Databricks API emulator — a single-developer tool for testingdatabricks-sdk/Terraform code against real SQL, real Delta Lake, and real Job execution,
without paying for cloud compute.
Quick Start
# Option 1: PyPI
pip install minilake
minilake --port 8000
# Option 2: GitHub Container Registry
docker run -p 8000:8000 ghcr.io/dmux/minilake:latest
# Option 3: Clone and build
git clone https://github.com/dmux/minilake && cd minilake
docker compose up -d
# Verify (any option)
curl http://localhost:8000/_minilake/health
No account, no API key, no sign-up — and the container image downloads nothing at runtime:
DuckDB's delta extension and the Delta / Unity Catalog Spark jars are baked in at build
time, so it works air-gapped (details).
Then point any Databricks client at it:
from databricks.sdk import WorkspaceClient
w = WorkspaceClient(host="http://localhost:8000", token="dev")
w.catalogs.create(name="vendas")
More in Getting Started.
Documentation
| Getting Started | Install, first catalog and query, internal endpoints |
| Configuration | Every environment variable, persistence, HTTPS/TLS |
| Databricks SDK | Unity Catalog, warehouses, SQL and jobs from Python |
| Terraform & Asset Bundles | The provider, and bundle deploy / bundle run |
| Spark & Delta Lake | Real Delta files, real Spark jobs, spark.table() by name |
| MCP Server | 67 tools for LLM agents — examples, tool reference, troubleshooting |
| Testing & development | Running the suite, adding an API group |
| Releases & CI/CD | How a tag becomes a published image |
| Feature status | Endpoint-by-endpoint status and design rationale |
Supported Services
| Service | Status | Notes |
|---|---|---|
| Unity Catalog (catalogs, schemas, tables, volumes) | ✅ Real | Each catalog = its own DuckDB database (ATTACH), native catalog.schema.table addressing |
| EXTERNAL Delta Tables | ✅ Real | Real Delta files; INSERT/UPDATE/DELETE via a generated Spark job, reads via delta_scan() |
| SQL Statement Execution | ✅ Real | Real DuckDB; JSON_ARRAY/ARROW_STREAM/CSV, INLINE/EXTERNAL_LINKS |
| SQL Warehouses | ✅ Real | Full CRUD + lifecycle |
| Jobs | ✅ Real | Sibling Docker container execution (Spark) or subprocess fallback; real DAG scheduling (depends_on/run_if); sql_task.file |
| Workspace | ✅ Real | File-backed notebook/script storage; raw-bytes workspace-files sync powers databricks bundle deploy / bundle run |
| DBFS & Files API | ✅ Real | File-backed storage, chunked upload |
| Secrets | ✅ Real | Real CRUD; values only resolvable inside job env vars, never via direct API (matches real Databricks) |
| Clusters | ✅ Real state machine | CRUD + timed lifecycle transitions; no real Spark compute (by design) |
| Permissions | ✅ Real CRUD | Single-user "allow-all" default (by design — see Gaps) |
| Identity (SCIM) | ✅ Static | Fake current-user endpoint |
Persistence (MINILAKE_PERSIST=1) |
✅ Real | JSON snapshot on shutdown, restored on startup |
| Unity Catalog protocol for Spark | ✅ Real | spark.table("cat.sch.tbl") resolves against minilake — see Spark & Delta Lake |
| JupyterLab + PySpark + Delta (optional) | ✅ Real | docker compose --profile notebook up |
MCP Server (optional, MINILAKE_MCP=1) |
✅ Real | 67 tools + resources + prompts at /mcp — see MCP Server |
| Secrets ACLs, Repos/Git, multi-language notebooks, DBT/pipeline tasks, Model Registry, Vector Search, Dashboards | 🚫 Not implemented | Returns 501 NOT_IMPLEMENTED |
Known Gaps
These are deliberate, not oversights — minilake targets one developer running it
locally, not a shared or multi-tenant server:
- No real authentication — any token is accepted; there's only ever one real user.
- No access-control enforcement — the Permissions API is real CRUD but always allow-all,
so a test that passes here says nothing about grants in a real workspace. - No real Spark compute for Clusters — state machine only; real compute happens through
Jobs' sibling containers instead. - Single process, no HA — and DuckDB's single-writer model means concurrent load
contends on locks. - Uneven test coverage —
jobs.py,sql_statements.pyandunity_catalog.pyare
covered mostly on happy paths, not edge cases. - Secrets ACLs not implemented — scope/secret CRUD is real, ACL endpoints aren't.
Contributing
See CONTRIBUTING.md for the project structure, how to add a new API
group, and the PR checklist.
License
MIT — see LICENSE.
Reviews (0)
Sign in to leave a review.
Leave a reviewNo results found