deskmind

agent
Security Audit
Fail
Health Warn
  • License — License: Apache-2.0
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Low visibility — Only 5 GitHub stars
Code Fail
  • rm -rf — Recursive force deletion command in .github/workflows/app.yml
  • rm -rf — Recursive force deletion command in app/build.sh
  • rm -rf — Recursive force deletion command in app/runtime.sh
  • rm -rf — Recursive force deletion command in app/tests/e2e/helper.sh
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Small enough to run on your Mac. Smart enough to ask. An open-source computer use agent for macOS, with local models (MLX) that see, decide and act. 得心,应手。

README.md

DeskMind 得心 · 得心,应手。

DeskMind · 得心

Small enough to run on your Mac. Smart enough to ask.

Open-source computer use for your Mac. DeskMind reads the screen, works out the next step and acts, all with small models running on your Mac. When a task could mean two things, it asks you instead of guessing.

中文 · Website · Docs · Download for Mac · Models · Discussions · Roadmap

Watch the 56-second demo on deskmind.dev: a real recording with the released model. Two orders match "Lisa Wong", so it asks which one before writing.

DeskMind asks: "Lisa Wong" is on more than one line, which one should I use? The answer field says 09-27.

How it was built, and what went wrong on the way: 3 weeks, 20 training rounds, $600. If DeskMind is useful or interesting to you, a star on this repo helps other people find it.

What makes it different

A small model, on your Mac. A 0.8B model decides each step and hands the unsure ones to a 4B. Both run on your Mac: no cloud round-trip, no per-step bill. The 0.8B decides in about 0.5 s; the 4B answers in about 3.6 s (median decision times).

System One: choices, not guesses. Each step is a multiple-choice question. The model scores every option instead of writing text, so every option gets a probability. Unsure steps go to the 4B or to you, and any agent can call it through POST /v1/systemone. In 39 real-desktop runs it never said "done" when the task was not done.

Open, from eyes to hands. Eyes, Brain, Hands and the Mac app are open source, along with Bench, which grades them. Every result below comes with its sample size, so you can reproduce it.

Get started

Use the Mac app. Download DeskMind for Mac (macOS 15+, Apple Silicon, signed and notarized). It ships the G18b release models. On first run it downloads the models (about 5.3 GB) and walks you through the permissions: Install the app.

Or run the model yourself. This runs the 4B alone (release G18b). It answers one step of a desktop task; it does not drive the desktop by itself.

# Terminal 1: get Brain and the model, then serve it
git clone https://github.com/deskmind-ai/brain && cd brain && uv sync --extra mlx
uv run hf download deskmind/brain-4b --revision g18b-q8 --local-dir models/brain-4b
uv run deskmind-brain-serve --predictor mlx:models/brain-4b --port 8793 --two-stage

The 4B is about 4.5 GB to download. If the download fails with a CAS Client Error, rerun it with HF_HUB_DISABLE_XET=1 in front. In mainland China, ModelScope carries the same files:
uvx modelscope download --model gxcsoccer/brain-4b --revision g18b-q8 --local-dir models/brain-4b.

# Terminal 2 (in the same brain folder): ask for the next step of a real desktop task
curl -s localhost:8793/v1/systemone -H 'Content-Type: application/json' -d @examples/request.json

Success looks like a typed decision with a probability for every option, e.g.
{"answers": {"operation": {"choice": "CLICK", "probabilities": {"CLICK": 0.96, "OPEN": 0.005, …}}, "click_target": {…}}}.
A probability is the model's weighting of the options, not a guarantee that the step is right.

The router (0.8B → 4B), as released: also download deskmind/brain-0.8b --revision g18b-q8 into models/brain-0.8b (0.8 GB), then serve
uv run deskmind-brain-serve --predictor mlx:models/brain-0.8b --escalate-to mlx:models/brain-4b --two-stage --port 8796.
The threshold (0.96) ships with the weights; each reply adds a routing record such as {"by": "strong", "reason": "low_conf", "fast_conf": 0.956}.

Step by step, with the full reply explained: Quickstart. To drive the desktop, add Hands.

Components

Role
Eyes finds the target on screen 4B visual grounder, for apps without an accessibility tree
Brain decides the next step 0.8B and 4B, MLX, 0.8B → 4B routing, /v1/systemone
Hands observes and acts on macOS accessibility and vision modes, budgets, cancellation
App (this repository) brings it to your Mac native app with a background helper; download
Bench checks what really happened sandbox desktop tasks with strict final-state graders

Models: huggingface.co/deskmind (brain-0.8b, brain-4b, eyes-4b). Website: deskmind.dev. Docs: deskmind.dev/docs.

In this repository

Path What
app/ The Mac app's source (Swift) and its build; build it yourself
docs/ How the app, its helper and the local model servers fit together
brand/ Logo, Xiaofang and social images; rules in BRAND.md
.github/workflows/app.yml CI: every app change is built and tested; a version tag is signed, notarized and drafted as a release
ROADMAP.md What we are working on next

Mac app releases are published here: Releases.

Results, with the sample size

What Setting Result
Real-desktop tasks Bench v25, 13 tasks × 3 runs, strict graders; router G18b (0.8B → 4B, 8-bit, threshold 0.96), through the app, one M4 Pro (48 GB) 39/39 passed; 0 false "done"
Decision time the same 39 runs, 208 decisions median 0.48 s when the 0.8B answers (about 30% of steps), 3.6 s when the 4B answers (about 70%); 2.85 s overall, slowest 5% 9.82 s. Per decision, not per task
Decision quality JevBench v1.4.2, 231 public items Brain 4B 0.835 · Brain 0.8B 0.723 · router 0.797; no sealed score yet
Visual grounding ScreenSpot-Pro, 1,581 items, one pass Eyes 4B 67.7% on a GPU (bf16, native resolution; base model 64.8%); 50.9% as the Mac app runs it (4-bit MLX, ≤ 2 MP)
  • Small sample. 13 tasks on one Mac with a Chinese system language. Runs cluster by task (almost every task passes 3/3 or 0/3), so the effective sample is closer to 13 tasks than to 39 runs.
  • One task was not clean. In all 3 runs of the Chinese exact-text task the file was right, but the model never said "done" and used its full step budget. The grader checks the final state, so these count as passes.
  • G18b traded some general judgement for the desktop. The 4B's JevBench public score fell from 0.866 (G14) to 0.835, mostly on the hard tier. The previous release, G14, passed 36/39 on the same bench.
  • Eyes' GPU score is not the app's setting. The app runs a 4-bit MLX conversion at up to 2 MP, which has no benchmark score yet.

Our own runs. Methods and full tables: Brain results · Bench reference · Eyes results · Results and limits.

Still hard

  • Copying long tables (more than about four rows) or only the rows that match a condition.
  • Filling a form from a photographed receipt.
  • Steps handed to the 4B take a few seconds, and most steps go to the 4B today.

What we are working on next: ROADMAP.

Your Mac does the thinking

Inference runs locally by default. The models download once, from Hugging Face or ModelScope; after that, deciding a step needs no network. A cloud model is opt-in: the app never calls one, but if you point the router's escalation tier at one yourself, the steps routed to it go to that service. Apps that DeskMind drives, such as a web page or a music app, still talk to their own servers.

Contribute

  • Questions and design discussion: Discussions.
  • Cross-component problems, reproduction reports and project direction: issues here, for every component (labelled area: …). Pull requests go to the repository that holds the code.
  • How to help, and what a good report contains: Contributing. A reproduction that disagrees with our numbers is welcome.
  • Security problems: not in public; see SECURITY.md.

Meet Xiaofang · 小方

Xiaofang, the DeskMind companion, with a task done

The frame from our logo, come to life. Artwork in brand/, rules in BRAND.md.

License

  • Code in this repository (the Mac app under app/ and the CI): Apache-2.0; see NOTICE.
  • Text and documentation in this repository: CC BY 4.0.
  • Not covered by either licence: the DeskMind and 得心 names, the DeskMind logo, the Xiaofang (小方) character and the other files under brand/. Their use is governed by BRAND.md.
  • Eyes, Brain, Hands and Bench are Apache-2.0 in their own repositories; see each one's LICENSE and NOTICE.
  • Model weights, base models and datasets follow their own terms.

Reviews (0)

No results found