dsh-docs

agent
Security Audit
Fail
Health Pass
  • License — License: MIT
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Community trust — 10 GitHub stars
Code Fail
  • fs module — File system access in package.json
  • exec() — Shell command execution in scripts/build-runtime-win32-x64.mjs
  • process.env — Environment variable access in scripts/build-runtime-win32-x64.mjs
  • network request — Outbound network request in scripts/build-runtime-win32-x64.mjs
Permissions Pass
  • Permissions — No dangerous permissions requested

No AI report is available for this listing yet.

SUMMARY

Fully local document intelligence for DeepSeek Harness. Parse PDF, Office files, images, and scanned documents with offline OCR. | DeepSeek Harness 全本地文档智能插件,支持 PDF、Office、图片与离线 OCR

README.md

dsh-docs

dsh-docs — local document intelligence for DeepSeek Harness: PDF, DOCX, XLSX, PPTX, Markdown, HTML, CSV, OCR, and text

中文 | Installation prompt

dsh-docs gives your DeepSeek Harness agent real document intelligence —
entirely on your own machine. Hand it a PDF, Word, Excel, or PowerPoint file
and get back clean Markdown, plain text, or structured JSON; hand it a scanned
page or image and a fully offline OCR pipeline reads it for you. No Docker, no
HTTP service, no API keys, and no document ever leaves your disk.

It ships a pinned, self-contained Python + Xberg
runtime with offline Tesseract language data (English and Simplified Chinese),
delivering complete PDF/Office/OCR coverage on Windows x64 out of the box. The
native Xberg Node binding serves as a lightweight non-OCR fallback on any
platform, and every file read stays confined to folders you explicitly
authorize.

The published package and plugin id use the dsh-doc spelling and the tools
use dshdoc_*; they were renamed from the initial dsh-docling / docling_*
release.

One-prompt install

No local checkout or build toolchain is needed. Paste the following prompt into
a running DSH session (for example dsh web) in your own project folder. The
Harness agent installs the published npm package, downloads the pinned offline
OCR runtime, and configures the plugin in one go. The only prerequisite is a
working dsh CLI on Node ^22.19 or >= 24; every runtime step is plain
Node.js, so any shell works: cmd, PowerShell, pwsh, or Git Bash.

Install the dsh-doc plugin into my DSH web profile, end to end. Do every
step yourself in the terminal and verify the result.

1. Install the published plugin package:
   dsh plugin --profile web add dsh-doc
2. Windows x64 only — download the prebuilt offline OCR runtime. The script
   verifies the pinned archive SHA-256, then verifies every extracted file
   against the bundled manifest:
   node <home>/.dsh/profiles/web/node_modules/dsh-doc/scripts/fetch-runtime-win32-x64.mjs <home>/.dsh/runtimes/dshdoc-runtime-win32-x64
   Replace <home> with my absolute home directory in this and every later step.
   On any other platform, skip this step and use engine: node below.
3. Edit <home>/.dsh/profiles/web/cordis.patch.yml. Preserve every existing
   entry and add or update this one:
   - id: dsh-doc
     config:
       engine: python
       runtimeDir: <home>/.dsh/runtimes/dshdoc-runtime-win32-x64
       defaultOcr: true
       maxOutputChars: 32000
   The session workspace is readable automatically; add allowedLocalRoots only
   for extra persistent directories such as a shared document vault.
   If you skipped step 2, use `engine: node` and `defaultOcr: false` instead
   and omit runtimeDir.
4. Verify with `dsh --profile web --dump-config` that the composed dsh-doc
   entry carries exactly this config, then report the result and remind me to
   restart `dsh web` so I can call dshdoc_health.

Hard constraints: never install, start, or configure Docling Serve, Docker,
containers, or any remote document-conversion service; never configure a
downloadable OCR backend or allow a model download.

After the agent finishes and you restart dsh web, check the engine with
dshdoc_health, then parse any file below your workspace with
dshdoc_extract. INSTALL.md documents the same flow plus the
manual procedure step by step.

Supported and tested inputs

  • PDF, DOCX, XLSX, PPTX, Markdown, HTML, CSV, and text
  • PNG, JPEG, TIFF, WebP, and scanned PDFs through local OCR
  • Markdown, plain text, or JSON-shaped Tool Results

The integration suite generates binary documents outside the repository and
verifies PDF, DOCX, XLSX, PPTX, PNG OCR, and scanned-PDF OCR. Test any other
Xberg-supported input against your own corpus before enabling it in production.

Quick start with dsh web

Install the published package into the web profile:

dsh plugin --profile web add dsh-doc

Windows x64 — fetch the prebuilt offline OCR runtime into a stable directory
outside node_modules (so plugin upgrades never delete it):

node ~/.dsh/profiles/web/node_modules/dsh-doc/scripts/fetch-runtime-win32-x64.mjs ~/.dsh/runtimes/dshdoc-runtime-win32-x64

Expand ~ to your absolute home directory everywhere, including inside the
YAML below. On other platforms, skip the runtime and use engine: node with
defaultOcr: false.

Add the plugin entry to the web profile's cordis.patch.yml:

- id: dsh-doc
  config:
    engine: python
    runtimeDir: ~/.dsh/runtimes/dshdoc-runtime-win32-x64
    maxFileBytes: 52428800
    maxOutputChars: 32000
    # Safe here because the configured runtime carries the local language packs.
    defaultOcr: true
    defaultTableMode: accurate
    defaultOutputFormat: md

The session workspace is readable without further configuration. Add
allowedLocalRoots only for extra persistent directories the model may read
beyond the workspace, and set allowWorkspaceFiles: false only for
allowlist-only lockdown deployments.

Restart dsh web, then ask it to read a local document:

Read ./reports/annual-report.pdf and give me the three main risks.
Extract the tables from ./financials.xlsx.
Read the text from ./scanned-invoice.png.

Only paths below the session workspace or allowedLocalRoots are readable.
Relative paths resolve against the DSH session workspace, not the directory
from which dsh web was started.

Offline embedded Python runtime (Windows x64)

Most users fetch the prebuilt, hash-pinned runtime from the GitHub Release:

node ./scripts/fetch-runtime-win32-x64.mjs

To audit and rebuild from source instead, run:

node ./scripts/build-runtime-win32-x64.mjs

Either command creates a gitignored .dsh-runtime/runtime-win32-x64 directory
containing CPython 3.11.9, xberg==1.0.14, and pinned eng / chi_sim
Tesseract data. Every downloaded file is SHA-256 validated; the artifact
contains a manifest, NOTICE, and SPDX inventory. It does not alter a global
Python installation. Run node ./scripts/verify-runtime-win32-x64.mjs before
pointing a profile at a copied runtime artifact.

Point the plugin at the runtime:

- id: dsh-doc
  config:
    engine: python
    runtimeDir: <absolute path to the runtime directory>

The Python worker receives only a byte snapshot, display name, MIME type, and
conversion options over stdio. It never receives a user path or URL. It runs
offline, refuses missing OCR language packs, and disables document-derived OCR
caching. dshdoc_health reports the available OCR languages. See the runtime
guide
.

Node-only fallback

Set engine: node only when you need PDF/Office/text parsing without the
embedded Python runtime. Its defaultOcr is false. To enable Node OCR, set
tessdataPath to a reviewed local directory containing every requested
<language>.traineddata pack; missing data returns ENGINE_OCR_UNAVAILABLE
instead of downloading a model. The Python runtime above is the supported
complete offline OCR path.

Tools

Tool Purpose
dshdoc_health Report readiness of the selected local engine.
dshdoc_convert_file Parse an allowlisted local file.
dshdoc_extract Preferred local-file convenience tool.
dshdoc_convert_url Compatibility stub that returns UNSUPPORTED_URL.

HTTP(S) input is detected only to reject it safely. Download a remote document
through a reviewed workflow into an allowed local root, then parse that file.
The plugin never forwards a URL to Xberg or Python, avoiding redirect and
DNS-rebinding risks.

page_range uses inclusive, one-based page numbers for Markdown and plain-text
results. JSON output deliberately retains the complete structured document.

The conversion tools also accept an optional per-request ocr_languages array
(for example ["chi_sim", "eng"]) to override the configured language set for
one document. Results report OCR: applied / OCR: not used when the engine
exposes whether the OCR pipeline contributed; enabling OCR never replaces a
healthy embedded PDF text layer.

Configuration

Field Default Meaning
engine auto node, python, or auto; auto selects configured embedded Python, otherwise Node Xberg.
runtimeDir unset Absolute embedded-runtime directory.
pythonCommand unset Trusted Python executable for a managed runtime.
pythonWorkerPath shipped worker Absolute Python worker override.
tessdataPath runtime ocr/tessdata Absolute bundled Tesseract language-data directory.
ocrBackend auto auto or tesseract; both select the pinned local Tesseract backend.
ocrLanguages every bundled pack Ordered local OCR language packs. Unset uses every .traineddata pack in the configured runtime.
timeoutMs 120000 Per-conversion deadline.
maxFileBytes 52428800 Authorized input-size cap.
allowedLocalRoots [] Extra absolute non-root directories the model may read, beyond the session workspace.
allowWorkspaceFiles true Implicitly authorize the session workspace (session cwd) as a readable root; set false for allowlist-only lockdown.
defaultOcr false OCR default for images and scans; enable it only with a configured local tessdata runtime.
defaultTableMode accurate fast or accurate PDF table behavior.
defaultOutputFormat md md, text, or json.
maxOutputChars 32000 Maximum result returned to the model.
debug false Logs safe engine metadata only.

Older baseUrl, apiKey, enableRemoteUrls, and allowPrivateUrls profile
fields are accepted only for migration; they do not enable a remote engine.

Security model

  • Paths are realpathed and checked against every configured root and, unless
    allowWorkspaceFiles is disabled, the session workspace. Traversal,
    symlink escapes, filesystem roots, non-files, and oversized inputs fail.
  • The authorized descriptor is read once into a snapshot before parsing, so a
    later path replacement cannot change the parsed bytes.
  • The Node and Python engines accept bytes only. The plugin creates no listener,
    URL fetcher, container, or external parser service.
  • OCR is Tesseract-only in this release. All requested language packs are read
    from the configured local artifact; missing packs fail closed rather than
    triggering a model download.
  • The descriptor opened for parsing must have the same device/inode identity as
    the post-open allowlisted path, blocking file replacement between authorization
    and the byte snapshot.
  • Results are bounded before becoming Tool Results. JSON is limited using the
    same pretty representation shown to the model.

Development

pnpm install
pnpm lint
pnpm typecheck
pnpm test
pnpm build
pnpm pack --pack-destination .pack

Tests create temporary documents only. They cover native Xberg, the Python
stdio worker, local OCR data, Cordis ToolRuntime, and local DSH AgentLoop
context injection.

Licenses

This project is MIT. Xberg 1.0.14 is MIT. The optional Windows runtime contains
CPython (PSF-2.0) and tessdata_fast language data (Apache-2.0), with exact
sources, hashes, and notices recorded in its generated artifact.

Reviews (0)

No results found