dsh-docs
Health Gecti
- License — License: MIT
- Description — Repository has a description
- Active repo — Last push 0 days ago
- Community trust — 10 GitHub stars
Code Basarisiz
- fs module — File system access in package.json
- exec() — Shell command execution in scripts/build-runtime-win32-x64.mjs
- process.env — Environment variable access in scripts/build-runtime-win32-x64.mjs
- network request — Outbound network request in scripts/build-runtime-win32-x64.mjs
Permissions Gecti
- Permissions — No dangerous permissions requested
Bu listing icin henuz AI raporu yok.
Fully local document intelligence for DeepSeek Harness. Parse PDF, Office files, images, and scanned documents with offline OCR. | DeepSeek Harness 全本地文档智能插件,支持 PDF、Office、图片与离线 OCR
dsh-docs
dsh-docs gives your DeepSeek Harness agent real document intelligence —
entirely on your own machine. Hand it a PDF, Word, Excel, or PowerPoint file
and get back clean Markdown, plain text, or structured JSON; hand it a scanned
page or image and a fully offline OCR pipeline reads it for you. No Docker, no
HTTP service, no API keys, and no document ever leaves your disk.
It ships a pinned, self-contained Python + Xberg
runtime with offline Tesseract language data (English and Simplified Chinese),
delivering complete PDF/Office/OCR coverage on Windows x64 out of the box. The
native Xberg Node binding serves as a lightweight non-OCR fallback on any
platform, and every file read stays confined to folders you explicitly
authorize.
The published package and plugin id use the dsh-doc spelling and the tools
use dshdoc_*; they were renamed from the initial dsh-docling / docling_*
release.
One-prompt install
No local checkout or build toolchain is needed. Paste the following prompt into
a running DSH session (for example dsh web) in your own project folder. The
Harness agent installs the published npm package, downloads the pinned offline
OCR runtime, and configures the plugin in one go. The only prerequisite is a
working dsh CLI on Node ^22.19 or >= 24; every runtime step is plain
Node.js, so any shell works: cmd, PowerShell, pwsh, or Git Bash.
Install the dsh-doc plugin into my DSH web profile, end to end. Do every
step yourself in the terminal and verify the result.
1. Install the published plugin package:
dsh plugin --profile web add dsh-doc
2. Windows x64 only — download the prebuilt offline OCR runtime. The script
verifies the pinned archive SHA-256, then verifies every extracted file
against the bundled manifest:
node <home>/.dsh/profiles/web/node_modules/dsh-doc/scripts/fetch-runtime-win32-x64.mjs <home>/.dsh/runtimes/dshdoc-runtime-win32-x64
Replace <home> with my absolute home directory in this and every later step.
On any other platform, skip this step and use engine: node below.
3. Edit <home>/.dsh/profiles/web/cordis.patch.yml. Preserve every existing
entry and add or update this one:
- id: dsh-doc
config:
engine: python
runtimeDir: <home>/.dsh/runtimes/dshdoc-runtime-win32-x64
defaultOcr: true
maxOutputChars: 32000
The session workspace is readable automatically; add allowedLocalRoots only
for extra persistent directories such as a shared document vault.
If you skipped step 2, use `engine: node` and `defaultOcr: false` instead
and omit runtimeDir.
4. Verify with `dsh --profile web --dump-config` that the composed dsh-doc
entry carries exactly this config, then report the result and remind me to
restart `dsh web` so I can call dshdoc_health.
Hard constraints: never install, start, or configure Docling Serve, Docker,
containers, or any remote document-conversion service; never configure a
downloadable OCR backend or allow a model download.
After the agent finishes and you restart dsh web, check the engine withdshdoc_health, then parse any file below your workspace withdshdoc_extract. INSTALL.md documents the same flow plus the
manual procedure step by step.
Supported and tested inputs
- PDF, DOCX, XLSX, PPTX, Markdown, HTML, CSV, and text
- PNG, JPEG, TIFF, WebP, and scanned PDFs through local OCR
- Markdown, plain text, or JSON-shaped Tool Results
The integration suite generates binary documents outside the repository and
verifies PDF, DOCX, XLSX, PPTX, PNG OCR, and scanned-PDF OCR. Test any other
Xberg-supported input against your own corpus before enabling it in production.
Quick start with dsh web
Install the published package into the web profile:
dsh plugin --profile web add dsh-doc
Windows x64 — fetch the prebuilt offline OCR runtime into a stable directory
outside node_modules (so plugin upgrades never delete it):
node ~/.dsh/profiles/web/node_modules/dsh-doc/scripts/fetch-runtime-win32-x64.mjs ~/.dsh/runtimes/dshdoc-runtime-win32-x64
Expand ~ to your absolute home directory everywhere, including inside the
YAML below. On other platforms, skip the runtime and use engine: node withdefaultOcr: false.
Add the plugin entry to the web profile's cordis.patch.yml:
- id: dsh-doc
config:
engine: python
runtimeDir: ~/.dsh/runtimes/dshdoc-runtime-win32-x64
maxFileBytes: 52428800
maxOutputChars: 32000
# Safe here because the configured runtime carries the local language packs.
defaultOcr: true
defaultTableMode: accurate
defaultOutputFormat: md
The session workspace is readable without further configuration. AddallowedLocalRoots only for extra persistent directories the model may read
beyond the workspace, and set allowWorkspaceFiles: false only for
allowlist-only lockdown deployments.
Restart dsh web, then ask it to read a local document:
Read ./reports/annual-report.pdf and give me the three main risks.
Extract the tables from ./financials.xlsx.
Read the text from ./scanned-invoice.png.
Only paths below the session workspace or allowedLocalRoots are readable.
Relative paths resolve against the DSH session workspace, not the directory
from which dsh web was started.
Offline embedded Python runtime (Windows x64)
Most users fetch the prebuilt, hash-pinned runtime from the GitHub Release:
node ./scripts/fetch-runtime-win32-x64.mjs
To audit and rebuild from source instead, run:
node ./scripts/build-runtime-win32-x64.mjs
Either command creates a gitignored .dsh-runtime/runtime-win32-x64 directory
containing CPython 3.11.9, xberg==1.0.14, and pinned eng / chi_sim
Tesseract data. Every downloaded file is SHA-256 validated; the artifact
contains a manifest, NOTICE, and SPDX inventory. It does not alter a global
Python installation. Run node ./scripts/verify-runtime-win32-x64.mjs before
pointing a profile at a copied runtime artifact.
Point the plugin at the runtime:
- id: dsh-doc
config:
engine: python
runtimeDir: <absolute path to the runtime directory>
The Python worker receives only a byte snapshot, display name, MIME type, and
conversion options over stdio. It never receives a user path or URL. It runs
offline, refuses missing OCR language packs, and disables document-derived OCR
caching. dshdoc_health reports the available OCR languages. See the runtime
guide.
Node-only fallback
Set engine: node only when you need PDF/Office/text parsing without the
embedded Python runtime. Its defaultOcr is false. To enable Node OCR, settessdataPath to a reviewed local directory containing every requested<language>.traineddata pack; missing data returns ENGINE_OCR_UNAVAILABLE
instead of downloading a model. The Python runtime above is the supported
complete offline OCR path.
Tools
| Tool | Purpose |
|---|---|
dshdoc_health |
Report readiness of the selected local engine. |
dshdoc_convert_file |
Parse an allowlisted local file. |
dshdoc_extract |
Preferred local-file convenience tool. |
dshdoc_convert_url |
Compatibility stub that returns UNSUPPORTED_URL. |
HTTP(S) input is detected only to reject it safely. Download a remote document
through a reviewed workflow into an allowed local root, then parse that file.
The plugin never forwards a URL to Xberg or Python, avoiding redirect and
DNS-rebinding risks.
page_range uses inclusive, one-based page numbers for Markdown and plain-text
results. JSON output deliberately retains the complete structured document.
The conversion tools also accept an optional per-request ocr_languages array
(for example ["chi_sim", "eng"]) to override the configured language set for
one document. Results report OCR: applied / OCR: not used when the engine
exposes whether the OCR pipeline contributed; enabling OCR never replaces a
healthy embedded PDF text layer.
Configuration
| Field | Default | Meaning |
|---|---|---|
engine |
auto |
node, python, or auto; auto selects configured embedded Python, otherwise Node Xberg. |
runtimeDir |
unset | Absolute embedded-runtime directory. |
pythonCommand |
unset | Trusted Python executable for a managed runtime. |
pythonWorkerPath |
shipped worker | Absolute Python worker override. |
tessdataPath |
runtime ocr/tessdata |
Absolute bundled Tesseract language-data directory. |
ocrBackend |
auto |
auto or tesseract; both select the pinned local Tesseract backend. |
ocrLanguages |
every bundled pack | Ordered local OCR language packs. Unset uses every .traineddata pack in the configured runtime. |
timeoutMs |
120000 |
Per-conversion deadline. |
maxFileBytes |
52428800 |
Authorized input-size cap. |
allowedLocalRoots |
[] |
Extra absolute non-root directories the model may read, beyond the session workspace. |
allowWorkspaceFiles |
true |
Implicitly authorize the session workspace (session cwd) as a readable root; set false for allowlist-only lockdown. |
defaultOcr |
false |
OCR default for images and scans; enable it only with a configured local tessdata runtime. |
defaultTableMode |
accurate |
fast or accurate PDF table behavior. |
defaultOutputFormat |
md |
md, text, or json. |
maxOutputChars |
32000 |
Maximum result returned to the model. |
debug |
false |
Logs safe engine metadata only. |
Older baseUrl, apiKey, enableRemoteUrls, and allowPrivateUrls profile
fields are accepted only for migration; they do not enable a remote engine.
Security model
- Paths are realpathed and checked against every configured root and, unless
allowWorkspaceFilesis disabled, the session workspace. Traversal,
symlink escapes, filesystem roots, non-files, and oversized inputs fail. - The authorized descriptor is read once into a snapshot before parsing, so a
later path replacement cannot change the parsed bytes. - The Node and Python engines accept bytes only. The plugin creates no listener,
URL fetcher, container, or external parser service. - OCR is Tesseract-only in this release. All requested language packs are read
from the configured local artifact; missing packs fail closed rather than
triggering a model download. - The descriptor opened for parsing must have the same device/inode identity as
the post-open allowlisted path, blocking file replacement between authorization
and the byte snapshot. - Results are bounded before becoming Tool Results. JSON is limited using the
same pretty representation shown to the model.
Development
pnpm install
pnpm lint
pnpm typecheck
pnpm test
pnpm build
pnpm pack --pack-destination .pack
Tests create temporary documents only. They cover native Xberg, the Python
stdio worker, local OCR data, Cordis ToolRuntime, and local DSH AgentLoop
context injection.
Licenses
This project is MIT. Xberg 1.0.14 is MIT. The optional Windows runtime contains
CPython (PSF-2.0) and tessdata_fast language data (Apache-2.0), with exact
sources, hashes, and notices recorded in its generated artifact.
Yorumlar (0)
Yorum birakmak icin giris yap.
Yorum birakSonuc bulunamadi