DeepSeek Harness plugin

dsh-knowledge-sorenabt

Cherry Studio-style knowledge base system for DeepSeek Harness (DSH): bases, documents, chunking, embeddings (OpenAI-compatible / Ollama / local / lexical fallback), retrieval, model-facing tools

Jump to install

Source facts

Repository
Soren-ABT/dsh-knowledge
Latest update
Aug 20, 2026
Category
Docs & Rendering
GitHub stars
10
Format
plugin
Catalog evidence
Upstream dsh.bundle evidence
Evidence path
package.json#dsh.bundle
Checked against
0.1.0-rc.8
Upstream check date
2026-08-20

This evidence comes from the upstream catalog. This site has not installed, run, or security-reviewed the plugin.

Install

Start with a prompt that asks an agent to review the GitHub repository and source. Switch to the command if you want to install it yourself.

Copy this prompt into DSH, Codex, or another agent and ask it to review the GitHub repository and source first.

Do not install or run any commands yet. Read this plugin's GitHub repository, README, and relevant source code. Then answer the questions below clearly and directly so I can decide whether it fits my needs:

1. What is this plugin, and what problem does it solve?
2. Who is it for, and what are its typical use cases?
3. How is it used after installation? Include one minimal example.
4. What known limitations or privacy, security, compatibility, or maintenance risks does it have?
5. Give a clear recommendation: recommend, conditionally recommend, or do not recommend, with reasons.

Distinguish statements documented by the repository, inferences from source code, and unknowns. If evidence is insufficient, say so explicitly. Do not guess or simply repeat the README.

GitHub: https://github.com/Soren-ABT/dsh-knowledge
Plugin: dsh-knowledge-sorenabt
Author: Soren-ABT

Check the source files

Read the README and other files from this plugin directory before installing.

File explorer4 files
README.en.mdSource · read only
README language

<div align="center">

dsh-knowledge

A knowledge base plugin for DSH

English · 中文

![npm version](https://www.npmjs.com/package/dsh-knowledge) ![npm downloads](https://www.npmjs.com/package/dsh-knowledge) ![Node.js >= 22](https://nodejs.org) ![TypeScript](https://www.typescriptlang.org/) ![SQLite](https://www.sqlite.org/) ![License: AGPL-3.0](LICENSE)

A knowledge base system as a standalone, open-source bundle plugin for DeepSeek Harness (DSH): bases (with groups) and documents, text chunking, embeddings (OpenAI-compatible / Ollama / local model / lexical fallback), retrieval, model-facing tools, and a browser management panel.

</div>

---

Features

  • Bases, groups & documents: create/delete/rename bases, documents and groups (grouped sidebar navigation with collapsible sections, move-to-group, create/rename/delete group); add files, URLs, or whole directories (pure text via the knowledge_add_document model tool; the picker caps one selection at 20, drag has no count limit, ≤22MB per file — imports run in a 5-way concurrent background pool); directory import (recursive scan of txt / md / csv / html / json / pdf / docx / doc / pptx / ppt / xlsx / xls / epub, imported as a drillable folder tree); same-name conflict dialog (rename all / replace); content-hash dedup; document preview (inline PDF viewer + text/chunk preview with truncation for huge files); per-document ✓ ready badge, live import status, and relative updated time. Imports run in the background — each file row appears the moment parsing starts, folders show importing while any descendant is processing, and errors surface as toasts.
  • Scanned-document OCR (local engine): scanned PDFs, vector-drawn PDFs without a text layer, and images are recognized automatically via full-page rendering + PaddleOCR PP-OCRv5 (≈21MB models with an 18k-entry Chinese dictionary, one-click download in Settings → Local Models; pages are rendered with the mupdf WASM engine), falling back to Tesseract on failure; 1-bit rasters (JBIG2/CCITT fax-style scans) are handled correctly, and PDFs with corrupt or per-glyph text layers automatically switch to the rendering path; recognized text is chunked, embedded, and searchable like any other document.
  • Per-base configuration: each base can pick its own embedding provider/model (including the local model), rerank model (remote API or the local bge-reranker), chunk size, Top K, semantic chunking, conflict strategy, and URL auto-refresh — unset fields inherit the global config; reindex one document or the whole base after a config change.
  • Optional MinerU remote processing: PDFs can be handed to a MinerU service (formulas, tables, layout reconstructed as Markdown) by entering a MinerU API Key in the base settings (global or per-base); without a key the local parse chain runs instead.
  • Embeddings & retrieval: pluggable providers — any OpenAI-compatible /embeddings endpoint, a local Ollama server, or an in-process local model (transformers.js, default onnx-community/Qwen3-Embedding-0.6B-ONNX) — with hybrid retrieval (BM25 + vector + Reciprocal Rank Fusion), rerank (Jina / SiliconFlow / Cohere v2 style APIs or a local cross-encoder), MMR diversity, search modes (auto/hybrid/vector/lexical), a score threshold, and multi-query retrieval (extraQueries — paraphrased/translated variants widen recall). A lexical BM25 fallback (CJK bigram + latin word) keeps it working with zero configuration; the recall test shows per-hit scores, latency, copies full Markdown citations (quote + source line), and keeps a replayable search history. knowledge_search returns a citations array for the model to quote verbatim.
  • Smart chunking: heading-aware chunking (preserves the Markdown heading path, code-fence protected) with the document title + heading injected as retrieval context for better recall; chunkSize/chunkOverlap are token budgets (converted per-document via the measured chars-per-token ratio, aligned with Cherry Studio) and long blocks are cut with Cherry's break-point scoring (heading 100 → code edge 80 → rule 60 → paragraph 20 → sentence punctuation 8 → list 5 → bare newline 1, distance-decayed inside the window), so CJK and Latin corpora window at comparable sizes; optional semantic chunking (embed paragraph segments, merge adjacent similar ones — no extra embedding pass) and token-budget refinement (oversized chunks split at sentence/comma/space boundaries).
  • Index management: reindex on demand (re-chunk + re-embed after changing chunk size or the provider), batched embedding, and statistics (docs / chunks / chars / tokens / embedded).
  • Model tools: 14 tools — search (with citations), list/create/delete bases, add/list/delete documents, import URL, refresh URL, stats, get document, deep-read a document (paged text slices or regex grep), reindex one document or a whole base.
  • Management panel: NOT in Settings — a sidebar-foot "Knowledge" action (beside Settings) opens a frame-wide Cherry Studio-style page: left grouped navigator with base cards, right content column with stat chips, an "updated at" header, add-source popover, a table-style source list (checkbox + name/type/status/updated-at columns with multi-select bulk reindex/delete), per-base settings panel (document processing / image captioning / embedding / rerank / Top K / advanced), recall test, and toasts.
  • Local model manager (in Settings): a Settings → "Local Models" page (via the settings.section slot) with Cherry Studio-style cards for the embedding, rerank, and OCR models: name/subtitle, a ready badge, download / retry / remove actions, and a live download progress bar. Downloaded models become selectable as the embedding provider, and scanned PDFs get OCR automatically. The model cache directory is configurable (native folder picker, open-in-file-manager, and one-click migration of downloaded models to a new location — dot-dirs are skipped and the target is never copied into itself, no blind C-drive growth), and the page includes Ollama management (pull with progress + cancel, installed-model cards with sizes, per-pull progress cards that survive panel close/reopen, embedding/vision recommendations).
  • Persistence: business state (bases, documents, runtime config) is durable via DSH's storageDomain seam (the json backend shipped by the web profile); chunks live in a dedicated SQLite file (<DSH_HOME>/storages/knowledge-chunks.sqlite, configurable via chunkStorePath) so put/delete stay O(1) as data grows — one row per chunk, embeddings as float32 BLOBs. Lexical search runs an FTS5 trigram index and the vector lane uses a resident per-base cache (Float32Array, exact invalidation); nothing loads into memory at open beyond it, so resident memory does not scale with the corpus. One-time migrations on first start after upgrade convert the legacy JSON unit and the previous bundle layout. Falls back to in-memory when no storage backend is available.

---

Architecture

One bundle mounts three plugin rows plus two dedicated worker threads that carry all inference (Cherry's own-worker posture — a native/WASM crash can never take down the host):

Plugin / threadPlatformRole
knowledge (ctx.knowledge)hostCore engine: storage domain, chunking, embeddings, retrieval, OCR scheduling, /knowledge/* HTTP surface
tool-knowledgehost14 model tools consuming ctx.knowledge
ui-knowledgeclientSidebar-foot entry (sidebar.footer.action) + frame-wide Cherry Studio-style panel (shell.overlay), calling the host over same-origin fetch
embed-worker (worker thread)hosttransformers.js local embedding inference (the ~600MB model never enters the host process)
ocr-worker (worker thread)hostPage rendering (mupdf WASM) + PaddleOCR / Tesseract recognition (onnxruntime, OpenCV.js, tesseract workers all isolated in-thread)

Data model: business state lives in the storage domain knowledge (version 0) — bases and documents tables plus a global slot for runtime config overrides; chunks (each may carry an embedding vector) live in the plugin-owned SQLite chunk store, one row per chunk.

---

Install

The package declares dsh.bundle.patch, so dsh plugin add registers it automatically:

dsh plugin --profile <name> add dsh-knowledge          # from npm
dsh plugin --profile <name> add file:/path/to/dsh-knowledge
dsh plugin --profile <name> add ./dsh-knowledge-0.3.3.tgz

> pnpm 10+ build allowlist (required): the plugin's dependencies onnxruntime-node, sharp, protobufjs, and tesseract.js ship postinstall scripts that pnpm refuses to run by default and exits non-zero — dsh plugin add then stops before registering the bundle, so the plugin never activates. Add this to the profile's pnpm-workspace.yaml before installing, then run the add: > > ``yaml > allowBuilds: > onnxruntime-node: true > sharp: true > protobufjs: true > tesseract.js: true > ` > > (Platform binaries are already bundled inside the npm packages, so skipping these scripts does not hurt functionality — but pnpm treats the refusal as an error, so approving them is the clean route. If you only add the config after a first failed add, just run the add again: the package is already in node_modules` and the rerun registers it.)

Restart the web service for the host half and refresh the page for the client panel.

> The plugin installs at the profile level (dsh plugin runs pnpm inside the profile directory), so the same three commands work identically whether DSH was installed from npm or run from a fresh source checkout — no plugin source, checkout links, or DSH builds are involved.

Zero-basis install (only DSH present)

Everything beyond the add command — local embedding, OCR, scanned-document recognition, hybrid retrieval, the management panel — ships with the plugin; there are no personal configs or external services to reproduce. Four prerequisites:

1. Node.js ≥ 22 and pnpm ≥ 10 on PATH (DSH itself already assumes these; dsh plugin add shells out to pnpm and will tell you if it is missing). 2. Write the allowBuilds block before adding (above — the only required pre-step; the profile's pnpm-workspace.yaml is generated on first init, just append those four lines). 3. Model-download network reachability: - Local embedding model (≈585MB, Qwen3-Embedding-0.6B): mirrors apply automatically in China; set hfEndpoint in the panel (base settings → Advanced, or Settings → Local Models) or the HF_ENDPOINT env var for a custom endpoint. - OCR model (≈21MB, PaddleOCR): downloads from hf-mirror.com by default; users outside China should set the same hfEndpoint field to https://huggingface.co — both OCR and embedding downloads then use it. - Skipping the downloads is fine (remote OpenAI-compatible / Ollama embeddings + text PDFs), only local vectorization and scan recognition are unavailable. 4. Platforms: Windows / macOS (Apple Silicon) / Linux x64+arm64 are fully supported; Intel Mac (macOS x64) cannot run local embedding or local OCR (no darwin-x64 onnxruntime binary) — use a remote embedding provider there.

First use after install: download the embedding model in Settings → "Local Models" (or point a base at a remote provider), download the OCR model if you need scan recognition — afterwards the feature set matches this repository's dev machine exactly.

---

Compatibility

  • DSH version: developed and verified against deepseek-harness commit 47f943859b (public npm-plugin-ecosystem era, 2026-08). The plugin declares no peer dependencies — the DSH host injects cordis / zod / storage etc. as externals — so a newer DSH source checkout installs without resolution errors; if a newer DSH release breaks something, report an issue with the DSH commit you run.
  • Node.js: ^22.19.0 || >=24.0.0 (the same floor DSH itself requires — the chunk store uses Node's built-in node:sqlite, which DSH's own session storage also uses).
  • Platforms: Windows / macOS (Apple Silicon) / Linux x64+arm64 fully supported. Legacy .doc / .ppt / .xls parsing uses @firecrawl/anydoc (per-platform native binaries); @napi-rs/canvas's Windows platform package is declared as an optionalDependency and is skipped on other platforms. Intel Mac (darwin-x64): no onnxruntime binary exists for that platform — local embedding and local OCR are unavailable; use a remote embedding provider.
  • First-run network: enabling embeddingProvider: local downloads model weights from Hugging Face on first use (cache under localModelCacheDir); the OCR model (≈21MB) downloads from Settings → Local Models. Both honor the panel's hfEndpoint field or the HF_ENDPOINT env var (OCR defaults to hf-mirror.com; users outside China can switch to huggingface.co).

---

Configuration

Deployment defaults live in the knowledge row of cordis.patch.yml; the panel's Settings can override them at runtime (persisted in the storage domain):

FieldDefaultNotes
embeddingProvidernoneopenai / ollama / local (in-process transformers.js) / none
embeddingBaseUrl''e.g. https://api.openai.com/v1 or http://127.0.0.1:11434 (unused by local)
embeddingModel''e.g. text-embedding-3-small; for local, a Hugging Face repo id (default onnx-community/Qwen3-Embedding-0.6B-ONNX)
embeddingApiKey''optional; KNOWLEDGE_API_KEY env var also works
rerankModel / rerankBaseUrl / rerankApiKey''rerank model (empty = disabled), Jina / SiliconFlow / Cohere v2 style APIs
smartChunktrueheading/paragraph-aware chunking; off = separator only
chunkSeparator\n\nblock boundary when smart chunking is off (\n allowed)
chunkSize800per-chunk token budget (converted to characters with the document's measured chars-per-token ratio; aligned with Cherry Studio)
chunkOverlap100overlap between consecutive chunks (token budget, converted likewise)
topK6results per search (1–50)
searchModeautoauto / hybrid / vector / lexical
similarityThreshold0drop results below this score (0–1)
mmrDiversity0MMR diversity (0–1, 0 = off)
rrfVectorWeight1relative weight of the vector lane in RRF hybrid fusion (0.1–5, 1 = balanced)
embeddingBatchSize32texts per embedding request
siblingChunks1neighbour chunks (±N) attached to each hit as context (0–3, 0 = off)
semanticChunkfalsesemantic chunking: embed paragraph segments, merge adjacent similar ones (per-base overridable)
semanticChunkThreshold0.75merge threshold for semantic chunking (0–1)
chunkTokenLimit0per-chunk token budget (0 = off); oversized chunks split at sentence/comma/space boundaries
conflictStrategyrenamesame-name file import: rename (auto _1 suffix) / replace / keep
urlRefreshHours0auto-refresh interval for URL documents (hours, 0 = off)
imageCaptionProvideroffPDF figure captioning: off / openai (vision-compatible API) / ollama (local VLM)
imageCaptionModel''captioning model id (e.g. qwen2.5vl, gpt-4o-mini)
imageCaptionBaseUrl''captioning API root; empty = embedding base URL (openai) or http://127.0.0.1:11434 (ollama)
imageCaptionApiKey''captioning API key (openai provider)
hfEndpoint''Hugging Face endpoint (download mirror for embedding and OCR models); empty = embeddings use the transformers default, OCR uses hf-mirror.com
documentProcessorProviderbuiltinPDF document processing: builtin (local parsing + optional OCR) / mineru (remote MinerU service)
mineruApiKey''MinerU API Key (needed in mineru mode; global or per-base)
mineruApiHost''MinerU service host; empty = official https://mineru.net
localModelCacheDir''local-model cache root; empty = <DSH_HOME>/cache/dsh-knowledge/local-models (~/.dsh when DSH_HOME is unset)
chunkStorePath''chunk SQLite file; empty = <DSH_HOME>/storages/knowledge-chunks.sqlite

Chunk data is stored in a dedicated SQLite file rather than the domain KV store: on the web profile's JSON backend every record write rewrites the whole unit file, which made deletion and import cost seconds-to-minutes as data grew. The SQLite store makes each put/delete a single statement, FTS5 trigram full-text search (BM25) plus a brute-force vector scan at query time, and bounded reads — resident memory does not grow with the corpus. A one-time migration on first start after an upgrade moves chunks out of a legacy JSON unit into the SQLite store (idempotent; duplicate rows from interrupted migrations are dropped).

> Every field can be overridden per base in the panel's Settings view (empty = inherit global); API keys are stored in plain text on the local machine.

Local (in-process) embeddings and OCR

With embeddingProvider: local, the host runs embeddings in a dedicated worker thread via @huggingface/transformers (+ onnxruntime) — no external service needed. The default model is onnx-community/Qwen3-Embedding-0.6B-ONNX (1024 dims); set embeddingModel to any ONNX embedding repo id on Hugging Face. The first use downloads the weights from the Hub (cached under `$DSH_HOME/cach