dsh-document-reader
一个 DeepSeek Harness (DSH) 网页端插件:在聊天输入框里直接读取 PDF / Word (.docx) / PPT (.pptx) / 纯文本与代码 文件,把提取出的文字填入输入框草稿。
A DeepSeek Harness (DSH) web plugin that reads PDF / Word (.docx) / PPT (.pptx) / plain-text & code files straight into the composer, by file picker or by dragging a file anywhere on the page.
功能 / Features
- 上传读取:输入框工具栏出现「文档」按钮,点击可选择文件(支持多选)。
- 整页拖拽:把文档拖到页面任意位置,出现「释放以读取文档」遮罩,松手即读取。
- 格式支持:
.pdf(服务端pdfjs-dist精确解析)、.docx、.pptx、以及txt / md / csv / json / yaml / 代码等纯文本。 - 输出上限:单文件 20 万字符,超出自动截断。
- Upload: a "文档" button appears in the composer tool row; click to pick files (multi-select supported).
- Drag & drop: drag a document anywhere on the page and release to read it.
- Formats:
.pdf(parsed server-side withpdfjs-dist),.docx,.pptx, and plain text liketxt / md / csv / json / yaml / code. - Output cap: 200k chars per file, truncated with a note.
架构 / Architecture
- 客户端 (
lib/client.js):注册到conversation.input.left插槽,负责 UI、拖拽、以及.docx/.pptx(ZIP + XML)和纯文本的前端解析。 - 服务端 (
lib/index.js):注册POST /read-document/parse路由,用 Mozillapdfjs-dist解析 PDF(正确支持对象流、xref stream、CID 字体、ToUnicode CMap)。
- Client (
lib/client.js): registers into theconversation.input.leftslot; owns the UI, drag-and-drop, and client-side parsing of.docx/.pptx(ZIP + XML) and plain text. - Server (
lib/index.js): registersPOST /read-document/parse, which parses PDFs with Mozillapdfjs-dist(correctly handles object streams, xref streams, CID fonts, ToUnicode CMaps).
安装 / Install
> 需要先安装 dsh 并可用 pnpm。
# 1. 安装到 web 配置(会自动注册到 dsh.profile.bundles)
dsh plugin --profile web add github:Yun-tech123/dsh-document-reader
# 2. 重启 web 服务
dsh web安装后刷新页面,输入框工具栏即出现「文档」按钮。
dsh plugin 需要本机可用 pnpm;若未安装,可先 corepack enable 或 npm i -g pnpm。
手动安装 / Manual install
1. 把本仓库克隆到本地,dsh plugin --profile web add <本地路径>。 2. 确认 package.json 的 dsh.profile.bundles 里已包含 dsh-document-reader(dsh plugin 会自动加入)。 3. 重启 dsh web,刷新页面。
限制 / Limitations
- PDF 为文本提取:纯扫描件(图片型 PDF) 无文本层,无法提取。
- 老式二进制
.doc / .xls / .ppt暂不支持,请先另存为.docx / .pptx。 - 需要较新浏览器(
.docx/.pptx用DecompressionStream,Chrome/Edge/Firefox 2022+ 均可)。
- PDF is text extraction: pure scans (image-only PDFs) have no text layer and cannot be extracted.
- Legacy binary
.doc / .xls / .pptare not supported; save them as.docx / .pptxfirst. - Requires a modern browser (uses
DecompressionStreamfor.docx/.pptx; Chrome/Edge/Firefox 2022+).
License
MIT