DeepSeek Harness 插件

dsh-pdf-reader

DSH PDF 读取插件:注册 read_pdf 工具,用 pdfjs-dist 提取 PDF 文本(路径限定工作区内,纯本地解析)。(英文原文)

跳到安装方式

来源信息

GitHub 仓库
ralfsqual/dsh-pdf-reader
最近更新
2026年8月15日
分类
工具与能力
GitHub stars
0
载体类型
plugin
目录证据
上游声明已找到 dsh.bundle
证据路径
package.json#dsh.bundle
核对版本
0.1.0-rc.8
上游核对日期
2026-08-20

该证据由上游目录提供。本站没有安装、运行或安全审核这个插件。

安装

默认先复制一段 Prompt,让 Agent 读 GitHub 仓库和源码;需要自己装时再切到命令。

复制这段 Prompt,发给 DSH、Codex 或其他 Agent,让它先读 GitHub 仓库和源码。

请先不要安装或执行任何命令。阅读这个插件的 GitHub 仓库、README 和关键源码,然后用清楚、直接的方式回答以下问题,帮助我判断它是否适合我的需求:

1. 这个插件是什么,解决什么问题;
2. 适合哪些用户和典型使用场景;
3. 安装后如何使用,并给出一个最小使用示例;
4. 有哪些已知限制,以及隐私、安全、兼容性或维护风险;
5. 给出“推荐 / 有条件推荐 / 不推荐”的明确建议和理由。

请区分仓库明确说明、根据源码推断和未知信息。证据不足时请明确说明,不要猜测或照抄 README。

GitHub:https://github.com/ralfsqual/dsh-pdf-reader
插件名:dsh-pdf-reader
作者:ralfsqual

检查来源文件

安装前先看这个插件目录里的 README 和其他文件。

文件资源管理器3 个文件
README.md来源说明 · 只读预览

dsh-pdf-reader

A DeepSeek Harness plugin that adds a read_pdf tool, extracting text from PDF files page by page. Fully local parsing — no upload, no network access.

为 DeepSeek Harness 增加 read_pdf 工具,逐页提取 PDF 文本。纯本地解析,不上传、不联网。

Install / 安装

dsh plugin --profile web add github:ralfsqual/dsh-pdf-reader

Or install from a local checkout / 或本地安装:

dsh plugin --profile web add /path/to/dsh-pdf-reader

Restart dsh web afterwards. The read_pdf tool becomes available to the agent in new sessions.

重启 dsh web 后生效,新会话中 Agent 即可调用 read_pdf

Usage / 使用

Tell the agent to read a PDF, or combine it with an @ file mention (e.g. via dsh-at-file):

Read @docs/spec.pdf and summarize the requirements.
读取 @docs/report.pdf 并总结要点。

Tool parameters

ParameterTypeDescription
pathstringPDF path, relative to the current working directory or an absolute path inside the workspace. / 相对当前工作目录或工作区内绝对路径
maxPagesnumberMax pages to extract, default 500; 0 = no limit. / 最多提取页数
maxCharsnumberMax characters to return, default 120000. / 最多返回字符数

How it works / 工作原理

  • Registered as a model tool via defineTool (@deepseek-ai/dsh-tools).
  • Parses the PDF with pdfjs-dist (2.6.347), extracting text per page with --- page N --- separators.
  • Output is a structured object (pages, extractedPages, truncated, empty, text) with a readable render.

Security / 安全

  • Workspace-confined: path must resolve inside the current working directory; escape attempts (e.g. ../) are rejected. / 路径限定工作区内,越界拒绝。
  • Read-only: the tool only reads the file, never writes, never executes commands. / 只读,绝不写入或执行命令。
  • Local only: parsing happens entirely on the host; no network requests are made. / 纯本地解析,无任何网络请求。
  • Password-protected, corrupted, or non-PDF files produce clear error messages.

Prerequisites / 环境要求

  • DeepSeek Harness with a web (or any agent-capable) profile.
  • Node.js >= 22.19.
  • Works on Windows, macOS, and Linux (host process only; no browser/native code).

Limitations / 已知限制

  • Extracts text layers only. Scanned/image-only PDFs contain no extractable text (the tool reports empty; OCR is out of scope). / 仅提取文本层,扫描件需 OCR。
  • Paths must not contain a leading @ when combined with @-mention plugins. / 与 @ 引用插件合用时路径不能以 @ 开头。
  • Very large PDFs are bounded by maxPages/maxChars; output is truncated with a truncated: true marker.

License

MIT — see [LICENSE](./LICENSE).