DeepSeek Harness 插件

dsh-image-reader

Give DeepSeek Harness agents the ability to read images directly: a model-facing read_image tool that answers questions about an image through any OpenAI-compatible vision endpoint.(英文原文)

跳到安装方式

来源信息

GitHub 仓库
zcXie777/dsh-image-reader
最近更新
2026年8月14日
分类
视觉与多模态
GitHub stars
3
载体类型
plugin
目录证据
上游声明已找到 dsh.bundle
证据路径
package.json#dsh.bundle
核对版本
0.1.0-rc.8
上游核对日期
2026-08-20

该证据由上游目录提供。本站没有安装、运行或安全审核这个插件。

安装

默认先复制一段 Prompt,让 Agent 读 GitHub 仓库和源码;需要自己装时再切到命令。

复制这段 Prompt,发给 DSH、Codex 或其他 Agent,让它先读 GitHub 仓库和源码。

请先不要安装或执行任何命令。阅读这个插件的 GitHub 仓库、README 和关键源码,然后用清楚、直接的方式回答以下问题,帮助我判断它是否适合我的需求:

1. 这个插件是什么,解决什么问题;
2. 适合哪些用户和典型使用场景;
3. 安装后如何使用,并给出一个最小使用示例;
4. 有哪些已知限制,以及隐私、安全、兼容性或维护风险;
5. 给出“推荐 / 有条件推荐 / 不推荐”的明确建议和理由。

请区分仓库明确说明、根据源码推断和未知信息。证据不足时请明确说明,不要猜测或照抄 README。

GitHub:https://github.com/zcXie777/dsh-image-reader
插件名:dsh-image-reader
作者:zcXie777

检查来源文件

安装前先看这个插件目录里的 README 和其他文件。

文件资源管理器3 个文件
README.md来源说明 · 只读预览

dsh-image-reader

Give a text-only DeepSeek Harness agent the ability to read images directly: one model-facing read_image tool that asks any OpenAI-compatible vision endpoint about an image by its workspace path.

Why

DeepSeek Harness is "everything is a plugin". This bundle mounts a single tool so the model can look at a screenshot, diagram, or photograph and answer questions about it, instead of only ever reasoning over text.

Verification status

  • Verified locally: npm run typecheck, npm run build, and npm test (16 tests) all pass.
  • Not yet verified: a real end-to-end read against a live vision endpoint. The request/response logic is covered by a mocked-fetch unit test, but the plugin has not been smoke-tested inside a running dsh profile against a real multimodal model. Do that once with a real VISION_API_KEY before relying on it.

Install

git clone https://github.com/zcXie777/dsh-image-reader.git
cd dsh-image-reader
npm install
npm run build          # lib/ is not committed; build once after cloning
cd ..
dsh plugin --profile web add "$PWD/dsh-image-reader"
dsh plugin --profile headless add "$PWD/dsh-image-reader"
dsh --profile web --dump-config | grep image-reader

Restart a running Web profile after installing.

Configure

provider.baseUrl and provider.model are required; the plugin never assumes a vendor. Override them in the profile patch row with the same id:

- id: image-reader
  config:
    provider:
      baseUrl: https://api.openai.com/v1
      model: gpt-4o-mini
      apiKeyEnv: VISION_API_KEY
    lang: zh
    timeoutMs: 60000
    maxImageBytes: 10485760
    allowedDirs: []

Set the key in the environment before starting the profile:

export VISION_API_KEY=sk-...

Use

In a conversation, point the model at an image path and ask:

read_image image="screenshot.png" query="What error is shown in this dialog?"
read_image image="diagram.png"

Configuration fields

FieldDefaultContract
provider.baseUrl— (required)OpenAI-compatible chat/completions base URL
provider.model— (required)Multimodal model name
provider.apiKeyEnvVISION_API_KEYEnvironment variable holding the API key
langzhAnswer language: zh or en
timeoutMs60000Whole-request deadline, 1000–600000 ms
maxImageBytes10485760Encoded-byte limit per image
allowedDirs[]Extra realpath-resolved input roots; the workspace is always allowed

Security

  • Inputs resolve against the workspace and allowedDirs through realpath, so a symlink cannot escape the fence.
  • Images are size-limited and extension-checked before upload.
  • The key is read from the environment per call, never stored in config.

Development

npm install
npm run typecheck
npm run build

Publish

Tag the repo with the dsh-plugin topic so it is discoverable, and publish to npm when ready.

License

MIT