DeepSeek Harness 插件

dsh-bundle-vision

Vision bundle + plugin for DeepSeek Harness: the describe_image tool reads local images and asks any configured multimodal route, with zero core changes(英文原文)

跳到安装方式

来源信息

GitHub 仓库
skillre/dsh-bundle-vision
最近更新
2026年8月15日
分类
工具与能力
GitHub stars
0
载体类型
plugin
目录证据
上游声明已找到 dsh.bundle
证据路径
package.json#dsh.bundle
核对版本
0.1.0-rc.8
上游核对日期
2026-08-20

该证据由上游目录提供。本站没有安装、运行或安全审核这个插件。

安装

默认先复制一段 Prompt,让 Agent 读 GitHub 仓库和源码;需要自己装时再切到命令。

复制这段 Prompt,发给 DSH、Codex 或其他 Agent,让它先读 GitHub 仓库和源码。

请先不要安装或执行任何命令。阅读这个插件的 GitHub 仓库、README 和关键源码,然后用清楚、直接的方式回答以下问题,帮助我判断它是否适合我的需求:

1. 这个插件是什么,解决什么问题;
2. 适合哪些用户和典型使用场景;
3. 安装后如何使用,并给出一个最小使用示例;
4. 有哪些已知限制,以及隐私、安全、兼容性或维护风险;
5. 给出“推荐 / 有条件推荐 / 不推荐”的明确建议和理由。

请区分仓库明确说明、根据源码推断和未知信息。证据不足时请明确说明,不要猜测或照抄 README。

GitHub:https://github.com/skillre/dsh-bundle-vision
插件名:dsh-bundle-vision
作者:skillre

检查来源文件

安装前先看这个插件目录里的 README 和其他文件。

文件资源管理器3 个文件
README.md来源说明 · 只读预览

dsh-bundle-vision

A zero-core-change vision capability for DeepSeek Harness, shipped as one installable npm package that is both a profile bundle and a plugin:

  • the plugin registers the model-facing describe_image tool;
  • the bundle patch mounts that plugin on any profile.

The tool reads a local PNG/JPEG/WebP/GIF file, commits the bytes through the shipped attachment service, and asks the named multimodal route about it in one direct LLM request (provider / model are tool arguments). The result is text only — no image block ever enters the calling session, so a text-only main model gains vision without any change to dsh itself.

How it works against the shipped seams

Everything the tool uses already ships with every dsh profile:

  • ctx.fs (bounded byte read, session-workspace resolution) — filesystem capability;
  • ctx.attachments (saveImage, image limits, magic-byte validation) — durable image storage;
  • ctx.llm (resolveModelInfo + stream) with the pi-ai multi-provider adapter — the multimodal request itself.

The one per-deployment prerequisite is the same as for any vision use of dsh: the multimodal model must declare image input in the llm-pi-ai settings section, e.g.:

llm-pi-ai:
  providers:
    my-vision:
      apiKeyEnv: MY_VISION_API_KEY
      api: openai-completions
      baseURL: https://example.invalid/v1
      models:
        - id: my-vision-model
          input: [text, image]

(On releases whose Models page has the input-modality control, the same declaration is one dropdown.)

Install (installed dsh, no source checkout)

From the npm registry:

dsh plugin --profile <name> add dsh-bundle-vision

Or from a packed tarball (e.g. before the first publish, or for a pinned version):

dsh plugin --profile <name> add ./dsh-bundle-vision-0.1.0.tgz

dsh plugin forwards to pnpm inside the profile directory and reconciles the profile's bundle layers automatically — a dependency declaring dsh.bundle joins the layer stack. Restart dsh <name>; the tool registers for every agent (profile-root registrations are visible to all preset scopes).

Use

Ask the main model, naming the route:

> Use describe_image with file_path /path/to/photo.jpg, provider my-vision, model my-vision-model, and prompt "OCR the text in this image".

The main model supplies provider/model per call — configure several multimodal routes and switch per call, no profile edits. Refusals name the failing gate (unknown extension, deployment media types, a route whose model does not declare image input, a missing file, a type-mismatched file, an errored stream).

Version floor

The package declares >=0.1.0-rc.6 peer dependencies on the dsh seam packages. Bump the floor if a later release is required.

Development

npm install          # dev deps resolve the seam packages from the npm registry
npm run typecheck    # tsc --noEmit over src + tests
npm test             # vitest
npm run build        # tsdown (lib/index.js) + tsc declarations (lib/types)
npm pack             # the installable tarball