DeepSeek Harness plugin

dsh-vision-tool

Image routing for text-only models in DeepSeek Harness: a global analyze_image tool (Kimi vision) plus automatic rewriting of pasted images into attachment references when the active model cannot s...

Jump to install

Source facts

Repository
visail/dsh-vision-tool
Latest update
Aug 14, 2026
Category
Vision & Multimodal
GitHub stars
1
Format
plugin
Catalog evidence
Upstream dsh.bundle evidence
Evidence path
package.json#dsh.bundle
Checked against
0.1.0-rc.8
Upstream check date
2026-08-20

This evidence comes from the upstream catalog. This site has not installed, run, or security-reviewed the plugin.

Install

Start with a prompt that asks an agent to review the GitHub repository and source. Switch to the command if you want to install it yourself.

Copy this prompt into DSH, Codex, or another agent and ask it to review the GitHub repository and source first.

Do not install or run any commands yet. Read this plugin's GitHub repository, README, and relevant source code. Then answer the questions below clearly and directly so I can decide whether it fits my needs:

1. What is this plugin, and what problem does it solve?
2. Who is it for, and what are its typical use cases?
3. How is it used after installation? Include one minimal example.
4. What known limitations or privacy, security, compatibility, or maintenance risks does it have?
5. Give a clear recommendation: recommend, conditionally recommend, or do not recommend, with reasons.

Distinguish statements documented by the repository, inferences from source code, and unknowns. If evidence is insufficient, say so explicitly. Do not guess or simply repeat the README.

GitHub: https://github.com/visail/dsh-vision-tool
Plugin: dsh-vision-tool
Author: visail

Check the source files

Read the README and other files from this plugin directory before installing.

File explorer3 files
README.mdSource · read only

dsh-vision-tool

Image routing for text-only models in DeepSeek Harness (DSH).

Text-only models (e.g. deepseek-v4-flash) cannot see images. This bundle gives them eyes in two coordinated steps:

1. vision-prompt — shadows POST /api/session.prompt. When the active session model does not support image input, pasted images are persisted as content-addressed attachments and rewritten in place into text prompts that carry the full attachment reference JSON. Any other request (no image, or a model that already supports images) is forwarded unchanged, with the official /api trust fence (DNS rebinding / cross-site defense) reimplemented. 2. vision-tool — registers a global analyze_image tool. The model calls it with the attachment reference (or a local file path); the tool routes the image to a vision model and returns its text answer.

Verified end-to-end: paste an image → attachment persisted → model calls analyze_image → vision model describes the image.

Install

Requires the dsh CLI (see package and install a plugin).

# from npm (once published)
dsh plugin --profile <name> add dsh-vision-tool

# straight from a git host (pin a commit; no build step needed — pure ESM)
dsh plugin --profile <name> add github:<you>/dsh-vision-tool#<sha>

# or from a local tarball
pnpm pack
dsh plugin --profile <name> add ./dsh-vision-tool-0.1.0.tgz

The bundle's cordis.patch.yml inserts two rows: vision-tool and vision-prompt. Restart the profile afterwards:

dsh --profile <name>

Verify the layer landed without booting:

dsh --profile <name> --dump-config   # look for the "# == dsh-vision-tool" layer

Configuration

Credential (required)

The tool resolves KIMI_CODE_API_KEY — from $DSH_HOME/.credentials.yaml or an environment variable of the same name. Get the key from your Kimi Code subscription page (sk-kimi- prefix).

# $DSH_HOME/.credentials.yaml
KIMI_CODE_API_KEY: sk-kimi-...

Switching vision models

The defaults target kimi-for-coding at https://api.kimi.com/coding/v1. Override the row in your profile's cordis.patch.yml (a patch replaces the whole config, so restate every key you keep):

- id: vision-tool
  name: dsh-vision-tool
  config:
    baseURL: https://api.kimi.com/coding/v1
    model: kimi-for-coding
    apiKeyEnv: KIMI_CODE_API_KEY
    maxImageBytes: 20971520
    timeoutMs: 120000

> Note: kimi-for-coding only accepts temperature: 1 (anything else is > rejected with HTTP 400). The tool hard-codes temperature: 1 as its default > and is not configurable for this model. Other OpenAI-compatible vision > endpoints generally work as long as they accept temperature: 1.

Supported inputs

  • attachment — full reference JSON injected by the paste-rewrite mechanism

({"attachmentId":"sha256:...","mediaType":...,"bytes":N,"width":N,"height":N}). Pass it verbatim; do not strip fields.

  • path — local image file (absolute, or relative to the session cwd).
  • Formats: png / jpg / jpeg / webp / gif. Local files up to

maxImageBytes (default 20 MB). Attachments are bounded by the harness attachment store limits.

Security

  • vision-prompt reimplements the official /api trust fence: loopback /

trustedHosts host check, sec-fetch-site and Origin checks.

  • Request bodies are capped at 160 MB (413 otherwise), matching the harness

http-bridge default.

  • Any failure degrades to passthrough — the original request is forwarded

unchanged, never swallowed or mangled.

  • The tool only reads the attachment you reference and your configured

credential; it never stores prompt or image content beyond the attachment store the harness itself maintains.

Diagnostics

Both plugins append to $DSH_HOME/vision-trace.log:

handle: rewrite result = REWRITTEN
vision-tool: execute: resolve KIMI_CODE_API_KEY -> source=file len=72

License

[MIT](./LICENSE)