DeepSeek Harness 插件

dsh-tool-vision

DeepSeek Harness 外置视觉模型插件:inspect_image 把本地图片或 http(s) 图片 URL 发给任意 OpenAI 兼容端点,视觉模型看图的文字回答直接带回对话;附 Web UI 设置栏。(英文原文)

跳到安装方式

来源信息

GitHub 仓库
Scorp1o117/dsh-tool-vision
最近更新
2026年8月21日
分类
模型与服务商
GitHub stars
6
载体类型
plugin
目录证据
上游声明已找到 dsh.bundle
证据路径
package.json#dsh.bundle
核对版本
0.1.0-rc.8
上游核对日期
2026-08-20

该证据由上游目录提供。本站没有安装、运行或安全审核这个插件。

安装

默认先复制一段 Prompt,让 Agent 读 GitHub 仓库和源码;需要自己装时再切到命令。

复制这段 Prompt,发给 DSH、Codex 或其他 Agent,让它先读 GitHub 仓库和源码。

请先不要安装或执行任何命令。阅读这个插件的 GitHub 仓库、README 和关键源码,然后用清楚、直接的方式回答以下问题,帮助我判断它是否适合我的需求:

1. 这个插件是什么,解决什么问题;
2. 适合哪些用户和典型使用场景;
3. 安装后如何使用,并给出一个最小使用示例;
4. 有哪些已知限制,以及隐私、安全、兼容性或维护风险;
5. 给出“推荐 / 有条件推荐 / 不推荐”的明确建议和理由。

请区分仓库明确说明、根据源码推断和未知信息。证据不足时请明确说明,不要猜测或照抄 README。

GitHub:https://github.com/Scorp1o117/dsh-tool-vision
插件名:dsh-tool-vision
作者:Scorp1o117

检查来源文件

安装前先看这个插件目录里的 README 和其他文件。

文件资源管理器3 个文件
README.md来源说明 · 只读预览

dsh-tool-vision

![中文文档](README.zh.md)

GitHub: Scorp1o117/dsh-tool-vision · npm: dsh-tool-vision

![Enhancement Suite](https://github.com/Scorp1o117/dsh-enhancement-suite) ![npm](https://www.npmjs.com/package/dsh-enhancement-suite)

Part of the DeepSeek Harness Enhancement Suite — Vision · Soul/Persona · Long-term Memory · Plugin Marketplace.

External vision model for DeepSeek Harness.

DSH 0.1.1 adds native image input for DeepSeek's vision catalog. This plugin remains useful when you want a separate OpenAI-compatible vision endpoint, pixel-level image tools, screenshots, or a text-model bridge. The harness derives every model request strictly from the session log (llm/stream requests must equal the durable derivation — the agent-loop invariant), so the bridge keeps its conversion inside that durable path:

1. inspect_image tool — sends an image (local file, or http(s) URL) to any OpenAI-compatible /chat/completions endpoint that supports image_url content parts, and returns the vision model's textual answer into the agent loop. 2. Image bridge (v0.2.1) — pasted images are turned into inspect_image hints before they enter the durable log, on the agent/pre-step waterfall (the one seam where the harness lets a plugin replace the messages of a proposed step). Images already logged by an older version are repaired lazily with a surface replace on the session's first pre-step. Only models listed in multimodalModels receive image blocks directly; a model's declared inputModalities are never consulted, because profiles routinely declare input: [text, image] on text-only models just to pass the harness's prompt-admission check.

  • Zero dependencies beyond the dsh SDK — works with any compatible endpoint:

OpenAI GPT-4o, Qwen-VL (DashScope), GLM-4V (Zhipu), Moonshot, Gemini compatible endpoints, local Ollama, etc.

  • Registered on the global tools layer: every agent in the process can

call inspect_image.

  • Web UI settings section (v0.3.0): Settings → 视觉模型 edits the

tool-vision namespace (API endpoint, write-only key, model, bridge options) in settings.yaml; changes hot-apply without a restart. The API key lives in settings.yaml, not the profile patch. Mount by package name (name: 'dsh-tool-vision') so the web client bundle is discovered.

Install

Mount in a profile patch ($DSH_HOME/profiles/<name>/cordis.patch.yml):

- insert:
    - id: tool-vision
      name: 'dsh-tool-vision'     # after: pnpm add dsh-tool-vision in the profile
      config:
        baseURL: 'https://api.openai.com/v1'
        apiKeyEnv: 'VISION_API_KEY'
        model: 'gpt-4o-mini'

Or load it from a local path without npm:

    - id: tool-vision
      name: './plugins/dsh-tool-vision/index.js'

Config

FieldDefaultMeaning
baseURLhttps://api.openai.com/v1OpenAI-compatible API base URL.
apiKey''API key (takes precedence over env).
apiKeyEnvVISION_API_KEYEnv var holding the key.
modelgpt-4o-miniVision model id.
maxTokens1024Max output tokens.
timeoutMs60000Per-request timeout.
maxImageBytes10MBLargest accepted local image.
descriptiondefaultTool description shown to the model.
bridgeTextOnlytrueBridge pasted images to text hints on models that cannot see images.
bridgeExportDirtempExport dir for bridged images (os.tmpdir()/dsh-vision-bridge).
multimodalModels[]Model ids that receive image blocks directly (e.g. mimo-v2.5).
bridgePreviewtrueInline preview for bridged images: thumbnail above the hint text in the user bubble (click to zoom).
bridgePreviewScanIntervalMs2000Fallback scan interval for the preview scanner (ms); 0 disables the fallback.
bridgePreviewHideHinttrueHide the bridged hint text once the preview image has loaded (kept on failure — safe degradation).
bridgeAutoImagetrueWhile the bridge is on, report image input capability for every model to the host admission gate, so pasted images are accepted on text-only models without hand-editing provider configs.

Image bridge setup

1. (Optional, usually not needed) If bridgeAutoImage is disabled, declare image input on the models you paste images onto, so the harness admits image messages (pi-ai style): ``yaml llm-pi-ai: providers: your-provider: models: - id: deepseek-v4-flash input: [text, image] ` 2. List genuinely multimodal models in the plugin config so they receive image blocks untouched: `yaml - id: tool-vision name: 'dsh-tool-vision' config: multimodalModels: ['mimo-v2.5', 'grok-4.5'] ``

Then pasting an image while on a text-only model stores a hint like [User sent an image, exported to: <path>. Inspect it with the inspect_image tool...] in the transcript (the pasted image no longer renders as pixels in that message), and the agent inspects it through the configured vision endpoint.

> Why not llm/stream? The harness freezes every request and the agent-loop > invariant fails any request whose messages diverge from the session-log > derivation (log-reconstruction desync), and this cordis waterfall's > next() cannot replace request arguments. The agent/pre-step waterfall is > the supported seam: its decision messages become the durable log, so the > invariant stays satisfied.

Key resolution order: config.apiKeyprocess.env[apiKeyEnv]process.env.OPENAI_API_KEY.

Bridge image preview (v0.4.0)

On text-only models, pasted images become [User sent an image...] hint text in the transcript. With bridgePreview enabled (default), the browser half renders those hints as inline thumbnails in the display layer only:

  • Thumbnail + lightbox: click to zoom full-screen; click anywhere or

press Esc to close;

  • Immediate + fallback: new messages are handled by a MutationObserver;

history is back-filled by a periodic scan (interval via bridgePreviewScanIntervalMs);

  • Hide the hint (P2): with bridgePreviewHideHint on, the hint text is

hidden once the image has loaded, leaving just the image; on load failure the text stays (safe degradation — never "no image AND no text");

  • Precise identification: bridged hints carry an invisible prefix marker

(\u200b[bridge]), so ordinary user text that happens to contain "exported to:" is never misidentified;

  • Display-layer red line: persisted messages, the transcript, the

model-facing text and the inspect_image chain are untouched.

Preview images are served by the same-origin loopback route /plugins/dsh-tool-vision/image: read-only access to the bridge export directory, localhost-only Host, image extensions only, ≤ 20MB per file, path-traversal protected.

Tool: inspect_image

ArgRequiredMeaning
pathImage path (absolute, or relative to the current workspace) or http(s) URL.
questionOptional specific question about the image.
detailauto / low / high resolution hint.

Example endpoints (baseURL):

  • OpenAI: https://api.openai.com/v1gpt-4o, gpt-4o-mini
  • Alibaba DashScope (Qwen-VL): https://dashscope.aliyuncs.com/compatible-mode/v1qwen-vl-plus, qwen-vl-max
  • Zhipu (GLM-4V): https://open.bigmodel.cn/api/paas/v4glm-4v-flash (free tier), glm-4v-plus
  • Moonshot (Kimi): https://api.moonshot.cn/v1moonshot-v1-8k-vision-preview
  • Ollama local: http://localhost:11434/v1llama3.2-vision (no key)

> Note for users > - This plugin is a standard profile bundle (dsh.bundle.patch): > dsh plugin --profile web add dsh-tool-vision installs and mounts it in > one step — no manual cordis.patch.yml edits needed. > - Settings changes hot-apply (no restart needed). > - Version 0.6.3 and newer require DSH 0.1.0-rc.7 or newer and are tested > against 0.1.0-rc.7, 0.1.0-rc.8, and 0.1.1-rc.1. > - DSH 0.1.0-rc.6 users must pin dsh-tool-vision@0.6.1, the last release > carrying the legacy settings-allowlist compatibility patch.

Pixel-level vision tools (v0.6.0, ported from dsh-vision-router)

14 vision_* tools driven by the same configured endpoint as inspect_image (baseURL/apiKey/model) — no provider chain, no local models, no extra settings:

ToolPurpose
vision_describeImage Q&A / multi-image comparison (optional structured JSON)
vision_groundLocate a target and return its ORIGINAL-pixel bounding box
vision_detectEnumerate elements (buttons, inputs, icons…) with numbered boxes
vision_cropCrop a pixel region to a PNG artifact
vision_pixel_diffPer-pixel comparison: ratio, worst regions, heatmap, report
vision_colorsDominant-color quantization for palette matching
vision_ocrVerbatim text transcription (letters only — not scene analysis)
vision_long_screenshot_ocrChunked long-screenshot transcription into Markdown
vision_tracePotrace vectorization into colored SVG (worker-thread, safe)
vision_extract_foregroundSolid-background removal → transparent PNG
vision_html_screenshotHeadless render of a local .html (network blocked)
vision_screenshotDesktop capture (privacy-gated: enable desktopScreenshot in settings; Win: PowerShell / macOS: screencapture / Linux: import/scrot)
vision_presentPublish a generated image to the user via the host attachment store
vision_materializeCopy an attachment/local image into the workspace as a real path

Quality & safety details:

  • Content-hash cache keyed by endpoint+model+image+question (no stale

answers across model switches, failures are never cached).

  • Uniform 4MP downscale before every model call; oversized inputs are

rejected with a clear error (stat pre-check, 20MB cap on both file and attachment paths).

  • Rate-limit / 5xx auto-retry with Retry-After-aware backoff; endpoint

content-safety rejections are surfaced as VISION_CONTENT_FILTERED instead of a generic backend error.

  • Long-OCR bounds: 120s total budget, 40-chunk cap, cancellation checks,

stop-on-first-backend-failure.

  • Path containment for relative inputs; artifacts land in

<workspace>/.dsh-tool-vision/.

Requires sharp / potrace / puppeteer-core (declared as optional dependencies: a failed platform install never blocks the plugin; missing ones degrade lazily with an install hint and never break other tools).

vision_screenshot is privacy-sensitive and therefore not registered by default — set desktopScreenshot: true in the tool-vision settings to enable desktop capture.

Limitations

  • A bridged image enters the conversation as a text hint (a transcript, not

pixels) — pixel-precise in-context reasoning is not available to text-only models; the vision model's description comes back through inspect_image.

  • Images are base64-transferred; mind privacy and size limits.
  • Independent of the dsh-llm routing/retry system; failures return clear

errors to the agent.

License

MIT — bridge preview & integration: xing666173. Pixel vision tools ported from dsh-vision-router (© ysr666, MIT) with gratitude.