DeepSeek Harness plugin

dsh-eye-vision

dsh-eye-vision: give text-only DeepSeek Harness models eyes — image understanding via any OpenAI-compatible multimodal API. Fork of dsh-free-vision with custom-provider fixes and an

Jump to install

Source facts

Repository
AlloyPlane/dsh-eye-vision
Latest update
Aug 17, 2026
Category
Models & Providers
GitHub stars
1
Format
plugin
Catalog evidence
Upstream dsh.bundle evidence
Evidence path
package.json#dsh.bundle
Checked against
0.1.0-rc.8
Upstream check date
2026-08-20

This evidence comes from the upstream catalog. This site has not installed, run, or security-reviewed the plugin.

Install

Start with a prompt that asks an agent to review the GitHub repository and source. Switch to the command if you want to install it yourself.

Copy this prompt into DSH, Codex, or another agent and ask it to review the GitHub repository and source first.

Do not install or run any commands yet. Read this plugin's GitHub repository, README, and relevant source code. Then answer the questions below clearly and directly so I can decide whether it fits my needs:

1. What is this plugin, and what problem does it solve?
2. Who is it for, and what are its typical use cases?
3. How is it used after installation? Include one minimal example.
4. What known limitations or privacy, security, compatibility, or maintenance risks does it have?
5. Give a clear recommendation: recommend, conditionally recommend, or do not recommend, with reasons.

Distinguish statements documented by the repository, inferences from source code, and unknowns. If evidence is insufficient, say so explicitly. Do not guess or simply repeat the README.

GitHub: https://github.com/AlloyPlane/dsh-eye-vision
Plugin: dsh-eye-vision
Author: AlloyPlane

Check the source files

Read the README and other files from this plugin directory before installing.

File explorer3 files
README.mdSource · read only

dsh-eye-vision

Give text-only DeepSeek Harness models eyes — image understanding, OCR, and UI analysis through any OpenAI-compatible multimodal API.

Fork of dsh-free-vision (MIT) v1.0.1, with:

  • custom-provider fixCUSTOM_MODEL_NAME is now forwarded, so arbitrary OpenAI-compatible endpoints (GPT-4o, Qwen-VL, GLM-4V, youtu-vita, vLLM, Ollama…) actually start
  • allowed-directories whitelist — the vision engine can read images from your configured workspace roots, not just the engine CWD and home directory

How it works

You send an image path
  → agent calls the image_understand tool
    → plugin spawns the luma-mcp vision engine (child process)
      → engine calls your multimodal API
        → text description comes back to the text-only model

The main model never needs image input support. Paste the image path, get answers.

Features

  • image_understand tool registered on ctx.tools, visible to every session in the profile
  • Any OpenAI-compatible endpoint via custom provider — bring your own multimodal API
  • Free-tier providers built in: qwen (Qwen3-VL-Flash), volcengine (Doubao), siliconflow (DeepSeek-OCR)
  • Multi-crop for large images (detail preservation)
  • Allowed-directories whitelist (LUMA_ALLOWED_DIRS) — read images from your workspace
  • Proxy vars stripped for direct mainland-China API access
  • Live settings — save via the settings route, no restart needed

Installation

# from a local checkout
dsh plugin --profile web add D:/xd/dsh-eye-vision

# once published
dsh plugin --profile web add dsh-eye-vision

Restart dsh web. The tool appears as image_understand.

Configuration

Settings file: ~/.dsh/free-vision.json (same path as upstream for drop-in compatibility):

{
  "modelProvider": "custom",
  "baseURLs": { "custom": "https://your-api.example.com/v1" },
  "modelName": "your-vision-model",
  "apiKey": "sk-...",
  "allowedDirs": "D:/workspace",
  "toolName": "image_understand"
}

Or use environment variables (fallback chain: settings file > env):

ProviderKey envBase URL env
customCUSTOM_API_KEYCUSTOM_BASE_URL + CUSTOM_MODEL_NAME
qwenDASHSCOPE_API_KEYQWEN_BASE_URL
volcengineVOLCENGINE_API_KEYVOLCENGINE_BASE_URL
siliconflowSILICONFLOW_API_KEYSILICONFLOW_BASE_URL
zhipuZHIPU_API_KEYZHIPU_BASE_URL
hunyuanHUNYUAN_API_KEYHUNYUAN_BASE_URL

allowedDirs: semicolon/comma-separated extra roots the engine may read images from (default: engine CWD + home directory).

Usage

看图:D:/path/to/screenshot.png
OCR:D:/path/to/document.png
UI 分析:D:/path/to/design.png (task_type: ui)

Tool arguments: image_source (local path / http(s) URL / data URI), prompt, task_type (auto|general|ocr|ui|debug|describe). PNG/JPG/WebP/GIF up to ~10 MB.

Security

  • API keys live only in the settings file or environment variables — never in this repo
  • Images are sent only to the endpoint you configured
  • Proxy environment variables are deliberately stripped from the engine child process
  • A gitguard-style pre-push scan is recommended before publishing forks

Development

cd dsh-eye-vision
pnpm install   # installs luma-mcp engine + MCP SDK
pnpm test

The engine patches in scripts/patch-luma.mjs re-apply automatically on install (idempotent, pinned to luma-mcp 1.7.1).

License

MIT — see [LICENSE](LICENSE). Upstream: dsh-free-vision (MIT) by FuzzySoul.