DeepSeek Harness plugin

dsh-ext-vision-proxy

External vision proxy plugin for DeepSeek Harness, enabling text-only models (DeepSeek-V3 / R1) to analyze and understand images via OpenAI-compatible vision APIs.

Jump to install

Source facts

Repository
jin123-alpha/dsh-ext-vision-proxy
Latest update
Aug 19, 2026
Category
Models & Providers
GitHub stars
1
Format
plugin
Catalog evidence
Upstream dsh.bundle evidence
Evidence path
package.json#dsh.bundle
Checked against
0.1.0-rc.8
Upstream check date
2026-08-20

This evidence comes from the upstream catalog. This site has not installed, run, or security-reviewed the plugin.

Install

Start with a prompt that asks an agent to review the GitHub repository and source. Switch to the command if you want to install it yourself.

Copy this prompt into DSH, Codex, or another agent and ask it to review the GitHub repository and source first.

Do not install or run any commands yet. Read this plugin's GitHub repository, README, and relevant source code. Then answer the questions below clearly and directly so I can decide whether it fits my needs:

1. What is this plugin, and what problem does it solve?
2. Who is it for, and what are its typical use cases?
3. How is it used after installation? Include one minimal example.
4. What known limitations or privacy, security, compatibility, or maintenance risks does it have?
5. Give a clear recommendation: recommend, conditionally recommend, or do not recommend, with reasons.

Distinguish statements documented by the repository, inferences from source code, and unknowns. If evidence is insufficient, say so explicitly. Do not guess or simply repeat the README.

GitHub: https://github.com/jin123-alpha/dsh-ext-vision-proxy
Plugin: dsh-ext-vision-proxy
Author: jin123-alpha

Check the source files

Read the README and other files from this plugin directory before installing.

File explorer4 files
README.mdSource · read only
README language

dsh-ext-vision-proxy

English | 简体中文

> External vision proxy plugin for DeepSeek Harness, enabling text-only LLMs (such as DeepSeek-V3 / DeepSeek-R1) to inspect, analyze, and understand images via OpenAI-compatible vision model APIs.

---

📸 Screenshots

1. Settings Panel (Configuration & Provider Management)

![Settings Preview](docs/assets/settings-preview.png)

2. Chat Composer Switch (Pill-shaped Toggle & Model Selector)

![Composer Preview](docs/assets/composer-preview.png)

---

🌟 Key Features

  • 🖼️ Multimodal Power for Text-Only Models: Allows text-only LLMs to autonomously invoke the vision_describe tool to inspect and answer questions about images, screenshots, diagrams, and attachments.
  • 🔘 Native Composer Switch: Integrates seamlessly into the chat input bar (conversation.input.left) with a pill-shaped toggle switch and an interactive model dropdown selector.
  • ⚙️ Dedicated Settings & Connectivity Testing: Provides an intuitive configuration panel in the Settings page supporting custom Base URLs, API Keys, vision model filtering, and one-click connectivity testing.
  • 🎛️ 4-State Session Matrix:

1. Multimodal LLM + Switch ON: vision_describe tool is visible. Model can choose native or delegated recognition. Prompts confirmation dialog upon image attachment. 2. Multimodal LLM + Switch OFF: Tool is stripped. Model processes images natively via standard DSH image messaging with zero proxy interference. 3. Text-only LLM + Switch ON: Tool is visible. Images are admitted without rejection; the LLM automatically invokes vision_describe with the attachment ID or file path. 4. Text-only LLM + Switch OFF: Tool is invisible. Behavior is 100% identical to not having the plugin installed.

  • 🔍 Intelligent Attachment & Image Resolution: Resolves content-addressed hashes (sha256:...), local file paths, HTTP/HTTPS URLs, and automatically detects MIME types using binary magic bytes.

---

🏗️ Architecture

  • Host (Node.js): src/index.ts

- Registers the vision_describe tool schema and execution handler. - Dynamically filters tool exposure via the Cordis system-prompt/assemble waterfall hook based on session state. - Exposes REST API endpoints on webServer for settings, model probing, and session state sync.

  • Client (React / Browser): src/client/

- Registers the settings.section slot (VisionSettingsSection) for provider configuration. - Registers the conversation.input.left slot (VisionComposerButton) for composer toggling and fast model switching.

---

📦 Installation & Setup

Option 1: Link in DSH Profile (Recommended for Development)

Add the dependency and bundle entry in your ~/.dsh/profiles/web/package.json:

{
  "dependencies": {
    "dsh-ext-vision-proxy": "link:/path/to/dsh-ext-vision-proxy"
  },
  "dsh": {
    "profile": {
      "bundles": [
        "@deepseek-ai/dsh-base",
        "@deepseek-ai/dsh-web-app",
        "dsh-ext-vision-proxy"
      ]
    }
  }
}

Option 2: Build & Pack

# Install dependencies and build bundles
npm install
npm run build

# Pack tarball
npm pack

---

🚀 Getting Started

1. Start DSH Web: ``bash dsh web ` 2. Configure Vision API: - Open DSH Web in your browser (http://127.0.0.1:3080). - Navigate to Settings -> 视觉代理 (Vision Proxy). - Enter your OpenAI-compatible Vision API Base URL (e.g., https://api.openai.com/v1 or local endpoint) and API Key. - Click 获取模型列表 (Fetch Models) and select your desired default vision model (e.g., gpt-4o, gemini-2.5-flash, qwen-vl-max). - Click 测试连通性 (Test Connection) to verify API connectivity, then save settings. 3. Chat & Inspect Images: - In any conversation with DeepSeek-V3 or DeepSeek-R1, turn the 视觉代理 switch ON. - Upload or paste an image and ask questions (e.g., "What is shown in this image?"). - The LLM will call vision_describe` behind the scenes and respond with detailed answers.

---

📄 License

[MIT](LICENSE) © 2026 Fan Yuejin