DeepSeek Harness plugin

dsh-plugin-vision-toolkit

Vision toolkit for DeepSeek Harness -- glance, ground, detect, crop CLI tools for text-only agents to understand images

Jump to install

Source facts

Repository
YYTbit/dsh-plugin-vision-toolkit
Latest update
Aug 13, 2026
Category
Vision & Multimodal
GitHub stars
1
Format
plugin
Catalog evidence
Upstream dsh.bundle evidence
Evidence path
package.json#dsh.bundle
Checked against
0.1.0-rc.8
Upstream check date
2026-08-20

This evidence comes from the upstream catalog. This site has not installed, run, or security-reviewed the plugin.

Install

Start with a prompt that asks an agent to review the GitHub repository and source. Switch to the command if you want to install it yourself.

Copy this prompt into DSH, Codex, or another agent and ask it to review the GitHub repository and source first.

Do not install or run any commands yet. Read this plugin's GitHub repository, README, and relevant source code. Then answer the questions below clearly and directly so I can decide whether it fits my needs:

1. What is this plugin, and what problem does it solve?
2. Who is it for, and what are its typical use cases?
3. How is it used after installation? Include one minimal example.
4. What known limitations or privacy, security, compatibility, or maintenance risks does it have?
5. Give a clear recommendation: recommend, conditionally recommend, or do not recommend, with reasons.

Distinguish statements documented by the repository, inferences from source code, and unknowns. If evidence is insufficient, say so explicitly. Do not guess or simply repeat the README.

GitHub: https://github.com/YYTbit/dsh-plugin-vision-toolkit
Plugin: dsh-plugin-vision-toolkit
Author: YYTbit

Check the source files

Read the README and other files from this plugin directory before installing.

File explorer3 files
README.mdSource · read only

dsh-plugin-vision-toolkit

Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images.

What it does

Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them.

Tools

  • glance -- describe, ask about, or OCR an image
  • ground -- locate a specific element (returns bounding box)
  • detect -- find all instances of an element kind
  • crop -- cut a region from an image

Install

dsh plugin --profile your-profile add dsh-plugin-vision-toolkit

Configuration

Set environment variables:

export VISION_API_KEY=sk-xxx           # Vision API key (falls back to DEEPSEEK_API_KEY)
export VISION_BASE_URL=https://...     # API endpoint (falls back to DEEPSEEK_BASE_URL)
export VISION_MODEL=deepseek-vl2       # Vision model name

Usage examples

# Describe an image
glance screenshot.png

# Ask a question
glance screenshot.png -q "What error is shown?"

# OCR
glance screenshot.png --ocr

# Find a button
ground screenshot.png "the login button"
# Output: 450,820,620,870

# Find all buttons
detect screenshot.png "buttons"

# Crop a region
crop screenshot.png 450,820,620,870 button.png

How it works

The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which:

1. Reads the image file 2. Encodes it as base64 3. Sends it to the vision API with a prompt 4. Returns the text response

The agent never sees raw pixels -- it gets text descriptions it can reason about.

Supported vision providers

  • DeepSeek VL (deepseek-vl2, deepseek-vl2.5)
  • OpenAI GPT-4V / GPT-4o
  • Any OpenAI-compatible multimodal endpoint

License

MIT -- YYTbit