dsh-plugin-vision-toolkit
Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images.
What it does
Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them.
Tools
glance-- describe, ask about, or OCR an imageground-- locate a specific element (returns bounding box)detect-- find all instances of an element kindcrop-- cut a region from an image
Install
dsh plugin --profile your-profile add dsh-plugin-vision-toolkitConfiguration
Set environment variables:
export VISION_API_KEY=sk-xxx # Vision API key (falls back to DEEPSEEK_API_KEY)
export VISION_BASE_URL=https://... # API endpoint (falls back to DEEPSEEK_BASE_URL)
export VISION_MODEL=deepseek-vl2 # Vision model nameUsage examples
# Describe an image
glance screenshot.png
# Ask a question
glance screenshot.png -q "What error is shown?"
# OCR
glance screenshot.png --ocr
# Find a button
ground screenshot.png "the login button"
# Output: 450,820,620,870
# Find all buttons
detect screenshot.png "buttons"
# Crop a region
crop screenshot.png 450,820,620,870 button.pngHow it works
The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which:
1. Reads the image file 2. Encodes it as base64 3. Sends it to the vision API with a prompt 4. Returns the text response
The agent never sees raw pixels -- it gets text descriptions it can reason about.
Supported vision providers
- DeepSeek VL (deepseek-vl2, deepseek-vl2.5)
- OpenAI GPT-4V / GPT-4o
- Any OpenAI-compatible multimodal endpoint
License
MIT -- YYTbit