DeepSeek Harness 插件

dsh-plugin-vision-toolkit

Vision toolkit for DeepSeek Harness -- glance, ground, detect, crop CLI tools for text-only agents to understand images(英文原文)

跳到安装方式

来源信息

GitHub 仓库
YYTbit/dsh-plugin-vision-toolkit
最近更新
2026年8月13日
分类
视觉与多模态
GitHub stars
1
载体类型
plugin
目录证据
上游声明已找到 dsh.bundle
证据路径
package.json#dsh.bundle
核对版本
0.1.0-rc.8
上游核对日期
2026-08-20

该证据由上游目录提供。本站没有安装、运行或安全审核这个插件。

安装

默认先复制一段 Prompt,让 Agent 读 GitHub 仓库和源码;需要自己装时再切到命令。

复制这段 Prompt,发给 DSH、Codex 或其他 Agent,让它先读 GitHub 仓库和源码。

请先不要安装或执行任何命令。阅读这个插件的 GitHub 仓库、README 和关键源码,然后用清楚、直接的方式回答以下问题,帮助我判断它是否适合我的需求:

1. 这个插件是什么,解决什么问题;
2. 适合哪些用户和典型使用场景;
3. 安装后如何使用,并给出一个最小使用示例;
4. 有哪些已知限制,以及隐私、安全、兼容性或维护风险;
5. 给出“推荐 / 有条件推荐 / 不推荐”的明确建议和理由。

请区分仓库明确说明、根据源码推断和未知信息。证据不足时请明确说明,不要猜测或照抄 README。

GitHub:https://github.com/YYTbit/dsh-plugin-vision-toolkit
插件名:dsh-plugin-vision-toolkit
作者:YYTbit

检查来源文件

安装前先看这个插件目录里的 README 和其他文件。

文件资源管理器3 个文件
README.md来源说明 · 只读预览

dsh-plugin-vision-toolkit

Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images.

What it does

Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them.

Tools

  • glance -- describe, ask about, or OCR an image
  • ground -- locate a specific element (returns bounding box)
  • detect -- find all instances of an element kind
  • crop -- cut a region from an image

Install

dsh plugin --profile your-profile add dsh-plugin-vision-toolkit

Configuration

Set environment variables:

export VISION_API_KEY=sk-xxx           # Vision API key (falls back to DEEPSEEK_API_KEY)
export VISION_BASE_URL=https://...     # API endpoint (falls back to DEEPSEEK_BASE_URL)
export VISION_MODEL=deepseek-vl2       # Vision model name

Usage examples

# Describe an image
glance screenshot.png

# Ask a question
glance screenshot.png -q "What error is shown?"

# OCR
glance screenshot.png --ocr

# Find a button
ground screenshot.png "the login button"
# Output: 450,820,620,870

# Find all buttons
detect screenshot.png "buttons"

# Crop a region
crop screenshot.png 450,820,620,870 button.png

How it works

The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which:

1. Reads the image file 2. Encodes it as base64 3. Sends it to the vision API with a prompt 4. Returns the text response

The agent never sees raw pixels -- it gets text descriptions it can reason about.

Supported vision providers

  • DeepSeek VL (deepseek-vl2, deepseek-vl2.5)
  • OpenAI GPT-4V / GPT-4o
  • Any OpenAI-compatible multimodal endpoint

License

MIT -- YYTbit