DeepSeek Harness plugin

dsh-tool-eyes

Local vision 'eyes' for DeepSeek Harness (DSH): screen tool (capture screen or image -> local OpenAI-compatible VLM description) and ocr tool (Windows built-in OCR, zero model / GPU / cloud).

Jump to install

Source facts

Repository
go-farther-and-farther/dsh-tool-eyes
Latest update
Aug 15, 2026
Category
Models & Providers
GitHub stars
2
Format
plugin
Catalog evidence
Upstream dsh.bundle evidence
Evidence path
package.json#dsh.bundle
Checked against
0.1.0-rc.8
Upstream check date
2026-08-20

This evidence comes from the upstream catalog. This site has not installed, run, or security-reviewed the plugin.

Install

Start with a prompt that asks an agent to review the GitHub repository and source. Switch to the command if you want to install it yourself.

Copy this prompt into DSH, Codex, or another agent and ask it to review the GitHub repository and source first.

Do not install or run any commands yet. Read this plugin's GitHub repository, README, and relevant source code. Then answer the questions below clearly and directly so I can decide whether it fits my needs:

1. What is this plugin, and what problem does it solve?
2. Who is it for, and what are its typical use cases?
3. How is it used after installation? Include one minimal example.
4. What known limitations or privacy, security, compatibility, or maintenance risks does it have?
5. Give a clear recommendation: recommend, conditionally recommend, or do not recommend, with reasons.

Distinguish statements documented by the repository, inferences from source code, and unknowns. If evidence is insufficient, say so explicitly. Do not guess or simply repeat the README.

GitHub: https://github.com/go-farther-and-farther/dsh-tool-eyes
Plugin: dsh-tool-eyes
Author: go-farther-and-farther

Check the source files

Read the README and other files from this plugin directory before installing.

File explorer3 files
README.mdSource · read only

dsh-tool-eyes

Local vision "eyes" for DeepSeek Harness (DSH).

Give your text-only agent eyes with two model-facing tools:

  • screen — capture the screen (or describe an existing image file) through a

local OpenAI-compatible vision endpoint (llama.cpp with --mmproj, LM Studio, Ollama, ...) and return the vision model's text description.

  • ocr — extract ALL text verbatim with the **Windows built-in OCR

engine**: zero model, zero GPU, zero cloud, milliseconds.

screen  =  screenshot / image  ->  local VLM  ->  text description (understanding)
ocr     =  screenshot / image  ->  Windows OCR -> verbatim text          (extraction)

Why

DeepSeek's chat-completions line is text-only. Instead of switching your whole conversation to a vision model, keep the text brain and add eyes as tools:

  • Private by default — point screen at a local endpoint and images never

leave your machine.

  • Cheap — a 0.8B–4B local VLM is plenty for describing screens; ocr costs

nothing at all.

  • Honest by default — the screen prompt tells the VLM to *describe only

what is visible* and never guess app/game/character names unless confirmed by on-screen text (this measurably cuts small-model name hallucination).

Requirements

  • Windows 10/11 (the ocr tool uses WinRT OCR; screen uses .NET for capture)
  • Node.js >= 22.19, DeepSeek Harness >= 0.1.0-rc.6
  • screen additionally needs any OpenAI-compatible VLM endpoint, e.g.:

- llama.cpp: llama-server -m model.gguf --mmproj mmproj.gguf --port 1235 - LM Studio (loaded vision model), Ollama, or any OpenAI-compatible gateway

Install

> This package is published on GitHub only (not on npm).

dsh plugin --profile web add https://github.com/go-farther-and-farther/dsh-tool-eyes

Then restart dsh web. The screen and ocr tools appear in the agent's toolkit automatically.

Manual install (offline / from source)

Copy this package into the profile's node_modules, then register it in $DSH_HOME/profiles/<profile>/cordis.patch.yml:

- insert:
    - id: tool-eyes
      name: 'dsh-tool-eyes'
      config:
        baseUrl: http://127.0.0.1:1235/v1
        model: ''
        timeoutMs: 180000

Configuration

Plugin config (all optional):

keydefaultmeaning
baseUrlhttp://127.0.0.1:1235/v1OpenAI-compatible endpoint for screen
model''model id to send; empty lets the server decide (llama.cpp serves one model)
timeoutMs180000hard cap for one capture call
captureScriptbundled capture.ps1override path to an alternate capture script

Override in your profile's cordis.patch.yml (id-targeted):

- id: tool-eyes
  name: 'dsh-tool-eyes'
  config:
    baseUrl: http://127.0.0.1:1235/v1
    model: qwen3.5-4b
    timeoutMs: 120000

Usage

In a conversation, the agent can now:

  • screen — "what is on my screen?", "describe this image file", with an

optional prompt to focus on a region or detail.

  • ocr — "read all the text on screen", "transcribe this error dialog".

Both accept an optional image path; without it they capture the screen. The bundled PowerShell scripts can also be run standalone:

powershell -NoProfile -ExecutionPolicy Bypass -File lib\capture.ps1 -Prompt "..." -BaseUrl http://127.0.0.1:1235/v1
powershell -NoProfile -ExecutionPolicy Bypass -File lib\ocr.ps1 -Image C:\path\x.png

Privacy

  • ocr is fully local (WinRT OCR, no network).
  • screen sends the captured image to the configured baseUrl. Point it at a

local endpoint (llama.cpp / LM Studio / Ollama) to keep images on your machine.

Related

transcription of attached images in the chat input (Chatbox-style), so you don't need to give file paths. Pairs well with this plugin.

Development

npm test    # node --test tests/

License

MIT