DeepSeek Harness 视觉工具箱——给纯文本 agent 装上眼睛
Vision toolkit for DeepSeek Harness -- give text-only agents eyes
安装
dsh plugin --profile web add github:yytbit/dsh-plugin-vision-toolkitGitHub 源码安装:首次需按提示配置 allowBuilds 构建授权后重试
安装与环境配置指引、插件开发教程见 DSH 中文社区文档 ↗
安装即在你的机器上以你的权限运行第三方代码——它可读写文件、使用凭据、访问网络,DSH 的工具审批不会为插件代码加沙箱。「检测到 manifest」仅代表发现 dsh.bundle / dsh.plugin 清单,不构成兼容性或安全审查;安装前请审阅源码,不熟悉的插件先在不含密钥的环境试用。
README
Vision toolkit for DeepSeek Harness -- give text-only agents the ability to see images.
What it does
Provides CLI tools that call a vision API (DeepSeek VL, GPT-4V, or any OpenAI-compatible endpoint) to describe, locate, detect, and crop elements from images. Registered as a dsh skill so agents know when and how to use them.
Tools
glance-- describe, ask about, or OCR an imageground-- locate a specific element (returns bounding box)detect-- find all instances of an element kindcrop-- cut a region from an image
Install
dsh plugin --profile your-profile add dsh-plugin-vision-toolkit
Configuration
Set environment variables:
export VISION_API_KEY=sk-xxx # Vision API key (falls back to DEEPSEEK_API_KEY)
export VISION_BASE_URL=https://... # API endpoint (falls back to DEEPSEEK_BASE_URL)
export VISION_MODEL=deepseek-vl2 # Vision model name
Usage examples
# Describe an image
glance screenshot.png
# Ask a question
glance screenshot.png -q "What error is shown?"
# OCR
glance screenshot.png --ocr
# Find a button
ground screenshot.png "the login button"
# Output: 450,820,620,870
# Find all buttons
detect screenshot.png "buttons"
# Crop a region
crop screenshot.png 450,820,620,870 button.png
How it works
The plugin registers a skill in the system prompt that teaches the agent about the vision tools. When the agent encounters an image (user pastes one, references a screenshot, etc.), it calls the appropriate CLI tool which:
- Reads the image file
- Encodes it as base64
- Sends it to the vision API with a prompt
- Returns the text response
The agent never sees raw pixels -- it gets text descriptions it can reason about.
Supported vision providers
- DeepSeek VL (deepseek-vl2, deepseek-vl2.5)
- OpenAI GPT-4V / GPT-4o
- Any OpenAI-compatible multimodal endpoint
License
MIT -- YYTbit
原始 README: https://github.com/YYTbit/dsh-plugin-vision-toolkit/blob/master/README.md ↗
同类插件
查看全部 →
modlens
为纯文本模型架起视觉桥梁:粘贴图片,输出结构化 JSON 证据(OCR、版面、语义)。

dsh-vision-toolkit
让纯文本模型更好地做视觉任务:带意图的图片问答、长截图 OCR、UI 还原等。

dsh-vision-router
为纯文本 Agent 提供视觉能力:内置免 Key 视觉链 + 像素级视觉工具(看图问答、定位、裁剪、像素对比、取色、OCR、矢量化、抠图、截图);粘贴图片即可用。

dsh-vision-complete
给 DeepSeek 补上「眼睛和耳朵」的多模态视觉插件:看图 / OCR / 物体检测 / 视频理解 / 语音转写 / 截图直读,一键安装(DSH 插件)。

dsh-media-skills
面向纯文本模型的免费视觉桥与生图:粘贴读图、GLM-4V-Flash 与 Gemini 引擎故障转移、modlens 同款结构化证据输出,并自动播种免费视觉模型路由。

dsh-vision-opencode
给纯文本主模型加可配置识图模型:vision_read_image 工具、输入框识图模型选择器,以及纯文本路由的图片自动转文字。