dsh-image-reader
by zcxie777
给纯文本 DeepSeek Harness agent 直接读图的能力:一个面向模型的 read image 工具,按需询问任意 OpenAI 兼容视觉端点
Give a text-only DeepSeek Harness agent the ability to read images directly : one model-facing read image tool that asks any OpenAI-compatible vision endpoint about an image by…
安装
dsh plugin --profile web add github:zcxie777/dsh-image-readerGitHub 源码安装:首次需按提示配置 allowBuilds 构建授权后重试
安装与环境配置指引、插件开发教程见 DSH 中文社区文档 ↗
安装即在你的机器上以你的权限运行第三方代码——它可读写文件、使用凭据、访问网络,DSH 的工具审批不会为插件代码加沙箱。「检测到 manifest」仅代表发现 dsh.bundle / dsh.plugin 清单,不构成兼容性或安全审查;安装前请审阅源码,不熟悉的插件先在不含密钥的环境试用。
README
Give a text-only DeepSeek Harness agent the ability to read images directly: one model-facing read_image tool that asks any OpenAI-compatible vision endpoint about an image by its workspace path.
Why
DeepSeek Harness is "everything is a plugin". This bundle mounts a single tool so the model can look at a screenshot, diagram, or photograph and answer questions about it, instead of only ever reasoning over text.
Verification status
- Verified locally:
npm run typecheck,npm run build, andnpm test(16 tests) all pass. - Not yet verified: a real end-to-end read against a live vision endpoint. The request/response logic is covered by a mocked-fetch unit test, but the plugin has not been smoke-tested inside a running dsh profile against a real multimodal model. Do that once with a real
VISION_API_KEYbefore relying on it.
Install
git clone https://github.com/zcXie777/dsh-image-reader.git
cd dsh-image-reader
npm install
npm run build # lib/ is not committed; build once after cloning
cd ..
dsh plugin --profile web add "$PWD/dsh-image-reader"
dsh plugin --profile headless add "$PWD/dsh-image-reader"
dsh --profile web --dump-config | grep image-reader
Restart a running Web profile after installing.
Configure
provider.baseUrl and provider.model are required; the plugin never assumes a vendor. Override them in the profile patch row with the same id:
- id: image-reader
config:
provider:
baseUrl: https://api.openai.com/v1
model: gpt-4o-mini
apiKeyEnv: VISION_API_KEY
lang: zh
timeoutMs: 60000
maxImageBytes: 10485760
allowedDirs: []
Set the key in the environment before starting the profile:
export VISION_API_KEY=sk-...
Use
In a conversation, point the model at an image path and ask:
read_image image="screenshot.png" query="What error is shown in this dialog?"
read_image image="diagram.png"
Configuration fields
| Field | Default | Contract |
|---|---|---|
provider.baseUrl |
— (required) | OpenAI-compatible chat/completions base URL |
provider.model |
— (required) | Multimodal model name |
provider.apiKeyEnv |
VISION_API_KEY |
Environment variable holding the API key |
lang |
zh |
Answer language: zh or en |
timeoutMs |
60000 |
Whole-request deadline, 1000–600000 ms |
maxImageBytes |
10485760 |
Encoded-byte limit per image |
allowedDirs |
[] |
Extra realpath-resolved input roots; the workspace is always allowed |
Security
- Inputs resolve against the workspace and
allowedDirsthroughrealpath, so a symlink cannot escape the fence. - Images are size-limited and extension-checked before upload.
- The key is read from the environment per call, never stored in config.
Development
npm install
npm run typecheck
npm run build
Publish
Tag the repo with the dsh-plugin topic so it is discoverable, and publish to npm when ready.
License
MIT
原始 README: https://github.com/zcXie777/dsh-image-reader/blob/main/README.md ↗
同类插件
查看全部 →
modlens
为纯文本模型架起视觉桥梁:粘贴图片,输出结构化 JSON 证据(OCR、版面、语义)。

dsh-vision-toolkit
让纯文本模型更好地做视觉任务:带意图的图片问答、长截图 OCR、UI 还原等。

dsh-vision-router
为纯文本 Agent 提供视觉能力:内置免 Key 视觉链 + 像素级视觉工具(看图问答、定位、裁剪、像素对比、取色、OCR、矢量化、抠图、截图);粘贴图片即可用。

dsh-vision-complete
给 DeepSeek 补上「眼睛和耳朵」的多模态视觉插件:看图 / OCR / 物体检测 / 视频理解 / 语音转写 / 截图直读,一键安装(DSH 插件)。

dsh-media-skills
面向纯文本模型的免费视觉桥与生图:粘贴读图、GLM-4V-Flash 与 Gemini 引擎故障转移、modlens 同款结构化证据输出,并自动播种免费视觉模型路由。

dsh-vision-opencode
给纯文本主模型加可配置识图模型:vision_read_image 工具、输入框识图模型选择器,以及纯文本路由的图片自动转文字。