dsh-image-reader

by zcxie777

视觉与多模态github收录于 08-23

给纯文本 DeepSeek Harness agent 直接读图的能力:一个面向模型的 read image 工具,按需询问任意 OpenAI 兼容视觉端点

Give a text-only DeepSeek Harness agent the ability to read images directly : one model-facing read image tool that asks any OpenAI-compatible vision endpoint about an image by…

安装

dsh plugin --profile web add github:zcxie777/dsh-image-reader

GitHub 源码安装:首次需按提示配置 allowBuilds 构建授权后重试

安装与环境配置指引、插件开发教程见 DSH 中文社区文档 ↗

安装即在你的机器上以你的权限运行第三方代码——它可读写文件、使用凭据、访问网络,DSH 的工具审批不会为插件代码加沙箱。「检测到 manifest」仅代表发现 dsh.bundle / dsh.plugin 清单,不构成兼容性或安全审查;安装前请审阅源码,不熟悉的插件先在不含密钥的环境试用。

README

目录

Give a text-only DeepSeek Harness agent the ability to read images directly: one model-facing read_image tool that asks any OpenAI-compatible vision endpoint about an image by its workspace path.

Why

DeepSeek Harness is "everything is a plugin". This bundle mounts a single tool so the model can look at a screenshot, diagram, or photograph and answer questions about it, instead of only ever reasoning over text.

Verification status

  • Verified locally: npm run typecheck, npm run build, and npm test (16 tests) all pass.
  • Not yet verified: a real end-to-end read against a live vision endpoint. The request/response logic is covered by a mocked-fetch unit test, but the plugin has not been smoke-tested inside a running dsh profile against a real multimodal model. Do that once with a real VISION_API_KEY before relying on it.

Install

git clone https://github.com/zcXie777/dsh-image-reader.git
cd dsh-image-reader
npm install
npm run build          # lib/ is not committed; build once after cloning
cd ..
dsh plugin --profile web add "$PWD/dsh-image-reader"
dsh plugin --profile headless add "$PWD/dsh-image-reader"
dsh --profile web --dump-config | grep image-reader

Restart a running Web profile after installing.

Configure

provider.baseUrl and provider.model are required; the plugin never assumes a vendor. Override them in the profile patch row with the same id:

- id: image-reader
  config:
    provider:
      baseUrl: https://api.openai.com/v1
      model: gpt-4o-mini
      apiKeyEnv: VISION_API_KEY
    lang: zh
    timeoutMs: 60000
    maxImageBytes: 10485760
    allowedDirs: []

Set the key in the environment before starting the profile:

export VISION_API_KEY=sk-...

Use

In a conversation, point the model at an image path and ask:

read_image image="screenshot.png" query="What error is shown in this dialog?"
read_image image="diagram.png"

Configuration fields

Field Default Contract
provider.baseUrl — (required) OpenAI-compatible chat/completions base URL
provider.model — (required) Multimodal model name
provider.apiKeyEnv VISION_API_KEY Environment variable holding the API key
lang zh Answer language: zh or en
timeoutMs 60000 Whole-request deadline, 1000–600000 ms
maxImageBytes 10485760 Encoded-byte limit per image
allowedDirs [] Extra realpath-resolved input roots; the workspace is always allowed

Security

  • Inputs resolve against the workspace and allowedDirs through realpath, so a symlink cannot escape the fence.
  • Images are size-limited and extension-checked before upload.
  • The key is read from the environment per call, never stored in config.

Development

npm install
npm run typecheck
npm run build

Publish

Tag the repo with the dsh-plugin topic so it is discoverable, and publish to npm when ready.

License

MIT

原始 README: https://github.com/zcXie777/dsh-image-reader/blob/main/README.md ↗