dsh-tool-describe-image

by soli0x4ea

0 视觉与多模态github未核验到 manifest收录于 08-16

本地视觉工具 — 给 DeepSeek Harness 的主模型「借一双眼睛」:传入图片路径,调用本地 LM Studio 视觉模型( qwen3vl4b ),返回图片的文字描述,主模型据此理解图片。

A local vision tool — lends the DeepSeek Harness main model a pair of eyes: pass in an image path, it calls a local LM Studio vision model (qwen3vl4b), returns a textual description of the image, and the main model understands the image from that.

安装

dsh plugin --profile web add github:soli0x4ea/dsh-tool-describe-image

GitHub 源码安装:首次需按提示配置 allowBuilds 构建授权后重试

安装与环境配置指引、插件开发教程见 DSH 中文社区文档 ↗

安装即在你的机器上以你的权限运行第三方代码——它可读写文件、使用凭据、访问网络,DSH 的工具审批不会为插件代码加沙箱。「检测到 manifest」仅代表发现 dsh.bundle / dsh.plugin 清单,不构成兼容性或安全审查;安装前请审阅源码,不熟悉的插件先在不含密钥的环境试用。

README

目录

本地视觉工具 — 给 DeepSeek Harness 的主模型「借一双眼睛」:传入图片路径,调用本地 LM Studio 视觉模型(qwen3vl4b),返回图片的文字描述,主模型据此理解图片。

图片只发往 127.0.0.1——不出网,隐私安全。

License: MIT DeepSeek Harness

功能

  • 主模型不声明视觉能力时,通过工具「借眼睛」:读图 → base64 → 本地视觉模型 → 文字描述;
  • 微信图片链路天然衔接:收到「[微信媒体] 已保存至 <路径>」后自动调用,回复图片内容;
  • MIME 自动识别(PNG/JPEG/WebP/GIF,magic bytes);
  • 超限保护(>15MB 拒绝并提示);超时明确报错(不编造图片内容)。

依赖

挂载

- insert:
    - id: tool-describe-image
      name: '@deepseek-ai/dsh-tool-describe-image'

配置

字段 默认 说明
baseURL http://127.0.0.1:1234/v1 LM Studio OpenAI 兼容端点
model qwen3vl4b 视觉模型 id
timeoutMs 90000 单次识别超时
defaultPrompt 通用描述指令 识别指令

工具

describe_image:path(必填,图片绝对路径)+ 可选 prompt(自定义识别指令)→ 返回图片文字描述。

开发

pnpm install
pnpm run typecheck
pnpm run build   # tsc 类型 + tsdown bundle → lib/

License

MIT © 2026 soli0x4ea

原始 README: https://github.com/soli0x4ea/dsh-tool-describe-image/blob/main/README.md ↗