dsh-voice
by zhuiyueya
DeepSeek Harness 的语音——给纯文本 DeepSeek 装上耳朵和嘴
Voice for DeepSeek Harness — give text-only DeepSeek ears and a mouth.
安装
dsh plugin --profile web add github:zhuiyueya/dsh-voiceGitHub 源码安装:首次需按提示配置 allowBuilds 构建授权后重试
安装与环境配置指引、插件开发教程见 DSH 中文社区文档 ↗
安装即在你的机器上以你的权限运行第三方代码——它可读写文件、使用凭据、访问网络,DSH 的工具审批不会为插件代码加沙箱。「检测到 manifest」仅代表发现 dsh.bundle / dsh.plugin 清单,不构成兼容性或安全审查;安装前请审阅源码,不熟悉的插件先在不含密钥的环境试用。
README
DeepSeek 的对话 API 是纯文本的——既听不懂音频、也说不出话。dsh-voice 在输入/输出边界把声音桥接成文本,模型全程只看到文本,却获得完整的语音回路:
🎤 语音 → 文本 → DeepSeek(纯文本)→ 文本 → 🔊 语音
和
dsh-vision-bridge是同一思路(把图片在进模型前转成文本)——只是这次补的是听觉,这块 DeepSeek Harness 的多模态空白还没人填。
✨ 功能
| 层级 | 作用 | |
|---|---|---|
| 🎤 | 语音输入(STT)— Web UI | 输入框工具行里的麦克风按钮。点击说话,转写结果通过浏览器 Web Speech API 直接写入输入框。 |
| 🔊 | 朗读(TTS)— Web UI | 每条 assistant 回复上的朗读按钮。点击用 speechSynthesis 把这条回答读出来。 |
| 📄 | voice_transcribe 工具 |
把附件的音频文件(wav/mp3/m4a/ogg/webm/flac)转成文字,走任意 Whisper 兼容的 /audio/transcriptions 端点。 |
| 🗣️ | voice_speak 工具 |
把文本合成语音文件,走任意 OpenAI 兼容的 /audio/speech 端点。 |
- ✅ Web UI 零 API key —— 纯浏览器语音,开箱即用。
- ✅ 不换模型 —— DeepSeek 仍是纯文本,语音在边缘处理。
- ✅ 后端可配 —— 指向本地 whisper.cpp / Kokoro,即可做到完全免费、免 key。
🧭 工作原理
┌──────────────────────────────────────────────────────────────┐
│ dsh Web GUI │
│ │
│ 你说话 ──🎤 SpeechRecognition──► 文本 ──► 输入框 │
│ │
│ 回复文本 ──🔊 speechSynthesis──► 你听到 │
└──────────────────────────────────────────────────────────────┘
│ ▲
│ 文本(STT) │ 文本(TTS)
▼ │
┌──────────────────────────────────────────────────────────────┐
│ DeepSeek(纯文本模型) │
└──────────────────────────────────────────────────────────────┘
附件音频 ── voice_transcribe(Whisper 兼容)──► 文本 ──► 模型
模型想说话 ── voice_speak(OpenAI 兼容 TTS)──► 音频文件
📦 安装
# 从本地目录安装
dsh plugin --profile web add "file:/path/to/dsh-voice"
# 或(发布到 npm 后)
dsh plugin --profile web add dsh-voice
激活是自动的:包内自带 bundle patch(cordis.patch.yml)并声明了 dsh.bundle.patch,dsh plugin add 会自动把它注册进 profile 的 bundles。
然后重启 dsh web(或等待 HMR)。你应在输入框看到 🎤、在每条回复上看到 🔊。
⚙️ 配置
🎤 麦克风按钮需要配置 voice.stt.apiBase(浏览器录音后发给宿主上的 Whisper 兼容后端)。🔊 朗读无需任何配置(浏览器 speechSynthesis)。如需自定义朗读的语言/语速/音调,编辑 lib/client.js 顶部的常量(TTS_LANG / TTS_RATE / TTS_PITCH)。
settings.yaml:
voice:
stt: # 麦克风按钮 + voice_transcribe 工具
enabled: true
apiBase: "" # 麦克风必需。示例:
# 硅基流动 SiliconFlow:https://api.siliconflow.cn/v1
# 本地 whisper.cpp:http://127.0.0.1:8080/v1
apiKeyEnv: VOICE_STT_API_KEY
model: whisper-1
language: "" # zh / en / ... ;空 = 自动检测
tts: # voice_speak 工具
enabled: true
apiBase: "" # 留空 = https://api.openai.com/v1
apiKeyEnv: VOICE_TTS_API_KEY
model: tts-1
voice: alloy # alloy/echo/fable/onyx/nova/shimmer,或本地服务的 voice id
format: mp3
为什么麦克风需要后端:Chrome 内置的
SpeechRecognition会把音频上传到 Google,国内无法访问(会报识别出错:network)。dsh-voice 改用MediaRecorder录音,再通过你自己的 Whisper 兼容后端转写。两个免费、免 key 的选择:
- 硅基流动 SiliconFlow(国内可访问、有免费额度)——
apiBase: https://api.siliconflow.cn/v1,模型FunAudioLLM/SenseVoiceSmall或whisper-1。- 本地 whisper.cpp(完全离线)——
apiBase: http://127.0.0.1:8080/v1(无需 key)。
🧰 Agent 工具
| 工具 | 参数 | 返回 |
|---|---|---|
voice_transcribe |
path(音频文件)、language? |
{ text, language } |
voice_speak |
text、outPath?、voice? |
{ path, bytes } |
🗂 项目结构
dsh-voice/
├── package.json # 双半插件:host(main)+ browser(client)
├── cordis.patch.yml # bundle 激活层
├── lib/
│ ├── index.js # host 半:settings + voice_transcribe/voice_speak 工具
│ ├── client.js # browser 半:🎤 / 🔊 按钮
│ └── types/
│ ├── index.d.ts
│ └── client/index.d.ts
├── README.md # 英文版
└── README.zh-CN.md # 本文件(中文版)
🗺 路线图
- 把浏览器 UI 的语言 / 语速 / 音调 / 自动朗读接入
voice:设置页(目前为代码内常量) -
autoRead:回复完成后自动朗读 - 内置免费
edge-tts后端(无需 OpenAI key) - 基于
@xenova/transformers的本地 Whisper STT - 长文本分句朗读、流式打断
🙏 致谢
设计时参考了这些「给其他 agent 补语音」的成熟方案:
- slopus/happy(~23k★)—— 实时语音交互形态
- mbailey/voicemode(~1.3k★)—— Claude Code 语音模式
- caiovicentino/claude-call —— 本地 Whisper STT + edge-tts、免 key
- edge-tts —— 免费微软 Edge 神经语音
- ggerganov/whisper.cpp / OpenAI Whisper —— 语音识别
- hexgrad/kokoro —— 本地神经 TTS
📄 License
原始 README: https://github.com/zhuiyueya/dsh-voice/blob/main/README.zh-CN.md ↗
同类插件
查看全部 →
dsh
将 DeepSeek Harness 的生命周期状态、错误与审批请求,桥接到本地运行的 OpenPets 桌面伙伴。

dsh-qqbot
让 QQ Bot 接入 DeepSeek Harness(dsh)的官方插件

dsh-open-in-vscode
从 Web GUI 一键在 VS Code 中打开工作区目录。

dsh-notification
回合完成桌面通知,按结果分控 + 关键词过滤。

dsh-notifier
DSH 统一通知推送与远程控制:一个 `notify()` API 打通 25+ 渠道(Telegram / 钉钉 / 飞书 / 企业微信 / QQ 机器人 / WxPusher / PushPlus / Server 酱 / Bark / Discord / Slack / ntfy / webhook 等),timeSensitive / active / passive 分级路由并重试;五通道反向审批(Telegram 按钮 / 飞书卡片 / QQ / WxPusher / 微信 iLink);QQ/钉钉/飞书官方扫码登录;本地 Web 管理台;多 agent 路由;系统桌面通知——以及**手机指挥中心**:在手机上发 `!status` / `!stop` / `!retry` 遥控 agent,通知带可操作按钮(查看结果 / 重试 / 日志,点击回调 agent)。密钥脱敏、工具限流、零运行时依赖。

chatccc
飞书(Lark)或微信(WeChat)聊天控制 DeepSeek Harness / Claude Code / Cursor / Codex / CCC Agent