dsh-tool-bandit-search
by siruignaw-sys
DeepSeek Harness插件,替换标准网络搜索工具,使用上下文多臂老虎机学习搜索策略。
A DeepSeek Harness plugin that replaces the standard web search tool with a search tool that learns which search strategy to use through a contextual multi-armed bandit,…
安装
dsh plugin --profile web add github:siruignaw-sys/dsh-tool-bandit-searchGitHub 源码安装:首次需按提示配置 allowBuilds 构建授权后重试
安装与环境配置指引、插件开发教程见 DSH 中文社区文档 ↗
安装即在你的机器上以你的权限运行第三方代码——它可读写文件、使用凭据、访问网络,DSH 的工具审批不会为插件代码加沙箱。「检测到 manifest」仅代表发现 dsh.bundle / dsh.plugin 清单,不构成兼容性或安全审查;安装前请审阅源码,不熟悉的插件先在不含密钥的环境试用。
README
A DeepSeek Harness plugin that replaces the standard web_search tool with a search tool that learns which search strategy to use through a contextual multi-armed bandit, instead of relying on a single hardcoded approach.
Why
Every web search has a tradeoff: a fast, narrow query gets you an answer quickly, but a broader, multi-angle query gets you better coverage at the cost of latency. Hardcoding one strategy means always overpaying for simple questions or always underdelivering on complex ones. This plugin lets the tool discover, from real usage, which strategy tends to pay off — and keeps adapting as conditions change.
How it works
The search tool has two internal strategies ("arms"):
quick— a single search query, capped at 5 results. Fast, good for simple factual lookups.thorough— three query variants (the original plus two reframed angles) run in parallel and merged/deduplicated, capped at 10 results. Slower, better for open-ended or multi-perspective questions.
On every call, the plugin uses Thompson sampling to pick an arm: each arm has a Beta(α, β) distribution representing its estimated reward, the plugin samples from both distributions, and whichever sample is higher gets used. This naturally balances exploration (trying the less-proven arm occasionally) against exploitation (favoring the arm that's performed better so far).
After the call, a continuous reward in [0, 1] is computed from two components, weighted equally:
- Quality — how many results came back, relative to that arm's own maximum (so a 5-of-5 "quick" result is scored the same as a 10-of-10 "thorough" result — neither arm is structurally favored by its own result cap).
- Speed — how fast the call completed, calibrated against realistic search latency.
That reward updates the chosen arm's Beta distribution (α += reward, β += 1 − reward), so the bandit's beliefs shift a little after every single call — no separate training phase, no manual tuning.
The model never sees the two arms directly. It just calls search(query); the plugin decides internally which strategy to run.
Example output
[bandit-search] arm=quick reward=1.000 durationMs=4393 resultCount=5 stats={"quick":{"alpha":2,"beta":1},"thorough":{"alpha":1,"beta":1}}
[bandit-search] arm=thorough reward=0.854 durationMs=8481 resultCount=10 stats={"quick":{"alpha":2,"beta":1},"thorough":{"alpha":1.85,"beta":1.15}}
[bandit-search] arm=quick reward=0.000 durationMs=5777 resultCount=0 stats={"quick":{"alpha":2,"beta":2},"thorough":{"alpha":1.85,"beta":1.15}}
Each log line shows which arm was picked, the reward it earned, and the running Beta parameters for both arms — you can watch the bandit's confidence shift in real time as it accumulates evidence.
Install
dsh plugin --profile web add github:siruignaw-sys/dsh-tool-bandit-search
For local development against a cloned/edited copy instead:
dsh plugin --profile web add link:/absolute/path/to/dsh-tool-bandit-search
Either way, restart the Web UI (a fresh pnpm dsh web / dsh web, not just a new chat) after installing — bundle installs only take effect on the next boot, and the plugin's system-prompt instruction steering the model toward search over the built-in web_search tool only applies to sessions started after that.
Requirements
Runs on top of dsh's native ctx.web search service — no separate API key needed beyond whatever search provider your dsh profile already has configured (e.g. dsh-web-search-deepseek).
Known limitations
- Bandit state is in-memory and resets on every restart. Persisting it via
ctx.storage(which dsh already exposes) is a natural next step. - Reward is a heuristic, not a measure of actual answer quality — it captures result count and latency, not whether the results were relevant or correct. A stronger version might score reward against whether the model's final answer actually used the returned sources.
- The model can still issue multiple
searchcalls per turn even whenthoroughis already broadening internally — the plugin optimizes strategy per call, not the model's own multi-call behavior. - Built and tested against dsh's developer preview; the plugin/tool APIs may change before a stable release.
License
MIT
原始 README: https://github.com/siruignaw-sys/dsh-tool-bandit-search/blob/main/README.md ↗
同类插件
查看全部 →
dsh-anchored-standard
两阶段 DeepSeek Harness 预设:先 Minimal 对齐的 bootstrap,再切完整 Standard 工具(Project2 98/99)

PicGo-Core
极致的图片上传引擎,CLI 与 API 双支持

awesome-deepseek-harness
DeepSeek Harness(DSH)及其优秀社区插件的精选指南。

awesome-deepseek-harness
DeepSeek Harness (DSH)生态系统:来自dsh-external/hub和公共dsh-plugin主题的精选插件、工具和基础设施。

AI-Novel-Writer
本地优先 AI 小说创作工作台,提供 Windows/macOS 桌面版与 DeepSeek Harness 插件开发预览,支持角色、大纲、章节蓝图、审稿修稿和本地模型。

mcp-for-stata
MCP-for-Stata:把 Stata 集成进你的 agent