dsh-eval-regression
by aryswisnu
DeepSeek Harness 的小型确定性回归评估插件
A small, deterministic regression-evaluation plugin for DeepSeek Harness.
安装
dsh plugin --profile web add github:aryswisnu/dsh-eval-regressionGitHub 源码安装:首次需按提示配置 allowBuilds 构建授权后重试
安装与环境配置指引、插件开发教程见 DSH 中文社区文档 ↗
安装即在你的机器上以你的权限运行第三方代码——它可读写文件、使用凭据、访问网络,DSH 的工具审批不会为插件代码加沙箱。「检测到 manifest」仅代表发现 dsh.bundle / dsh.plugin 清单,不构成兼容性或安全审查;安装前请审阅源码,不熟悉的插件先在不含密钥的环境试用。
README
A small, deterministic regression-evaluation plugin for DeepSeek Harness.
It registers evaluate_golden_output, a model-callable tool that compares supplied candidate output against required and forbidden fragments. It does not call a model, persist data, or claim semantic correctness. Its job is repeatable pass/fail evidence, not vibes-based architecture in a trench coat.
Why
Agent changes routinely regress answers that appear superficially acceptable. A stable corpus of expected fragments gives a cheap, transparent signal for release smoke tests and replayed transcripts:
- required fragments catch omissions
- forbidden fragments catch known bad claims or unsafe fallbacks
- per-case reports make failures reviewable
- deterministic scoring is suitable for CI thresholds
Install as a DSH plugin
dsh plugin --profile <profile> add github:aryswisnu/dsh-eval-regression
The package is a DSH bundle. Its cordis.patch.yml registers the tool automatically after the profile's base tool runtime.
For local development:
git clone https://github.com/aryswisnu/dsh-eval-regression.git
cd dsh-eval-regression
npm install
npm run build
dsh plugin --profile <profile> add .
Run a version-controlled suite in CI
The plugin also ships a small CLI. It reads a JSON suite, prints an evaluation report to stdout, exits 0 when every case passes, exits 1 when any case fails, and exits 2 for invalid input or usage errors.
{
"suite": "release-smoke",
"cases": [
{
"id": "grounded-answer",
"actual": "The result is 42. Source: benchmark.csv",
"includes": ["42", "Source:"],
"excludes": ["I cannot verify"]
}
]
}
npx dsh-eval-regression suites/release-smoke.json
# or, from this repository:
npm run evaluate -- suites/release-smoke.json
The report includes total passed and failed cases, a 0..1 score, and case-level missing or forbidden fragments. This makes the evaluation corpus ordinary, reviewable source code and makes a failed expectation fail the CI job.
Tool example
{
"suite": "release-smoke",
"cases": [
{
"id": "grounded-answer",
"actual": "The result is 42. Source: benchmark.csv",
"includes": ["42", "Source:"],
"excludes": ["I cannot verify"]
}
]
}
The canonical result includes total passed and failed cases, a 0..1 score, and each case's missing or forbidden fragments.
Boundaries
This is intentionally a narrow deterministic evaluator. It does not replace model-quality review, factual grounding, tool execution checks, or snapshot replay. Use it as one gate in an evaluation harness, then add stronger signals where the product needs them.
Development
npm install
npm test
npm run typecheck
npm run build
MIT License.
原始 README: https://github.com/aryswisnu/dsh-eval-regression/blob/main/README.md ↗
同类插件
查看全部 →
k8e
k8e.sh — 开源 Agentic AI 沙箱矩阵

hol-guard
开源AI代理防病毒:运行时拦截风险工具、秘密访问、提示注入、恶意软件包、MCP服务器、插件和技能。

anolisa
ANOLISA(Agentic Nexus Operating Layer & Interface System Architecture):具备运行时、安全性、可观测性和 Tokenless 响应压缩能力的 Agentic OS,可降低 Token 使用量与成本。

mobius
首个自我演进的开源 Agent OS:连接你的团队、AI agent、设备与算力

deepseek-harness-desktop
DeepSeek Harness Tauri 桌面版 | Only 5mb installer, zero environment setup. Windows / macOS / Linux.

open-managed-agents
开源Claude管理代理API实现和自托管Claude标签式代理运行时。即插即用;在Cloudflare Workers/Durable Objects或Node.js上运行。Apache 2.0。