dsh-llm-longcat

by ffyuuu

模型与账号接入github收录于 08-21

LongCat-2.0 模型适配器:100 万上下文、二元思考模式、工具调用,基于 OpenAI 兼容接口的流式输出。

LongCat-2.0 LLM adapter: 1M context, binary thinking mode, tool calling, and streaming over the OpenAI-compatible endpoint.

安装

dsh plugin --profile web add github:ffyuuu/dsh-llm-longcat

GitHub 源码安装:首次需按提示配置 allowBuilds 构建授权后重试

安装与环境配置指引、插件开发教程见 DSH 中文社区文档 ↗

安装即在你的机器上以你的权限运行第三方代码——它可读写文件、使用凭据、访问网络,DSH 的工具审批不会为插件代码加沙箱。「检测到 manifest」仅代表发现 dsh.bundle / dsh.plugin 清单,不构成兼容性或安全审查;安装前请审阅源码,不熟悉的插件先在不含密钥的环境试用。

README

目录

LongCat adapter for the DeepSeek Harness LLM seam.

Adds LongCat-2.0 as a model provider: 1M context, thinking mode, tool calling.

Features

  • Thinking mode — recognizes LongCat's reasoning_content field and translates it into harness ReasoningBlocks
  • Tool calling — full function-calling support, with arguments kept a raw JSON string end to end
  • Multi-turn — replays reasoning_content on tool-call turns, as thinking-mode passback requires
  • Streaming — SSE with the usage-before-finish ordering the harness relies on
  • Credential seam — the key resolves per request from ctx.credentials or the environment; no secret in any config file

Supported models

Model Context Max output Notes
LongCat-2.0 1,048,576 131,072 text-only; thinking + tool calling

Facts from GET /openai/v1/models/LongCat-2.0, the only documented endpoint that reports supported_parameters. Tool calling is not mentioned on the chat-completions doc page and is only visible there.

Install

dsh plugin --profile default add github:ffyuuu/dsh-llm-longcat
export LONGCAT_API_KEY=...   # create one at https://longcat.chat/platform/api_keys

Installing a bundle lets the package's install scripts run on your machine, outside the sandbox the agent runs under. Pin a commit so a later push cannot change what executes:

dsh plugin --profile default add github:ffyuuu/dsh-llm-longcat#3dcb3b1b5870ba52baab053453bdbb28826e5f13

Then pick LongCat-2.0 in the model selector. The key may also be stored through the Web UI's Models page instead of the environment.

If dsh itself will not install

At the time of writing, installing the harness can fail before any plugin is reached, with either ETARGET … dsh-typert-protocol@^0.1.0-rc.8 or an npm heap exhaustion. That is an upstream packaging state, not this plugin: @deepseek-ai/dsh published 0.1.0-rc.8 while several packages it depends on stopped at 0.1.0-rc.7, and because the manifests use caret ranges, ^0.1.0-rc.7 still resolves up into the missing rc.8. npm then backtracks over an unsatisfiable graph until it runs out of memory.

Pinning every @deepseek-ai/* package to an exact 0.1.0-rc.7 through npm overrides avoids the drift. Nothing in this plugin needs changing either way — it declares >=0.1.0-rc.7 and works against whichever of those the host ends up with.

Config

- id: llm-longcat
  name: dsh-llm-longcat
  config:
    apiKeyEnv: LONGCAT_API_KEY   # default; resolved per request, never a literal key
    baseURL: https://api.longcat.chat/openai/v1  # optional; $LONGCAT_BASE_URL then the public API
    thinking: enabled            # optional deployment policy; `disabled` locks every request to off
    reasoningEffort: high        # optional; off | high — LongCat's switch is binary
    maxTokens: 131072            # optional per-request output cap
    defaultContextWindow: 1048576
    streamIdleTimeoutMs: 300000  # optional; five-minute default
    retryPolicy:                 # optional; omission uses bounded normal defaults
      mode: normal
      maxRetries: 3
    models:
      - id: LongCat-2.0
        contextWindow: 1048576

A llm-longcat: section in $DSH_HOME/settings.yaml overrides any field without a restart: base URL, catalog, request defaults, and idle budget all take effect on the next request, while an in-flight stream keeps the facts it started with.

Reasoning is binary, deliberately

LongCat controls thinking with thinking: {type: enabled|disabled} and does not accept OpenAI's top-level reasoning_effort — its supported_parameters lists the former and omits the latter. There is therefore no low/medium/high gradient to map, and this adapter offers exactly two levels rather than advertising controls that would collapse onto the same two request bodies:

Selected effort Wire body
high ("Thinking") {"thinking": {"type": "enabled"}}
off {"thinking": {"type": "disabled"}}
(none named) resolves from config; still explicit

off serializes an explicit disabled rather than omitting the field — omitting it would hand the decision to LongCat's server-side default, which is not what selecting Off should mean. Requesting low, medium, or max fails with UNSUPPORTED_REASONING_EFFORT before any network I/O.

Wire-format notes

  • Tool-call deltas repeat id and name as explicit null. LongCat sends them on the opening delta and then null (not omitted) on every continuation, so a naive !== undefined guard blanks the assembled call's name. Verified on live traffic; pinned by a regression test.
  • Streaming only, with stream_options.include_usage always on. Usage may arrive attached to the finish chunk or as a trailing usage-only chunk; both are deferred to [DONE] so usage always precedes finish.
  • The first thinking-mode delta can be an empty string — it must not open a reasoning block.
  • Reasoning passback: on assistant turns that carried tool calls, reasoning_content is serialized back into history; on tool-call-free turns it is dropped (ignored anyway — saves tokens).
  • Assistant content is always a string, never null: the message is durable session history, and a null there would make later turns replay a body the endpoint can reject.
  • Cache accounting: prompt_tokens_details.cached_tokens maps to cacheReadTokens and is subtracted out of inputTokens to keep the harness's disjoint-count convention.

Errors

Non-2xx responses throw LlmError with stable codes. LongCat documents a dedicated 402 for exhausted token quota and puts insufficient_quota on 403, where most OpenAI-compatible providers use 429 — both are classified as QUOTA before the auth and rate-limit buckets, so a depleted balance is never reported as a bad key or retried as a transient rate limit.

Condition Code
402, or quota detail at any status QUOTA_EXCEEDED
401 / 403 AUTH
429 RATE_LIMIT
400 with context-overflow detail CONTEXT_WINDOW_EXCEEDED
other 400 INVALID_REQUEST
5xx SERVER
no [DONE] / bad JSON STREAM_CLOSED / MALFORMED_RESPONSE

A completed stream that opened no content blocks becomes a finish error with EMPTY_RESPONSE, which the shipped retry policy treats as retryable.

Tests

npm run typecheck   # against the published @deepseek-ai/dsh-llm types
npm test            # 30 unit tests over serialize + translate
npm run build       # emits lib/ and lib/types/
npm run test:e2e    # real API, needs LONGCAT_API_KEY, spends a few hundred tokens

test:e2e drives the built adapter's own serialize → SSE → translate pipeline against api.longcat.chat, so it verifies what the plugin actually sends rather than a hand-written approximation. It is what caught the null-name delta bug.

Limitations

  • No image input. LongCat-2.0 reports modality: text->text, so image content is refused before sending, naming the model.
  • No stop sequences. stop is absent from supported_parameters; passing one fails with UNSUPPORTED_OPTION rather than silently running past it.
  • Reasoning is binary — no low/medium/high gradient exists to map.

License

MIT

原始 README: https://github.com/ffyuuu/dsh-llm-longcat/blob/main/README.md ↗

同类插件

查看全部 →
模型与账号接入Anionex

agent-vision-toolkit

为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode

查看详情
928github+08-16
模型与账号接入toby-bridges

api-relay-audit

从 DeepSeek Harness 对 AI API 中转站和 LLM 代理运行本地安全审计,生成 Markdown 报告,覆盖提示词注入、模型替换信号、工具调用改写、错误泄漏、流完整性和按 profile 启用的 Web3 风险。

查看详情
791github+08-21
模型与账号接入ZJU-LLMs

OpenStory

✨ OpenStory 现已支持 DeepSeek Harness 插件! 现在可以通过 dsh-openstory 将 OpenStory 多智能体推演接入 DeepSeek Harness,让 agent 直接启动模拟、查看角色、下达指令并逐回合推进故事。查看 DSH 插件配置与使用指南。

查看详情
377github+08-17
模型与账号接入pulseaiclub

phi

pi的编码代理 ∞ 提供者、子代理、hashline编辑和权限门

查看详情
89github+08-16
模型与账号接入anysearch-team

anysearch-dsh

DeepSeek Harness(DSH)的 AnySearch 网络搜索提供方与高级搜索工具。

查看详情
79github+08-17
模型与账号接入kuangre123

codex-switch

Codex Switch 是一个 macOS 工具,一键配置 Codex 的自定义 API,同时保留官方 OpenAI 登录。保存后 Codex 的模型选择器里只会出现你选的那个 provider 的模型。也支持 Claude Code 的官方 / 自定义 API 切换。Codex Switch is a lightweight helper for configuring multiple coding-agent API routes. For Codex, it keeps Official OpenAI and a custom API provider configured in parallel, registers the custom model in Codex's mod

查看详情
67github+08-16