Skills Plugins MCP Prompt Model 导航 博客 资讯 我的中心
模型与提供方 #deepseek-harness#dsh-plugin

dsh-llm-gate

DSH llm/stream 流程的按提供商并发门:将多余模型请求放入 FIFO 队列,避免单槽本地服务器超时。

d3vmeh @d3vmeh ⬇ 1 ★ 1 main

安装

dsh plugin --profile web add github:d3vmeh/dsh-llm-gate
下载安装清单

需要可复现安装时,可在仓库后追加 #commit 固定提交。

DSH llm/stream 流程的按提供商并发门:将多余模型请求放入 FIFO 队列,避免单槽本地服务器超时。

该插件未提供要点说明,请参考仓库 README。

deepseek-harnessdsh-plugin
  1. 安装并启动 DeepSeek Harness:npx @deepseek-ai/dsh web
  2. 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
  3. 用 dsh plugins list 确认已安装,必要时重启 Harness 生效

插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。

代码仓库github.com/d3vmeh/dsh-llm-gate
许可证MIT
主要语言main
下载量1
GitHub 星标1
最近推送2026-08-29
收录日期2026-09-19
分类模型与提供方

事实信息来自公开插件目录快照(2026-10-01),介绍文案由本站再加工。

以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。

# dsh-llm-gate

Per-provider concurrency gate for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) model requests.

If a provider can only serve a fixed number of requests at once (e.g a local `llama-server` with `--parallel 1`), every extra request is deferred by the server with nothing sent back. The client cannot tell "waiting for a slot" from "dead", and Node HTTP layer times out after 300 seconds with `terminated`. In practice this happens when there is overlap between a subagent and the main agent or compaction and the agent.

This plugin holds surplus requests inside dsh instead. A request waits in a FIFO queue before any HTTP request is made so no timeout is running while it waits. When a slot frees, the next request is dispatched.

## Install

```
dsh plugin --profile web add dsh-llm-gate
```

Then configure the providers to gate in `~/.dsh/profiles/web/cordis.patch.yml`:

```yaml
- id: llm-gate
  config:
    providers:
      llamacpp:
        maxConcurrent: 1
        maxQueued: 16
        queueTimeoutMs: 3600000
```

The provider key is the route name from your `llm-pi-ai.providers` (or other adapter) settings. Providers not listed are not gated. Restart `dsh web` and open a new session.

Check the composed config with `dsh --profile web --dump-config`.

## Settings

| Setting | Required | Meaning |
|---|---|---|
| `maxConcurrent` | yes | Requests allowed in flight to this provider. For llama.cpp, match `--parallel`. |
| `maxQueued` | no | Requests allowed to wait. Beyond this, a request fails at once with `QUEUE_FULL`. Default: unlimited. |
| `queueTimeoutMs` | no | Longest a request may wait for a slot before failing with `QUEUE_TIMEOUT`. Default: wait indefinitely. |

Queue failures end the turn with the code shown. They are not retried by `dsh-llm-retry`.

## What you will see

The plugin prints a line to the dsh terminal only when a request has to wait:

```
llm-gate: llamacpp session=a61e6e40 queued (depth 1)
llm-gate: llamacpp session=a61e6e40 dispatched after 5730ms
```

`purpose=compaction` or `purpose=session-title` is added for auxiliary requests. Requests that get a slot immediately print nothing.

## Notes

- This gate serializes requests so it does not make a single-slot server faster. For parallelizing, give llama.cpp more slots (`--parallel 2 --kv-unified`) and raise `maxConcurrent` to match.
- Waiting time is not counted by the adapter's `streamIdleTimeoutMs` because the adapter is not called until the slot is acquired. You still need `streamIdleTimeoutMs` large enough for your prompt processing time (see the `llm-pi-ai` provider settings).
- A queued request is cancelled through its abort signal. Dropping the stream without aborting leaves the request queued until a slot frees, at which point it dispatches and is closed immediately.
- Requires the `llm` service; hooks the `llm/stream` waterfall, so it covers every model request in the host: agents, subagents, compaction, and title generation.

## License

MIT

数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。