dsh-llm-gate
DSH llm/stream 流程的按提供商并发门:将多余模型请求放入 FIFO 队列,避免单槽本地服务器超时。
d3vmeh
@d3vmeh
⬇ 1
★ 1
main
安装
dsh plugin --profile web add github:d3vmeh/dsh-llm-gate
需要可复现安装时,可在仓库后追加 #commit 固定提交。
DSH llm/stream 流程的按提供商并发门:将多余模型请求放入 FIFO 队列,避免单槽本地服务器超时。
该插件未提供要点说明,请参考仓库 README。
deepseek-harnessdsh-plugin
- 安装并启动 DeepSeek Harness:
npx @deepseek-ai/dsh web - 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
- 用 dsh plugins list 确认已安装,必要时重启 Harness 生效
插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。
| 代码仓库 | github.com/d3vmeh/dsh-llm-gate |
| 许可证 | MIT |
| 主要语言 | main |
| 下载量 | 1 |
| GitHub 星标 | 1 |
| 最近推送 | 2026-08-29 |
| 收录日期 | 2026-09-19 |
| 分类 | 模型与提供方 |
事实信息来自公开插件目录快照(2026-10-01),介绍文案由本站再加工。
以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。
# dsh-llm-gate
Per-provider concurrency gate for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness) model requests.
If a provider can only serve a fixed number of requests at once (e.g a local `llama-server` with `--parallel 1`), every extra request is deferred by the server with nothing sent back. The client cannot tell "waiting for a slot" from "dead", and Node HTTP layer times out after 300 seconds with `terminated`. In practice this happens when there is overlap between a subagent and the main agent or compaction and the agent.
This plugin holds surplus requests inside dsh instead. A request waits in a FIFO queue before any HTTP request is made so no timeout is running while it waits. When a slot frees, the next request is dispatched.
## Install
```
dsh plugin --profile web add dsh-llm-gate
```
Then configure the providers to gate in `~/.dsh/profiles/web/cordis.patch.yml`:
```yaml
- id: llm-gate
config:
providers:
llamacpp:
maxConcurrent: 1
maxQueued: 16
queueTimeoutMs: 3600000
```
The provider key is the route name from your `llm-pi-ai.providers` (or other adapter) settings. Providers not listed are not gated. Restart `dsh web` and open a new session.
Check the composed config with `dsh --profile web --dump-config`.
## Settings
| Setting | Required | Meaning |
|---|---|---|
| `maxConcurrent` | yes | Requests allowed in flight to this provider. For llama.cpp, match `--parallel`. |
| `maxQueued` | no | Requests allowed to wait. Beyond this, a request fails at once with `QUEUE_FULL`. Default: unlimited. |
| `queueTimeoutMs` | no | Longest a request may wait for a slot before failing with `QUEUE_TIMEOUT`. Default: wait indefinitely. |
Queue failures end the turn with the code shown. They are not retried by `dsh-llm-retry`.
## What you will see
The plugin prints a line to the dsh terminal only when a request has to wait:
```
llm-gate: llamacpp session=a61e6e40 queued (depth 1)
llm-gate: llamacpp session=a61e6e40 dispatched after 5730ms
```
`purpose=compaction` or `purpose=session-title` is added for auxiliary requests. Requests that get a slot immediately print nothing.
## Notes
- This gate serializes requests so it does not make a single-slot server faster. For parallelizing, give llama.cpp more slots (`--parallel 2 --kv-unified`) and raise `maxConcurrent` to match.
- Waiting time is not counted by the adapter's `streamIdleTimeoutMs` because the adapter is not called until the slot is acquired. You still need `streamIdleTimeoutMs` large enough for your prompt processing time (see the `llm-pi-ai` provider settings).
- A queued request is cancelled through its abort signal. Dropping the stream without aborting leaves the request queued until a slot frees, at which point it dispatches and is closed immediately.
- Requires the `llm` service; hooks the `llm/stream` waterfall, so it covers every model request in the host: agents, subagents, compaction, and title generation.
## License
MIT
数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。