Skills Plugins MCP Prompt Model 导航 博客 资讯 我的中心
工具与能力 #deepseek-harness#dsh#dsh-plugin#llm#rate-limit#ai-agent

dsh-llm-rate-limit

LLM API rate limiting, concurrency control, queuing, and adaptive cooldown for DeepSeek Harness

asong6824 @asong6824 ⬇ 1 ★ 0 main

安装

dsh plugin --profile web add github:asong6824/dsh-llm-rate-limit
下载安装清单

需要可复现安装时,可在仓库后追加 #commit 固定提交。

LLM API rate limiting, concurrency control, queuing, and adaptive cooldown for DeepSeek Harness

该插件未提供要点说明,请参考仓库 README。

deepseek-harnessdshdsh-pluginllmrate-limitai-agent
  1. 安装并启动 DeepSeek Harness:npx @deepseek-ai/dsh web
  2. 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
  3. 用 dsh plugins list 确认已安装,必要时重启 Harness 生效

插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。

代码仓库github.com/asong6824/dsh-llm-rate-limit
许可证MIT
主要语言main
下载量1
GitHub 星标0
最近推送2026-08-20
收录日期2026-09-19
分类工具与能力

事实信息来自公开插件目录快照(2026-10-03),介绍文案由本站再加工。

以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。

# dsh-llm-rate-limit

[![npm](https://img.shields.io/npm/v/dsh-llm-rate-limit)](https://www.npmjs.com/package/dsh-llm-rate-limit)
[![downloads](https://img.shields.io/npm/dm/dsh-llm-rate-limit)](https://www.npmjs.com/package/dsh-llm-rate-limit)
[![CI](https://github.com/Asong6824/dsh-llm-rate-limit/actions/workflows/ci.yml/badge.svg)](https://github.com/Asong6824/dsh-llm-rate-limit/actions/workflows/ci.yml)
[![license](https://img.shields.io/npm/l/dsh-llm-rate-limit)](LICENSE)

English | [中文](README.zh.md)

A DeepSeek Harness (DSH) plugin that prevents avoidable API rate-limit errors by pacing LLM requests before they reach the provider. It provides per-provider RPM limits, optional token budgets, concurrency control, bounded FIFO queuing, and adaptive cooldown for DeepSeek API, Volcengine Ark, and other DSH providers.

Use it when parallel agents, subagents, retries, or background requests are producing HTTP 429 errors, provider throttling, or traffic bursts.

## Install from npm

Install the latest release into the Web profile:

```sh
dsh plugin --profile web add dsh-llm-rate-limit
dsh web
```

Pin a version for reproducible environments:

```sh
dsh plugin --profile web add dsh-llm-rate-limit@0.1.1
```

Install separately for Headless:

```sh
dsh plugin --profile headless add dsh-llm-rate-limit
```

GitHub installation is also supported:

```sh
dsh plugin --profile web add github:Asong6824/dsh-llm-rate-limit#v0.1.1
```

The bundled default protects `deepseek-official` with 30 requests per minute, burst 1, two concurrent requests, and a bounded queue.

## Features

- Provider-scoped requests-per-minute token buckets with configurable burst capacity.
- Optional estimated-token-per-minute budgets with actual-usage reconciliation.
- Concurrency limits and bounded FIFO queues with timeout and cancellation.
- Adaptive cooldown for provider error codes, HTTP statuses, and `Retry-After`.
- Explicit auxiliary-request shedding so background traffic does not block primary work.
- Durable admission wait/start events for DSH session diagnostics.
- Clean lifecycle disposal without abandoning queued or active requests.
- Retry-aware admission: every `dsh-llm-retry` attempt is admitted independently; this plugin never retries requests itself.

## Configure DeepSeek and Ark

Override the complete `llm-rate-limit` config in `$DSH_HOME/profiles//cordis.patch.yml`:

```yaml
- id: llm-rate-limit
  config:
    providers:
      deepseek-official:
        requests: { perMinute: 30, burst: 1 }
        maxConcurrentRequests: 2
        queue: { maxSize: 100, maxWaitMs: 300000, auxiliary: reject }
        cooldown:
          codes: [RATE_LIMIT, SERVER]
          statuses: [429, 529]
          initialDelayMs: 500
          maxDelayMs: 60000
          maxProviderDelayMs: 3600000
          jitterRatio: 0.1
      volcengine-ark-coding:
        requests: { perMinute: 30, burst: 1 }
        maxConcurrentRequests: 2
        queue: { maxSize: 100, maxWaitMs: 300000, auxiliary: reject }
```

Provider keys must exactly match `GenerateOptions.provider`. Optional token limiting adds:

```yaml
tokens:
  perMinute: 1000000
  burst: 200000
  estimatedOutputTokens: 8192
  imageTokens: 1024
```

`tokens.burst` must be large enough for one complete request estimate. Omit `tokens` when a provider should have RPM and concurrency control without a local token ceiling.

## How it works

Before each provider call, the plugin reserves request capacity, estimated token capacity, and a concurrency slot. Requests without capacity wait in FIFO order. Provider throttling responses activate a shared cooldown; successful responses reconcile estimated tokens with actual usage. The state is process-local and resets when DSH restarts.

The plugin deliberately does not provide distributed quotas, automatic retries, or provider failover.

## Compatibility and links

- Requires DeepSeek Harness `0.1.0-rc.8` or newer and Node.js `22.19` or newer.
- [npm package](https://www.npmjs.com/package/dsh-llm-rate-limit)
- [GitHub releases](https://github.com/Asong6824/dsh-llm-rate-limit/releases)
- [DSH plugins topic](https://github.com/topics/dsh-plugin)
- [Machine-readable summary](llms.txt)

## Development

```sh
pnpm install
pnpm run check
```

MIT

数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。