界面增强
#crash-recovery#deepseek-harness#deepseek-harness-plugin#deepseek-harness-plugins#dsh#dsh-plugin
dsh-plugin-watchdog
Restart your ongoing agent sessions once on server restart in DeepSeek Harness.
davidgereb
@davidgereb
⬇ 2
★ 0
main
安装
dsh plugin --profile web add github:davidgereb/dsh-plugin-watchdog
需要可复现安装时,可在仓库后追加 #commit 固定提交。
Restart your ongoing agent sessions once on server restart in DeepSeek Harness.
该插件未提供要点说明,请参考仓库 README。
crash-recoverydeepseek-harnessdeepseek-harness-plugindeepseek-harness-pluginsdshdsh-plugin
- 安装并启动 DeepSeek Harness:
npx @deepseek-ai/dsh web - 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
- 用 dsh plugins list 确认已安装,必要时重启 Harness 生效
插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。
| 代码仓库 | github.com/davidgereb/dsh-plugin-watchdog |
| 许可证 | MIT |
| 主要语言 | main |
| 下载量 | 2 |
| GitHub 星标 | 0 |
| 最近推送 | 2026-08-15 |
| 收录日期 | 2026-09-19 |
| 分类 | 界面增强 |
事实信息来自公开插件目录快照(2026-10-01),介绍文案由本站再加工。
以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。
# dsh-plugin-watchdog
**Host-side crash/restart watchdog for dsh.** When the dsh server process dies —
a crash, OOM, `kill -9`, or a restart — every running agent is interrupted and
the work simply stops. The persisted session logs survive, but nothing picks the
work back up until a human reopens the sessions. This plugin closes that gap:
- **No browser needed.** It is a host-only Cordis plugin; it runs inside the dsh
server and works whether the GUI is open, closed, or never opened.
- **On restart it checks recently active sessions** and resumes the ones whose
work was genuinely interrupted — not the ones that were deliberately stopped.
- **It picks up where they left off**: it resumes the persisted session via
`ctx.agents.resume`, then either **re-arms the active goal** (the goal-round
driver queues the next `` itself) or **wakes the agent with a
"continue where you left off" follow-up** prompt. Crash-recovery closers
(`TOOL_OUTCOME_UNKNOWN` / `TOOL_NOT_STARTED`) are already in the resumed
history, so the model verifies uncertain side effects before retrying.
## Install
The plugin is a host-only Cordis plugin resolved from your dsh profile's
package store. The profile's `dsh plugin` command forwards to pnpm, so any
pnpm-supported git spec works.
### From GitHub
One block — installs the plugin (and its git-hosted dependency
`dsh-lib-context-injection`) and registers the loader row:
```bash
dsh plugin --profile web add github:davidgereb/dsh-plugin-watchdog --config.block-exotic-subdeps=false
PATCH="${DSH_HOME:-$HOME/.dsh}/profiles/web/cordis.patch.yml"
grep -q "name: dsh-plugin-watchdog" "$PATCH" 2>/dev/null || cat >> "$PATCH" <<'EOF'
- insert:
- id: watchdog
name: dsh-plugin-watchdog
EOF
```
> **pnpm ≥ 11 note.** The `--config.block-exotic-subdeps=false` flag lifts
> pnpm 11's default ban on git-hosted *transitive* dependencies — this
> plugin's dependency
> [`dsh-lib-context-injection`](https://github.com/davidgereb/dsh-lib-context-injection)
> is resolved that way. No config files to edit.
### From a local checkout
```bash
# 1. build (regenerates lib/index.js + lib/client.js from src/)
node scripts/build.js
# 2. make the local package resolvable by the profile
dsh plugin --profile web link /path/to/dsh-plugin-watchdog
# 3. same for the lib dependency
dsh plugin --profile web link /path/to/dsh-lib-context-injection
```
Then add the loader row to the profile patch file
(`$DSH_HOME/profiles/web/cordis.patch.yml`):
```yaml
# ── Watchdog: resume sessions interrupted by a server crash/restart ──
# Host-only plugin (no browser half). On boot it resumes recently-active
# sessions whose agents were interrupted by an unclean server stop and
# continues their work (goal re-arm or a continue follow-up).
- insert:
- id: watchdog
name: dsh-plugin-watchdog
```
The profile patch is read at server start, so the row takes effect on the **next
`dsh web` start** — which is also the first real test of the watchdog.
## How it decides what to resume
### 1. It tracks sessions while the server runs
A small persisted registry (`$DSH_HOME/storages/watchdog.json`) records every
session the watchdog sees and when it last did anything:
- every durable `session/event` touches the session's `lastActiveAt`;
- `agent/status` → `running` (and a periodic poll of `ctx.agents.list()`) also
records `lastRunningAt` and clears `cleanStop`.
### 2. A graceful stop is a deliberate stop
On a graceful shutdown (Ctrl+C / SIGTERM — the plugin's effect disposer and the
loader's `exit` event both run), every tracked session is marked `cleanStop`.
A crash or `kill -9` runs neither, so `cleanStop` stays `false`.
### 3. The boot scan resumes only real pending work
After the loader tree settles, the watchdog looks at every tracked session that
is **recently active** (`lastActivity` within `staleMs`, default 24 h), **not
cleanly stopped** (unless `resumeAfterGracefulShutdown`), and **not already
live**. It also does a one-time disk scan of the sessions directory for
recently-touched logs it has never seen (so it also helps sessions that were
running before the plugin was installed).
For each candidate it reads the session log **non-mutatingly**
(`sessionPersistence.inspect`) and checks for durable pending work:
| Signal | Means | Present when… |
|---|---|---|
| `interrupted-turn` | a `turn/end { interrupted }` closer is in the log — the persistence layer only emits these when it closed a turn the crashed process left open | crash mid-turn |
| `pending-inbox` | the durable `agent/inbox/spliced` projection is non-empty | prompts queued but never claimed |
| `active-goal` | the latest `goal/change` is phase `active` with `roundsStarted < maxGoalRounds` | automatic continuation was pending |
Deliberately stopped sessions show none of these: their turns close normally,
their inboxes are cleared, and a user-cancelled goal is durably **paused** by
the round driver. One race is covered explicitly: a goal session cancelled right
before the crash still reads phase `active` (the pause mutation had not landed)
— if the **only** signal is the goal and the log's **last** turn ended in a
user-abort, it is treated as deliberately stopped and left alone.
### 4. Resume mechanics
- `ctx.agents.resume({ resumeSessionId })` — the watchdog owns the returned
handle and disposes it on unload.
- Goal-only sessions are **re-armed** (`ctx.goals.resume` with the live
`{ id, revision }`) and left to the goal-round driver; they get **no**
follow-up prompt (the driver queues the next round itself).
- Anything else gets a `followup()` continuation prompt explaining the
interruption (configurable via `continuePrompt`).
- Subagent sessions are never resumed on their own — their parent drives them.
- Per-boot cap (`maxResumePerBoot`, default 8) prevents a resume storm; each
candidate's failure is contained (logged, next candidate proceeds).
## Config (per-entry `config:` in cordis.patch.yml)
| Key | Default | Meaning |
|---|---|---|
| `staleMs` | `86400000` (24 h) | Only sessions active within this window are resume candidates. |
| `bootDelayMs` | `1500` | Delay after activation before the boot scan starts. |
| `loaderSettleTimeoutMs` | `20000` | Max wait for the loader tree to settle before scanning. |
| `loaderSettlePollMs` | `250` | Poll interval while waiting. |
| `heartbeatIntervalMs` | `10000` | Running-agent heartbeat scan (unref'd timer). |
| `persistDebounceMs` | `300` | Debounce for registry writes. |
| `resumeAfterGracefulShutdown` | `false` | Also resume sessions after a *graceful* server stop — but only sessions that were actively **running** right up to the stop (see `gracefulWindowMs`). |
| `gracefulWindowMs` | `60000` | With `resumeAfterGracefulShutdown`, the window around the shutdown time in which a session's last activity must fall to be retriggered. Sessions that finished earlier stay asleep. |
| `resumeGoals` | `true` | Re-arm active goals with remaining capacity. |
| `maxResumePerBoot` | `8` | Safety cap on resumes per boot scan. |
| `scanUntracked` | `true` | Disk-scan recently-active sessions not yet in the registry. |
| `scanUntrackedMax` | `10` | Max untracked sessions inspected per boot. |
| `continuePrompt` | `""` | Custom continuation prompt; empty uses the short generated one. |
| `fallbackProvider` | `"deepseek"` | Provider used when a resumed session has no provider in its log. |
| `fallbackModel` | `"deepseek-v4-flash"` | Model used when a resumed session has no model in its log (headless resumes have no browser to install the model selection, so without a fallback the persona's `{{model}}` is empty and the first turn fails). |
| `resumeOnlyWhenNoClient` | `false` | Opt-in: skip auto-resume while a browser client is connected. Defaults to false — every interrupted session is resumed whether or not a browser is open. (A mid-turn kill can leave a truncated tool-call text block in the log; that is handled by the truncated-block note on resume, not by skipping the resume.) |
On resume, the watchdog also inspects the log tail for a **truncated tool-call
block** (an unclosed `` / ` The resume mechanics (route derivation, `installModelSelection` wrap, and the
> follow-up message injection) live in the shared
> [`dsh-lib-context-injection`](https://github.com/davidgereb/dsh-lib-context-injection)
> library, which dsh-plugin-cost-lens also uses for its scheduled/off-peak
> sends.
Browser half (`lib/client.js`): the Settings → Plugins card reading/writing
`/watchdog/api`.
## Caveats
- **Resume is a fresh continuation, not mid-turn time travel.** A crash-mid-turn
session is closed by the persistence layer's crash repair (synthetic
`tool/result` + `turn/end { interrupted }`), and the model continues from
there — it sees the repair markers and our follow-up prompt, and verifies
uncertain side effects before retrying. There is deliberately no
partial-turn rewind (that is a persistence-layer limitation, not this
plugin's).
- **Single-writer sessions.** Only one dsh process may write a session log at a
time. Do not run two servers against the same `$DSH_HOME` and expect both to
resume the same sessions; the losing `resume` is caught and logged.
- **Deliberate restarts are treated as deliberate stops** by default; flip
`resumeAfterGracefulShutdown` if you want graceful restarts to continue too.
- **The watchdog only tracks what it has seen.** The untracked disk scan covers
sessions that ran before install, but it is bounded (`scanUntrackedMax`) and
mtime-based — very large old logs are not scanned.
## Testing
```bash
# install the git dependency (dsh-lib-context-injection) so lib/index.js resolves:
npm install
node scripts/build.js
node scripts/smoke-test.mjs # pure event-analysis logic (30 checks)
node scripts/mock-boot-test.mjs # full apply() orchestration against a mock ctx (20 checks)
```
> **Compatibility.** Tested against **dsh `0.1.0-rc.6`** on Node.js
> **v24.19.0** (dsh web profile). Older or newer dsh releases may change the
> internals this plugin hooks into — check the changelog before upgrading.
---
> **⚠️ AI-generated, provided as-is.** This project was written with the
> assistance of an AI. It is provided **AS IS** without warranty of any kind,
> express or implied. The author cannot be held responsible for any damage,
> data loss, or misbehaviour that results from using it. Use at your own risk.
数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。