Skills Plugins MCP Prompt Model 导航 博客 资讯 我的中心
工具与能力 #dsh-plugin

dsh-client-vision (tool-vision)

屏幕截图与外部视觉识别:take_screenshot、list_windows、analyze_image、view_image 四个工具,可配置 GPT 视觉通道(gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra),API Key 经凭据服务存储,带设置卡片;view_image 在 Web 对话中显示截图,模型上下文只保留文字。

ankye @ankye ⬇ 1 ★ 1 main

安装

dsh plugin --profile web add github:ankye/dsh-client-vision
下载安装清单

需要可复现安装时,可在仓库后追加 #commit 固定提交。

屏幕截图与外部视觉识别:take_screenshot、list_windows、analyze_image、view_image 四个工具,可配置 GPT 视觉通道(gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra),API Key 经凭据服务存储,带设置卡片;view_image 在 Web 对话中显示截图,模型上下文只保留文字。

该插件未提供要点说明,请参考仓库 README。

dsh-plugin
  1. 安装并启动 DeepSeek Harness:npx @deepseek-ai/dsh web
  2. 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
  3. 用 dsh plugins list 确认已安装,必要时重启 Harness 生效

插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。

代码仓库github.com/ankye/dsh-client-vision/tree/main/packages/tool-vision
许可证MIT
主要语言main
下载量1
GitHub 星标1
最近推送2026-09-10
收录日期2026-09-19
分类工具与能力

事实信息来自公开插件目录快照(2026-10-01),介绍文案由本站再加工。

以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。

# dsh-client-vision

English | [中文](README.zh.md)

Give your DeepSeek Harness agent **eyes**. `dsh-client-vision` is a screen-capture + external image-recognition plugin for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness): the agent takes a screenshot (or points at any image), hands it to a vision-capable model through a pluggable channel, and gets back plain text it can actually act on — **no multimodal model required**.

## Compatibility

This revision requires DeepSeek Harness core `0.1.2-alpha.5` or later within the `0.1.x` line. It uses the Settings service API introduced in that core release.

## Why you want it

- **DeepSeek can't see — now it can.** The harness model has no image input. This plugin runs the whole "look" outside the model and returns text the agent can reason about, exactly like Codex's semantic vision tool.
- **Capture anything, any way.** `fullscreen` / `window` (with live window enumeration) / `region` / `interactive` — grab the browser, a game window, or one corner of the screen.
- **Multi-channel by design.** Tools are decoupled from recognition backends. The `gpt` channel ships ready to use; adding Claude, Gemini, or a local model is **one `analyze()` implementation + one registry line** — the three tools never change.
- **Secret-safe.** The API key lives in the harness `credentials` store (`VISION_GPT_API_KEY`) — never in settings files, logs, or the conversation transcript.
- **Every preset, out of the box.** Mounted on the host plane, so `code`, `standard`, `cordis`, `minimal` — every agent sees the tools. No preset switching.
- **Ready to ship.** Prebuilt bundles included; three install paths (drop into the monorepo / `pnpm publish` / tarball).
- **Smart payloads.** Large captures are auto-downscaled and re-encoded (≤1568px JPEG q80) before they leave the machine.

## Capabilities

### Tools

| Tool | What it does |
|---|---|
| `take_screenshot` | Capture the screen: `fullscreen` (primary display), `window` (by id from `list_windows`), `region` (x, y, width, height), `interactive` (user selection), `android` (adb device/emulator), or `ios` (booted simulator). Returns the PNG path + dimensions. |
| `list_windows` | Enumerate on-screen windows (`id`, `app`, `title`) — macOS `CGWindowList`, Windows `Get-Process` main handles, Linux X11 (`wmctrl`/`xprop`) — pick the browser or game window to capture. |
| `analyze_image` | Submit an image (a path, or the most recent screenshot) to the configured vision channel and return a plain-text description. |
| `view_image` | One-shot "look at this": capture the screen (or use `image_path`) and recognize it through the active channel. The screenshot is rendered as an image card in the Web conversation, while the model context receives only the plain-text description — the image bytes never enter the model context. |

### Platforms

| Platform | Capture backend | Window enumeration | Extra requirements |
|---|---|---|---|
| macOS | `screencapture` (system) | Swift `CGWindowList` | Screen Recording permission on first use |
| Windows | PowerShell `System.Drawing` (system) | `Get-Process` main window handles | PowerShell `System.Drawing` |
| Linux | ImageMagick `import` | `wmctrl` + `xprop` | ImageMagick (`convert`/`identify`), `wmctrl`, `x11-utils` |

`mode=interactive` (system selection UI) is macOS-only; on Windows and Linux
use `mode=region` with explicit coordinates.

### Device capture

| Mode | What it captures | Requirements |
|---|---|---|
| `android` | A connected Android device or emulator screen | `adb` on PATH with a device online (`adb devices`); works from any host. With several devices online, pass `device=`. |
| `ios` | The booted iOS simulator | macOS host with Xcode (`xcrun simctl`) |

### Settings (`vision` namespace)

Configured in **Settings → Plugins → Plugin configuration → Vision**:

| Field | Meaning |
|---|---|
| Endpoint (`baseUrl`) | Domain + optional path prefix; `/chat/completions` is appended. e.g. `https://api.example.com/v1` |
| Channel | The active recognition backend (currently `gpt`). |
| Model | `gpt-5.5` / `gpt-5.6-sol` / `gpt-5.6-terra` |
| API key | Stored through the harness `credentials` service as `VISION_GPT_API_KEY`; the literal never leaves your machine. |

### Channels

| Channel | Backend | Model | API key |
|---|---|---|---|
| `gpt` | OpenAI-compatible `/chat/completions` | `gpt-5.5` / `gpt-5.6-sol` / `gpt-5.6-terra` | required (e.g. `VISION_GPT_API_KEY`) |
| `zhipu` | Zhipu GLM-4V, OpenAI-compatible `/chat/completions` | `glm-4v-plus` / `glm-4v-flash` | required (e.g. `VISION_ZHIPU_API_KEY`) |
| `ollama` | local Ollama `/api/chat` (default `http://localhost:11434`) | `llava` / `llava-llama3` / `bakllava` / `moondream` / `qwen2-vl` / `minicpm-v` (or any installed vision model) | none |

Pick the channel in **Settings → Plugins → Vision**; the model dropdown follows
the channel and the API-key control is hidden for `ollama`. For `ollama` the
base URL defaults to `http://localhost:11434` and the model to `llava` when
left blank.

### Multi-channel architecture

```
model → analyze_image(image, prompt)
          │  reads vision.channel
          ▼
  channels//analyze()        ← one implementation per backend
          │
  gpt:    POST {baseUrl}/chat/completions   (image_url data URL)
  claude / gemini / local: …    ← add yours here
```

Adding a channel is deliberately small:

```ts
// src/channels//index.ts
export async function myAnalyze(ctx, call): Promise {
  // call.imageB64, call.mime, call.prompt, call.config, call.signal
  return await fetchYourVisionApi(...)
}
```

```ts
// src/channels/index.ts — one registry line
export const channels = {
  gpt: { label: 'GPT', analyze: gptAnalyze },
  myChannel: { label: 'My Channel', analyze: myAnalyze },
}
```

The tools (`take_screenshot` / `list_windows` / `analyze_image`) and their schemas never change.

## Installation (official — no repo modification)

`dsh plugin add` installs the packages into your profile; each package declares `dsh.bundle`, so the rows mount automatically — no patch rows, no repo edits.

### Prerequisites

- DeepSeek Harness core `0.1.2-alpha.5` or later within the `0.1.x` line, plus `dsh` and `pnpm` on PATH.

### 1. Get the packages (pick one)

**a. From this repository (recommended until published to npm):**

```sh
dsh plugin --profile web add \
  file:/path/to/dsh-client-vision/packages/tool-vision \
  file:/path/to/dsh-client-vision/packages/ui-vision
```

**b. Tarball:**

```sh
cd packages/tool-vision && npm pack
cd packages/ui-vision   && npm pack
dsh plugin --profile web add file:/path/to/deepseek-ai-dsh-tool-vision-0.1.0-rc.7.tgz \
                            file:/path/to/deepseek-ai-dsh-client-ui-vision-0.1.0-rc.7.tgz
```

**c. npm registry (after publishing):**

```sh
dsh plugin --profile web add @deepseek-ai/dsh-tool-vision @deepseek-ai/dsh-client-ui-vision
```

> A `[WARN] Issues with peer dependencies` message is expected and safe to ignore — the peers come from your deployment's own bundles at runtime.

### 2. Verify

```sh
node -e "console.log(JSON.stringify(require(process.env.HOME + '/.dsh/profiles/web/package.json').dsh.profile.bundles))"
# should list dsh-tool-vision and dsh-client-ui-vision
```

### 3. Restart + configure

Restart the harness, then **Settings → Plugins → Plugin configuration → Vision**: set the endpoint, model, and your own API key (`VISION_GPT_API_KEY`), save.

### 4. Verify

Ask the agent to "look at the screen" — it should call `take_screenshot` → `analyze_image` and describe what it sees.

### Uninstall

```sh
dsh plugin --profile web remove @deepseek-ai/dsh-tool-vision @deepseek-ai/dsh-client-ui-vision
```

## Alternative: build inside a harness fork

If you run a **fork** of deepseek-harness (not the official deployment), you can drop the packages into the monorepo instead:

```sh
cp -R packages/tool-vision /packages/vision/tool-vision
cp -R packages/ui-vision   /packages/client/ui-vision
```

Then add both to `apps/cli/package.json` (`workspace:^`), add `./packages/vision/tool-vision` to `tsconfig.host.json` and `./packages/client/ui-vision` to `tsconfig.client.json`, `pnpm install`, build (`tsdown` host + client passes), and restart.

## Quick start

1. Restart the harness.
2. The tool catalog now includes `take_screenshot` / `list_windows` / `analyze_image`.
3. Open **Settings → Plugins → Plugin configuration → Vision**, set the endpoint, model, and your own API key, and save.
4. Ask the agent to "look at the screen" — it will screenshot and describe what it sees.

## Development

- This repository is a **source distribution**: the peer packages (`@deepseek-ai/dsh-tools`, …) resolve from your deployment. `lib/` ships prebuilt, so `npm pack` works immediately.
- The `tsconfig.json` files are standalone; the harness monorepo's build pipeline (including the client-bundle `tsdown.config.ts`) applies in Option A.
- **Never commit secrets.** The API key stays in each machine's `.credentials.yaml`.

## License

MIT

数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。