dsh-client-vision (tool-vision)
屏幕截图与外部视觉识别:take_screenshot、list_windows、analyze_image、view_image 四个工具,可配置 GPT 视觉通道(gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra),API Key 经凭据服务存储,带设置卡片;view_image 在 Web 对话中显示截图,模型上下文只保留文字。
安装
dsh plugin --profile web add github:ankye/dsh-client-vision
需要可复现安装时,可在仓库后追加 #commit 固定提交。
屏幕截图与外部视觉识别:take_screenshot、list_windows、analyze_image、view_image 四个工具,可配置 GPT 视觉通道(gpt-5.5 / gpt-5.6-sol / gpt-5.6-terra),API Key 经凭据服务存储,带设置卡片;view_image 在 Web 对话中显示截图,模型上下文只保留文字。
该插件未提供要点说明,请参考仓库 README。
- 安装并启动 DeepSeek Harness:
npx @deepseek-ai/dsh web - 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
- 用 dsh plugins list 确认已安装,必要时重启 Harness 生效
插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。
| 代码仓库 | github.com/ankye/dsh-client-vision/tree/main/packages/tool-vision |
| 许可证 | MIT |
| 主要语言 | main |
| 下载量 | 1 |
| GitHub 星标 | 1 |
| 最近推送 | 2026-09-10 |
| 收录日期 | 2026-09-19 |
| 分类 | 工具与能力 |
事实信息来自公开插件目录快照(2026-10-01),介绍文案由本站再加工。
以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。
# dsh-client-vision
English | [中文](README.zh.md)
Give your DeepSeek Harness agent **eyes**. `dsh-client-vision` is a screen-capture + external image-recognition plugin for [DeepSeek Harness](https://github.com/deepseek-ai/deepseek-harness): the agent takes a screenshot (or points at any image), hands it to a vision-capable model through a pluggable channel, and gets back plain text it can actually act on — **no multimodal model required**.
## Compatibility
This revision requires DeepSeek Harness core `0.1.2-alpha.5` or later within the `0.1.x` line. It uses the Settings service API introduced in that core release.
## Why you want it
- **DeepSeek can't see — now it can.** The harness model has no image input. This plugin runs the whole "look" outside the model and returns text the agent can reason about, exactly like Codex's semantic vision tool.
- **Capture anything, any way.** `fullscreen` / `window` (with live window enumeration) / `region` / `interactive` — grab the browser, a game window, or one corner of the screen.
- **Multi-channel by design.** Tools are decoupled from recognition backends. The `gpt` channel ships ready to use; adding Claude, Gemini, or a local model is **one `analyze()` implementation + one registry line** — the three tools never change.
- **Secret-safe.** The API key lives in the harness `credentials` store (`VISION_GPT_API_KEY`) — never in settings files, logs, or the conversation transcript.
- **Every preset, out of the box.** Mounted on the host plane, so `code`, `standard`, `cordis`, `minimal` — every agent sees the tools. No preset switching.
- **Ready to ship.** Prebuilt bundles included; three install paths (drop into the monorepo / `pnpm publish` / tarball).
- **Smart payloads.** Large captures are auto-downscaled and re-encoded (≤1568px JPEG q80) before they leave the machine.
## Capabilities
### Tools
| Tool | What it does |
|---|---|
| `take_screenshot` | Capture the screen: `fullscreen` (primary display), `window` (by id from `list_windows`), `region` (x, y, width, height), `interactive` (user selection), `android` (adb device/emulator), or `ios` (booted simulator). Returns the PNG path + dimensions. |
| `list_windows` | Enumerate on-screen windows (`id`, `app`, `title`) — macOS `CGWindowList`, Windows `Get-Process` main handles, Linux X11 (`wmctrl`/`xprop`) — pick the browser or game window to capture. |
| `analyze_image` | Submit an image (a path, or the most recent screenshot) to the configured vision channel and return a plain-text description. |
| `view_image` | One-shot "look at this": capture the screen (or use `image_path`) and recognize it through the active channel. The screenshot is rendered as an image card in the Web conversation, while the model context receives only the plain-text description — the image bytes never enter the model context. |
### Platforms
| Platform | Capture backend | Window enumeration | Extra requirements |
|---|---|---|---|
| macOS | `screencapture` (system) | Swift `CGWindowList` | Screen Recording permission on first use |
| Windows | PowerShell `System.Drawing` (system) | `Get-Process` main window handles | PowerShell `System.Drawing` |
| Linux | ImageMagick `import` | `wmctrl` + `xprop` | ImageMagick (`convert`/`identify`), `wmctrl`, `x11-utils` |
`mode=interactive` (system selection UI) is macOS-only; on Windows and Linux
use `mode=region` with explicit coordinates.
### Device capture
| Mode | What it captures | Requirements |
|---|---|---|
| `android` | A connected Android device or emulator screen | `adb` on PATH with a device online (`adb devices`); works from any host. With several devices online, pass `device=`. |
| `ios` | The booted iOS simulator | macOS host with Xcode (`xcrun simctl`) |
### Settings (`vision` namespace)
Configured in **Settings → Plugins → Plugin configuration → Vision**:
| Field | Meaning |
|---|---|
| Endpoint (`baseUrl`) | Domain + optional path prefix; `/chat/completions` is appended. e.g. `https://api.example.com/v1` |
| Channel | The active recognition backend (currently `gpt`). |
| Model | `gpt-5.5` / `gpt-5.6-sol` / `gpt-5.6-terra` |
| API key | Stored through the harness `credentials` service as `VISION_GPT_API_KEY`; the literal never leaves your machine. |
### Channels
| Channel | Backend | Model | API key |
|---|---|---|---|
| `gpt` | OpenAI-compatible `/chat/completions` | `gpt-5.5` / `gpt-5.6-sol` / `gpt-5.6-terra` | required (e.g. `VISION_GPT_API_KEY`) |
| `zhipu` | Zhipu GLM-4V, OpenAI-compatible `/chat/completions` | `glm-4v-plus` / `glm-4v-flash` | required (e.g. `VISION_ZHIPU_API_KEY`) |
| `ollama` | local Ollama `/api/chat` (default `http://localhost:11434`) | `llava` / `llava-llama3` / `bakllava` / `moondream` / `qwen2-vl` / `minicpm-v` (or any installed vision model) | none |
Pick the channel in **Settings → Plugins → Vision**; the model dropdown follows
the channel and the API-key control is hidden for `ollama`. For `ollama` the
base URL defaults to `http://localhost:11434` and the model to `llava` when
left blank.
### Multi-channel architecture
```
model → analyze_image(image, prompt)
│ reads vision.channel
▼
channels//analyze() ← one implementation per backend
│
gpt: POST {baseUrl}/chat/completions (image_url data URL)
claude / gemini / local: … ← add yours here
```
Adding a channel is deliberately small:
```ts
// src/channels//index.ts
export async function myAnalyze(ctx, call): Promise {
// call.imageB64, call.mime, call.prompt, call.config, call.signal
return await fetchYourVisionApi(...)
}
```
```ts
// src/channels/index.ts — one registry line
export const channels = {
gpt: { label: 'GPT', analyze: gptAnalyze },
myChannel: { label: 'My Channel', analyze: myAnalyze },
}
```
The tools (`take_screenshot` / `list_windows` / `analyze_image`) and their schemas never change.
## Installation (official — no repo modification)
`dsh plugin add` installs the packages into your profile; each package declares `dsh.bundle`, so the rows mount automatically — no patch rows, no repo edits.
### Prerequisites
- DeepSeek Harness core `0.1.2-alpha.5` or later within the `0.1.x` line, plus `dsh` and `pnpm` on PATH.
### 1. Get the packages (pick one)
**a. From this repository (recommended until published to npm):**
```sh
dsh plugin --profile web add \
file:/path/to/dsh-client-vision/packages/tool-vision \
file:/path/to/dsh-client-vision/packages/ui-vision
```
**b. Tarball:**
```sh
cd packages/tool-vision && npm pack
cd packages/ui-vision && npm pack
dsh plugin --profile web add file:/path/to/deepseek-ai-dsh-tool-vision-0.1.0-rc.7.tgz \
file:/path/to/deepseek-ai-dsh-client-ui-vision-0.1.0-rc.7.tgz
```
**c. npm registry (after publishing):**
```sh
dsh plugin --profile web add @deepseek-ai/dsh-tool-vision @deepseek-ai/dsh-client-ui-vision
```
> A `[WARN] Issues with peer dependencies` message is expected and safe to ignore — the peers come from your deployment's own bundles at runtime.
### 2. Verify
```sh
node -e "console.log(JSON.stringify(require(process.env.HOME + '/.dsh/profiles/web/package.json').dsh.profile.bundles))"
# should list dsh-tool-vision and dsh-client-ui-vision
```
### 3. Restart + configure
Restart the harness, then **Settings → Plugins → Plugin configuration → Vision**: set the endpoint, model, and your own API key (`VISION_GPT_API_KEY`), save.
### 4. Verify
Ask the agent to "look at the screen" — it should call `take_screenshot` → `analyze_image` and describe what it sees.
### Uninstall
```sh
dsh plugin --profile web remove @deepseek-ai/dsh-tool-vision @deepseek-ai/dsh-client-ui-vision
```
## Alternative: build inside a harness fork
If you run a **fork** of deepseek-harness (not the official deployment), you can drop the packages into the monorepo instead:
```sh
cp -R packages/tool-vision /packages/vision/tool-vision
cp -R packages/ui-vision /packages/client/ui-vision
```
Then add both to `apps/cli/package.json` (`workspace:^`), add `./packages/vision/tool-vision` to `tsconfig.host.json` and `./packages/client/ui-vision` to `tsconfig.client.json`, `pnpm install`, build (`tsdown` host + client passes), and restart.
## Quick start
1. Restart the harness.
2. The tool catalog now includes `take_screenshot` / `list_windows` / `analyze_image`.
3. Open **Settings → Plugins → Plugin configuration → Vision**, set the endpoint, model, and your own API key, and save.
4. Ask the agent to "look at the screen" — it will screenshot and describe what it sees.
## Development
- This repository is a **source distribution**: the peer packages (`@deepseek-ai/dsh-tools`, …) resolve from your deployment. `lib/` ships prebuilt, so `npm pack` works immediately.
- The `tsconfig.json` files are standalone; the harness monorepo's build pipeline (including the client-bundle `tsdown.config.ts`) applies in Option A.
- **Never commit secrets.** The API key stays in each machine's `.credentials.yaml`.
## License
MIT
数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。