dsh-vision-bridge
为纯文本 DSH 会话按需提供视觉能力:图片变为标记,vision_describe 工具只把图片和问题发给 OpenAI 兼容的视觉模型。
sfyyy
@sfyyy
⬇ 2
★ 5
main
安装
dsh plugin --profile web add github:sfyyy/dsh-vision-bridge
需要可复现安装时,可在仓库后追加 #commit 固定提交。
为纯文本 DSH 会话按需提供视觉能力:图片变为标记,vision_describe 工具只把图片和问题发给 OpenAI 兼容的视觉模型。
该插件未提供要点说明,请参考仓库 README。
agentdeepseekdeepseek-harnessdshimage-understandingllm
- 安装并启动 DeepSeek Harness:
npx @deepseek-ai/dsh web - 在终端执行上面的安装命令(CLI 会解析插件并核验来源)
- 用 dsh plugins list 确认已安装,必要时重启 Harness 生效
插件以当前 dsh 进程的权限运行,安装时可能执行代码。请先通读仓库源码与许可证,确认无破坏性命令与越权访问;本站只做索引,不对第三方插件安全性作担保。
| 代码仓库 | github.com/sfyyy/dsh-vision-bridge |
| 许可证 | MIT |
| 主要语言 | main |
| 下载量 | 2 |
| GitHub 星标 | 5 |
| 最近推送 | 2026-08-16 |
| 收录日期 | 2026-09-19 |
| 分类 | 工具与能力 |
事实信息来自公开插件目录快照(2026-10-03),介绍文案由本站再加工。
以下为插件仓库 README 全文(原始内容,由公开目录抓取整理)。
# @dsh-extension/dsh-vision-bridge
> On-demand vision for text-only DeepSeek Harness (DSH) sessions.
[](https://developer.mozilla.org/en-US/docs/Web/JavaScript)
[](https://www.typescriptlang.org/)
[](https://www.npmjs.com/package/@dsh-extension/dsh-vision-bridge)
[](https://www.npmjs.com/package/@dsh-extension/dsh-vision-bridge)
[](https://github.com/sfyyy/dsh-vision-bridge)
[](./LICENSE)
[](https://www.npmjs.com/package/@deepseek-ai/dsh)
[中文文档](./README.zh-CN.md) · [npm](https://www.npmjs.com/package/@dsh-extension/dsh-vision-bridge)
A DSH plugin that gives a **text-only DeepSeek session on-demand multimodal capability**: the session stays on its text model for every turn, and only when the model actually needs to look at pixels — a screenshot, an uploaded image, a diagram, a chart — does it call the `vision_describe` tool, which sends **only the image(s) + a focused question** to an OpenAI-compatible vision model.
- **No long context ever reaches the vision model** — a 300k-token conversation history is never sent; each vision call is just image + question, keeping cost minimal.
- **Session log and UI keep the original images** — only the model *input* is rewritten to text markers.
- **Bring your own vision endpoint** — any OpenAI-compatible `/v1/chat/completions` service (OpenAI, DeepSeek, Gemini proxy, local vLLM/One-API, …).
## ❤️ Sponsors
> [Want to appear here?](mailto:sfyyy@users.noreply.github.com) — sponsor this project with an API donation.
[图片: xiaoyaoapi]
🎉 Thanks to xiaoyaoapi for donating their API to this project! xiaoyaoapi is an OpenAI-compatible AI API aggregation gateway for developers, built on New API with a unified admin dashboard. It offers unified key management, transparent usage tracking, and multi-channel access to mainstream large models under a single endpoint — letting developers integrate leading LLM services at lower cost and with greater convenience, ready to use as the vision endpoint of this plugin.
## How it works
```text
User / tool produces an image ──► image stays in the session and UI
│
▼ (model input layer)
image is rewritten to a text marker
(marker carries the attachment id and hints
the model to call vision_describe)
│
▼
text model calls vision_describe(attachmentIds / paths, question)
│
▼
vision model (receives only image + question) → text answer
→ returned as a normal tool result
```
- The `agent/pre-step` hook records every image attachment that appears in the session (user uploads **and** tool-produced screenshots, including ones nested inside `tool-result`), building an attachment index that `vision_describe` uses to resolve bytes by id.
- `session.deriveMessages()` is wrapped so that **no text-model request ever contains image blocks** (the native DeepSeek adapter rejects them); images are replaced by text markers. The session event log and UI keep showing the original images.
- DSH's built-in `llm-pi-ai` builds the OpenAI multimodal request; the plugin only maintains a single `vision-bridge` provider route and does not re-implement a protocol adapter.
### Image admission
DSH Web runs an image-capability check before a message enters the agent, based on the current DeepSeek model. This plugin keeps an admission bypass so images can enter the session first; the marker rewrite then guarantees no text-model request carries image blocks. When the plugin is disabled or uninstalled, the native admission check is restored.
Check `GET /_dsh/vision-bridge/settings` for the live `value.admissionBypass` and dependency-service status.
## Installation
Install from the npm registry (not a local checkout) — one command:
```sh
# if you already have the `dsh` CLI on PATH:
dsh plugin --profile web add @dsh-extension/dsh-vision-bridge
# or, if you have been using npx all along:
npx @deepseek-ai/dsh@0.1.0-rc.6 plugin --profile web add @dsh-extension/dsh-vision-bridge
```
> The `--profile` flag targets the profile you boot (`web` is the browser UI profile). Omit it or adapt it if your profile has a different name.
>
> After a new client bundle is added, restart `dsh web` once so the UI picks it up.
## Configuration
Configure it in **Settings → Vision Bridge** (DSH Web), or edit `~/.dsh/vision-bridge.json`:
```json
{
"enabled": true,
"baseUrl": "https://api.openai.com/v1",
"apiKey": "sk-xxxx",
"apiKeyEnv": "",
"model": "gpt-5.6-terra"
}
```
- `baseUrl` accepts an API root, a `.../v1` base, or a full `.../chat/completions` URL (the plugin normalizes it).
- `apiKey` and `apiKeyEnv` are mutually exclusive. A directly entered key is synced to the DSH credential store and referenced as `DSH_VISION_BRIDGE_API_KEY`.
- The plugin maintains exactly one `vision-bridge` route inside DSH's `llm-pi-ai.providers` and never touches other providers.
- `enabled: false` disables the whole chain: no tool registration, no image rewriting, no admission bypass (native behavior restored).
**Precedence (highest wins):** Settings page (with schema defaults) → environment variables → config file.
**Environment overrides:** `DSH_VISION_BRIDGE_BASE_URL`, `DSH_VISION_BRIDGE_API_KEY`, `DSH_VISION_BRIDGE_API_KEY_ENV`, `DSH_VISION_BRIDGE_MODEL`, `DSH_VISION_BRIDGE_ENABLED`.
## `vision_describe` tool
- **Arguments**
- `attachmentIds`: image attachment ids from the current conversation (shaped like `sha256:...`), one or several;
- `paths`: absolute local image file paths (`png`/`jpeg`/`webp`/`gif`) — use either or both, **1–4 images in total**;
- `question`: required — a focused, specific question about the image(s).
- **Behavior**: resolves the images → sends image(s) + question to the vision model via the `vision-bridge` route → returns the text answer as a tool result.
- **Multi-image comparison** is supported: put several images in the same user message.
- Attachment ids must come from the current conversation (user uploads or tool output); `paths` go through DSH's sandbox-aware file service.
## Verify
```sh
npm test
```
The suite covers: image-marker rewriting (including nested `tool-result`), both id- and path-based resolution, event-log attachment indexing, full-chain shutdown when disabled, and text-only sessions staying untouched.
## Development
From a local checkout:
```sh
dsh plugin inject /path/to/dsh-vision-bridge
```
## Search keywords
`deepseek` · `deepseek-harness` · `dsh` · `plugin` · `vision` · `multimodal` · `vision-language-model` · `VLM` · `image understanding` · `screenshot` · `OCR` · `image analysis` · `OpenAI-compatible` · `text-only-llm` · `on-demand vision` · `LLM agent`
## License
[MIT](./LICENSE)
数据来源:公开的 DeepSeek Harness 插件目录与各插件 GitHub 仓库。本站为独立第三方目录,与 DeepSeek、幻方(High-Flyer)及插件作者均无隶属或背书关系。