seek://map/anionex-dsh-vision-toolkit
DSH Vision Toolkit
#vision #screenshots #ui-restoration #ocr #tool-integration
TL;DR
Native DeepSeek Harness integration of an upstream agent-vision-toolkit: ships, per its README, ten specialised vision tools plus a bundled vision-skills Skill that teaches the agent when to inspect, ground, OCR, crop, trace, or compare pixels; targets text-only models with a free shared vision quota and a self-bootstrapping Python runtime.
The entry-level summary the upstream README leads with, quoted: "the plugin provides 10 tools" plus a bundled Skill that turns text-only DeepSeek Harness agents into something closer to a multimodal worker, anchored on the project's own upstream toolkit. Worth a slot on the map because vision is a gap in the broader DSH tooling landscape and this one ships a coherent set of primitives instead of a single caption endpoint.
Facts
- install
- dsh plugin --profile web add @anionex/dsh-vision-toolkit
- license
- MIT
Key points
- Native DSH profile integration: installable into web, headless and desktop profiles via dsh plugin add and surfaced under Settings → Vision Toolkit
- Per its README: ten tools, independently callable and composable into workflows; coordinates always use original-image pixels so grounding can feed directly into cropping, tracing, or downstream automation
- Bundled vision-skills Skill carries upstream playbooks covering long-screenshot OCR, UI/graphic/sketch restoration, structured extraction and GUI-from-screenshot operation
- Shipped as the DeepSeek Harness integration of the upstream Anionex/agent-vision-toolkit, which holds the underlying visual-task methodology
- Self-described in its own README as the ecosystem's first comprehensive vision plugin; not adopted as a fact here
- MIT licensed; npm package @anionex/dsh-vision-toolkit is the install target
FAQ
How is this different from dsh-tool-describe-image?
describe-image is a single-tool wrapper that forwards attached images to any OpenAI-compatible vision endpoint you configure. per its README, DSH Vision Toolkit ships ten specialised tools plus a Skill, brings its own free vision quota, and runs an isolated Python runtime for image processing — different scope and configuration surface.
Does it need my own API key?
The README documents a free shared vision service that works immediately after install with a daily quota; an API key is only needed if you want to switch providers in Settings.
Will it replace a multimodal model?
It is designed to make text-only models handle visual tasks via tool calls and the vision-skills workflow. README describes it as a parallel interaction mode rather than a replacement.
Any install gotchas on DSH Desktop?
README notes that the marketplace install in DSH Desktop 2.0.1 has known issues — the documented reliable path is opening DSH Terminal from the tray and running the dsh plugin --profile desktop add command manually.
Official references