Anionex/agent-vision-toolkit
1,125 stars · Last commit 2026-08-27
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, with optional seamless integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode
README preview
<p align="center"> <img src="assets/hero.png" alt="agent-vision-toolkit — Give text-only LLM agents eyes." width="100%"> </p> <div align="center"> # agent-vision-toolkit <a href="https://trendshift.io/repositories/99395?utm_source=trendshift-badge&utm_medium=badge&utm_campaign=badge-trendshift-99395" target="_blank" rel="noopener noreferrer"><img src="https://trendshift.io/api/badge/trendshift/repositories/99395/daily?language=Python" alt="Anionex%2Fagent-vision-toolkit | Trendshift" width="250" height="55"/></a> [](https://x.com/anion_ex) [](https://github.com/Anionex/agent-vision-toolkit/stargazers) [](https://github.com/Anionex/agent-vision-toolkit/forks) [](https://github.com/Anionex/agent-vision-toolkit/blob/main/LICENSE) [](https://agentskills.io) [](https://github.com/Anionex/agent-vision-toolkit/tree/main/extensions) [](https://github.com/Anionex/agent-vision-toolkit/tree/main/bin) **What it thinks is what it sees — give any text-only coding agent eyes: image Q&A, long-screenshot OCR, frontend UI restoration, and GUI automation, as a vision toolkit plus a skill, with optional drop-in integration for Codex, Claude Code, Pi, Oh My Pi, and OpenCode.**