analyze_images
One screenshot or an ordered batch ... resolved from paths, file:// URLs, markdown links, inline data or OpenCode caches ... returned as structured UI JSON: layout, components, colors, spacing, accessibility & more.
omni-vision-pro is a production-grade MCP server that lets text-only models ... DeepSeek, Codex, OpenCode, Claude, VS Code & more ... see screenshots, read source trees, and inspect ZIP archives. Free local OCR out of the box. Optional Gemini & OpenAI vision.
What it does
omni-vision-pro turns your text-only AI assistant into a multimodal powerhouse ... no new tool, no copy-paste folders, no cloud dependency.
One screenshot or an ordered batch ... resolved from paths, file:// URLs, markdown links, inline data or OpenCode caches ... returned as structured UI JSON: layout, components, colors, spacing, accessibility & more.
A clean file tree followed by safe text contents from any file or folder ... with node_modules, .git, binaries and secret files automatically skipped so the model only sees what matters.
Inspect any .zip entirely in memory ... virtual tree plus readable contents, with path-traversal, ZIP-bomb, encryption and duplicate checks. Nothing is ever extracted to disk.
How it works
A tiny local server sits next to your AI client and speaks the Model Context Protocol over stdio. Three steps, fully automatic.
Absolute & relative paths, file:// URLs, markdown image links, inline data URLs, OpenCode caches ... every reference is found and prepared.
AI vision first ... Gemini or OpenAI ... with automatic local Tesseract OCR fallback. Images are resized, EXIF-corrected and sent at low detail.
Ordered, delimited context blocks ... one result per input index. Failed images get their own result; nothing crashes, nothing is lost.
One-click setup
Paste one command into your AI chat ... or run it yourself. Detects clients, installs a pinned runtime, writes configs, and verifies all three tools.
One command detects your AI clients, installs a private version-pinned runtime and writes the MCP config. Setup is a one-time job.
Add a Gemini or OpenAI key later through a hidden prompt or a clipboard import. Switching providers never touches your client config.
Reload your AI client and ask it to analyze a screenshot. No cloud key? Local Tesseract OCR handles it, free and private.
$ npx --yes omni-vision-pro@latest setup --yes
Vision providers
Start free with OCR. Add a Gemini or OpenAI key whenever you want ... your choice, switchable any time, without reinstalling.
Local, private, instant ... works with zero API keys and zero internet. The safest fallback for every mode.
Powerful visual understanding. Defaults to gemini-1.5-flash with fallback; choose a current model via GEMINI_MODEL.
gpt-4o-mini at low detail ... EXIF-corrected, compressed to a 1024px box with sharp to keep costs down.
Install omni-vision-pro today and give your AI the context it's been missing.