ScreenVision Skill
Prefer the fixed, read-only files for agent-facing checks and searches. They do not require shell execution or network permission prompts.
Agent workflow
- Check service health by reading
screenvision_status.jsonin this skill folder. Treatstatus=runningas ready. If the file is missing or stale, use the fixed fallback command:python screenvision_cli.py health- Treat
status=runningas ready. - On connection failure or a stopped service, run
setup.batfrom this skill folder. - For
status=initializing, wait and retry. The first model export can take 5-20 minutes. - For
status=error, reportmessageto the user.
- Treat
- Remember the current screen:
python capture_once.py - Search past screens by reading
screen_memory.jsonlin this skill folder and matching the user's natural-language request against each record'ssummary. The newest records are at the end. If the file is unavailable, use the fixed fallback command:python screenvision_cli.py search "python error"
Do not use curl, PowerShell web cmdlets, browser tools, or direct HTTP URLs for
health checks or searches. Do not inspect the ChromaDB files directly.
MCP tools (offline-native, preferred for MCP-capable hosts)
screenvision_mcp.py is a stdio MCP server exposing the same capabilities as
native tools. It needs no skill instructions, no internet, and works while the
network is fully disconnected (the only online step ever is the one-time model
download in setup.bat).
Tools provided:
| Tool | Purpose |
| --- | --- |
| screenvision_status | Health: backend state / device / model cache readiness |
| screenvision_capture | Screenshot -> local VLM summary -> index (auto-starts backend) |
| screenvision_search | Natural-language memory search (offline JSONL fallback) |
| screenvision_recent | List recent memories (pure file read, always offline) |
| screenvision_start | Start the local backend |
Register it with any MCP client using mcp_config.example.json in this folder
(copy the mcpServers block). All traffic is stdio + 127.0.0.1 loopback.
Unified CLI
python screenvision_cli.py health # service status
python screenvision_cli.py capture # capture + summarize + index
python screenvision_cli.py search "query" # search memories (--limit N)
python screenvision_cli.py recent [N] # newest N memories
python screenvision_cli.py start # start the backend
All commands print JSON and work offline. search/recent fall back to the
local screen_memory.jsonl log when the backend is not running.
Setup
Run the Windows setup once:
setup.bat
The setup installs Python and dependencies when required, starts server.py and the Ctrl+Shift+M hotkey client, then waits for model readiness.
Optional silent autostart: run start_hidden.vbs to start both the server and hotkey client.
Files
screenvision/
|-- SKILL.md
|-- setup.bat
|-- wait_ready.ps1
|-- server.py # FastAPI backend (OpenVINO VLM + ChromaDB)
|-- screenvision_mcp.py # stdio MCP server (offline-native agent tools)
|-- screenvision_core.py # shared logic for MCP server + CLI
|-- screenvision_stdio.py # Windows-safe UTF-8 print/JSON helpers
|-- screenvision_cli.py # Unified offline CLI: health/capture/search/recent/start
|-- mcp_config.example.json # Generic MCP client registration snippet
|-- screenvision_status.json # Agent-readable service status
|-- screen_memory.jsonl # Agent-readable screen summaries
|-- capture_once.py # One-shot screen capture and indexing (legacy entry)
|-- auto_capture.py # Ctrl+Shift+M listener
|-- start_hidden.vbs
|-- requirements.txt
|-- skill.json
|-- screen_memory_db/ # Created at runtime
`-- ov_model_cache/ # Created after first model export
Notes
- Device policy (automatic): probe OpenVINO
available_devices, then load in NPU → GPU → CPU order. Missing hardware is skipped; a failed compile falls through to the next candidate. Active device is reported in/healthandscreenvision_status.jsonasdevice, plusavailable_devices/device_candidates. - Force a device with
SCREENVISION_DEVICE=NPU|GPU|CPU(or a concrete id likeGPU.0) when startingserver.py. - Console / log output is UTF-8 safe on Windows (
PYTHONUTF8,PYTHONIOENCODING=utf-8:replace, sharedscreenvision_stdiohelpers, and UTF-8FileHandlerin the hotkey client) so Unicode in messages or VLM summaries cannot crash under cp1252. - Cache the INT4 OpenVINO model under
ov_model_cache/. - Allow about 30 GB of free disk space for the first model download and conversion.
- Offline behavior: MCP (stdio) and backend (loopback) never touch the internet.
The one-time model download in
setup.batis the only online step; after the cache exists everything runs disconnected. - Environment variables honored by the MCP server / CLI / backend:
SCREENVISION_BASE_URL(defaulthttp://127.0.0.1:8000),SCREENVISION_MEMORY_LOG(override the memory JSONL path),SCREENVISION_NO_AUTOSTART=1(disable backend auto-start), andSCREENVISION_DEVICE(optional device override; see above).
微信扫一扫