← 返回 Skill 列表
extension
分类: 开发与工程API Key 暂未确认

屏幕视觉感知技能

一款专为 Intel AI PC 打造的本地化屏幕视觉感知技能。

person作者: showdaihubModelScope

ScreenVision Skill

Prefer the fixed, read-only files for agent-facing checks and searches. They do not require shell execution or network permission prompts.

Agent workflow

  1. Check service health by reading screenvision_status.json in this skill folder. Treat status=running as ready. If the file is missing or stale, use the fixed fallback command:
    python screenvision_cli.py health
    
    • Treat status=running as ready.
    • On connection failure or a stopped service, run setup.bat from this skill folder.
    • For status=initializing, wait and retry. The first model export can take 5-20 minutes.
    • For status=error, report message to the user.
  2. Remember the current screen:
    python capture_once.py
    
  3. Search past screens by reading screen_memory.jsonl in this skill folder and matching the user's natural-language request against each record's summary. The newest records are at the end. If the file is unavailable, use the fixed fallback command:
    python screenvision_cli.py search "python error"
    

Do not use curl, PowerShell web cmdlets, browser tools, or direct HTTP URLs for health checks or searches. Do not inspect the ChromaDB files directly.

MCP tools (offline-native, preferred for MCP-capable hosts)

screenvision_mcp.py is a stdio MCP server exposing the same capabilities as native tools. It needs no skill instructions, no internet, and works while the network is fully disconnected (the only online step ever is the one-time model download in setup.bat).

Tools provided:

| Tool | Purpose | | --- | --- | | screenvision_status | Health: backend state / device / model cache readiness | | screenvision_capture | Screenshot -> local VLM summary -> index (auto-starts backend) | | screenvision_search | Natural-language memory search (offline JSONL fallback) | | screenvision_recent | List recent memories (pure file read, always offline) | | screenvision_start | Start the local backend |

Register it with any MCP client using mcp_config.example.json in this folder (copy the mcpServers block). All traffic is stdio + 127.0.0.1 loopback.

Unified CLI

python screenvision_cli.py health          # service status
python screenvision_cli.py capture         # capture + summarize + index
python screenvision_cli.py search "query"  # search memories (--limit N)
python screenvision_cli.py recent [N]      # newest N memories
python screenvision_cli.py start           # start the backend

All commands print JSON and work offline. search/recent fall back to the local screen_memory.jsonl log when the backend is not running.

Setup

Run the Windows setup once:

setup.bat

The setup installs Python and dependencies when required, starts server.py and the Ctrl+Shift+M hotkey client, then waits for model readiness.

Optional silent autostart: run start_hidden.vbs to start both the server and hotkey client.

Files

screenvision/
|-- SKILL.md
|-- setup.bat
|-- wait_ready.ps1
|-- server.py               # FastAPI backend (OpenVINO VLM + ChromaDB)
|-- screenvision_mcp.py     # stdio MCP server (offline-native agent tools)
|-- screenvision_core.py    # shared logic for MCP server + CLI
|-- screenvision_stdio.py   # Windows-safe UTF-8 print/JSON helpers
|-- screenvision_cli.py     # Unified offline CLI: health/capture/search/recent/start
|-- mcp_config.example.json # Generic MCP client registration snippet
|-- screenvision_status.json # Agent-readable service status
|-- screen_memory.jsonl     # Agent-readable screen summaries
|-- capture_once.py         # One-shot screen capture and indexing (legacy entry)
|-- auto_capture.py         # Ctrl+Shift+M listener
|-- start_hidden.vbs
|-- requirements.txt
|-- skill.json
|-- screen_memory_db/       # Created at runtime
`-- ov_model_cache/         # Created after first model export

Notes

  • Device policy (automatic): probe OpenVINO available_devices, then load in NPU → GPU → CPU order. Missing hardware is skipped; a failed compile falls through to the next candidate. Active device is reported in /health and screenvision_status.json as device, plus available_devices / device_candidates.
  • Force a device with SCREENVISION_DEVICE=NPU|GPU|CPU (or a concrete id like GPU.0) when starting server.py.
  • Console / log output is UTF-8 safe on Windows (PYTHONUTF8, PYTHONIOENCODING=utf-8:replace, shared screenvision_stdio helpers, and UTF-8 FileHandler in the hotkey client) so Unicode in messages or VLM summaries cannot crash under cp1252.
  • Cache the INT4 OpenVINO model under ov_model_cache/.
  • Allow about 30 GB of free disk space for the first model download and conversion.
  • Offline behavior: MCP (stdio) and backend (loopback) never touch the internet. The one-time model download in setup.bat is the only online step; after the cache exists everything runs disconnected.
  • Environment variables honored by the MCP server / CLI / backend: SCREENVISION_BASE_URL (default http://127.0.0.1:8000), SCREENVISION_MEMORY_LOG (override the memory JSONL path), SCREENVISION_NO_AUTOSTART=1 (disable backend auto-start), and SCREENVISION_DEVICE (optional device override; see above).