<essential_principles> <principle name="Catalog First, Crawl Second"> This skill has two halves that share one knowledge base:
- Discover & catalog — periodically scan GitHub for "collection-class" repos (scrapers, API collectors, agent skills/MCP, datasets/awesome-lists) and store them in
references/tool-catalog.{md,json}. - Match & crawl — when the user names a target, recommend tools from the catalog via progressive disclosure, then start crawling.
Always check the catalog freshness before recommending. If stale (> N days), surface that to the user.
</principle>
<principle name="Chinese-Social First">
The headline use case is Chinese social/e-commerce platforms. When the user names one (小红书/抖音/B站/微博/知乎/贴吧/快手/公众号/视频号/淘宝/京东/拼多多/豆瓣/雪球…), route directly to the fast path (workflows/match-and-crawl.md Step 3B) via references/chinese-social-platforms.md — never show the generic category menu. The general five-category catalog is the fallback breadth for every other target.
Compliance gate (non-negotiable): most Chinese-platform tools are reverse-engineered and violate platform ToS. Always show the compliance reminder before tool selection, re-confirm intent before the first network request, never help evade anti-bot / risk-control systems, and default to low request rates + anonymized personal data. </principle>
<principle name="Progressive Disclosure"> Never dump the whole catalog. Show a **category menu** first, then drill into tool cards within the chosen category, then load the matching workflow only after the user picks a tool. Load references lazily — only what the current step needs. </principle> <principle name="Five Canonical Categories"> Every repo in the catalog is tagged with exactly one primary category:| Tag | Means |
|-----|-------|
| web-scraper | Static HTML / simple HTTP fetch (BeautifulSoup, httpx, Selectolax) |
| dynamic-scraper | JS-rendered pages, SPAs (Playwright, Selenium, Crawl4AI) |
| api-collector | REST/GraphQL endpoints, SDK-driven pulls, ETL pipelines |
| agent-skill | Claude/GPT agent skills, MCP servers, tool-use frameworks |
| dataset | Public datasets, awesome-lists, curated resource repos |
Categories drive both discovery keywords and the matching menu. </principle>
<principle name="Catalog Is the Source of Truth"> - `references/tool-catalog.json` — structured data, the canonical source. - `references/tool-catalog.md` — human-readable view, **regenerated from JSON** by `scripts/build_catalog_md.py`. Never hand-edit the markdown. - `references/discovery-log.md` — append-only run history (when, what, how many, errors). </principle> <principle name="Safety & Boundaries"> - Respect `robots.txt`, rate limits, and ToS. Default to authenticated GitHub API calls (higher limits) when scanning repos. - Never store credentials in the repo. Read tokens from env vars (`GITHUB_TOKEN`) or the `gh` CLI keyring. - Always confirm scope with the user before the first network request against a new target domain. </principle> <principle name="LLM Judging Is Optional & Safely Guarded"> Discovery *can* use an LLM (`LLM_API_KEY`) to judge candidates and reassign categories, but **pruning is never the model's call alone** — deterministic collection-signal and human-curation guards defend it. Full mechanism in `references/llm-judging.md`. Not needed for the match-and-crawl path. </principle> </essential_principles> <intake> On invocation, determine the user's intent. Most messages fall into one of:- "抓小红书 / 抖音评论 / 微博热搜…" (Chinese platform named) — fast path: platform-tagged shortlist + compliance gate, straight to a tool.
- "Find/discover tools" — refresh the catalog, search GitHub for collection-class repos.
- "I want to scrape/fetch X" (any other target) — match a tool from the catalog to a target and start crawling.
- "Browse the catalog" — show categories / tool cards without scraping.
- "Schedule / automate" — set up the periodic refresh hook.
If the message is ambiguous, ask one minimal question to disambiguate, then route. </intake>
<routing> Map user intent to a workflow:| User says | Route to |
|-----------|----------|
| any Chinese platform name (小红书/抖音/B站/微博/淘宝…, full alias table in references/chinese-social-platforms.md) | workflows/match-and-crawl.md (Step 3B fast path) |
| "refresh", "discover", "find new repos", "update catalog" | workflows/discover-catalog.md |
| "I want to scrape/fetch/collect X", "抓 X 数据" | workflows/match-and-crawl.md |
| "show me the catalog", "what tools do we have", "browse" | workflows/browse-catalog.md |
| "schedule", "automate refresh", "定时", "cron" | workflows/schedule-refresh.md |
After routing, the workflow tells you which references to load. Do not preload everything. </routing>
<reference_index>
All in references/:
tool-catalog.json— canonical structured catalog (data source)tool-catalog.md— human-readable catalog (generated, do not edit)discovery-log.md— append-only run logcategory-keywords.md— search keywords + GitHub topic mapping per categorychinese-social-platforms.md— canonical platform registry: alias → key map, tag conventions, per-platform crawl characteristics (load whenever a Chinese platform is suspected)repo-schema.md— the JSON schema every catalog entry must satisfyrate-limit-guide.md— GitHub API quota, pagination, retry/backoff patternsllm-judging.md— optional LLM judge mechanism + safe-pruning guards (load only when running/tuning discovery) </reference_index>
<workflows_index>
All in workflows/:
| Workflow | Purpose |
|----------|---------|
| discover-catalog.md | Scan GitHub, dedupe, update catalog JSON + MD |
| match-and-crawl.md | Progressive-disclosure matching → tool selection → crawl |
| browse-catalog.md | Read-only category/card view, no network crawling |
| schedule-refresh.md | Install a periodic refresh hook (cron / Task Scheduler) |
| tools/*.md | Per-tool crawl workflows — loaded by match-and-crawl.md Step 5 when a catalog entry's workflow_file points here. Each is a concrete crawl recipe (spec → skeleton → pre-flight → run → validate → report). See tools/mediacrawler.md for the compliance-gated pattern. |
Existing per-tool workflows: tools/scrapy.md, tools/playwright.md, tools/crawl4ai.md, tools/mediacrawler.md.
</workflows_index>
<success_criteria> This skill works when:
- A Chinese-platform request hits the fast path and sees the compliance reminder before any tool card.
- The catalog contains real, recently-verified entries across all five categories.
tool-catalog.mdandtool-catalog.jsonstay in sync (MD regenerated from JSON).- A user asking "I want to scrape X" gets a category menu → tool card → workflow in ≤ 2 turns.
- Discovery runs are idempotent and logged in
discovery-log.md. - No credentials are committed; tokens come from env or
ghkeyring. </success_criteria>
Scan to join WeChat group