Crawl WeChat Articles
This skill extracts content from WeChat public account articles using the crawl4ai library. WeChat articles require special handling because they check the User-Agent header, render content dynamically, and use lazy-loading for images.
安装方法
-
先问使用者用的是什么 AI 工具(如 Claude Code、Codex、Cursor),回答记作 X
-
只安装给 X,不安装或链接其他任何工具:
npx -y --registry=https://registry.npmmirror.com skills add https://modelscope.cn/skills/sunny31221/crawl-wechat -a X -y(
--registry为 npmmirror 镜像,保证国内不翻墙也能装;-a X表示只装给 X 这一个工具) X 不在上述命令支持范围时,按该工具自身的技能机制安装,装完确认 X 能识别本技能即可 -
装完后询问是否现在使用本技能
When to use
- User provides a
mp.weixin.qq.com/s/...URL and wants its content - User asks to scrape/crawl/extract/read a WeChat (微信) article
- User wants to batch-process multiple WeChat article links
- User needs the article in markdown or structured format
Setup (run once before first use)
Before running the script, ensure dependencies are installed:
pip install crawl4ai aiohttp && crawl4ai-setup
If crawl4ai is already importable and the browser backend is ready, skip this step. When the script fails with ModuleNotFoundError or browser-related errors, run the commands above to fix it.
How it works
Run the bundled script to crawl a WeChat article:
python <skill-dir>/scripts/crawl_wechat.py <URL> [--download-images] [--save-html] [--save-markdown] [--output-dir DIR]
The script outputs a JSON summary to stdout and optionally saves the full HTML and/or markdown to files.
Key technical details
-
User-Agent spoofing: The script sets
MicroMessenger/8.0.43in the UA string so WeChat serves the full article instead of a "please open in WeChat" block. -
Dynamic wait: Uses
wait_for="css:#js_content"to ensure the article body has fully rendered before scraping. -
Lazy-image fix: WeChat uses
data-srcfor lazy-loaded images. The script injects JS to copydata-src→srcso the markdown generator can pick up real image URLs. -
Structured extraction: Uses
JsonCssExtractionStrategywith a schema targeting WeChat's DOM structure (#activity-namefor title,#js_namefor author,#publish_timefor date,#js_contentfor body). -
Clean markdown with images: Uses
DefaultMarkdownGeneratorto produce readable markdown. SVG placeholder images and data-URI artifacts are cleaned out, preserving only real article images inline with the text. -
UI noise stripped automatically: WeChat pages leak widgets into the markdown after the body (赞赏/留言/投诉 dialogs, 互动身份 prompts) and reading-mode prompts at the top (在小说阅读器中沉浸阅读). The script cuts everything from the "预览时标签不可点" anchor onward and filters known UI lines, so the saved markdown contains only the article.
-
Image hotlink protection: WeChat images on
mmbiz.qpic.cnblock requests with non-QQ referrers. Use--download-imagesto download all images locally with the correct Referer header, automatically replacing remote URLs with local paths in both HTML and markdown output. -
Reading template:
--save-markdownwrites the body into a reading template — cover image, title,> 来源:作者 | 时间 | [原文](url), thenAI 摘要 / 关键要点 / 我的笔记placeholders, then the full body under## 原文.
Extracted fields
| Field | Description |
|----------------|------------------------------------|
| title | Article title |
| author | Public account name |
| publish_time | Publication timestamp |
| account_desc | Account description/bio |
| markdown | Clean markdown of article body |
| html | Raw HTML of article body |
| url | Final URL after any redirects |
Example usage
Single article with images downloaded locally:
python <skill-dir>/scripts/crawl_wechat.py "https://mp.weixin.qq.com/s/xxx" --download-images --save-markdown --output-dir ./output
Archive one article as 对标素材:
python <skill-dir>/scripts/crawl_wechat.py "https://mp.weixin.qq.com/s/xxx" --archive-dir "AI第二大脑/02_规律研究/对标内容采集/公众号" --output-dir <临时目录>
入库模式:对标素材(--archive-dir)
把文章直接落成对标素材文件夹,供下游规律提取使用。产出与小红书 / YouTube / 视频号素材同构:
对标内容采集/公众号/
└── NNN_标题_文章ID/
├── 01_原始素材/
│ ├── 01-封面图/01.png 第一张图的副本(封面单独留一份,供以后做封面规律提取)
│ └── 02-正文图片/01.png … 按正文出现顺序重命名
├── 02_笔记信息.md 字段表(编号 / 文章ID / 来源 / 公众号 / 发布时间 / 原文链接 / 入库时间)+ 正文原文 + 图片清单
└── 04_文案结构化梳理.md 中心思想 + 核心论点 + 逐段梳理 + 结构规律提炼
约定:
- 编号取目标目录已有最大编号 +1(三位补零),不假设从 001 开始
- 图片按正文出现顺序命名,正文里的图片引用同步改成
01_原始素材/02-正文图片/NN.ext,保留图片在原文中的位置;第一张另复制一份进01_原始素材/01-封面图/ - 去重:同一
原文链接已入库时跳过,不重复建文件夹 - 抓取失败不落盘:正文为空(被拦或页面改版)时不建文件夹,只在 stderr 说明原因
- 标题兜底:分享型页面(
common_share.html)没有标题元素,取正文首行做标题 - 单张图片下载失败时保留远程链接,不丢这篇文章
批量入库时逐条调用,条与条之间间隔几秒,避免触发风控。入库不填 AI 摘要——素材要的是原文,摘要只在阅读模板里生成。
写结构化梳理(入库模式必做)
脚本落盘后,再出一份 04_文案结构化梳理.md,供下游提炼规律:
# 文案结构化梳理
> 来源:<编号> <标题>([[02_笔记信息.md]])
> 梳理日期:<YYYY-MM-DD>
> 说明:本文件做结构化梳理:中心思想、核心论点、逐段小标题 + 一句话总结。总结保留原文的具体论据、数据与比喻,不删减关键信息;完整原文见 02 号文件。
---
## 一、中心思想
<一段话说清这篇文章主张什么;下面附一句可直接引用的一句话说概括>
## 二、核心论点(N 条)
1. **<论点>**:<展开,保留原文的数据、例子、比喻>
## 三、逐段梳理
### <原文小标题>|<一句话总结>
- **一句话总结**:<这一段讲了什么>
- **原文示例**:<引原句,可核对>
- **可复用点**:<这条做法在哪能用>
## 四、结构规律提炼
<3-5 条,讲这篇文章的骨架怎么搭的:几段式、每段承担什么、开头与结尾各做什么>
要点:每一句都要能追回 02_笔记信息.md 的正文;不新增原文没有的数据或例子。
Fill in the summary (required step)
After the script saves the markdown, immediately read the saved file's ## 原文 section and:
- Replace the
## AI 摘要placeholder with a 3-5 sentence summary of the article's core argument and conclusion. - Replace the
## 关键要点placeholder with 3-6 bullet-point key takeaways (facts, data, frameworks the reader can act on). - Leave
## 我的笔记untouched — the user fills it in after reading.
Do this in the same turn as the extraction; never hand the user a file with empty [待 AI 生成] placeholders.
For programmatic use in Python:
from crawl_wechat import crawl_wechat_article
import asyncio
article = asyncio.run(crawl_wechat_article(
"https://mp.weixin.qq.com/s/...",
images_dir="./output/images", # download images locally
))
print(article["title"])
print(article["markdown"]) # images reference local paths
Limitations
- Requires a valid, non-expired WeChat article URL — cannot search or list articles from an account
- High-frequency crawling may trigger WeChat's anti-bot measures (CAPTCHAs, IP blocks)
- Some temporary share links expire after a period
微信扫一扫