Crawl WeChat Articles
This skill extracts content from WeChat public account articles using the crawl4ai library. WeChat articles require special handling because they check the User-Agent header, render content dynamically, and use lazy-loading for images.
安装方法
- 先问使用者用的是什么 AI 工具,选项按这个顺序列:1. WorkBuddy 2. Claude Code 3. Codex 4. 其他(用户直接说工具名),回答记作 X
- 只安装给 X,不安装或链接其他任何工具:
-
X 是 WorkBuddy:把本技能下载后放进
~/.workbuddy/skills/(WorkBuddy 的技能目录;skills CLI 不支持 WorkBuddy,由你直接放置) -
X 是其他工具:
npx -y --registry=https://registry.npmmirror.com skills add https://modelscope.cn/skills/sunny31221/crawl-wechat-buheliwa -a X -y(不带
-a会安装给本机所有检测到的工具,禁止;--registry为 npmmirror 镜像,保证国内不翻墙也能拉到 CLI) -
X 不在上述范围时,按该工具自身的技能机制安装,装完确认 X 能识别本技能即可
-
- 装完后询问是否现在使用本技能
来源与署名
本技能的公众号爬取引擎(scripts/crawl_wechat.py)来自开源项目 wechat-article-crawler(作者:gxcsoccer,GitHub: https://github.com/gxcsoccer/wechat-article-crawler ,MIT License)。本技能在其基础上按训练营标准素材规格做了修改与扩展(对标素材入库模式、建夹脚本、字段表按平台分套、封面取页面元数据真封面等)。感谢原作者。
When to use
- User provides a
mp.weixin.qq.com/s/...URL and wants its content - User asks to scrape/crawl/extract/read a WeChat (微信) article
- User wants to batch-process multiple WeChat article links
- User needs the article in markdown or structured format
Setup (run once before first use)
Before running the script, ensure dependencies are installed:
pip install crawl4ai aiohttp && crawl4ai-setup
If crawl4ai is already importable and the browser backend is ready, skip this step. When the script fails with ModuleNotFoundError or browser-related errors, run the commands above to fix it.
采回自己号的文章后:去掉模板文案
从自己公众号的发布链接采回文章时,正文前后会带着排版模板的固定文案(顶部关注引导、END 分隔、底部品牌区),它们不是文章内容,转到小红书、视频脚本等下游产物时必须去掉。跑:
python3 scripts/strip_template_copy.py <文章.md> [--dry-run]
它读排版技能自带模板的 待替换素材.md(模板固定文案的唯一真相源,取 模板/ 里的第一套,跟排版脚本用同一套),按排版脚本同样的方式把键值拼成完整句子,再拿整行精确比对——不直接用清单里的碎片(如 。、关注),避免误删正常正文。采回到下游还要重排 Markdown:抓回来的段落之间没有空行,> 引用 会把后面紧跟的段落整段吸进引用块;01~07 这类小节编号也是模板自动生成的,要还原成 ## 标题。
采集别人的公众号对标素材时不跑这一步——那是人家的模板文案,属于原文的一部分。
少花调用次数
有些链路按请求次数计费(第三方兼容接口的每日额度多是这样),不按 token。一次模型回复里发多个工具调用,只算 1 次请求;串行一次发一个,就一次一个数。同一个任务排得好坏,能差三倍次数。
- 能串的串成一条命令:
python3 a.py && python3 b.py,中间结果用$(...)接住。互不依赖的命令别分成两次 Bash。 - 无依赖的一轮发完:抓取、下载、读文件这些互不依赖的动作,放在同一条回复里一起发。
- 脚本能批量就别逐条调:先看脚本的
--help,支持一次处理多条或整个目录的,就一次跑完。 - 输出要过滤:脚本输出配
head/grep再看。全量输出进上下文既拖速度,也容易把关键信息淹掉。 - 图片只存档,不进模型:下载下来的封面和正文图直接落位,不要用 Read 之类的读图能力打开看。图片进 API 会以 base64 文本计费,一张原图约十万 token,此后每轮请求都带着重发。确实要看图时,先跑
scripts/shrink.py压小,一次最多两张。
How it works
Run the bundled script to crawl a WeChat article:
python <skill-dir>/scripts/crawl_wechat.py <URL> [--download-images] [--save-html] [--save-markdown] [--output-dir DIR]
The script outputs a JSON summary to stdout and optionally saves the full HTML and/or markdown to files.
Key technical details
-
User-Agent spoofing: The script sets
MicroMessenger/8.0.43in the UA string so WeChat serves the full article instead of a "please open in WeChat" block. -
Dynamic wait: Uses
wait_for="css:#js_content"to ensure the article body has fully rendered before scraping. -
Lazy-image fix: WeChat uses
data-srcfor lazy-loaded images. The script injects JS to copydata-src→srcso the markdown generator can pick up real image URLs. -
Structured extraction: Uses
JsonCssExtractionStrategywith a schema targeting WeChat's DOM structure (#activity-namefor title,#js_namefor author,#publish_timefor date,#js_contentfor body). -
Clean markdown with images: Uses
DefaultMarkdownGeneratorto produce readable markdown. SVG placeholder images and data-URI artifacts are cleaned out, preserving only real article images inline with the text. -
UI noise stripped automatically: WeChat pages leak widgets into the markdown after the body (赞赏/留言/投诉 dialogs, 互动身份 prompts) and reading-mode prompts at the top (在小说阅读器中沉浸阅读). The script cuts everything from the "预览时标签不可点" anchor onward and filters known UI lines, so the saved markdown contains only the article.
-
Image hotlink protection: WeChat images on
mmbiz.qpic.cnblock requests with non-QQ referrers. Use--download-imagesto download all images locally with the correct Referer header, automatically replacing remote URLs with local paths in both HTML and markdown output. -
Reading template:
--save-markdownwrites the body into a reading template — cover image, title,> 来源:作者 | 时间 | [原文](url), thenAI 摘要 / 关键要点 / 我的笔记placeholders, then the full body under## 原文.
Extracted fields
| Field | Description |
|----------------|------------------------------------|
| title | Article title |
| author | Public account name |
| publish_time | Publication timestamp |
| account_desc | Account description/bio |
| markdown | Clean markdown of article body |
| html | Raw HTML of article body |
| url | Final URL after any redirects |
Example usage
Single article with images downloaded locally:
python <skill-dir>/scripts/crawl_wechat.py "https://mp.weixin.qq.com/s/xxx" --download-images --save-markdown --output-dir ./output
Archive one article as 对标素材:
python <skill-dir>/scripts/crawl_wechat.py "https://mp.weixin.qq.com/s/xxx" --archive-dir "AI第二大脑/规律研究/对标内容采集/公众号" --output-dir <临时目录>
入库模式:对标素材(--archive-dir)
把文章直接落成对标素材文件夹,供下游规律提取使用。产出与小红书 / YouTube / 视频号素材同构:
对标内容采集/公众号/
└── NNN_标题_文章ID/
├── 01_原始素材/
│ ├── 01-封面图/cover.jpg 文章真封面(页面元数据取;拿不到就不建此夹)
│ └── 02-正文图片/01.png … 按正文出现顺序重命名
├── 02_内容信息.md 字段表(编号 / 文章ID / 来源 / 公众号 / 发布时间 / 原文链接 / 入库时间)+ 正文原文 + 图片清单(用 `assets/内容信息模板.md`)
└── 04_文案结构化梳理.md 中心思想 + 核心论点 + 逐段梳理 + 结构规律提炼(用 `assets/文案结构化梳理模板.md`)
约定:
- 编号取目标目录已有最大编号 +1(三位补零),不假设从 001 开始
- 图片按正文出现顺序命名,正文里的图片引用同步改成
01_原始素材/02-正文图片/NN.ext,保留图片在原文中的位置;封面从页面元数据(msg_cdn_url)取文章真封面存01_原始素材/01-封面图/cover.jpg——公众号第一张正文图不是封面,不许拿正文图冒充;元数据拿不到封面时就不建封面夹 - 去重:同一
原文链接已入库时跳过,不重复建文件夹 - 抓取失败不落盘:正文为空(被拦或页面改版)时不建文件夹,只在 stderr 说明原因
- 标题兜底:分享型页面(
common_share.html)没有标题元素,取正文首行做标题 - 单张图片下载失败时保留远程链接,不丢这篇文章
批量入库时逐条调用,条与条之间间隔几秒,避免触发风控。入库不填 AI 摘要——素材要的是原文,摘要只在阅读模板里生成。
写结构化梳理(入库模式必做)
脚本落盘后,再出一份 04_文案结构化梳理.md,供下游提炼规律(模板见 assets/文案结构化梳理模板.md):
# 文案结构化梳理
> 来源:<编号> <标题>([[02_内容信息.md]])
> 梳理日期:<YYYY-MM-DD>
> 说明:本文件做结构化梳理:中心思想、核心论点、逐段梳理。总结保留原文的具体论据、数据与比喻——**写到「不读原文也能知道这段说了什么、能吸收这段的知识」**;完整原文见 02 号文件,本文件只做概括与归类,不改写、不新增原文没有的东西。
---
## 一、中心思想
<一段话说清这篇文章主张什么:讲了什么、结论是什么、对谁说的。不评价、不引申>
一句话概括:**<原文里最有代表性的一句总结;原文没有就自己凝练一句,不加引号>**
## 二、核心论点(N 条)
1. **<论点>**:<展开,保留原文的数据、例子、比喻>
2. **<论点>**:<…>
## 三、逐段梳理
### <原文小标题(原文没有就按自然结构拟一个)>
**一句话总结**:<保留该段的核心主张 + 具体论据、数据、比喻、例子,写到不读原文也能知道这段说了什么;不整段照抄>
#### <原文小节里的子问题(有就保留,没有可省略这一层)>
**一句话总结**:<…>
### <下一段小标题>
**一句话总结**:<…>
## 四、结构规律提炼
<3-5 条,讲这篇文章的骨架怎么搭的:几段式、每段承担什么、开头与结尾各做什么>
要点:每一句都要能追回 02_内容信息.md 的正文;不新增原文没有的数据或例子。
Fill in the summary (required step)
After the script saves the markdown, immediately read the saved file's ## 原文 section and:
- Replace the
## AI 摘要placeholder with a 3-5 sentence summary of the article's core argument and conclusion. - Replace the
## 关键要点placeholder with 3-6 bullet-point key takeaways (facts, data, frameworks the reader can act on). - Leave
## 我的笔记untouched — the user fills it in after reading.
Do this in the same turn as the extraction; never hand the user a file with empty [待 AI 生成] placeholders.
For programmatic use in Python:
from crawl_wechat import crawl_wechat_article
import asyncio
article = asyncio.run(crawl_wechat_article(
"https://mp.weixin.qq.com/s/...",
images_dir="./output/images", # download images locally
))
print(article["title"])
print(article["markdown"]) # images reference local paths
不做什么
- 不识别链接属于哪个平台——那是编排器的事
- 不建目录、不提规律、不沉淀卡片
- 不跑归档 CSV——「已入库 ✓」标记由本技能写,归档由编排器批量采完后统一跑
- 抓不到正文就不落盘——只报原因,不拿页面上的其他内容凑一份
抓不到正文时的退路
公众号有反爬(验证页、限流)时,抓取会失败。这时候可以让 AI 走腾讯元宝取正文:把链接交给腾讯元宝,让它输出全文,再按同一套素材规格落盘——素材规格不变,只是换一条取正文的路径。
Limitations
- Requires a valid, non-expired WeChat article URL — cannot search or list articles from an account
- High-frequency crawling may trigger WeChat's anti-bot measures (CAPTCHAs, IP blocks)
- Some temporary share links expire after a period
Scan to join WeChat group