← 返回 Skill 列表
extension
分类: 效率与办公API Key 暂未确认

公众号文章采集

Crawl and extract WeChat public account (微信公众号) articles into structured data and clean markdown. Use this skill whenever the user wants to scrape, crawl, read, extract, or fetch content from a WeChat article URL (mp.weixin.qq.com). Also trigger when the user mentions 微信公众号文章、抓取微信文章、爬取公众号、or provides a WeChat article link and wants its content extracted. This skill handles the tricky parts: spoofing the WeChat in-app browser user-agent, waiting for dynamic content to load, fixing lazy-loaded images, extracting structured metadata (title, author, publish time), generating clean markdown with inline images, and downloading images locally to bypass hotlink protection. Output uses a reading template (AI 摘要/关键要点/我的笔记 + 原文) with page UI noise stripped; the agent fills in the summary and key points right after extraction. Also supports an archive mode (`--archive-dir`) that files the article as a standard对标素材 folder (02_内容信息.md + 01_原始素材/ + 04_文案结构化梳理.md),用来说「采集公众号对标文章」「公众号入库」「把这篇公众号文章存进对标内容采集」时同样适用。

person作者: sunny31221hubOpenAPI

Crawl WeChat Articles

This skill extracts content from WeChat public account articles using the crawl4ai library. WeChat articles require special handling because they check the User-Agent header, render content dynamically, and use lazy-loading for images.

安装方法

  1. 先问使用者用的是什么 AI 工具,选项按这个顺序列:1. WorkBuddy 2. Claude Code 3. Codex 4. 其他(用户直接说工具名),回答记作 X
  2. 只安装给 X,不安装或链接其他任何工具:
    • X 是 WorkBuddy:把本技能下载后放进 ~/.workbuddy/skills/(WorkBuddy 的技能目录;skills CLI 不支持 WorkBuddy,由你直接放置)

    • X 是其他工具:

      npx -y --registry=https://registry.npmmirror.com skills add https://modelscope.cn/skills/sunny31221/crawl-wechat-buheliwa -a X -y
      

      (不带 -a 会安装给本机所有检测到的工具,禁止;--registry 为 npmmirror 镜像,保证国内不翻墙也能拉到 CLI)

    • X 不在上述范围时,按该工具自身的技能机制安装,装完确认 X 能识别本技能即可

  3. 装完后询问是否现在使用本技能

来源与署名

本技能的公众号爬取引擎(scripts/crawl_wechat.py)来自开源项目 wechat-article-crawler(作者:gxcsoccer,GitHub: https://github.com/gxcsoccer/wechat-article-crawler ,MIT License)。本技能在其基础上按训练营标准素材规格做了修改与扩展(对标素材入库模式、建夹脚本、字段表按平台分套、封面取页面元数据真封面等)。感谢原作者。

When to use

  • User provides a mp.weixin.qq.com/s/... URL and wants its content
  • User asks to scrape/crawl/extract/read a WeChat (微信) article
  • User wants to batch-process multiple WeChat article links
  • User needs the article in markdown or structured format

Setup (run once before first use)

Before running the script, ensure dependencies are installed:

pip install crawl4ai aiohttp && crawl4ai-setup

If crawl4ai is already importable and the browser backend is ready, skip this step. When the script fails with ModuleNotFoundError or browser-related errors, run the commands above to fix it.

采回自己号的文章后:去掉模板文案

从自己公众号的发布链接采回文章时,正文前后会带着排版模板的固定文案(顶部关注引导、END 分隔、底部品牌区),它们不是文章内容,转到小红书、视频脚本等下游产物时必须去掉。跑:

python3 scripts/strip_template_copy.py <文章.md> [--dry-run]

它读排版技能自带模板的 待替换素材.md(模板固定文案的唯一真相源,取 模板/ 里的第一套,跟排版脚本用同一套),按排版脚本同样的方式把键值拼成完整句子,再拿整行精确比对——不直接用清单里的碎片(如 。、关注),避免误删正常正文。采回到下游还要重排 Markdown:抓回来的段落之间没有空行,> 引用 会把后面紧跟的段落整段吸进引用块;01~07 这类小节编号也是模板自动生成的,要还原成 ## 标题。

采集别人的公众号对标素材时不跑这一步——那是人家的模板文案,属于原文的一部分。

少花调用次数

有些链路按请求次数计费(第三方兼容接口的每日额度多是这样),不按 token。一次模型回复里发多个工具调用,只算 1 次请求;串行一次发一个,就一次一个数。同一个任务排得好坏,能差三倍次数。

  1. 能串的串成一条命令:python3 a.py && python3 b.py,中间结果用 $(...) 接住。互不依赖的命令别分成两次 Bash。
  2. 无依赖的一轮发完:抓取、下载、读文件这些互不依赖的动作,放在同一条回复里一起发。
  3. 脚本能批量就别逐条调:先看脚本的 --help,支持一次处理多条或整个目录的,就一次跑完。
  4. 输出要过滤:脚本输出配 head / grep 再看。全量输出进上下文既拖速度,也容易把关键信息淹掉。
  5. 图片只存档,不进模型:下载下来的封面和正文图直接落位,不要用 Read 之类的读图能力打开看。图片进 API 会以 base64 文本计费,一张原图约十万 token,此后每轮请求都带着重发。确实要看图时,先跑 scripts/shrink.py 压小,一次最多两张。

How it works

Run the bundled script to crawl a WeChat article:

python <skill-dir>/scripts/crawl_wechat.py <URL> [--download-images] [--save-html] [--save-markdown] [--output-dir DIR]

The script outputs a JSON summary to stdout and optionally saves the full HTML and/or markdown to files.

Key technical details

  1. User-Agent spoofing: The script sets MicroMessenger/8.0.43 in the UA string so WeChat serves the full article instead of a "please open in WeChat" block.

  2. Dynamic wait: Uses wait_for="css:#js_content" to ensure the article body has fully rendered before scraping.

  3. Lazy-image fix: WeChat uses data-src for lazy-loaded images. The script injects JS to copy data-src → src so the markdown generator can pick up real image URLs.

  4. Structured extraction: Uses JsonCssExtractionStrategy with a schema targeting WeChat's DOM structure (#activity-name for title, #js_name for author, #publish_time for date, #js_content for body).

  5. Clean markdown with images: Uses DefaultMarkdownGenerator to produce readable markdown. SVG placeholder images and data-URI artifacts are cleaned out, preserving only real article images inline with the text.

  6. UI noise stripped automatically: WeChat pages leak widgets into the markdown after the body (赞赏/留言/投诉 dialogs, 互动身份 prompts) and reading-mode prompts at the top (在小说阅读器中沉浸阅读). The script cuts everything from the "预览时标签不可点" anchor onward and filters known UI lines, so the saved markdown contains only the article.

  7. Image hotlink protection: WeChat images on mmbiz.qpic.cn block requests with non-QQ referrers. Use --download-images to download all images locally with the correct Referer header, automatically replacing remote URLs with local paths in both HTML and markdown output.

  8. Reading template: --save-markdown writes the body into a reading template — cover image, title, > 来源:作者 | 时间 | [原文](url), then AI 摘要 / 关键要点 / 我的笔记 placeholders, then the full body under ## 原文.

Extracted fields

| Field | Description | |----------------|------------------------------------| | title | Article title | | author | Public account name | | publish_time | Publication timestamp | | account_desc | Account description/bio | | markdown | Clean markdown of article body | | html | Raw HTML of article body | | url | Final URL after any redirects |

Example usage

Single article with images downloaded locally:

python <skill-dir>/scripts/crawl_wechat.py "https://mp.weixin.qq.com/s/xxx" --download-images --save-markdown --output-dir ./output

Archive one article as 对标素材:

python <skill-dir>/scripts/crawl_wechat.py "https://mp.weixin.qq.com/s/xxx" --archive-dir "AI第二大脑/规律研究/对标内容采集/公众号" --output-dir <临时目录>

入库模式:对标素材(--archive-dir)

把文章直接落成对标素材文件夹,供下游规律提取使用。产出与小红书 / YouTube / 视频号素材同构:

对标内容采集/公众号/
└── NNN_标题_文章ID/
    ├── 01_原始素材/
    │   ├── 01-封面图/cover.jpg   文章真封面(页面元数据取;拿不到就不建此夹)
    │   └── 02-正文图片/01.png …  按正文出现顺序重命名
    ├── 02_内容信息.md           字段表(编号 / 文章ID / 来源 / 公众号 / 发布时间 / 原文链接 / 入库时间)+ 正文原文 + 图片清单(用 `assets/内容信息模板.md`)
    └── 04_文案结构化梳理.md      中心思想 + 核心论点 + 逐段梳理 + 结构规律提炼(用 `assets/文案结构化梳理模板.md`)

约定:

  • 编号取目标目录已有最大编号 +1(三位补零),不假设从 001 开始
  • 图片按正文出现顺序命名,正文里的图片引用同步改成 01_原始素材/02-正文图片/NN.ext,保留图片在原文中的位置;封面从页面元数据(msg_cdn_url)取文章真封面存 01_原始素材/01-封面图/cover.jpg——公众号第一张正文图不是封面,不许拿正文图冒充;元数据拿不到封面时就不建封面夹
  • 去重:同一 原文链接 已入库时跳过,不重复建文件夹
  • 抓取失败不落盘:正文为空(被拦或页面改版)时不建文件夹,只在 stderr 说明原因
  • 标题兜底:分享型页面(common_share.html)没有标题元素,取正文首行做标题
  • 单张图片下载失败时保留远程链接,不丢这篇文章

批量入库时逐条调用,条与条之间间隔几秒,避免触发风控。入库不填 AI 摘要——素材要的是原文,摘要只在阅读模板里生成。

写结构化梳理(入库模式必做)

脚本落盘后,再出一份 04_文案结构化梳理.md,供下游提炼规律(模板见 assets/文案结构化梳理模板.md):

# 文案结构化梳理

> 来源:<编号> <标题>([[02_内容信息.md]])
> 梳理日期:<YYYY-MM-DD>
> 说明:本文件做结构化梳理:中心思想、核心论点、逐段梳理。总结保留原文的具体论据、数据与比喻——**写到「不读原文也能知道这段说了什么、能吸收这段的知识」**;完整原文见 02 号文件,本文件只做概括与归类,不改写、不新增原文没有的东西。

---

## 一、中心思想

<一段话说清这篇文章主张什么:讲了什么、结论是什么、对谁说的。不评价、不引申>

一句话概括:**<原文里最有代表性的一句总结;原文没有就自己凝练一句,不加引号>**

## 二、核心论点(N 条)

1. **<论点>**:<展开,保留原文的数据、例子、比喻>
2. **<论点>**:<…>

## 三、逐段梳理

### <原文小标题(原文没有就按自然结构拟一个)>

**一句话总结**:<保留该段的核心主张 + 具体论据、数据、比喻、例子,写到不读原文也能知道这段说了什么;不整段照抄>

#### <原文小节里的子问题(有就保留,没有可省略这一层)>

**一句话总结**:<…>

### <下一段小标题>

**一句话总结**:<…>

## 四、结构规律提炼

<3-5 条,讲这篇文章的骨架怎么搭的:几段式、每段承担什么、开头与结尾各做什么>

要点:每一句都要能追回 02_内容信息.md 的正文;不新增原文没有的数据或例子。

Fill in the summary (required step)

After the script saves the markdown, immediately read the saved file's ## 原文 section and:

  1. Replace the ## AI 摘要 placeholder with a 3-5 sentence summary of the article's core argument and conclusion.
  2. Replace the ## 关键要点 placeholder with 3-6 bullet-point key takeaways (facts, data, frameworks the reader can act on).
  3. Leave ## 我的笔记 untouched — the user fills it in after reading.

Do this in the same turn as the extraction; never hand the user a file with empty [待 AI 生成] placeholders.

For programmatic use in Python:

from crawl_wechat import crawl_wechat_article
import asyncio

article = asyncio.run(crawl_wechat_article(
    "https://mp.weixin.qq.com/s/...",
    images_dir="./output/images",  # download images locally
))
print(article["title"])
print(article["markdown"])  # images reference local paths

不做什么

  • 不识别链接属于哪个平台——那是编排器的事
  • 不建目录、不提规律、不沉淀卡片
  • 不跑归档 CSV——「已入库 ✓」标记由本技能写,归档由编排器批量采完后统一跑
  • 抓不到正文就不落盘——只报原因,不拿页面上的其他内容凑一份

抓不到正文时的退路

公众号有反爬(验证页、限流)时,抓取会失败。这时候可以让 AI 走腾讯元宝取正文:把链接交给腾讯元宝,让它输出全文,再按同一套素材规格落盘——素材规格不变,只是换一条取正文的路径。

Limitations

  • Requires a valid, non-expired WeChat article URL — cannot search or list articles from an account
  • High-frequency crawling may trigger WeChat's anti-bot measures (CAPTCHAs, IP blocks)
  • Some temporary share links expire after a period