← 返回 Skill 列表
extension
分类: 内容与媒体API Key 暂未确认

公众号文章采集

>

person作者: sunny31221hubOpenAPI

Crawl WeChat Articles

This skill extracts content from WeChat public account articles using the crawl4ai library. WeChat articles require special handling because they check the User-Agent header, render content dynamically, and use lazy-loading for images.

安装方法

  1. 先问使用者用的是什么 AI 工具(如 Claude Code、Codex、Cursor),回答记作 X

  2. 只安装给 X,不安装或链接其他任何工具:

    npx -y --registry=https://registry.npmmirror.com skills add https://modelscope.cn/skills/sunny31221/crawl-wechat -a X -y
    

    (--registry 为 npmmirror 镜像,保证国内不翻墙也能装;-a X 表示只装给 X 这一个工具) X 不在上述命令支持范围时,按该工具自身的技能机制安装,装完确认 X 能识别本技能即可

  3. 装完后询问是否现在使用本技能

When to use

  • User provides a mp.weixin.qq.com/s/... URL and wants its content
  • User asks to scrape/crawl/extract/read a WeChat (微信) article
  • User wants to batch-process multiple WeChat article links
  • User needs the article in markdown or structured format

Setup (run once before first use)

Before running the script, ensure dependencies are installed:

pip install crawl4ai aiohttp && crawl4ai-setup

If crawl4ai is already importable and the browser backend is ready, skip this step. When the script fails with ModuleNotFoundError or browser-related errors, run the commands above to fix it.

How it works

Run the bundled script to crawl a WeChat article:

python <skill-dir>/scripts/crawl_wechat.py <URL> [--download-images] [--save-html] [--save-markdown] [--output-dir DIR]

The script outputs a JSON summary to stdout and optionally saves the full HTML and/or markdown to files.

Key technical details

  1. User-Agent spoofing: The script sets MicroMessenger/8.0.43 in the UA string so WeChat serves the full article instead of a "please open in WeChat" block.

  2. Dynamic wait: Uses wait_for="css:#js_content" to ensure the article body has fully rendered before scraping.

  3. Lazy-image fix: WeChat uses data-src for lazy-loaded images. The script injects JS to copy data-src → src so the markdown generator can pick up real image URLs.

  4. Structured extraction: Uses JsonCssExtractionStrategy with a schema targeting WeChat's DOM structure (#activity-name for title, #js_name for author, #publish_time for date, #js_content for body).

  5. Clean markdown with images: Uses DefaultMarkdownGenerator to produce readable markdown. SVG placeholder images and data-URI artifacts are cleaned out, preserving only real article images inline with the text.

  6. UI noise stripped automatically: WeChat pages leak widgets into the markdown after the body (赞赏/留言/投诉 dialogs, 互动身份 prompts) and reading-mode prompts at the top (在小说阅读器中沉浸阅读). The script cuts everything from the "预览时标签不可点" anchor onward and filters known UI lines, so the saved markdown contains only the article.

  7. Image hotlink protection: WeChat images on mmbiz.qpic.cn block requests with non-QQ referrers. Use --download-images to download all images locally with the correct Referer header, automatically replacing remote URLs with local paths in both HTML and markdown output.

  8. Reading template: --save-markdown writes the body into a reading template — cover image, title, > 来源:作者 | 时间 | [原文](url), then AI 摘要 / 关键要点 / 我的笔记 placeholders, then the full body under ## 原文.

Extracted fields

| Field | Description | |----------------|------------------------------------| | title | Article title | | author | Public account name | | publish_time | Publication timestamp | | account_desc | Account description/bio | | markdown | Clean markdown of article body | | html | Raw HTML of article body | | url | Final URL after any redirects |

Example usage

Single article with images downloaded locally:

python <skill-dir>/scripts/crawl_wechat.py "https://mp.weixin.qq.com/s/xxx" --download-images --save-markdown --output-dir ./output

Archive one article as 对标素材:

python <skill-dir>/scripts/crawl_wechat.py "https://mp.weixin.qq.com/s/xxx" --archive-dir "AI第二大脑/02_规律研究/对标内容采集/公众号" --output-dir <临时目录>

入库模式:对标素材(--archive-dir)

把文章直接落成对标素材文件夹,供下游规律提取使用。产出与小红书 / YouTube / 视频号素材同构:

对标内容采集/公众号/
└── NNN_标题_文章ID/
    ├── 01_原始素材/
    │   ├── 01-封面图/01.png      第一张图的副本(封面单独留一份,供以后做封面规律提取)
    │   └── 02-正文图片/01.png …  按正文出现顺序重命名
    ├── 02_笔记信息.md           字段表(编号 / 文章ID / 来源 / 公众号 / 发布时间 / 原文链接 / 入库时间)+ 正文原文 + 图片清单
    └── 04_文案结构化梳理.md      中心思想 + 核心论点 + 逐段梳理 + 结构规律提炼

约定:

  • 编号取目标目录已有最大编号 +1(三位补零),不假设从 001 开始
  • 图片按正文出现顺序命名,正文里的图片引用同步改成 01_原始素材/02-正文图片/NN.ext,保留图片在原文中的位置;第一张另复制一份进 01_原始素材/01-封面图/
  • 去重:同一 原文链接 已入库时跳过,不重复建文件夹
  • 抓取失败不落盘:正文为空(被拦或页面改版)时不建文件夹,只在 stderr 说明原因
  • 标题兜底:分享型页面(common_share.html)没有标题元素,取正文首行做标题
  • 单张图片下载失败时保留远程链接,不丢这篇文章

批量入库时逐条调用,条与条之间间隔几秒,避免触发风控。入库不填 AI 摘要——素材要的是原文,摘要只在阅读模板里生成。

写结构化梳理(入库模式必做)

脚本落盘后,再出一份 04_文案结构化梳理.md,供下游提炼规律:

# 文案结构化梳理

> 来源:<编号> <标题>([[02_笔记信息.md]])
> 梳理日期:<YYYY-MM-DD>
> 说明:本文件做结构化梳理:中心思想、核心论点、逐段小标题 + 一句话总结。总结保留原文的具体论据、数据与比喻,不删减关键信息;完整原文见 02 号文件。

---

## 一、中心思想

<一段话说清这篇文章主张什么;下面附一句可直接引用的一句话说概括>

## 二、核心论点(N 条)

1. **<论点>**:<展开,保留原文的数据、例子、比喻>

## 三、逐段梳理

### <原文小标题>|<一句话总结>

- **一句话总结**:<这一段讲了什么>
- **原文示例**:<引原句,可核对>
- **可复用点**:<这条做法在哪能用>

## 四、结构规律提炼

<3-5 条,讲这篇文章的骨架怎么搭的:几段式、每段承担什么、开头与结尾各做什么>

要点:每一句都要能追回 02_笔记信息.md 的正文;不新增原文没有的数据或例子。

Fill in the summary (required step)

After the script saves the markdown, immediately read the saved file's ## 原文 section and:

  1. Replace the ## AI 摘要 placeholder with a 3-5 sentence summary of the article's core argument and conclusion.
  2. Replace the ## 关键要点 placeholder with 3-6 bullet-point key takeaways (facts, data, frameworks the reader can act on).
  3. Leave ## 我的笔记 untouched — the user fills it in after reading.

Do this in the same turn as the extraction; never hand the user a file with empty [待 AI 生成] placeholders.

For programmatic use in Python:

from crawl_wechat import crawl_wechat_article
import asyncio

article = asyncio.run(crawl_wechat_article(
    "https://mp.weixin.qq.com/s/...",
    images_dir="./output/images",  # download images locally
))
print(article["title"])
print(article["markdown"])  # images reference local paths

Limitations

  • Requires a valid, non-expired WeChat article URL — cannot search or list articles from an account
  • High-frequency crawling may trigger WeChat's anti-bot measures (CAPTCHAs, IP blocks)
  • Some temporary share links expire after a period