Video Downloader
Overview
Use this skill to turn a video-platform URL into a local source-material folder containing the video file, post_caption.txt, audio.wav / audio.m4a / audio.mp3, transcript.txt, transcript.srt (when timing is available), and metadata.json.
Terminology:
post_caption.txt: platform publish text, title, description, and hashtags.transcript.txt: spoken in-video text from a verified page transcript, otherwise audio ASR.metadata.json: normalized platform metadata plus download and ASR status.
Output Folder Naming
Each download creates a folder with a human-readable name:
YY_MM_DD_标题摘要_平台_作者
Example: 26_01_01_示例视频_抖音_示例作者
Rules:
- Date — video publish date preferred, falls back to download date. Format:
YY_MM_DD. - 标题摘要 — title trimmed to ≤40 characters; newlines, hashtags (
#), mentions (@), emoji removed. - 平台 — Chinese display name:
抖音/B站/YouTube/小红书. - 作者 — falls back to
未知作者when unavailable. - Dedup — if the folder already exists, appends the last 5 characters of the video ID.
- Internal file names (
post_caption.txt,metadata.json, etc.) are not affected.
The current implemented providers are Douyin, Bilibili, YouTube, and Xiaohongshu. WeChat Channels is an explicit extension point; do not claim that provider works until its provider module has been implemented and tested. For WeChat Channels links, tell the user to first use the WeChat mini program kg百宝箱 to download the video from the 视频号 link, then continue with local ASR or downstream processing on the downloaded file.
Quick Start
For Douyin, first use the host Agent's supported browser and page-fetch tools as described in references/douyin-browser.md. Guest access is sufficient when the video plays; reuse an existing session if present. Ask the user to handle sign-in only when the page actually requires it. Never extract browser credentials to implement this route.
Pass the browser observation to the bundled CLI:
python3 scripts/download_video.py "https://www.douyin.com/video/1234567890123456789" \
--browser-capture ./browser-capture.json --output-dir ./downloads
If the page exposes a complete speech transcript or subtitle track, fetch it with the available page-fetch tool (such as web.fetch, when provided), normalize it as documented, and add --page-transcript ./page-transcript.json. A publish caption, chapter summary, comments or search snippet is not a transcript. If no valid page transcript is available, use the existing ASR route.
The CLI consumes these files; it does not launch or control a browser itself. If the host has no supported browser tool, explain that limitation and run the original public fallback routes:
python3 scripts/download_video.py "https://v.douyin.com/..." --output-dir ./downloads
For metadata/post-caption extraction without downloading the video:
python3 scripts/download_video.py "https://v.douyin.com/..." --output-dir ./downloads --metadata-only
For download without ASR:
python3 scripts/download_video.py "https://v.douyin.com/..." --output-dir ./downloads --asr none
For Chinese speech with local whisper.cpp (recommended):
python3 scripts/download_video.py "https://v.douyin.com/..." --output-dir ./downloads --asr whisper_cpp --asr-language Chinese
For Chinese speech with cloud ASR:
SILICONFLOW_API_KEY="..." python3 scripts/download_video.py "https://v.douyin.com/..." --output-dir ./downloads --asr siliconflow --asr-language Chinese
To bias ASR toward domain terms or a specific script:
python3 scripts/download_video.py "https://v.douyin.com/..." --output-dir ./downloads --asr-language Chinese --asr-prompt "请使用简体中文转写。关键词:视频中的项目名、术语和产品名称。"
Workflow
- Identify the platform from the URL.
- For Douyin, collect the browser observation and available page speech transcript first. Use
scripts/download_video.pyto download, validate and archive them. For other providers, call the CLI directly. - Inspect the output folder and report:
- video path, when downloaded
- post caption path
- audio path
- transcript path
- SRT path (timed page subtitles, whisper.cpp or openai-whisper)
- metadata path
- platform, item ID, author, duration, resolution, and any provider caveats
ASR Transcript
After a video is downloaded, a valid supplied Douyin page transcript is preferred and audio is extracted without ASR. Otherwise the CLI extracts audio with ffmpeg and transcribes it. Use --transcript-source asr to force speech recognition. --asr none skips both transcript import and audio extraction. Four backends are available, selected via --asr.
Backend Options
| Backend | Description |
|---------|-------------|
| auto (default) | Auto-detect: whisper.cpp > openai-whisper > SiliconFlow |
| whisper_cpp | Local whisper.cpp (whisper-cli), fastest on Apple Silicon |
| whisper | Local openai-whisper Python CLI |
| siliconflow | Cloud API via SiliconFlow SenseVoiceSmall |
| none | Skip audio extraction and transcription entirely |
--asr auto Priority
When --asr auto (the default) is used, the backend is selected in this order:
- whisper.cpp — if
WHISPER_CPP_BINpoints to an executable file, orwhisper-cliis available on PATH, andWHISPER_CPP_MODELpoints to an existing model file. - openai-whisper — if the
whisperCLI is available on PATH. - SiliconFlow — if
SILICONFLOW_API_KEYis set in the environment. - If none of the above are available, an error is returned. No software is auto-installed.
whisper.cpp (--asr whisper_cpp)
The recommended local backend for best performance on Apple Silicon.
Environment variables:
WHISPER_CPP_BIN— optional explicit path towhisper-cli; when unset, search PATHWHISPER_CPP_MODEL— required path to a local GGML model file
export WHISPER_CPP_BIN="/path/to/whisper-cli"
export WHISPER_CPP_MODEL="/path/to/ggml-large-v3-turbo.bin"
Do not assume a platform-specific default path. No software is installed automatically.
Output artifacts:
audio.wav(16 kHz PCM)transcript.txttranscript.srt(subtitles with timestamps)transcript.whisper_cpp.json(metadata)
The model is always the file path, not a model name — --asr-model is ignored for whisper_cpp.
openai-whisper (--asr whisper)
Output artifacts:
audio.m4atranscript.txttranscript.srttranscript.whisper.json
SiliconFlow (--asr siliconflow)
Requires SILICONFLOW_API_KEY. Calls https://api.siliconflow.cn/v1/audio/transcriptions.
Output artifacts:
audio.mp3transcript.txttranscript.siliconflow.json
SiliconFlow does not currently provide an SRT file; srt_path is returned as null.
Result Contract
Every ASR result includes status, backend, model, language, audio_path,
transcript_path, srt_path, raw_json_path, and error. A selected backend
that fails returns status: failed; auto does not silently retry another
backend after transcription has started.
Common Options
--asr-language— Language for transcription, e.g.Chinese,zh,English,auto. Default:auto.--asr-model— Model name. Default:auto(whisper.cpp ignores this; openai-whisper defaults tobase; SiliconFlow maps toFunAudioLLM/SenseVoiceSmall).--asr-prompt— Initial prompt passed to the ASR engine. Use it to request Simplified Chinese output and provide domain terms.--asr-max-seconds— Debug limit: transcribe only the first N seconds of audio.
Douyin Provider
The Agent first opens the requested video in its supported browser, observes the actual media source, and supplies a capture file. The CLI routes in this order:
- Browser-observed public media URL →
curl→ media validation; merge separate video/audio streams with FFmpeg when necessary. - H5 share-page metadata and public media endpoint.
- Anonymous
yt-dlp, with external yt-dlp configuration disabled. - Browser Cookie fallback only when the user explicitly authorizes it and
--douyin-cookies-from-browser <browser>is passed. It is off by default and may cause an OS keychain prompt. Do not add this flag just because another route failed.
Stop after the first successful, validated download. A stale media URL may be refreshed once in the browser; do not repeatedly retry login challenges. Missing browser capability is recorded as not_supplied, not as a successful browser attempt. --no-yt-dlp-fallback keeps browser and H5 routes enabled.
Use references/douyin-browser.md for capture and transcript schemas, source verification, expiry handling and profile/collection boundaries. Browser captures and page transcripts are runtime inputs outside the installed Skill, not distributable resources.
Browser-sourced media is saved as delivered by the platform; do not promise watermark removal or a specific resolution. Every successful Douyin download is checked with ffprobe. metadata.json records the successful method, Cookie use, validation and fallback history. Page transcript provenance is recorded separately from ASR.
Bilibili Provider
The Bilibili provider uses yt-dlp as the primary route for both metadata extraction and media download.
Important behavior:
- Support
bilibili.comandb23.tvlinks. - Store the Bilibili title plus description in
post_caption.txt. - Store
yt-dlpraw metadata and normalized fields inmetadata.json. - Download with
bv*+ba/band merge to mp4 when possible. - Retry with Chrome cookies when anonymous metadata extraction or download fails.
- Use
--metadata-onlyfor fast tests or when the user only needs title/description metadata.
YouTube Provider
The YouTube provider uses yt-dlp as the primary route for both metadata extraction and media download.
Important behavior:
- Support
youtube.comandyoutu.belinks. - Store the YouTube title plus description in
post_caption.txt. - Store
yt-dlpraw metadata and normalized fields inmetadata.json. - Download with
bv*+ba/band merge to mp4 when possible. - Use local Node as the yt-dlp JavaScript runtime when available.
- Retry with
--remote-components ejs:githubor Chrome cookies when the basic route fails. - Use
--metadata-onlyfor fast tests or when the user only needs title/description metadata.
Xiaohongshu Provider
The Xiaohongshu provider uses a three-tier download fallback:
- Anonymous yt-dlp — public download without any authentication (no cookies).
- Public direct URL — if metadata is available but anonymous yt-dlp can't download, extracts the highest-quality video stream URL from yt-dlp's formats metadata and downloads it via plain HTTP.
- Chrome cookies yt-dlp — last resort; requires local Chrome browser access.
Metadata extraction is always anonymous — cookies are never triggered just to get author info or metadata.
Important behavior:
- Support
xiaohongshu.com,xhslink.com, andxhslink.cnlinks. - Store the Xiaohongshu title plus note body in
post_caption.txt. - Store
yt-dlpraw metadata and normalized fields inmetadata.json. - Author nickname is taken from anonymous metadata; falls back to
未知作者when unavailable. - Remote assistant: if both public routes fail, reports "需要在电脑端授权 Chrome Cookie" immediately instead of hanging.
- Safety: no
cookies.txtis saved; cookie contents are never logged or included in output. metadata.jsonrecordsdownload.method,download.cookie_used,download.direct_url_source, anddownload.fallback_errorsfor full traceability.- Use
--metadata-onlyfor fast tests or when the user only needs title/note metadata.
WeChat Channels Provider
WeChat Channels is recognized but not implemented yet.
Current finding:
yt-dlpdoes not support the testedweixin.qq.com/sph/...share link.- The public web page can expose text, cover image, and QR-code flow, but did not expose a playable video URL in the tested case.
- Do not claim that this skill can directly download WeChat Channels videos until a provider has been implemented and tested.
Temporary workflow:
- Ask the user to open WeChat and search for the mini program
kg百宝箱. - In
kg百宝箱, paste the 视频号 video link and download the video there. - After the user has the downloaded video file locally, this skill can still be used for local ASR/transcript work if the file is passed through the ASR helper or a future local-file entrypoint.
Provider Extension Contract
Add new platforms by creating a module under scripts/providers/ and registering it in scripts/providers/__init__.py.
Each provider should expose:
PLATFORM: stable provider namesupports(url: str) -> boolfetch(url: str, output_root: Path, *, metadata_only: bool = False, **options) -> dict
Output folders should include the same artifact contract whenever possible:
metadata.jsonpost_caption.txtaudio.wav/audio.m4a/audio.mp3andtranscript.txtwhen ASR is enabledtranscript.srtwhen timed page subtitles or a timestamped ASR backend are used- video file when download is enabled
Reserved providers:
wechat_channels
When a reserved provider is detected but not implemented, say so plainly and do not fabricate a download result.
Safety
Download only material the user owns, has permission to download, or can lawfully archive for their intended use. Do not bypass DRM, paid access controls, private permissions, or platform restrictions for unauthorized redistribution.
Scan to join WeChat group