← 返回 Skill 列表
extension
分类: 内容与媒体API Key 暂未确认

md-to-video

Convert a Markdown document with images into a narrated MP4 video with subtitles, and export a companion document (Feishu or PDF)

person作者: keen2026hubgithub

MD-to-Video Skill

Convert a Markdown file (with embedded images) into a narrated MP4 video with AI-extracted image content, TTS voiceover, subtitles, and a companion document.

When to use

User says any of:

  • "把这个 MD 文档做成视频"
  • "将 Markdown 转成带配音的视频"
  • "make a video from this markdown"
  • "把这篇文章生成视频并导出文档"

Prerequisites

  • Node.js >= 18, npm
  • Python 3 + edge-tts (pip3 install edge-tts)
  • A Remotion project initialized (npx create-video@latest --yes --blank)
  • Optional: Feishu/Lark MCP connector for document export

Installation

# Install this skill into your Remotion project
npx skills add <your-github-username>/md-to-video

Full workflow (7 steps)

Step 1: Parse the Markdown

Read the Markdown file. Split content into "scenes":

  • Text scenes: Each paragraph or heading becomes a scene. Long paragraphs are split by sentence (Chinese: 。?!;) into chunks of ≤150 chars.
  • Image scenes: Each ![alt](path) becomes a scene. The image file is copied to public/images/<CompositionId>/. The alt text is used as a placeholder narration — it will be replaced in Step 2.

Parse rules:

  • Skip pure URL lines
  • Strip Markdown formatting (**bold**, [link](url), `code`)
  • Headings (#, ##) are treated as text content, not structural separators
  • Images referenced via local paths only (skip http:// URLs)

Save the parsed scenes to src/scenes/article-scenes.json:

[
  {"id": "scene-01", "title": "short title", "text": "narration text", "image": "images/MyComp/xxx.png"},
  {"id": "scene-02", "title": "short title", "text": "narration text"}
]

Step 2: AI vision extraction (CRITICAL)

The alt text of images is NOT enough. Images contain charts, diagrams, bullet points, and knowledge content that must be extracted by AI vision.

For each image scene:

  1. Read the image file using the Read tool (it supports image input)
  2. Have the AI analyze the image content and write a 2-4 sentence Chinese narration (≤150 chars) that:
    • Covers the core knowledge points visible in the image
    • Is written in a colloquial, voice-ready style
    • Does NOT start with "这张图片展示了" — just state the knowledge directly
  3. Update the scene's text field in article-scenes.json with the AI-extracted narration

Parallelization: When there are many images (>10), launch multiple agents in parallel. Each agent handles a batch of ~13 images. Use 3 agents max. Each agent writes its results to a temp JSON file (narration-batch-N.json), then a merge script combines them.

If the Read tool cannot process images (model lacks multimodal), fall back to:

  • macOS Vision framework via Swift: swift -e 'import Vision ...'
  • Tesseract OCR: tesseract image.png output -l chi_sim+eng

Step 3: Generate TTS voiceover

Use Edge TTS (free, high quality Chinese voices):

edge-tts --voice "zh-CN-YunxiNeural" --rate "+0%" --pitch "+0Hz" \
  --file "input.txt" --write-media "output.mp3"
  • Default voice: zh-CN-YunxiNeural (male, natural)
  • Alternative voices: zh-CN-XiaoxiaoNeural (female), zh-CN-YunyangNeural (male, steady)
  • Rate: +10% faster, -10% slower
  • Generate one MP3 per scene, save to public/voiceover/<CompositionId>/scene-01.mp3
  • Retry up to 3 times on failure (Edge TTS has occasional network hiccups)

IMPORTANT: Do NOT re-parse the Markdown after AI narrations are merged — this would overwrite the AI-extracted text with alt text again. Use a separate regenerate-voiceover.ts script that only reads article-scenes.json and generates audio.

Step 4: Build Remotion components

Scene component (src/scenes/Scene.tsx)

Each scene renders differently based on whether it has an image:

Text scene: gradient background + icon + title with Highlight + body text Image scene: gradient background + title with Highlight + image (with Ken Burns effect: slow zoom/pan)

Both scene types share:

  • useCurrentFrame() + interpolate() for all animations
  • Easing.bezier(0.16, 1, 0.3, 1) for smooth entrance
  • @remotion/rough-notation <Highlight> on the title
  • Bottom subtitle bar (see Step 5)
  • Bottom progress bar (0% → 100%)
  • <Audio src={staticFile(...)}> for voiceover

Composition (src/Composition.tsx)

  • Use calculateMetadata to compute total duration from audio file durations
  • Use <TransitionSeries> with fade() transitions (15 frames) between scenes
  • Total frames = sum(scene frames) - transition_frames × transition_count

Scene config (src/scenes/script.ts)

import scenesJson from "./article-scenes.json";

export interface Scene {
  id: string;
  title: string;
  text: string;
  bgGradient: string;  // assigned by index
  icon: string;         // assigned by index
  image?: string;
}

const GRADIENTS = [/* 8 gradient strings */];
const ICONS = ["📖", "💡", "🌐", "🔬", "🚀", "📊", "🎯", "✨"];

export const scenes: Scene[] = (scenesJson as RawScene[]).map((s, i) => ({
  ...s,
  bgGradient: GRADIENTS[i % GRADIENTS.length],
  icon: ICONS[i % ICONS.length],
}));

Step 5: Subtitles

Add a subtitle bar at the bottom of every scene:

  • Split scene.text by sentence endings (。?!;) into short phrases
  • Distribute phrases evenly across the scene's duration
  • Each phrase displays for durationInFrames / sentenceCount frames
  • No fade in/out — direct cut
  • Style: black semi-transparent background, white text, rounded corners
  • For image scenes: subtitle appears below the image
  • For text scenes: subtitle appears at the very bottom (text body is already in the middle)
const sentences = splitSentences(scene.text);
const sentenceDuration = durationInFrames / Math.max(sentences.length, 1);
const currentSentenceIndex = Math.min(
  Math.floor(frame / sentenceDuration),
  sentences.length - 1
);
const currentSentence = sentences[currentSentenceIndex] || scene.text;

Step 6: Render MP4

npx remotion render <CompositionId> out/article-video.mp4

Output: out/article-video.mp4

Step 7: Export companion document

Option A: Feishu document (if Lark MCP is connected)

  1. Read article-scenes.json for all scene data
  2. For each scene: write title (h2) + text content + image (if present)
  3. Separate scenes with horizontal rules
  4. Use Feishu MCP: lark-cli docs +create --doc-format xml
  5. Run from the public/ directory so image relative paths resolve correctly

Option B: Local PDF (fallback, no external dependencies)

  1. Generate an HTML file with all scenes (title + text + image embedded as base64)
  2. Use puppeteer to convert to PDF, or open in browser and "Print → Save as PDF"
  3. Save to out/article-knowledge.pdf

File structure

my-video/
├── src/
│   ├── Composition.tsx          # TransitionSeries + calculateMetadata
│   ├── Root.tsx
│   ├── scenes/
│   │   ├── script.ts            # loads article-scenes.json, assigns gradients/icons
│   │   ├── article-scenes.json  # auto-generated scene config
│   │   └── Scene.tsx            # scene renderer (text/image/subtitle)
│   └── ...
├── scripts/
│   ├── make-video-from-md.ts    # MD → scenes JSON + copy images + TTS (full pipeline)
│   ├── merge-narrations.ts      # merge AI vision narrations into scenes JSON
│   ├── regenerate-voiceover.ts  # re-generate TTS only (without overwriting JSON)
│   └── export-pdf.ts            # export companion document as PDF
├── public/
│   ├── voiceover/ArticleVideo/  # scene-01.mp3, scene-02.mp3, ...
│   └── images/ArticleVideo/     # copied images
└── out/
    └── article-video.mp4

Quick command reference

# Full pipeline: MD → scenes + images + TTS
npm run make-md -- "/path/to/article.md"

# Full pipeline + render MP4
npm run make-md -- "/path/to/article.md" --render

# Re-generate voiceover only (after AI narrations merged)
npx tsx scripts/regenerate-voiceover.ts --force

# Render MP4 only
npm run render

# Preview in Studio
npm run dev

# Change voice
EDGE_VOICE=zh-CN-XiaoxiaoNeural npm run make-md -- "/path/to/article.md" --force

# Export companion PDF
npx tsx scripts/export-pdf.ts

Key lessons learned

  1. Never re-parse MD after AI narrations are merged — make-video-from-md.ts overwrites article-scenes.json with alt text. Use regenerate-voiceover.ts for audio-only regeneration.

  2. Image content >> alt text — Always run AI vision extraction on images. The alt text is just a filename hint, not knowledge content.

  3. Edge TTS retry logic — Network failures happen. Always retry 3× with 1s delay.

  4. calculateMetadata with TransitionSeries — Total frames = sum(scene frames) - transition_frames × transition_count. Forgetting to subtract transition overlap causes audio sync issues.

  5. Subtitles: direct cut, no fade — Split text by sentence, display each for equal duration, no animation.

  6. Feishu document: Run lark-cli from public/ directory so image paths resolve. Use XML format for document creation.