← Back to skills
extension
Category: Data & AnalyticsNo API key required

书籍训练数据整理器

Extract and prepare local training or retrieval datasets from books, PDFs, scanned PDFs, DOCX, ODT, EPUB, HTML, CSV, JSON, JSONL, TXT, Markdown, and course transcripts. Use for requests such as "整理电脑里的 PDF 书籍", "提取扫描书并保留页码", "把讲义清洗分块去重", or exporting traceable JSONL/TXT datasets without generating another Agent Skill or uploading source material.

personAuthor: user_123d3840hubcommunity

Book Dataset Builder

Build a local, source-traceable text dataset from books and course materials. Keep every processing stage on the user's machine. Do not generate an expert Skill, publish a package, call a hosted AI API, or upload source content unless the user separately asks for it.

Workflow

  1. Confirm the user is authorized to process the source material.
  2. Create a dedicated output directory outside this installed skill.
  3. Run scripts/check_environment.py before the first extraction or whenever OCR fails.
  4. Run scripts/extract_documents.py to extract and normalize supported files.
  5. Read extraction-report.md. For scanned or low-text PDF pages, use local OCR when available and clearly report missing OCR dependencies.
  6. Spot-check representative pages against the originals when tables, formulas, vertical text, footnotes, or complex layouts matter.
  7. Run scripts/prepare_dataset.py to clean, chunk, deduplicate, and export the corpus.
  8. Run scripts/audit_dataset.py and fix errors before using the result for retrieval, continued pretraining, or later annotation.
  9. Read references/troubleshooting.md for failures and tuning, references/dataset-formats.md for downstream formats, and references/copyright-and-safety.md before distribution.

Commands

Resolve this skill directory first; commands assume it is the current directory.

python scripts/check_environment.py
python scripts/extract_documents.py <input-file-or-folder> --output <work-dir> --ocr auto
python scripts/prepare_dataset.py <work-dir> --output <dataset-dir> --target-chars 1800
python scripts/audit_dataset.py <dataset-dir>/dataset.jsonl --manifest <work-dir>/manifest.json

Supported inputs: .pdf, .docx, .odt, .epub, .html, .htm, .csv, .json, .jsonl, .txt, .md, .markdown. For audio or video courses, transcribe first and preserve timestamps in the transcript.

Example requests

  • “帮我整理电脑里的 PDF 书籍,导出成本地 JSONL。”
  • “提取这本扫描书,做中文 OCR,并保留每条数据对应的页码。”
  • “把这个文件夹里的课程讲义清洗、分块、去重。”
  • “检查这份书籍数据有没有乱码、重复和来源丢失。”
  • “数据块太碎了,帮我调大分块并重新生成。”

Local-only contract

  • Use pypdf for text-layer PDFs.
  • Use local pdftoppm and Tesseract OCR for scanned PDFs when installed.
  • Require no model API key.
  • Never send book text to OpenAI, Anthropic, DeepSeek, or another hosted service automatically.
  • Use AI semantic classification only when the user explicitly requests a later enrichment step and understands where the content will be processed.

Output

<work-dir>/
├── documents/              # normalized Markdown with page markers
├── manifest.json           # source hashes and extraction metadata
└── extraction-report.md    # errors and OCR warnings

<dataset-dir>/
├── dataset.jsonl           # traceable chunks
├── corpus.txt              # plain-text corpus
├── dataset-stats.json      # counts and duplicate statistics
└── rejected.jsonl          # empty/short/duplicate records

Each JSONL record includes a stable ID, text, original filename, page or section, character count, and source hash. Preserve these fields through later transformations.

Quality rules

  • Do not invent missing text or page numbers.
  • Preserve page boundaries before chunking.
  • Remove obvious extraction artifacts without changing meaning.
  • Keep tables and formulas in a loss-aware form; report when plain-text extraction is inadequate.
  • Deduplicate normalized identical chunks while retaining rejected-record reasons.
  • Keep raw extracted text separate from derived datasets so processing can be reproduced.
  • Call the output “training-ready text data” only after format and quality checks; it is not automatically labeled instruction-tuning data.