← 返回 Skill 列表
extension
分类: 数据与分析无需 API Key

书籍训练数据整理器

本地提取书籍、OCR 扫描件并整理成可追溯数据集

person作者: user_123d3840hubcommunity

Book Dataset Builder

Build a local, source-traceable text dataset from books and course materials. Keep every processing stage on the user's machine. Do not generate an expert Skill, publish a package, call a hosted AI API, or upload source content unless the user separately asks for it.

Workflow

  1. Confirm the user is authorized to process the source material.
  2. Create a dedicated output directory outside this installed skill.
  3. Run scripts/check_environment.py before the first extraction or whenever OCR fails.
  4. Run scripts/extract_documents.py to extract and normalize supported files.
  5. Read extraction-report.md. For scanned or low-text PDF pages, use local OCR when available and clearly report missing OCR dependencies.
  6. Spot-check representative pages against the originals when tables, formulas, vertical text, footnotes, or complex layouts matter.
  7. Run scripts/prepare_dataset.py to clean, chunk, deduplicate, and export the corpus.
  8. Run scripts/audit_dataset.py and fix errors before using the result for retrieval, continued pretraining, or later annotation.
  9. Read references/troubleshooting.md for failures and tuning, references/dataset-formats.md for downstream formats, and references/copyright-and-safety.md before distribution.

Commands

Resolve this skill directory first; commands assume it is the current directory.

python scripts/check_environment.py
python scripts/extract_documents.py <input-file-or-folder> --output <work-dir> --ocr auto
python scripts/prepare_dataset.py <work-dir> --output <dataset-dir> --target-chars 1800
python scripts/audit_dataset.py <dataset-dir>/dataset.jsonl --manifest <work-dir>/manifest.json

Supported inputs: .pdf, .docx, .odt, .epub, .html, .htm, .csv, .json, .jsonl, .txt, .md, .markdown. For audio or video courses, transcribe first and preserve timestamps in the transcript.

Example requests

  • “帮我整理电脑里的 PDF 书籍,导出成本地 JSONL。”
  • “提取这本扫描书,做中文 OCR,并保留每条数据对应的页码。”
  • “把这个文件夹里的课程讲义清洗、分块、去重。”
  • “检查这份书籍数据有没有乱码、重复和来源丢失。”
  • “数据块太碎了,帮我调大分块并重新生成。”

Local-only contract

  • Use pypdf for text-layer PDFs.
  • Use local pdftoppm and Tesseract OCR for scanned PDFs when installed.
  • Require no model API key.
  • Never send book text to OpenAI, Anthropic, DeepSeek, or another hosted service automatically.
  • Use AI semantic classification only when the user explicitly requests a later enrichment step and understands where the content will be processed.

Output

<work-dir>/
├── documents/              # normalized Markdown with page markers
├── manifest.json           # source hashes and extraction metadata
└── extraction-report.md    # errors and OCR warnings

<dataset-dir>/
├── dataset.jsonl           # traceable chunks
├── corpus.txt              # plain-text corpus
├── dataset-stats.json      # counts and duplicate statistics
└── rejected.jsonl          # empty/short/duplicate records

Each JSONL record includes a stable ID, text, original filename, page or section, character count, and source hash. Preserve these fields through later transformations.

Quality rules

  • Do not invent missing text or page numbers.
  • Preserve page boundaries before chunking.
  • Remove obvious extraction artifacts without changing meaning.
  • Keep tables and formulas in a loss-aware form; report when plain-text extraction is inadequate.
  • Deduplicate normalized identical chunks while retaining rejected-record reasons.
  • Keep raw extracted text separate from derived datasets so processing can be reproduced.
  • Call the output “training-ready text data” only after format and quality checks; it is not automatically labeled instruction-tuning data.