Book Dataset Builder
Build a local, source-traceable text dataset from books and course materials. Keep every processing stage on the user's machine. Do not generate an expert Skill, publish a package, call a hosted AI API, or upload source content unless the user separately asks for it.
Workflow
- Confirm the user is authorized to process the source material.
- Create a dedicated output directory outside this installed skill.
- Run
scripts/check_environment.pybefore the first extraction or whenever OCR fails. - Run
scripts/extract_documents.pyto extract and normalize supported files. - Read
extraction-report.md. For scanned or low-text PDF pages, use local OCR when available and clearly report missing OCR dependencies. - Spot-check representative pages against the originals when tables, formulas, vertical text, footnotes, or complex layouts matter.
- Run
scripts/prepare_dataset.pyto clean, chunk, deduplicate, and export the corpus. - Run
scripts/audit_dataset.pyand fix errors before using the result for retrieval, continued pretraining, or later annotation. - Read
references/troubleshooting.mdfor failures and tuning,references/dataset-formats.mdfor downstream formats, andreferences/copyright-and-safety.mdbefore distribution.
Commands
Resolve this skill directory first; commands assume it is the current directory.
python scripts/check_environment.py
python scripts/extract_documents.py <input-file-or-folder> --output <work-dir> --ocr auto
python scripts/prepare_dataset.py <work-dir> --output <dataset-dir> --target-chars 1800
python scripts/audit_dataset.py <dataset-dir>/dataset.jsonl --manifest <work-dir>/manifest.json
Supported inputs: .pdf, .docx, .odt, .epub, .html, .htm, .csv, .json, .jsonl, .txt, .md, .markdown. For audio or video courses, transcribe first and preserve timestamps in the transcript.
Example requests
- “帮我整理电脑里的 PDF 书籍,导出成本地 JSONL。”
- “提取这本扫描书,做中文 OCR,并保留每条数据对应的页码。”
- “把这个文件夹里的课程讲义清洗、分块、去重。”
- “检查这份书籍数据有没有乱码、重复和来源丢失。”
- “数据块太碎了,帮我调大分块并重新生成。”
Local-only contract
- Use
pypdffor text-layer PDFs. - Use local
pdftoppmand Tesseract OCR for scanned PDFs when installed. - Require no model API key.
- Never send book text to OpenAI, Anthropic, DeepSeek, or another hosted service automatically.
- Use AI semantic classification only when the user explicitly requests a later enrichment step and understands where the content will be processed.
Output
<work-dir>/
├── documents/ # normalized Markdown with page markers
├── manifest.json # source hashes and extraction metadata
└── extraction-report.md # errors and OCR warnings
<dataset-dir>/
├── dataset.jsonl # traceable chunks
├── corpus.txt # plain-text corpus
├── dataset-stats.json # counts and duplicate statistics
└── rejected.jsonl # empty/short/duplicate records
Each JSONL record includes a stable ID, text, original filename, page or section, character count, and source hash. Preserve these fields through later transformations.
Quality rules
- Do not invent missing text or page numbers.
- Preserve page boundaries before chunking.
- Remove obvious extraction artifacts without changing meaning.
- Keep tables and formulas in a loss-aware form; report when plain-text extraction is inadequate.
- Deduplicate normalized identical chunks while retaining rejected-record reasons.
- Keep raw extracted text separate from derived datasets so processing can be reproduced.
- Call the output “training-ready text data” only after format and quality checks; it is not automatically labeled instruction-tuning data.
Scan to join WeChat group