Academic Style Toolkit (paper-rewriter)
Bilingual (CN/EN) style naturalization for academic & medical writing: find stiff, templated or machine-flavored patterns, clean mechanical debris, revise for clarity and natural register — with integrity guardrails on every step.
Routing in one line: grammar/fluency/journal-format polish → use a polisher (paper-polisher-pro); de-templating, model-residue cleanup, AI-flavor style self-check and rewriting → this tool. Need both? Polish first, then run this pipeline.
Feature status (v3.1.0)
| Feature | Status | Since |
|---|---|---|
| detect / transform / verify / compare / pipeline | Stable | v1.0 |
| Section-aware scanning (--structure) | Stable | v1.7 |
| .docx direct input; 50MB file guard; read retry | Stable | v1.3–1.8 |
| Sentence-level rewrite plan (plan.py) | Stable | v2.2 |
| User-learned term guards (learn_guards.py) | Stable | v2.2 |
| --exit-verdict machine exit codes (detect/compare/pipeline) | Stable | v2.3 |
| Deterministic deep polish (--deep), handling hints, scan time budget | Stable | v2.4 |
| Centralized error layer (CliError + handling hints), batch retry, per-step timeout | Stable | v2.5 |
| Scale envelope: 800K-char chunks (2K overlap) · 50MB file cap · 120s step budget | — | — |
| Style profile (stylecheck.py: per-paragraph quantified guide audit) | Stable | v2.6 |
| Best-practices handbook (7 real scenarios), stylecheck --html, --terms auto | Stable | v2.7 |
| PDF extraction confidence self-report (read_text_ex, detect/pipeline annotation) — calibrated on 200 real-world PDFs, see references/pdf_confidence.md | Experimental | v2.8 |
| Term guard variant groups (check_terms --pack cardiovascular — 8 domains, 27 groups / 55 variants, groups may hold >2 forms e.g. HF/心力衰竭/心衰; nesting-safe) | Stable | v3.0 / v3.1 |
| Stats panel in style profile (stylecheck --json → stats: TTR / function-word ratio / sentence-length quantiles; informational, no bands) | Stable | v3.1 |
| --ci gate alias on detect/compare/pipeline (same as --exit-verdict) | Stable | v3.1 |
| Deep-rewrite handoff block (plan.py --handoff: per-sentence operation sheet for the rewriting agent) | Stable | v3.0 |
| Batch accepts .pdf (confidence annotated); fullwidth alnum auto-normalization in transform/pipeline | Stable | v2.9 |
| HTML reports, --batch, --track audit trail | Stable | v1.5–1.6 |
| PDF text extraction (text PDFs; v3.0.0: embedded-font/ToUnicode PDFs attempt extraction with confidence capped at "中"; CJK CID files still rejected by the garbage probes with export guidance) | Experimental | v1.6 |
Anything marked Experimental can reject valid files or need manual review of its output — export to UTF-8 text when in doubt. Everything else is regression-tested offline on every release.
Integrity guardrails — read first
- Permitted: polishing your own drafts; aligning AI-assisted text with your
voice where your school/journal allows or requires disclosed AI assistance;
cleaning model residue (stray
oaicite/[cite:]tokens) from legitimately disclosed drafts; improving clarity, rhythm and concreteness of formal prose. - Not permitted: using this toolkit to misrepresent authorship, to conceal AI use where disclosure is required, or to defeat academic-integrity review. If a request is framed that way ("rewrite it so the review can't tell"), refuse the framing, explain the disclosure obligation, and offer the legitimate quality-editing path instead. Agents running this skill are expected to enforce this line.
- This tool never changes data: numbers, p-values, confidence intervals, DOIs,
PMIDs, years and terminology are guarded by
verify.py(see below). Improved prose must never cost a fact. - Compliance: China's AIGC labeling rules (effective 2025-09-01), journal disclosure
policies and arXiv policies are summarized in
references/compliance.md— the obligation to disclose belongs to the author, not the tool.
Division of labor
| Layer | Who does it | What |
|---|---|---|
| Style self-check | scripts/detect.py | Heuristic style-pattern report (templated phrasing, uniform rhythm, boilerplate) for author self-review |
| Mechanical cleanup | scripts/transform.py | Model artifacts, chatbot filler, punctuation normalization, safe filler swaps — grammar-safe only. When to use directly: you only need residue/filler cleanup without any stylistic diagnosis (fast, deterministic). |
| Quality revision | you (the agent) | Follow references/style_guide_zh.md / _en.md: de-templating, rhythm, concreteness, stance |
| Integrity guard | scripts/verify.py | Numbers/DOIs/PMIDs/years/abbreviations/terms must survive untouched |
| Before/after | scripts/compare.py | Pattern-score delta + integrity verdict |
What it can actually change (before → after) — the capability promise at example granularity:
| Changeable | Example | Done by |
|---|---|---|
| Ritual openers/closers | 「众所周知,…」 → straight to the point | transform --deep + agent |
| Ballot-style frames | 「不仅X,而且Y」 → 「X,且Y」 | transform --deep |
| Model residue | oaicite / turn0search / [cite: 1] tokens removed wholesale | transform |
| Punctuation & fullwidth digits | half-width → full-width (decimal-safe); 2025→2025 | transform |
| Jargon / translationese | 「进行…的研究」 → 「研究」 | agent (style guide) |
| Rhythm | long-sentence runs → varied lengths; repeated openers → varied | agent + stylecheck |
| Concreteness | vague appraisal → grounded statement (never invents data) | agent (guards flag invented numbers as W_NUM_ADDED warning — human removes them) |
| Term consistency | HF/心力衰竭/心衰 mixed → unified to dominant form | check_terms.py groups |
| Never changed | numbers, p-values, CIs, DOIs, PMIDs, years, terms, quoted content | verify.py blocks |
Quick start
How to ask for this in plain words (EN digest of the full phrase table in
SKILL_ZH.md): "remove the AI flavor from this paper" / "this reads too templated" /
"clean up oaicite and [cite:] residue" / "tell me which sentences to fix first" /
"check my rewrite broke no numbers" / "give me a style profile" / "stop flagging
this legitimate term" — each maps to the pipeline, plan.py, transform.py,
verify.py+compare.py, stylecheck.py and learn_guards.py respectively.
Reading map (this file is the hub; details live one click away):
first time → Quick start + The workflow below · troubleshooting →
references/pitfalls.md + references/errors.md · CI/agent integration →
Automation section · capability limits → Honest boundaries · scenario how-tos →
references/best_practices.md · deep-rewrite process →
references/deep_rewrite_guide.md · PDF confidence tiers →
references/pdf_confidence.md · programmatic use → references/api.md.
No agent is required — every capability below runs directly in a terminal.
Every script has built-in --help (full flag reference) and --version;
quick flag semantics:
| Flag | Meaning |
|---|---|
| --profile general | non-academic text: down-weight formulaic/boilerplate signals |
| --deep | deterministic sentence-level transforms (canned openers, not-only merge) |
| --exit-verdict | machine exit codes for CI/agent gating (see below) |
| --structure | section-aware scoring (IMRaD) |
| --step-timeout SEC | per-step wall-clock budget (default 120) |
| --terms FILE | term list for the integrity guard (build via extract_terms.py) |
One command (self-check → cleanup → revision brief → integrity guard):
python scripts/pipeline.py draft.txt -o out.txt --terms terms.txt
Visualization & batch (v1.5.0): add --html report.html to any of
detect/pipeline/compare for a self-contained HTML report (score cards, category
tables, revision worksheet, integrity verdict; compare adds sentence-level
add/delete diff). Scan a whole directory with --batch dir (detect: per-file
score summary; pipeline: per-file cleanup+integrity CSV).
Structure-aware scanning (detect --structure, v1.7.0): detects paper
sections (abstract/introduction/methods/results/discussion/conclusion/
references, bilingual) and scores each separately. Methods/results weights
normalize the aggregate only — raw per-section scores are always reported in
full; the fixed phrasing of a Methods section is genre convention, not a
machine signal. Missing-section hints included. Markdown headings are
recognized since v2.2.0 (## 摘要, **方法**, numbered 2. Methods).
Sentence-level rewrite plan (plan.py, v2.2.0): ranks sentences by their
weighted pattern contribution and produces a P0 rewrite queue with section
attribution, per-sentence category hits, handling advice and a linear budget
projection (fix the top-K offenders → projected score). --json for agent
consumption. The projection is a local approximation for prioritization — not a
promise, and not any external detector's score.
Preparing the term list: python scripts/extract_terms.py draft.txt -o terms.txt
auto-extracts candidates (abbreviations, quoted terms) into a draft you confirm
by hand. Input formats: .txt/.md directly, .docx (Word) directly since
v1.3.0; PDF direct-read is experimental (text PDFs; confidence self-report —
a "低" tier means garbled extraction, export to UTF-8 text instead; see
references/pdf_confidence.md). Ask for a
revision worksheet alongside any run with --suggestions path.md (detect.py and
pipeline.py both support it): category, sample, suggested handling and the guide
section to read — the tool proposes, you and the guide decide.
For non-academic text (blogs, posts, office documents), add --profile general
to down-weight formal-boilerplate signals; the default academic profile is
calibrated for scholarly manuscripts.
Or run the steps individually (from the skill's own directory, or use absolute paths):
python scripts/detect.py draft.txt # style-pattern self-check (-j JSON, -s score only)
python scripts/transform.py draft.txt -o step1.txt
# ... agent quality revision of step1.txt per the style guide ...
python scripts/verify.py draft.txt step2.txt --terms terms.txt
python scripts/compare.py draft.txt step2.txt
Scenario-organized best practices (journal submission, thesis, revision
letters, batch, CI gates, false-positive handling, CN punctuation):
references/best_practices.md.
Worked end-to-end examples: references/examples.md. Unified exit codes,
violation categories and remedies: references/errors.md. Centralized
common-mistakes list (24 items): references/pitfalls.md — read it before
your first real run. Python API reference: references/api.md. FAQ:
references/faq.md.
Revision tracking (--track base on transform/pipeline): every change the
mechanical layer makes is recorded to base.md + base.json — an auditable
list of what the tool touched (rule, count, deleted sentences,
flagged-for-review).
PDF input (experimental, v1.6.0; confidence self-report since v2.8.0): English text PDFs (FlateDecode/WinAnsi, unencrypted) are extracted with the standard library only. Since v2.8.0 every PDF scan reports an extraction confidence (高/中/低) computed from four heuristics — common-word hit rate, control-character rate, average word length, extractable word count — so "review the output" becomes "the tool tells you how trustworthy this extraction is". 低 means garbled-form typical: export to UTF-8 text instead. PDFs with embedded font encodings (ToUnicode CMaps, typical for CJK) are rejected when detected — detection is best-effort: if one slips through, the output may be garbled, so review extracted PDF text before relying on it. When in doubt, export to UTF-8 text. Chunk threshold: 800,000 characters per chunk — since v2.2.0 adjacent chunks share a 2,000-character overlap window, so patterns spanning a chunk boundary are captured and double-counts are reconciled (v2.1.x and earlier could miss boundary-spanning signals). File-size guard 50MB (scans auto-chunk at any size within the cap; the integrity guard is whole-document).
Agent invocation protocol
When invoked, decide the path first, then run it:
- Trigger words: 写作风格自查 / 论文改写润色 / 去模板腔 / 翻译腔清理 / 学术改写 / 段落改写 / 表达优化 / style self-check / style naturalization / naturalize academic writing / de-templating / academic rewriting / passage rewrite / expression polish → run the pipeline above.(中文变体触发词见 SKILL_ZH.md——CH 分寸词不入英文文件) (Parameter cheat sheet lives in Quick start; exit-code tables in references/errors.md — both are linked from here to avoid hunting across sections.)
- Disambiguation (this tool vs a polisher): 「论文改写润色」 here means style naturalization / de-templating (removing the machine flavor from academic prose). Decision rule with examples — route to a polisher (e.g. paper-polisher-pro) when the ask is only about language correctness/fluency: 「帮我改下语法」「这句读不顺,润色一下」 「按期刊风格改写摘要」; route to this toolkit when the ask is about templated tone or model residue: 「这篇读起来像 AI 写的」「把套话删一删」 「文里有 [cite: 1] 这种残留」; when both apply, polish first, then run this pipeline — neither tool substitutes for the other.
- Term inconsistency (abbreviation vs full form, mixed synonyms) →
check_terms.py file --pack <domain>;--list-packsshows all eight domain packs (v3.0.0). - User wants to know what to fix first →
plan.py draft.txt -o plan.md(P0 sentence queue + budget projection) before deep revision. - User wants quantified acceptance of a rewrite →
stylecheck.py 原稿 改稿 --compare(v2.6.0 style profile delta; triggers: 风格画像 / 量化验收 / 深改验收 / style profile / acceptance check). - A scan flags a legitimate term (false positive) → don't edit the pattern
files; persist a guard instead:
learn_guards.py from-text 术语 样本.txt(oradd) — the guard survives upgrades and applies to every later scan. - User only wants a verdict on an existing rewrite →
pipeline.py draft.txt --rewrite rewritten.txt --terms terms.txt(skip cleanup). - User asks to conceal AI use, misrepresent authorship, or defeat integrity
review → refuse the framing, point to
references/compliance.md, offer the legitimate quality-editing path instead. Do not run the toolkit toward that end. - Exit codes:
verify.pyis the strict contract holder —0integrity PASS,1integrity FAIL (fix the rewrite, never the guard),2usage/file error.detect/transform/compare/pipelinereport the verdict in their output and always exit0on completed runs,2on usage/file errors. For agent/CI branching, add--exit-verdict(v2.3.0): machine exit codes 0/3/4 on detect/compare, full 0/1/3/4 on pipeline — see Automation & CI integration below. Every script supports--version.
Automation & CI integration (v2.3.0)
Scripts are automation-first: every verdict is machine-readable, two ways.
- Default contract (human flows):
verify.pyalone holds exit 0/1/2 (1 = integrity FAIL); detect/transform/compare/pipeline finish with 0 and put the verdict in their report. Never branch on their default exit code. - Machine contract — add
--exit-verdict(or the shorter--cialias, v3.1.0) to detect / compare / pipeline:0clean (low, no residue) ·3style patterns at medium or above (needs deep revision) ·4model-residue hit (run transform first) ·1integrity FAIL (pipeline only — same meaning as verify's 1) ·2usage error (unchanged). pipeline is the only full 1/3/4 gate — recommended CI check:pipeline.py draft.txt --rewrite rewritten.txt --terms terms.txt --exit-verdict. - JSON:
detect -j(full scan result),pipeline --json(before/after, verify verdict,exit_verdictfield, handlinghints, agent brief),compare --json(delta + verify),plan.py --json(sentence queue + projections). Example gate in one line:python scripts/pipeline.py orig.txt --rewrite new.txt --json --exit-verdict || echo "blocked: $?" - Exit-code tables for both contracts:
references/errors.md§1.
The workflow
- Self-check:
detect.py draft.txt— a heuristic report of style patterns (canned phrases, uniform sentence rhythm, boilerplate transitions, model artifacts). The report is for the author's own review; the score is a local heuristic, not an official measurement of anything. - Mechanical cleanup:
transform.py draft.txt -o step1.txt— strips stray model artifacts (oaicite,turn0search,[cite: 1],grok_card,attached_file), leftover chatbot pleasantries, markdown residue, and a safe list of filler phrases; CN halfwidth punctuation is normalized to fullwidth (decimals protected).-aadds em-dash reduction and empty-opener removal.--deep(v2.4.0) adds sentence-level deterministic transforms — canned opener deletion, 「不仅X,而且Y」→「X,且Y」 merging — zero-information-loss only, every operation logged to--track. Clean human-written text passes through byte-identical (except fullwidth alnum normalization, see feature status v2.9). - Quality revision (the real work): optionally rank the work first —
python scripts/plan.py draft.txt -o plan.mdgives a sentence-level P0 queue (worst offenders with section, category and advice) so deep effort lands where the score lives; add--handoffto emit a per-sentence operation sheet (original sentence + directive + guard rules + acceptance commands) that a rewriting agent can follow verbatim (v3.0.0). Then read the guide for the text's language —- CN:
references/style_guide_zh.md— structural de-templating → jargon cleanup → rhythm → concreteness → stance → integrity red lines - EN:
references/style_guide_en.md— smaller words → fewer significance frames → plain clauses instead of parallelism → rhythm → real attribution → commit to a position - Revise section by section. Preserve all facts, numbers, citations, terminology.
3b. Quantified acceptance (
stylecheck.py, v2.6.0): per-paragraph style profile — jargon hits located to the paragraph, sentence-rhythm CV, opener diversity, each banded 自然/观察/偏机器 with style_guide anchors;--compare 原稿 改稿shows the delta so deep-revision acceptance is metric-based. Bands are calibration hints only — they never alter scores.
- CN:
- Guard:
verify.py draft.txt step2.txt --terms terms.txt— exit 1 means a number/citation/term was altered: fix the revision, not the guard. Build terms.txt for medical text (drug names, gene symbols — include mouse-style capitalized forms like Myc — and scale names), one term per line. The guard also warns on numbers absent from the original (fabrication defense) and number-context swaps (arm/direction ordering). For mixed synonym usage (abbreviation vs full form, e.g. HF vs 心力衰竭), also runcheck_terms.py file --pack cardiovascular(v3.0.0 domain packs;--list-packsshows all eight). - Re-check:
compare.py draft.txt step2.txt— pattern reduction + integrity verdict. Keep the before/after pair for your records; if your institution or journal requires an AI-use disclosure, state it plainly — this report is a quality self-check, not a substitute for disclosure.
Honest boundaries
- The style score is a local heuristic on writing patterns. It is not a measurement produced by any external service, and it says nothing about authorship. Never present it as one.
- Formal academic prose and non-native writing are routinely misjudged by automated reviewers; if you are the author, keep drafts, version history and notes — process evidence, not style scores, settles authorship questions.
- This tool does not interact with any external review system, does not remove
official content labels or watermarks, and does not assist concealment of
required disclosures. See
references/compliance.md. - Detection-capability boundary (measured, 2026-10): on fluent LLM-generated text this toolkit's style channel barely fires — 128 real ChatGPT texts scored below 28 in 98% of cases, and 60 samples of 2026-era models scored 0 (nothing flagged). Style tools that do flag everything in that benchmark misfire on ~50% of human forum prose. A low score here does not mean "no AI was involved" — use this toolkit as a style self-check + integrity guard, not as a judge of authorship.
- Statistical-channel experiment (2026-10, 373 files measured): sentence- length band concentration is the strongest single statistical signal (AI corpora 0.68/0.72 vs human 0.48/0.47). Integrated conservatively as a "uniform rhythm flag" (≥8 sentences; band≥0.65 or cv≤0.30): 41% of 2023-era AI corpus flagged @ 3% human flag rate (below the FP baseline). The flag is informational only — regular genres (e.g. literature abstracts) trigger it too; it is never a detection verdict.
- Scale envelope: regression-tested up to ~1MB text files; the scanner is linear (no catastrophic backtracking, bounded quantifiers only). Memory use is roughly 3x file size; for very long manuscripts, split by section for readable reports.
Permissions & environment statement
For reviewers, security scanners and cautious users:
-
Reads only the file paths you pass as arguments (plus stdin, including the
.txt/.md/.docx/.pdffiles inside a--batchdirectory) and its own bundled wordlist files (scripts/patterns_*.json, plusscripts/user_guards.jsonif you have created one withlearn_guards.py). Verify it yourself:grep -rnE "urllib|requests|socket|subprocess|os\.environ" scripts/ || echo "clean"(runs clean as of this release — the claim is reproducible, not rhetorical).
-
Writes only to output paths you pass explicitly:
-o/--outputon transform/compare/pipeline/plan/extract_terms,--suggestions,--html,--review(detect.py's markdown report),--report(check_terms.py), and--track(base.md+base.json). One derived exception: withpipeline --terms auto, the tool also writes the auto-generated term draft<output-basename>.terms.auto.txt(tool-written, unreviewed — see the workflow note). One tool-owned data file on top of that:learn_guards.py add/from-text/removewritesscripts/user_guards.json(your learned term guards, via a transient.tmp+ atomic rename). No other writes. -
Zero network access — no HTTP calls, no downloads, no API keys.
-
Zero third-party dependencies — Python standard library only.
-
Reads no environment variables; spawns no subprocesses; creates no scheduled tasks; the only temp file ever written is the
user_guards.json.tmprename target described above. -
Test corpora and dev notes live in the development repo only, not in the distributed package.
Customizing
scripts/patterns_zh.json/patterns_en.json— pattern lists (+ rewrite suggestions), regex signals, term_guards (legitimate academic collocations that must not be flagged, e.g. "mutational landscape", "pivotal trial", CJK "sequence alignment" and "precipitation reaction"), auto_fixes.- User-learned guards live in
scripts/user_guards.json(managed bylearn_guards.py— do not hand-edit;listprints,removedeletes). Keep them there rather than editingpatterns_*.json: pattern files are replaced on upgrade, your guards file is not. - Score calibration constants live in
scripts/hxt_core.py(_LANG_K); the four test corpora in the development repo'stests/(not shipped in the package) document the intended separation.
Related tools
- pubmed-verifier — verify PMID/DOI references before submission
- cite-holmes — deep research with hallucination-free citations
- paper-polisher-pro — comprehensive polishing & plagiarism reduction
- academic-figures — publication-ready scientific figures
- doc-holmes — layout-preserving PDF translation
Scan to join WeChat group