Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

md2docx

Convert Markdown (.md) to Word (.docx) under a fixed SCI manuscript format lock — and, optionally, resolve [@key] citations against a Better-BibTeX .bib into numbered Vancouver references.

Built as a Claude Code skill, but the converter is a plain Node.js script you can run standalone.

Why

In-Word reference managers (EndNote / Zotero CWYW) and Pandoc's default DOCX output both drift away from strict journal formatting: built-in Heading 1/2/3 styles, colored themes, page breaks, inconsistent fonts. This tool makes the DOCX a build artifact of your Markdown + .bib, so the format never drifts and renumbering is automatic.

The format lock

Aspect Value
Font Times New Roman, every paragraph
Size 12pt — body, headings, table cells, captions
Color Black (#000000)
Line spacing 2.0 (double), global
Paragraph style Normal only; headings differ solely by bold=true
Built-in styles Never Heading 1/2/3, Title, Subtitle, themes
Page breaks Zero (Markdown --- is dropped, not promoted)
Tables Three-line only (top + bottom of header, bottom of last row)
Header / footer / page numbers None
Margins / page 1.0" all sides, US Letter portrait

The lock is embedded in scripts/md2docx_template.js by design — it is not exposed as per-file CLI flags, to prevent drift across a manuscript / supplement / cover-letter set.

Install

npm install -g docx

Requires Node.js. The script depends only on the global docx package.

Usage

NODE_PATH="$(npm root -g)" node scripts/md2docx_template.js input.md output.docx

Both arguments are positional; if output.docx is omitted it defaults to the input name with a .docx extension. The NODE_PATH prefix lets Node find the global docx module — omitting it causes Cannot find module 'docx', the single most common failure.

With citations

NODE_PATH="$(npm root -g)" node scripts/md2docx_template.js manuscript.md manuscript.docx --bib references.bib

With --bib, the converter pre-processes the Markdown through scripts/resolve_citations.js:

  1. Numbering — each unique [@key] gets a Vancouver number in order of first appearance; re-cited keys reuse it.
  2. Range collapse — [@a; @b; @c] with contiguous numbers (4,5,6) → ⁴⁻⁶; gapped lists stay comma-separated.
  3. Body replacement — [@key] → superscript number, adjacent citations merged.
  4. References regeneration — an existing ## References / ## Bibliography block is replaced with a freshly numbered Vancouver list; if absent, one is appended.
  5. Missing-key safety — an unknown [@key] is left literal and reported to stderr, never silently dropped.

Citation token grammar (subset of pandoc-citeproc):

[@kobayashi2003]                      single
[@wang2019; @sahoo2024]               multi (semicolon or comma)
[@a2024; @b2024; @c2024]              consecutive numbers auto-collapse to a range

Supported BibTeX fields (Better-BibTeX export defaults): title, author, journal/journaltitle/shortjournal, volume, number/issue, pages, year/date, doi, pmid, eprint+eprinttype=pubmed, and PMID: inside note. Output:

Author1 et al. Title. Journal. Year;Volume(Issue):Pages. PMID: NNNNNNNN.

What the converter handles

  • Headings #…###### → Normal-style bold (no size escalation)
  • **bold**, *italic*, `code` (Courier New)
  • <sup>…</sup> / <sub>…</sub>
  • Inline LaTeX $…$ → plaintext + italic, common patterns mapped to Unicode (\lambda→λ, \times→×, \log_2→log₂)
  • Pipe tables → three-line Word tables
  • --- → silently dropped (lock forbids page breaks)
  • [@key] tokens → resolved when --bib is supplied

Verify the output

python -c "
import zipfile, re, sys
xml = zipfile.ZipFile(sys.argv[1]).read('word/document.xml').decode('utf-8','ignore')
print('page-break tokens:', len(re.findall(r'<w:br[^/]*w:type=\"page\"', xml)))
print('pageBreakBefore  :', xml.count('w:pageBreakBefore'))
print('paragraph styles :', set(re.findall(r'w:pStyle w:val=\"([^\"]+)\"', xml)) or '{Normal default}')
" output.docx

A passing build reports 0, 0, and an empty / Normal-only style set.

Repository layout

SKILL.md                            Claude Code skill manifest (full reference)
scripts/
  md2docx_template.js               main converter (the format lock lives here)
  resolve_citations.js             [@key] → Vancouver numbering + references
  migrate_existing_manuscript.js   helper for converting legacy manuscripts
examples/
  sample.md / sample.bib            input demo
  resolved.md / sample.docx         resolved + rendered output

Customizing the lock

Edit the // LOCKED DEFAULTS block at the top of scripts/md2docx_template.js (BASE_FONT, BASE_SIZE, LINE_SPACING, …) and save — every subsequent conversion uses the new values. Do not add per-file overrides; that defeats the lock.

Notes (Windows / Chinese locale)

  • GBK encoding errors from a downstream validate.py are false positives (Python defaulting to GBK for Unicode em-dashes / curly quotes); the DOCX is structurally valid.
  • Dublin Core dcterms.xsd fetch failures in docProps/core.xml are a network-only external-schema issue every docx-js file triggers, not a document defect.

License

MIT

About

Markdown→DOCX converter enforcing the SCI manuscript format lock (TNR 12pt, double spacing, single Normal style, three-line tables) with optional Vancouver --bib citation resolution.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages