Turn any document into clean Markdown — in the browser, in one file, with real Persian/RTL support.
Built for the moment you hand a document to an AI chat and it answers "file type not supported" or swallows your whole context window.
Architecture · Conversion matrix · Persian/RTL notes · Deployment
Chat assistants are picky and expensive to feed:
| Problem | What happens in practice |
|---|---|
| Unsupported uploads | DOCX, oddly encoded CSV or an exported HTML page get rejected |
| Bloated context | A PDF pasted as raw text spends tokens on page furniture, watermarks and layout noise |
| Broken Persian | Copying Persian out of a PDF yields mirrored, disconnected letters |
Markdown is the format every model reads best: plain text, explicit structure, no markup tax. This tool converts locally, strips the noise and hands back a Markdown file that is smaller, cleaner and cheaper to send — without uploading anything anywhere.
- 19 conversion paths between TXT, Markdown, CSV, JSON, HTML, DOCX and PDF — see the matrix
- Text-based PDF extraction with heading detection, paragraph reflow and watermark removal
- Persian / Arabic pipeline: de-shaping of presentation forms, bidi reordering, half-space (ZWNJ) repair, tanwin restoration
- Batch mode: drop many files, convert in one pass, download individually or as a ZIP
- Live preview, syntax-highlighted editor, search, undo/redo and a session history panel
- Word / character / line counters plus an estimated reading time
- Auto-detection of file type and text encoding (UTF-8, UTF-16LE, Windows-1256, Windows-1252)
- Light / dark / high-contrast themes, keyboard navigation, ARIA labels, mobile-first layout
- One deployable file:
dist/worker.jsembeds the HTML, CSS and every module
git clone https://github.com/Shayan-alinezhad/cloner-document-converter.git
cd cloner-document-converter
npm start # dev server on http://localhost:8123 (nothing to install)Build and deploy the single-file worker:
npm run build # -> dist/worker.js
npm run preview # serve the built artifact exactly as Cloudflare will
npm run deploy # npx wrangler deployFull instructions, including the copy-paste dashboard route and Cloudflare Pages, live in docs/DEPLOYMENT.md.
No server, no database, no analytics. Files are read with the File API, converted in the tab
and written back with Blob + URL.createObjectURL. The only outbound requests fetch the
conversion libraries from jsDelivr, which is pinned by a strict Content-Security-Policy
(object-src 'none', form-action 'none', no unsafe-eval). Rendered HTML always passes
through DOMPurify.
src/ the application — plain ES modules, no build step required
index.html markup and panels
style.css design system, themes, responsive layers
script.js UI controller: events, editor, history, theming
converter.js conversion registry: every from -> to pair
markdown.js Markdown parsing, block model and highlighting
csv.js CSV parsing and the document CSV schema
json.js JSON parsing, flattening and document objects
docx.js DOCX -> HTML -> Markdown via Mammoth
pdf.js PDF text extraction and PDF generation
rtl.js Persian/Arabic de-shaping, bidi and half-space repair
utils.js DOM, file, encoding, ZIP and lazy library helpers
scripts/ dev server, worker build, worker preview
test/ dependency-free test files + runner
docs/ architecture, conversion matrix, RTL notes, deployment
public/_headers security headers for Cloudflare Pages deployments
npm testThe suite is deliberately boring Node: it exercises what is easy to break — half-space rules, tanwin restoration, mirrored punctuation, bidi row assembly and the watermark filter — using cases taken from real Persian PDFs.
Current Chrome, Edge, Firefox and Safari (including iOS). The app relies on ES modules,
async/await and dynamic import(); no polyfills are shipped.
Bug reports and pull requests are welcome — start with CONTRIBUTING.md. For a conversion bug, attach the input file (or a redacted excerpt) and the output you got.
MIT © 2026 Shayan Alinezhad — clonerr.ir · Telegram