xcb tools

PDF & DOCX to Markdown converter

PDF and DOCX to Markdown Converter

Drop a .pdf or .docx file here, or

Everything above runs locally in your browser — your file is never uploaded anywhere.

Why convert to Markdown before using an LLM?

Sending a PDF straight to a vision-based API can cost several times what the same content costs as plain text, because each page is processed as an image to preserve layout. Markdown keeps the structure that matters — headings, lists, tables — at a fraction of the token cost. See why your PDF costs more than it should for the full breakdown.

This tool doesn't do OCR: if a PDF is scanned (a photo or image of a page rather than real text), it will tell you clearly instead of returning an empty or broken result.

FAQ

Does this tool do OCR on scanned PDFs?

No. If a PDF is detected as scanned or image-based, the tool tells you directly rather than attempting extraction and returning nothing useful.

Can it handle password-protected PDFs?

Not yet. A password-protected PDF will show a clear error rather than a password prompt.

Does it support the older .doc format?

No, only modern .docx files. The legacy .doc binary format has no reliable browser-based conversion library, so it isn't supported.

Does this tool upload my file anywhere?

No. Both the PDF and DOCX conversion paths run entirely in your browser using WebAssembly and JavaScript. Your file is never sent to a server.

Do DOCX tables always convert to Markdown tables?

Only when the source Word document uses a genuine header-row table style. Tables without header-row styling are preserved as raw HTML inside the Markdown output rather than being mangled into an invalid table — check the result for tables from documents you didn't build yourself.