Get the text out of a PDF
Correct reading order even on two-column pages, with optional Markdown structure for headings, lists and emphasis.
- Nothing is uploaded
- No signup
- No watermark
- Works offline after the first visit
Questions
My PDF is a scan and nothing comes out. Why?
A scan is a picture of text, not text. Run it through OCR first to create a text layer, then extract from that.
Does it handle two-column papers?
Two-column pages are detected and read down each column rather than straight across. More complicated layouts — tables, sidebars, three columns — can still come out in the wrong order.