At a glance
| Input | |
|---|---|
| Output | Plain text, copied or downloaded as .txt |
| What it cannot read | A scanned page, which holds pictures of words rather than words |
| Where it runs | Entirely in your browser — no file is uploaded to a server |
| Price | Free, with no account and no sign-up |
| Offline | Works with no connection after the first visit |
It cannot read a scan
Extraction pulls out the text a PDF actually contains. A document produced from Word or a web page has that text; a scan or a photograph of a page is a picture, and contains no text at all.
The quick test: try selecting text in your PDF viewer. If you cannot, there is nothing to extract, and you need optical character recognition rather than this. That is a different job and this page does not pretend to do it.
Why the layout comes out strange
A PDF stores where each piece of text sits on the page, not that a passage is a paragraph or that two blocks are columns. Extraction reads them in the order they were written into the file, which for a two-column academic paper often means the left column and the right interleaved line by line.
Tables suffer similarly: cells come out as a stream with no structure. Expect to tidy anything with a complex layout, and expect straightforward prose to come through cleanly.
Frequently asked questions
Why did I get nothing back?
The PDF is a scan — a picture of a page with no text in it. If you cannot select the text in a viewer, there is nothing to extract, and you need OCR instead.
Why are my columns mixed together?
A PDF records where text sits, not that it forms columns. Extraction follows the order in the file, which interleaves them.
Are tables preserved?
No. Cells come out as a stream of text without their structure. Complex tables need tidying by hand.
Something wrong with this tool, or an idea for it? Tell us