TOOL

Convert PDF to text

Pull the text out of any PDF, scanned pages included, as plain text or Markdown. Free, no signup, and your documents never leave your device.

Drop files here or click to browse. PDF accepted, batches welcome

Drop a PDF, download its words as a plain .txt file or as Markdown. Ordinary PDFs finish in a moment; scanned ones take longer, because they are read rather than copied.

Two kinds of PDF, one tool. A PDF made by software carries its text inside the file, and that text is copied out exactly: instant, and accurate to the character. A scanned PDF is photographs of pages, and the tool reads each page with character recognition instead. It checks page by page and picks the right route on its own; a report with a scanned appendix gets each treatment where it belongs.

What recognition can do. Clean printed pages read well. Phone photos of documents read reasonably when the light was decent. Handwriting mostly does not read.

The reading engine fetches on first use - about 17MB, kept by your browser afterwards - and only when a scanned page needs it. Copying text out of ordinary PDFs downloads nothing extra.

Two formats, chosen before you drop the file. Plain text gives you the words, with paragraphs joined back up and pages marked. PDF to Markdown keeps what the layout meant: headings as headings, tables fenced with their columns still aligned, and a line the page held apart staying apart rather than dissolving into the paragraph above it.

For pages as pictures instead, PDF to JPG is the other direction.

The document is read on your device and is not sent anywhere. That is the point when it is a contract, a letter or a bank statement.

Frequently asked questions

Why did my PDF finish instantly when others take ages?

Because it carried its text already: software-made PDFs store their words inside the file, and copying them out is quick. The slow route is for scanned pages, which are pictures and have to be read character by character.

Can I get Markdown instead of plain text?

Yes. Pick Markdown before you drop the file and the download is a .md file. It carries the structure the plain-text version has to throw away: headings as headings, tables in a fenced block with their columns aligned, page breaks as rules. Useful if the text is heading for notes, a wiki, or anything that reads Markdown.

How does it know what a heading is?

By where the line sits on the page. In ordinary documents a title is often no larger than the body text around it, which makes size a poor test: what marks a heading is standing alone, stopping well short of the margin, carrying no sentence punctuation, and introducing the text beneath it. Anything ambiguous is left as an ordinary paragraph, which is the quieter mistake of the two.

The text came out jumbled. Why?

Layout is the usual culprit, though less often than it was. Two-column pages are detected and read one column at a time rather than straight across, and table rows are kept as rows instead of being dissolved into prose. Forms remain awkward, and a scan whose columns run together in the picture itself can still defeat the reading. The words are always there; only their arrangement is ever in doubt.

Does it read handwriting?

Mostly not. The recogniser is built for print; neat block capitals sometimes read, cursive rarely does. A typed document photographed clearly is a much better bet.