How to get text out of a scanned PDF

Published 20 August 2026

You have a PDF of something scanned - a contract, a payslip, a letter from a solicitor - and you need the words out of it. Ctrl+F finds nothing. Dragging across the text selects nothing. The document looks like text and behaves like a photograph, because that is what it is.

Two things share the name PDF

A PDF made by software - exported from Word, printed to PDF, generated by a payroll system - carries its text inside the file as text. Searching works, selecting works, and copying it out is instant and perfect.

A scanned PDF is photographs of paper in a PDF wrapper. There are no words in it, only pixels arranged in the shape of words. Every scanner, every phone scanning app and every fax-to-email system produces this kind.

Telling them apart takes two seconds. Try to select a line of text with the mouse. If a blue selection appears over the words, the text is there. If you get a rectangle over the whole page or nothing at all, it is a scan.

Getting words out of the second kind

Optical character recognition reads the shapes and works out which letters they are. Modern OCR on clean printed pages is accurate enough to be useful without proofreading every line, and it has known limits worth knowing before you start:

  • Clean print reads well. A flatbed scan at 300dpi of an ordinary printed page is the good case.
  • Phone photos read reasonably when the lighting was even and the page was flat.
  • Handwriting mostly does not read. Neat block capitals sometimes survive; cursive rarely does.
  • Layout is harder than letters. Two-column pages, tables and forms need the reading order rebuilt as well as the characters recognised, and that is where most tools quietly fall apart.

Where the document goes matters more than usual

Think about what people put through OCR: contracts, payslips, bank statements, medical letters, passports, tenancy agreements. Mainstream converters upload the file to a server, process it there, and ask you to trust a retention policy you have not read.

PDF to text does the reading inside the browser tab. The document is never sent anywhere, which is checkable rather than promised: open the network tab in your browser's developer tools and watch nothing leave while it works. The reading engine downloads once, about 17MB, and stays cached afterwards.

Getting a better result

  • Scan at 300dpi. Higher wastes time and adds noise; much lower and the recogniser starts guessing between similar letters.
  • Straighten the page before scanning it. A degree or two of skew is survivable, but ten will cost you accuracy.
  • Use greyscale or black and white, not colour, for ordinary printed documents.
  • Check the numbers by hand. OCR errors cluster in digits, where 1 and l and 8 and B look alike and no context exists to correct them.

What comes out

Plain text gives you the words with paragraphs joined back up and pages marked. Markdown keeps the structure the layout implied: headings as headings, tables in a block with their columns still lined up, and a line the page held apart staying apart instead of dissolving into the paragraph above it.

Pick Markdown if the text is heading for notes, a wiki, or anything that reads it. Pick plain text if you just need the words.

If the PDF turned out to be the first kind - real text inside - the same tool copies it straight out in a moment, no recognition involved.