Scanned PDF or text PDF? Why you cannot copy the text — and how to fix it

Last updated 24 September 2026

Two kinds of PDF that look the same

A PDF saved from Word, a website or an accounting program stores its words as text: each letter is a character the computer knows. A PDF made by a scanner or a phone camera stores a picture of the page. On screen they are indistinguishable; to a computer, the second one contains no words at all.

How to tell in ten seconds

Some PDFs are mixed — typed pages with a scanned signature page or an attached certificate — so check the page you actually need.

  • Try to select a single word. If a blue box covers the whole page instead, it is a picture.
  • Search for a word you can see (Ctrl+F, or ⌘F on a Mac). A scan finds nothing.
  • Zoom in a long way. Text stays sharp at any zoom; a scan turns soft or blocky.
  • Look at the size. Ten pages of text are usually a few hundred kilobytes; ten scanned pages are often several megabytes.

Why it matters

With a scanned PDF you cannot copy a paragraph, search for a name, or have a screen reader read it aloud. Converting it to Word produces a document containing a picture of each page, because there is no text to put in the document. And scans are large, which is why they are the files that break upload limits.

Getting the words out: text recognition (OCR)

OCR — optical character recognition — looks at the picture of each page and works out which letters are there. PDF OCR does this for a scanned PDF, and image to text does it for a photo or screenshot; both run in your browser and read fifteen languages, including Arabic, Chinese, Hindi and Bangla.

Recognition is very good on clean printed pages and never perfect. It reads what it sees, so the quality of the scan decides the quality of the result:

  • Scan at 300 DPI for small print; 200 DPI is enough for normal-sized text.
  • Keep the page flat and straight, in even light, with no shadow from your hand or phone.
  • Choose the right language. A page read as the wrong language comes out as the nearest-looking letters of the language chosen.
  • Handwriting is recognised poorly by every OCR engine; expect to type it yourself.
  • Always check names, numbers and dates by eye afterwards. A 0 read as an O or a 1 read as an l is the most common error, and on an official form it matters most.

Which tool for which PDF

More guides

This page in other languages: العربية · বাংলা · Deutsch · Español · فارسی · Français · עברית · हिन्दी · Bahasa Indonesia · Português · Русский · Türkçe · اردو · 中文