Last updated 24 September 2026
A PDF saved from Word, a website or an accounting program stores its words as text: each letter is a character the computer knows. A PDF made by a scanner or a phone camera stores a picture of the page. On screen they are indistinguishable; to a computer, the second one contains no words at all.
Some PDFs are mixed — typed pages with a scanned signature page or an attached certificate — so check the page you actually need.
With a scanned PDF you cannot copy a paragraph, search for a name, or have a screen reader read it aloud. Converting it to Word produces a document containing a picture of each page, because there is no text to put in the document. And scans are large, which is why they are the files that break upload limits.
OCR — optical character recognition — looks at the picture of each page and works out which letters are there. PDF OCR does this for a scanned PDF, and image to text does it for a photo or screenshot; both run in your browser and read fifteen languages, including Arabic, Chinese, Hindi and Bangla.
Recognition is very good on clean printed pages and never perfect. It reads what it sees, so the quality of the scan decides the quality of the result:
This page in other languages: العربية · বাংলা · Deutsch · Español · فارسی · Français · עברית · हिन्दी · Bahasa Indonesia · Português · Русский · Türkçe · اردو · 中文