Extract text from a scanned PDF

A scanned PDF is a stack of pictures — that is why selecting or searching its text does nothing, and why converting it to Word comes back empty. This reads the characters out of each page image and gives you real text.

Every page is rendered at a high resolution first and then recognised, with progress shown page by page. Long documents take a while: recognition is genuinely doing work on your device rather than queueing on a server.

Arabic, English and mixed Arabic/English documents are all supported, and Arabic output keeps its right-to-left direction in both the preview and the Word export.

How to use it

  1. Add a scanned PDF — one where selecting text does nothing.
  2. Choose the language: English, Arabic, or both for a mixed document.
  3. Press Extract. The recognition model downloads once on the first run and is cached afterwards.
  4. Copy the text, or download it as TXT or Word.

Supported formats and limits

Input
Scanned PDFs and photographed pages
Output
Text — copy, TXT or Word
Languages
English, Arabic, Arabic+English
First run
One-time model download, then cached offline
Where it runs
In your browser; the document is never uploaded

Not for: Handwriting — recognition is reliable on printed text and much less so on handwritten notes.

Frequently asked questions

Are my files uploaded to a server?
No. This conversion runs entirely inside your browser using your own device’s processing power. The file never leaves your computer or phone, which is why it also works when your connection drops.
How do I know if my PDF needs OCR?
Try selecting text in it. If nothing highlights, it is a scan and needs OCR; if text selects normally, the plain PDF-to-Word converter will do a better job.
Does the layout survive?
You get the text in reading order, not a reconstruction of the page design. Columns and tables come through as text rather than as formatted structures.

Related tools