How to extract text from a scanned PDF
If you have ever tried to select text in a PDF and could not, the file is most likely made of images of scanned or photographed pages rather than real text. That is common with old books, signed contracts and archived official letters. This is where optical character recognition (OCR) helps: the tool reads the image of each page and turns the characters into text you can copy, edit and search.
The tool runs entirely in your browser using the Tesseract engine. On first use, the engine and the selected language data are downloaded; each page of your file is then rendered as a high-resolution image and read on your own device, without sending your file to a server.
How to use it
- Drag a PDF onto Drag PDF file here, or click to choose it.
- From Text language, choose English, Arabic, or English + Arabic.
- Click Extract text.
- Follow the progress bar and status messages such as "Recognizing page 3 of 10".
- When finished, the text appears in a box below, split by headings like "--- Page 1 ---".
- Click Copy text to copy it, or Download as .txt to save it as extracted-text.txt.
What the tool offers
- Three language options for English, Arabic and mixed documents.
- All pages processed in order, with a clear separator between pages.
- A progress bar so you can follow long files.
- Plain text output you can copy or download.
Practical use cases
- Pulling the text from a scanned lease to quote a clause in an email.
- Turning pages of an old photographed book into text for notes or research.
- Copying details from an archived letter instead of retyping them.
- Making photographed printed lecture notes searchable.
Tips for better results
- Pick a single language when the document uses only one; it is faster and often more accurate. Use English + Arabic only for mixed documents.
- Results depend on scan quality: straight, well-lit pages with clear printed type work best, while handwriting and decorative fonts give weak results.
- Review Arabic output carefully, especially with diacritics or small fonts.
- Rotate upside-down or sideways pages before running OCR.
- Long files take time because every page is processed on your device; extract only the pages you need first.
- The output is text only; no searchable PDF is created.
For a single image rather than a PDF, use Image to text (OCR). If your PDF already has selectable text, PDF to Word is faster, and to fix page orientation first, use Rotate PDF.