Adwaty.

Extract Text from PDF (OCR)

Upload a scanned or image-based PDF, and extract its text in English, Arabic, or both — all processing happens in your browser.

extract text from pdfpdf ocr onlinescanned pdf to text

🔒 All processing happens inside your browser — your file is never uploaded to any server. Loading the OCR engine the first time may take a minute depending on your connection, and processing is relatively slower for large or multi-page files.

How to extract text from a scanned PDF

If you have ever tried to select text in a PDF and could not, the file is most likely made of images of scanned or photographed pages rather than real text. That is common with old books, signed contracts and archived official letters. This is where optical character recognition (OCR) helps: the tool reads the image of each page and turns the characters into text you can copy, edit and search.

The tool runs entirely in your browser using the Tesseract engine. On first use, the engine and the selected language data are downloaded; each page of your file is then rendered as a high-resolution image and read on your own device, without sending your file to a server.

How to use it

  1. Drag a PDF onto Drag PDF file here, or click to choose it.
  2. From Text language, choose English, Arabic, or English + Arabic.
  3. Click Extract text.
  4. Follow the progress bar and status messages such as "Recognizing page 3 of 10".
  5. When finished, the text appears in a box below, split by headings like "--- Page 1 ---".
  6. Click Copy text to copy it, or Download as .txt to save it as extracted-text.txt.

What the tool offers

  • Three language options for English, Arabic and mixed documents.
  • All pages processed in order, with a clear separator between pages.
  • A progress bar so you can follow long files.
  • Plain text output you can copy or download.

Practical use cases

  • Pulling the text from a scanned lease to quote a clause in an email.
  • Turning pages of an old photographed book into text for notes or research.
  • Copying details from an archived letter instead of retyping them.
  • Making photographed printed lecture notes searchable.

Tips for better results

  • Pick a single language when the document uses only one; it is faster and often more accurate. Use English + Arabic only for mixed documents.
  • Results depend on scan quality: straight, well-lit pages with clear printed type work best, while handwriting and decorative fonts give weak results.
  • Review Arabic output carefully, especially with diacritics or small fonts.
  • Rotate upside-down or sideways pages before running OCR.
  • Long files take time because every page is processed on your device; extract only the pages you need first.
  • The output is text only; no searchable PDF is created.

For a single image rather than a PDF, use Image to text (OCR). If your PDF already has selectable text, PDF to Word is faster, and to fix page orientation first, use Rotate PDF.

Frequently Asked Questions

Does the tool support Arabic files?

Yes, you can choose Arabic, English, or both together before starting extraction, depending on your file's text language.

Why is processing a bit slow?

The tool loads a full OCR engine (Tesseract) and runs it inside your browser — this takes longer than a dedicated server, especially for large or multi-page files, but your file stays completely private.

Is my file uploaded to any server?

Not at all — text recognition happens entirely inside your browser; the only thing downloaded is the OCR engine itself (generic files, unrelated to your personal file).

Does the tool create a searchable PDF?

No. It extracts the text and shows it for copying or downloading as a .txt file; the original PDF is not modified.

Can it read handwriting?

The engine is designed mainly for printed text, so results with handwriting are usually poor.

Which language option should I pick for a mixed document?

Choose English + Arabic so both scripts are recognised. For single-language documents, choose that language alone for faster results.

Do I need an internet connection?

Yes, the first time you use each language, to download the OCR engine and language data. Your PDF itself is processed on your device.