PDF OCR Text Extraction — Get Text from Scanned PDFs, Free
Extract text from scanned and typed PDFs: text layers are copied as-is, scanned pages are read with Tesseract OCR in 17 languages. Free with a Google sign-in; files deleted right after.
About PDF OCR Text Extraction
OCR (Optical Character Recognition) extracts text from images of text — scanned PDFs, photographed documents, faxes. The ZTools PDF OCR tool sends your PDF to our own processing server, which handles each page the fastest correct way: a page that already has a text layer is copied as-is, and a scanned page is read with the open-source Tesseract OCR engine in any of 17 languages (up to three at once). It takes PDFs up to 50 MB, 30 pages per run (pick a page range for longer files), shows each page with its method and OCR confidence, and deletes the PDF right after processing. Scanned pages take a few seconds each; it needs a free Google sign-in.
Use cases
- Digitise a scanned book / magazine. No embedded text in the PDF; OCR reads the page images. Output as searchable text.
- Process a fax-scanned contract. Faxes are image-only. OCR makes the text usable for downstream search / copy.
- Extract text from screenshots embedded in PDFs. A report with screenshots of code. OCR reads the code text.
- Multilingual document processing. Pick up to three of the 17 languages, such as English, Urdu and Arabic, for a document that mixes them.
How it works
- Sign in and drop a PDF. Up to 50 MB, not password-protected. Choose a page range for files longer than 30 pages.
- Pick the languages. Choose up to three of the 17 languages the scanned pages are written in.
- Our server reads each page. Pages with a text layer are copied exactly; scanned pages are rendered and read with Tesseract OCR. The server then deletes the PDF.
- Review and export. Each page shows whether it came from the text layer or OCR, with its confidence. Copy all the text or download it as a .txt file.
Examples
Input: 10-page scanned contract, English
Output: Text of all 10 pages, marked as OCR with a confidence score; a few seconds per page.
Input: Multilingual scan (English + French)
Output: Pick both languages and Tesseract reads both in one run.
Input: Handwritten notes
Output: Tesseract is built for print, so handwriting often comes out partly wrong. For handwriting, a handwriting-specific OCR service does better.
Frequently asked questions
How accurate is it?
Text-layer pages are exact. Clean printed scans read very accurately; noisy scans less so; handwriting poorly. The confidence score on each page shows where to proofread.
Is my PDF uploaded?
Yes. OCR runs on our own processing server, not in your browser, so the PDF goes there over HTTPS. The server deletes it right after processing and keeps only a request record with no document data.
Maximum PDF size?
Up to 50 MB and 30 pages per run. For longer files, run it again with the next page range.
Why do I need to sign in?
Because OCR runs on our server, a free Google sign-in and a quick bot check keep it available for everyone. Each account has a daily limit.
Pro tips
- For best accuracy, use the highest-resolution scan available. 300 DPI minimum.
- For multilingual docs, only pick the languages actually present — extra languages slow processing and can add errors.
- If your PDF has selectable text, every page comes back as "Text layer" in seconds, with no OCR needed.
- For handwriting, use a handwriting OCR service. Tesseract is print-optimised.
Reviewed by Ahsan Mahmood · Last updated 2026-09-11 · Part of ZTools.
For the full,
formatted version of this page, please enable JavaScript and reload
https://ztools.zaions.com/tools/pdf-ocr-text-extraction.