PDF Tools
PDF OCR
Extract text from scanned PDFs or image-based PDFs using OCR (Optical Character Recognition). Uses Tesseract.js which runs entirely in your browser.
100% private — processed in your browser only
Drop a scanned PDF to extract text
or click to browse · PDF only · English text
OCR is slow: Tesseract.js runs in your browser and takes 30–120 seconds per page. Multi-page PDFs take proportionally longer. Do not close the tab while processing.
How to use PDF OCR
1
Upload a scanned PDF
Drop a PDF that contains scanned images of text.
2
Wait for OCR
Each page is rendered and run through Tesseract.js OCR. This takes 30–120 seconds per page.
3
Copy or download the text
Copy the extracted text or download it as a .txt file.
Your privacy is protected
This tool runs entirely in your browser. Files you upload are never sent to any server — they are processed locally using JavaScript and immediately discarded when you close the page.
Frequently Asked Questions
How long does it take?
Tesseract.js runs entirely in your browser using WASM. Expect 30–120 seconds per page depending on your device.
Is my PDF uploaded?
No. OCR runs entirely in your browser. Nothing is sent to a server.
What languages are supported?
This tool recognizes English text. Multi-language support would require loading additional Tesseract language packs.
Why is my text garbled?
OCR quality depends on image resolution and clarity. Very low-resolution or skewed scans will produce poor results.