OCR your scanned PDFs
This tool runs Tesseract.js — an open-source OCR engine — entirely inside your browser. Each PDF page is rendered to a high-resolution image, then converted to text. Heavy on CPU and memory; lightning-private.
How to use
- Pick a scanned PDF
- Press Run OCR (engine downloads on first use)
- Wait — 5–20 sec per page
- Copy the text or download as a .txt file
Common questions
Does this make my PDF itself searchable, or just give me the text separately?
Just the text separately — the output is plain text you copy or download as a .txt file, not a modified PDF with a searchable text layer added back into it. If you specifically need a searchable PDF, you'd need to combine the extracted text with the original file yourself.
Does this work on languages other than English?
The default is English; accuracy on other languages depends on how well Tesseract's English-trained model handles that script, which is generally poor for non-Latin alphabets.
OCR accuracy depends on scan quality. Crisp 300+ dpi black-text-on-white scans give 95%+ accuracy; faded, skewed or handwritten pages are unreliable.