OCR your scanned PDFs

This tool runs Tesseract.js — an open-source OCR engine — entirely inside your browser. Each PDF page is rendered to a high-resolution image, then converted to text. Heavy on CPU and memory; lightning-private.

How to use

  1. Pick a scanned PDF
  2. Press Run OCR (engine downloads on first use)
  3. Wait — 5–20 sec per page
  4. Copy the text or download as a .txt file

Common questions

Does this make my PDF itself searchable, or just give me the text separately?

Just the text separately — the output is plain text you copy or download as a .txt file, not a modified PDF with a searchable text layer added back into it. If you specifically need a searchable PDF, you'd need to combine the extracted text with the original file yourself.

Does this work on languages other than English?

The default is English; accuracy on other languages depends on how well Tesseract's English-trained model handles that script, which is generally poor for non-Latin alphabets.

OCR accuracy depends on scan quality. Crisp 300+ dpi black-text-on-white scans give 95%+ accuracy; faded, skewed or handwritten pages are unreliable.