OCR PDF
Extract text — or build a searchable PDF — from scanned PDFs and images. On-device AI OCR for major languages, Tesseract for the rest. Nothing is uploaded.
Drop a scan or photo of a document
or click to browse (PDF or image files)
Extract text — or build a searchable PDF — from scanned PDFs and images. On-device AI OCR for major languages, Tesseract for the rest. Nothing is uploaded.
or click to browse (PDF or image files)
Free, with no page cap and no account. Nothing is uploaded: every page is processed inside this browser tab.
OCR here is not one library. Pages are routed to whichever engine handles them best.
For the major scripts, a modern neural recognition model does the work. It is markedly better than classical OCR on the things real documents actually throw at you: text in columns, mixed font sizes, tables, tightly set body copy, and pages photographed rather than scanned. Pages are straightened upright before recognition, which alone accounts for a good share of the accuracy difference — skew is the single biggest cause of garbled OCR output.
For everything else, a well-established classical engine covers a long tail of languages and scripts. Between them the language list runs to English, Hindi, Spanish, French, German, Portuguese, Italian, Japanese, Korean, Simplified Chinese, Arabic and Russian, plus paired options — English with Hindi, Tamil, Bengali, Telugu or Marathi — that read both scripts in a single pass with no accuracy cost to either.
The first run of a new language downloads its model, around 5 to 10 MB, and caches it. After that the language is available instantly, and offline.
Extracted text is handy. The searchable PDF is what changes how you can use the document.
It keeps your original page images exactly as they were — the scan still looks like the scan, with its signatures, stamps, letterhead and handwriting all visible — and writes an invisible layer of recognised text positioned over the words. Nothing looks different. Everything works differently.
You can search the document in any reader. You can select and copy a paragraph out of a scan. Screen readers can read it aloud. Your operating system's desktop search finally indexes it. And the rest of the toolkit starts working on it: PDF to Word converts it with full structure, PDF to Excel reads its tables, Edit PDF lets you click a word and retype it, and PDF to Text extracts it in one step.
This is the single most valuable operation you can perform on an archive of scans, and it is the step that turns a folder of unsearchable images into a usable library.
Making an archive searchable. Years of scanned invoices, contracts or correspondence that no search can currently see inside.
Getting data out of statements and bills. Recognise the text, then take the tables into a spreadsheet with PDF to Excel.
Quoting from a scanned book or paper without retyping the passage.
Preparing documents for legal or compliance review, where a bundle has to be searchable to be workable.
Reading a scan on a phone. Selectable text reflows and can be enlarged; an image cannot.
Accessibility. A scanned PDF is invisible to a screen reader until it has a text layer. This is often the difference between a document being usable and not.
Yes, entirely. The recognition model is downloaded once and cached, then all the pixel-to-text work happens in a WebAssembly worker inside this browser tab. Your pages are never transmitted — which is unusual for OCR, since most services upload every page to a server. You can disconnect from the network after the model has cached and it still runs.
It is built for printed and typeset text, which is what scanned documents overwhelmingly contain — invoices, receipts, books, tax papers, printed forms, statements and contracts. Handwriting recognition is a different class of model entirely, and keeping the scope to print is part of what allows this to run on your own device instead of shipping your scans to a server. Handwritten signatures and annotations stay visible in the searchable PDF exactly as they are, since the original page image is preserved.
Roughly 5 to 15 seconds per page on a mid-range laptop, slower on a phone. A 50-page document typically finishes in a few minutes — start it and leave the tab open. Splitting a very large document with Split PDF and running it in parts is a good way to work through an archive.
Yes — choose one of the paired options. English with Hindi, Tamil, Bengali, Telugu or Marathi are all available, and both scripts are read in the same pass with no accuracy penalty on either. These exist because Indian government and legal documents are routinely bilingual. For other combinations, split the document and run each part with its own language.
Yes — no account, no page limit, no watermark. Because the processing happens on your own device rather than on a server, there is no per-page cost to pass on, which is exactly why most OCR services meter it and this one does not.
No. Your original page images are kept exactly as they are, and the recognised text is written as an invisible layer positioned over the words. The document looks identical and behaves completely differently.
No OCR is perfect on a poor original, and accuracy tracks the quality of the scan closely. Raising the source resolution, straightening the page and evening out the lighting all help more than any setting. For a small number of corrections in the final document, Edit PDF lets you click a word and retype it.
Yes — images in PNG, JPG, BMP, WebP and TIFF are accepted directly. For phone photos of paper, running them through Image to PDF with the document enhancement turned on first evens out the lighting and improves recognition noticeably.
Privacy: OCR is CPU-heavy but entirely local — the Tesseract language model is downloaded once and cached. Your scanned pages, recognised text, and export TXT all stay on your device.