Scan → Searchable PDF

Make Scanned PDFs Searchable (and Bates-Numbered)

Make a searchable PDF from scans when a filing or discovery set has to be searched. OCR reads each page, and low-confidence text is reviewed before anything is built. The export is a two-layer PDF — original image on top, invisible corrected text underneath — with optional sequential Bates stamps.

No credit card Free credits Visual verification

How it works

  1. 1

    Upload the scanned PDFs or images — a batch queues automatically and processes in order.

  2. 2

    OCR extracts text with per-word coordinates; low-confidence regions are flagged for review, so misreads don't end up hidden in the text layer.

  3. 3

    Export each file as a two-layer searchable PDF — with Bates numbering and your own prefix stamped on every page if you need it.

Why OhMyOCR

The text layer is corrected, not raw

One-click OCR tools stamp raw output under the scan. Here the text layer includes your corrections, so the PDF searches on what the document actually says.

Bates numbering

Set your prefix, and stamps run sequentially across every page of every file in the export — continuous numbering for a whole production.

Low-confidence text is visible

If OCR misreads a name, a search for that name silently fails. Low-confidence regions are flagged with image snippets, and a human clears them before the PDF is built.

Zero-LLM option and auto-deletion

For confidentiality requirements: a deterministic no-AI extraction mode, and originals can auto-delete N days after processing.

A searchable PDF is only as good as its hidden text

The fear in discovery and records work isn't OCR quality in general — it's the one document that never comes back in a search. Any tool can stamp a text layer under a scan; the question is what lands in that layer. Raw OCR output, errors included, means a misread “not” or a name spelled one letter off simply won't match. The search runs, returns nothing, and nobody knows the document was missed.

OhMyOCR builds the text layer from text that was reviewed first. Low-confidence regions are flagged with image snippets before export; you clear them with keyboard-only review (J/K/Enter), and the corrected reading — not the raw guess — is what goes under the image. The PDF looks identical to the scan and searches on what the document actually says.

Bates numbering and batch production

The Export Center merges the whole batch into one two-layer PDF: your prefix, continuous numbering across every page of every file, applied as a legible corner stamp. Paired with the review pass, that's a defensible pipeline — scan in, flags cleared, corrected text layer, numbered pages out.

For confidentiality-sensitive work there are two controls: a deterministic zero-LLM extraction mode, and auto-deletion of originals N days after processing. Documents are never used for model training. We don't currently offer HIPAA BAAs, so medical records custodians should plan accordingly.

Frequently asked questions

What is a two-layer (sandwich) PDF?

Each page shows the original scanned image with an invisible text layer underneath. It looks identical to the scan but supports search, select and copy — the standard for e-filing and archives.

Can I add Bates numbers?

Yes. Turn on Bates numbering in the Export Center and set your prefix — stamps are applied to the corner of every page of the exported PDF, numbered per document.

Why review before exporting?

Because the hidden text is what gets searched. If OCR misread a name, that name silently fails to find anything. Flagged regions let you fix exactly those words before the text layer is built.

Does it handle photographs of documents, not just scans?

Yes. Phone photos work the same — pages are rendered and OCR runs with coordinates either way.

Is this suitable for confidential documents?

Documents run through our own pipeline, never used for training, and originals can auto-delete N days after processing. Note: we don't currently offer HIPAA BAAs.

Try it on your own file

Free to start, no credit card, results you can verify line by line.

Get started — free