Skip to main content
reference9 min read

OCRmyPDF puts searchable text beneath the scanned page image

Illustration for OCRmyPDF puts searchable text beneath the scanned page image
OCRmyPDF puts searchable text beneath the scanned page image

Before rescanning a perfectly legible document, check whether the problem is the page image or the absence of text behind it. A scanned PDF can look impeccable and still contain nothing a search box can find. It is a photograph wearing office clothes, and PDF is quite content to let the misunderstanding continue.

OCRmyPDF addresses that mismatch. It accepts PDFs or images, uses Tesseract OCR to recognize text, and places that text accurately beneath the visible page image. The resulting PDF can be searched and copied while retaining the scan as its visible content. By default, the project says it produces PDF/A and validates its input and output files. See the OCRmyPDF documentation.

The visible page is only half the file

At the level described by the project, an OCRmyPDF job has four main parts:

  1. It accepts a PDF or image and validates the input.
  2. If requested, it can rotate misoriented pages, deskew crooked pages, or clean page images before recognition.
  3. Tesseract recognizes text using the requested language packs.
  4. OCRmyPDF positions the recognized text beneath the image, may optimize embedded images, and validates the resulting PDF.

The hidden text layer is why a document can become searchable without being visibly retyped. It also explains why OCR errors may appear only when somebody searches or copies text. The scan can look flawless while the text underneath has developed its own interpretation of spelling.

OCRmyPDF requires external installations of Ghostscript and Tesseract in addition to the required Python version. OCRmyPDF itself is pure Python, but that describes its own code rather than the complete executable supply chain. The external dependencies did not receive the purity memo.

The project's promises about image handling are deliberately qualified. It keeps the exact resolution of original embedded images and, when possible, inserts OCR information as a lossless operation without disrupting other content. It also optimizes images and often produces a smaller file. "When possible" and "often" are doing useful engineering work here.

Build the command around the artifact you want

The shortest useful pattern names separate input and output files:


ocrmypdf --output-type pdfa input.pdf output.pdf

The README says PDF/A is already the default, but making --output-type pdfa explicit can still clarify the intended artifact in a script. PDF/A is designed for long-term storage, although creating a valid PDF/A file does not certify that every OCR word is correct. File validation and transcription accuracy answer different questions.

A more configured job can combine language selection, page correction, metadata, and an explicit worker count:


ocrmypdf -l eng+fra --rotate-pages --deskew --title "My PDF" --jobs 4 --output-type pdfa input_scanned.pdf output_searchable.pdf

Here, -l eng+fra requests English and French, --rotate-pages handles misrotated pages, and --deskew straightens crooked ones. --title "My PDF" changes output metadata. OCRmyPDF uses multiple CPU cores by default; --jobs 4 states a particular worker count instead of leaving that choice implicit.

OCRmyPDF can also turn an image into a single-page PDF:


ocrmypdf input.jpg output.pdf

The documented workflow operates on local input and output paths. It does not require a document upload, and the README explicitly says private data remains private. If a browser is used to inspect the result, that browser stage must also be local-only for the complete workflow to remain local. A browser tab is an interface, not a privacy guarantee.

Language selection is document configuration

OCRmyPDF relies on Tesseract language packs. Installing OCRmyPDF therefore does not establish that every language needed by a document is available. The language choice belongs with the document, not merely with the computer on which the command happens to run.

For a multilingual document, request multiple languages explicitly:


ocrmypdf -l eng+fra Bilingual-English-French.pdf Bilingual-English-French.pdf

On Debian or Ubuntu, the README gives this example for finding and installing language data:


apt-cache search tesseract-ocr
apt-get install tesseract-ocr-chi-sim

OCRmyPDF automatically uses the first Tesseract installation it finds on the PATH environment variable.

That selection rule matters for reproducibility. Two computers can have the same OCRmyPDF command and still select different Tesseract installations because their PATH order differs. Record the Tesseract environment alongside an automated job rather than treating OCRmyPDF's version number as the whole processing configuration.

Automation is local, but accuracy still needs inspection

OCRmyPDF is designed as a scriptable command-line program. It distributes work across available CPU cores, and the project says it scales to files containing thousands of pages. It also validates both ends of the operation, which is valuable when large jobs would otherwise produce an output file nobody inspects until months later.

Validation should not be confused with proofreading. A structurally valid PDF may still contain incorrectly recognized words. PDF validity is a low bar wearing a tie. For consequential documents, search representative phrases and copy text from representative pages after processing, including pages from each requested language.

OCRmyPDF also permits the input and output paths to be identical:


ocrmypdf myfile.pdf myfile.pdf

The README says this modifies the file only when processing succeeds. That is useful for automation, but it also means a successful run replaces the untouched document at that path. Use a separate output name whenever retaining the original matters.

Where OCRmyPDF goes sideways

The document is searchable, but one language is nonsense

Symptom: Text in one language searches or copies badly while another language works.

Cause: OCRmyPDF relies on installed Tesseract language packs, and the command did not request all languages present in the document.

What to do: Install the required language data and pass the appropriate languages with -l, combining them with + when necessary, as in -l eng+fra.

The output is still sideways or crooked

Symptom: OCR finishes, but some pages remain misrotated or their text lines visibly slope.

Cause: Adding an OCR layer does not itself request page rotation or deskewing. Those are separate processing choices.

What to do: Use --rotate-pages for misrotated pages and --deskew for crooked scans. Inspect the visible result as well as the searchable text.

OCRmyPDF selects an unexpected Tesseract installation

Symptom: Recognition behavior changes between computers even though the OCRmyPDF command is identical.

Cause: OCRmyPDF uses the first Tesseract installation found on PATH.

What to do: Control which installation appears first on PATH and record that environment with the job configuration.

Installing the Python package was not enough

Symptom: OCRmyPDF is present, but the processing stack cannot complete an OCR job.

Cause: Ghostscript and Tesseract are external requirements. OCRmyPDF being pure Python does not bundle those programs into existence.

What to do: Use the documented platform installation route, such as apt install ocrmypdf on Debian or Ubuntu or brew install ocrmypdf on macOS, and consult the installation documentation for other systems.

The output did not become smaller

Symptom: A valid searchable PDF is produced, but its size does not fall as expected.

Cause: The project says image optimization often produces a smaller file, not that every output will shrink. It also preserves embedded-image resolution and inserts OCR losslessly only when possible.

What to do: Treat reduced size as a possible benefit rather than the success criterion. Verify searchability, copy and paste behavior, visual fidelity, and validity independently of file size.

The successful in-place run consumed the original

Symptom: myfile.pdf now contains the processed result, and the untouched scan is no longer available at that path.

Cause: Using the same path for input and output intentionally modifies the file on success. "Only on success" is reassuring until success is precisely when the original was needed.

What to do: Write to a separate output file or preserve an independent copy before using the in-place form.

Validation passed, but copied text is wrong

Symptom: OCRmyPDF reports a usable output, yet copied words contain recognition mistakes.

Cause: The README says OCRmyPDF validates PDF input and output. It does not say that validation proves every recognized character matches the page image.

What to do: Test search and copy-paste behavior on representative pages. Keep the visible scan as the authority when the OCR layer and image disagree.

Licence notes

The supplied PyPI metadata does not state a licence. The GitHub source repository identifies the project licence as MPL-2.0. An unstated PyPI field should not be mistaken for an unlicensed project; it is a difference between metadata sources.

References


Important notice. Tap any item to read it in full.

Accuracy is not guaranteed

This article was produced with substantial automated assistance and is published without individual expert verification of every statement. It may contain errors, omissions, oversimplifications, or claims that were accurate when written and have since been superseded. Software, protocols, specifications and best practice in this field change quickly.

Verify before you rely on it

Treat this page as a starting point and a pointer to primary sources, never as an authority in itself. Before acting on anything here, check it against the official documentation, the original publication, or the vendor's own materials, which are linked in the references above. Where this page and a primary source disagree, the primary source is correct and this page is wrong.

No warranty

This content is provided "as is", without warranty of any kind, express or implied, including but not limited to warranties of accuracy, completeness, currency, merchantability, or fitness for a particular purpose.

No liability

To the fullest extent permitted by applicable law, pdfb2.io and its authors accept no liability for any loss or damage whatsoever, whether direct, indirect, incidental, consequential or otherwise, arising from use of or reliance on this article. This expressly includes lost time, lost data, damaged samples or specimens, wasted reagents or compute, failed experiments, equipment damage, and commercial loss.

Not professional advice

Nothing here constitutes professional, scientific, engineering, regulatory, safety or legal advice. You remain solely responsible for your own experimental design, safety assessment, regulatory compliance and data handling, and for any code you run or procedure you perform.

About the illustration

Any image accompanying this article is editorial and decorative. It was produced with generative AI, is not a technical diagram, is not to scale, and is not an accurate depiction of any structure, process or result. Do not read measurements, structures or relationships from it.

Third-party names and links

Product, project and organisation names are the property of their respective owners and are used for identification only. Their mention is not endorsement, affiliation or sponsorship in either direction. External links are provided for convenience and we neither control nor are responsible for third-party content.

Corrections

If you find an error, tell us and we will correct or withdraw the page.

ocrmypdfentity referencepdf

Ready to Try PDFb2?

Process your PDFs privately in your browser — 2 free downloads per day, no account needed. Your files never leave your device.

Try PDF Tools Free