Skip to content

Docs update for 2.0 - #5146

Open
jamie-lemon wants to merge 14 commits into
mainfrom
docs-update-for-2.0
Open

jamie-lemon wants to merge 14 commits into
mainfrom
docs-update-for-2.0

Conversation

@jamie-lemon

Copy link
Copy Markdown
Collaborator

No description provided.

Comment thread docs/ocr/index.rst Outdated
OCR Adaptors
------------

By default, PyMuPDF uses **Tesseract** and **OpenCV** for image pre-processing. If you need a different OCR engine — for higher accuracy, language support, or cloud-based processing — you can plug in a custom adaptor.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We do not use OpenCV anymore since some. For image pre-processing, we use our own utility that is based on a LightGBM model (Gradient Boosting Decision Tree) that has been trained on 70 thousand documents of DocLayNet's ground truth data.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@JorjMcKie Please advise, earlier under "How OCR is Triggered" we have a note:

For these heuristics to work, both a Tesseract installation and OpenCV must be available in your Python environment. If either is missing, no OCR is attempted.

Is this now simply:

For these heuristics to work, a Tesseract installation must be available in your Python environment. If it is missing, no OCR is attempted.

Comment thread docs/converting-files.rst Outdated
~~~~~~~~~~~~~~~~~

By utlilizing the :doc:`PyMuPDF4LLM API <pymupdf4llm/api>` we are able to convert PDF to a Markdown representation.
By utlilizing the :meth:`Document.to_markdown` method we are able to convert PDF to a Markdown representation.

@JorjMcKie JorjMcKie Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This conversion is not restricted to PDF:

By utlilizing the :meth:Document.to_markdown method we are able to convert any of the following document types to a Markdown representation:

  • PDFs
  • Image documents (because they are internally converted to 1-page PDFs)
  • Office documents opened using PyMuPDF-Office (because they are converted to PDFs via .to_pdf() before processing them).

Comment thread docs/document.rst Outdated
:meth:`Document.set_xml_metadata` PDF only: create or update document XML metadata
:meth:`Document.subset_fonts` PDF only: create font subsets
:meth:`Document.switch_layer` PDF only: activate OC configuration
:meth:`Document.to_json` PDF only: convert the document to JSON

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I know I communicated it differently once. But the restriction to PDF is exclusively connected to the fact that we currently require PDF to write back OCRed text to pages that have been determined to need OCR.
We are investigating how to avoid this restriction.
But still, while this result is pending, the following document types are also eligible:

  • Image documents (because we internally convert them to 1-page PDFs)
  • Office documents handled with PyMuPDF-Office (because we internally deal with the .to_pdf() output)

@JorjMcKie JorjMcKie left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have made a handful of comments...

Comment thread docs/ocr/index.rst

.. note::

For these heuristics to work, both a `Tesseract installation <installation_ocr>` and `OpenCV <https://pypi.org/project/opencv-python/>`_ must be available in your Python environment. If either is missing, no OCR is attempted.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

New suggested text:

For these heuristics to work, at least one of the supported default OCR engines, Tesseract or RapidOCR must be installed. Otherwise, the behavior is the same as if use_ocr=False was specified. If you have your own, non-default OCR engine, you must supply the respective plugin.

Comment thread docs/ocr/index.rst

If you want to use both OCR engines side-by-side, you can do so by implementing a custom OCR function which calls both OCR engines — one for bbox recognition (RapidOCR) and the other for text recognition (Tesseract) — and then combines their results.

This pre-made callable OCR function can be found in the ``PyMuPDF.ocr`` module as ``rapidtess_api.exec_ocr``.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
This pre-made callable OCR function can be found in the ``PyMuPDF.ocr`` module as ``rapidtess_api.exec_ocr``.
This pre-made default callable OCR function can be found in the ``PyMuPDF.ocr`` module as ``rapidtess_api.exec_ocr``.

Comment thread docs/ocr/index.rst
RapidOCR & Tesseract Side-by-Side
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

If you want to use both OCR engines side-by-side, you can do so by implementing a custom OCR function which calls both OCR engines — one for bbox recognition (RapidOCR) and the other for text recognition (Tesseract) — and then combines their results.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
If you want to use both OCR engines side-by-side, you can do so by implementing a custom OCR function which calls both OCR engines — one for bbox recognition (RapidOCR) and the other for text recognition (Tesseract) — and then combines their results.
If both OCR engines are installed -- Tesseract and one of RapidOCR / RapidOCR-ONNXRuntime -- they are automatically used together by default: RapidOCR for text line rectangle detection and Tesseract for text recognition in these rectangles. Experiments show that this approach delivers the best results and in a shorter time than RapidOCR alone.

Comment thread docs/ocr/index.rst
- Engines
- Notes
* - ``rapidocr_api.exec_ocr``
- RapidOCR

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
- RapidOCR
- RapidOCR or RapidOCR-ONNXRuntime

Comment thread docs/ocr/index.rst
- Notes
* - ``rapidocr_api.exec_ocr``
- RapidOCR
- Requires RapidOCR and ONNX Runtime

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
- Requires RapidOCR and ONNX Runtime

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants