Docs update for 2.0 - #5146
Docs update for 2.0#5146jamie-lemon wants to merge 14 commits into
Conversation
| OCR Adaptors | ||
| ------------ | ||
|
|
||
| By default, PyMuPDF uses **Tesseract** and **OpenCV** for image pre-processing. If you need a different OCR engine — for higher accuracy, language support, or cloud-based processing — you can plug in a custom adaptor. |
There was a problem hiding this comment.
We do not use OpenCV anymore since some. For image pre-processing, we use our own utility that is based on a LightGBM model (Gradient Boosting Decision Tree) that has been trained on 70 thousand documents of DocLayNet's ground truth data.
There was a problem hiding this comment.
@JorjMcKie Please advise, earlier under "How OCR is Triggered" we have a note:
For these heuristics to work, both a Tesseract installation and OpenCV must be available in your Python environment. If either is missing, no OCR is attempted.
Is this now simply:
For these heuristics to work, a Tesseract installation must be available in your Python environment. If it is missing, no OCR is attempted.
| ~~~~~~~~~~~~~~~~~ | ||
|
|
||
| By utlilizing the :doc:`PyMuPDF4LLM API <pymupdf4llm/api>` we are able to convert PDF to a Markdown representation. | ||
| By utlilizing the :meth:`Document.to_markdown` method we are able to convert PDF to a Markdown representation. |
There was a problem hiding this comment.
This conversion is not restricted to PDF:
By utlilizing the :meth:Document.to_markdown method we are able to convert any of the following document types to a Markdown representation:
- PDFs
- Image documents (because they are internally converted to 1-page PDFs)
- Office documents opened using PyMuPDF-Office (because they are converted to PDFs via
.to_pdf()before processing them).
| :meth:`Document.set_xml_metadata` PDF only: create or update document XML metadata | ||
| :meth:`Document.subset_fonts` PDF only: create font subsets | ||
| :meth:`Document.switch_layer` PDF only: activate OC configuration | ||
| :meth:`Document.to_json` PDF only: convert the document to JSON |
There was a problem hiding this comment.
I know I communicated it differently once. But the restriction to PDF is exclusively connected to the fact that we currently require PDF to write back OCRed text to pages that have been determined to need OCR.
We are investigating how to avoid this restriction.
But still, while this result is pending, the following document types are also eligible:
- Image documents (because we internally convert them to 1-page PDFs)
- Office documents handled with PyMuPDF-Office (because we internally deal with the
.to_pdf()output)
JorjMcKie
left a comment
There was a problem hiding this comment.
I have made a handful of comments...
|
|
||
| .. note:: | ||
|
|
||
| For these heuristics to work, both a `Tesseract installation <installation_ocr>` and `OpenCV <https://pypi.org/project/opencv-python/>`_ must be available in your Python environment. If either is missing, no OCR is attempted. |
There was a problem hiding this comment.
New suggested text:
For these heuristics to work, at least one of the supported default OCR engines, Tesseract or RapidOCR must be installed. Otherwise, the behavior is the same as if use_ocr=False was specified. If you have your own, non-default OCR engine, you must supply the respective plugin.
|
|
||
| If you want to use both OCR engines side-by-side, you can do so by implementing a custom OCR function which calls both OCR engines — one for bbox recognition (RapidOCR) and the other for text recognition (Tesseract) — and then combines their results. | ||
|
|
||
| This pre-made callable OCR function can be found in the ``PyMuPDF.ocr`` module as ``rapidtess_api.exec_ocr``. |
There was a problem hiding this comment.
| This pre-made callable OCR function can be found in the ``PyMuPDF.ocr`` module as ``rapidtess_api.exec_ocr``. | |
| This pre-made default callable OCR function can be found in the ``PyMuPDF.ocr`` module as ``rapidtess_api.exec_ocr``. |
| RapidOCR & Tesseract Side-by-Side | ||
| ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ | ||
|
|
||
| If you want to use both OCR engines side-by-side, you can do so by implementing a custom OCR function which calls both OCR engines — one for bbox recognition (RapidOCR) and the other for text recognition (Tesseract) — and then combines their results. |
There was a problem hiding this comment.
| If you want to use both OCR engines side-by-side, you can do so by implementing a custom OCR function which calls both OCR engines — one for bbox recognition (RapidOCR) and the other for text recognition (Tesseract) — and then combines their results. | |
| If both OCR engines are installed -- Tesseract and one of RapidOCR / RapidOCR-ONNXRuntime -- they are automatically used together by default: RapidOCR for text line rectangle detection and Tesseract for text recognition in these rectangles. Experiments show that this approach delivers the best results and in a shorter time than RapidOCR alone. |
| - Engines | ||
| - Notes | ||
| * - ``rapidocr_api.exec_ocr`` | ||
| - RapidOCR |
There was a problem hiding this comment.
| - RapidOCR | |
| - RapidOCR or RapidOCR-ONNXRuntime |
| - Notes | ||
| * - ``rapidocr_api.exec_ocr`` | ||
| - RapidOCR | ||
| - Requires RapidOCR and ONNX Runtime |
There was a problem hiding this comment.
| - Requires RapidOCR and ONNX Runtime | |
No description provided.