Skip to content

tests/test_tables.py: make the two new table tests hold with layout and pymupdf4llm imported - #5144

Merged
veget-able merged 1 commit into
mainfrom
table-tests-layout-default
Sep 29, 2026
Merged

veget-able merged 1 commit into
mainfrom
table-tests-layout-default

Conversation

@veget-able

@veget-able veget-able commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Two tests added in #5136 assumed that neither layout nor pymupdf4llm is imported. With the 2.0 setup, where import pymupdf also imports both, they fail. This PR changes the tests only; the table code is unchanged.

main now has layout-dependent expectations for these two tests (010f012, "Cope with regression"). This PR replaces them, so both tests check the same result with or without layout and pymupdf4llm, instead of expecting the caption as header or the split hyphens under layout.

test_table_header_is_computed_on_first_access()

With layout, the caption above the first table of chinese-tables.pdf is taken as an external header. This already happens with 1.28.2 and is not what the test checks. The test is about Table.header being computed lazily and only once, so it now uses the second table, whose header is its top row with or without layout.

test_find_tables_refine_splits_a_repeated_leading_header()

pymupdf4llm calls TOOLS.unset_quad_corrections(True) on import. After that, small glyphs such as "-" and "." are extracted as a line of their own, so "2024-01" comes back as "2024 01\n-". The synthetic dates now use "2024/01", which extracts the same either way. The row-count assertion also prints the extracted tables when it fails, to make any remaining difference visible.

Checked with

  • Windows 11 and Ubuntu 24.04
  • MuPDF 1.28.2 (wheel) and 1.28.6 (built from the 1.28.x branch)
  • pymupdf4llm 1.28.2 and main
  • each of: nothing imported, layout only, pymupdf4llm only, both

Both tests pass in all of these combinations. Rebased on current main, all 46 tests in tests/test_tables.py pass with the 2.0 defaults and with PYMUPDF_TEST_USE_LAYOUT=0 PYMUPDF_TEST_USE_4LLM=0.

…nd pymupdf4llm imported

test_table_header_is_computed_on_first_access() now uses the second table
of chinese-tables.pdf. With layout, the caption above the first table is
taken as an external header, which the test is not about; the second table
has its header as its top row with or without layout.

test_find_tables_refine_splits_a_repeated_leading_header() now writes
"2024/01" instead of "2024-01". pymupdf4llm sets unset_quad_corrections(True)
on import, and then "-" and "." are extracted as a line of their own. The
row-count assertion prints the extracted tables if it fails.

This replaces the layout-dependent expectations added in 010f012, so both
tests check the same result with or without layout and pymupdf4llm.
@veget-able
veget-able merged commit 336b261 into main Sep 29, 2026
3 checks passed
@veget-able
veget-able deleted the table-tests-layout-default branch September 29, 2026 00:25
@github-actions github-actions Bot locked and limited conversation to collaborators Sep 29, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants