
❌ This issue is not open for contribution. Visit Contributing guidelines to learn about the contributing process and how to find suitable issues.

Overview
Any URL whose HEAD response is text/html is sent to the page renderer, whatever the status. A server that rejects HEAD therefore fails, and a 404 URL "succeeds" as a snapshot of the error page. The retry-with-backoff adapter is replaced at the start of every run. YouTube thumbnails and links inside HTML5 zips are never re-fetched, even with --update.
Complexity: Medium
Target branch: main
Context
The Change
- Routing to the page renderer should only happen on a successful HEAD response. A failed HEAD should fall through to a plain GET.
- A 4xx or 5xx response should fail the file.
- Downloads should use one retry policy with backoff and 429/5xx retries, sized by
--download-attempts.
--update should re-fetch YouTube thumbnails and links inside zips.
How to Get There
Point DocumentFile at a URL that answers HEAD with a 405 text/html page and GET with a PDF.
Out of Scope
Acceptance Criteria
AI usage
I directed the regression hunt and decided which behaviour changes were intended; Claude Code compared v0.7.3 against main, reproduced each regression with scripts, and drafted this issue.
❌ This issue is not open for contribution. Visit Contributing guidelines to learn about the contributing process and how to find suitable issues.
Overview
Any URL whose HEAD response is
text/htmlis sent to the page renderer, whatever the status. A server that rejects HEAD therefore fails, and a 404 URL "succeeds" as a snapshot of the error page. The retry-with-backoff adapter is replaced at the start of every run. YouTube thumbnails and links inside HTML5 zips are never re-fetched, even with--update.Complexity: Medium
Target branch: main
Context
transfer.py:462-483routes on the HEADContent-Typewithout checking the status (af1a125, 8634293). A 405 on HEAD fails a PDF that GET would have served. A 404 becomes an HTML snapshot, and a dead<script src>inside a zip is rewritten to that snapshot.commands.py:67-72re-mounts a plainHTTPAdapter(max_retries=download_attempts)over theRetryadapter fromconfig.py:219-228. That adapter was added in 6cb6bf8 (Download external resources for archive content types (#233) #690) to replace the retriesdownloader.pyused to do, and thecommands.pyre-mount was missed.chefs.py:904andarchive_assets.py:180don't passskip_cache=config.UPDATE. v0.7.3 re-fetched YouTube thumbnails on every run.The Change
--download-attempts.--updateshould re-fetch YouTube thumbnails and links inside zips.How to Get There
Point
DocumentFileat a URL that answers HEAD with a 405text/htmlpage and GET with a PDF.Out of Scope
Acceptance Criteria
<script src>inside a zip is left as-is.--download-attempts.--update, a changed YouTube thumbnail and a changed image inside a zip are re-downloaded.AI usage
I directed the regression hunt and decided which behaviour changes were intended; Claude Code compared v0.7.3 against main, reproduced each regression with scripts, and drafted this issue.