Skip to content

HTTP downloads pick the wrong route or never refresh #778

Description

@rtibbles

❌ This issue is not open for contribution. Visit Contributing guidelines to learn about the contributing process and how to find suitable issues.

Overview

Any URL whose HEAD response is text/html is sent to the page renderer, whatever the status. A server that rejects HEAD therefore fails, and a 404 URL "succeeds" as a snapshot of the error page. The retry-with-backoff adapter is replaced at the start of every run. YouTube thumbnails and links inside HTML5 zips are never re-fetched, even with --update.

Complexity: Medium
Target branch: main

Context

The Change

  • Routing to the page renderer should only happen on a successful HEAD response. A failed HEAD should fall through to a plain GET.
  • A 4xx or 5xx response should fail the file.
  • Downloads should use one retry policy with backoff and 429/5xx retries, sized by --download-attempts.
  • --update should re-fetch YouTube thumbnails and links inside zips.

How to Get There

Point DocumentFile at a URL that answers HEAD with a 405 text/html page and GET with a PDF.

Out of Scope

Acceptance Criteria

  • A URL answering HEAD with 405 and GET with a PDF is stored as a PDF.
  • A 404 URL fails the file, and a dead <script src> inside a zip is left as-is.
  • A 503 followed by a 200 succeeds within --download-attempts.
  • With --update, a changed YouTube thumbnail and a changed image inside a zip are re-downloaded.

AI usage

I directed the regression hunt and decided which behaviour changes were intended; Claude Code compared v0.7.3 against main, reproduced each regression with scripts, and drafted this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions