Skip to content

Axum: add safe production diagnostics and trusted-proxy handling #397

Description

@ChristianPavilonis

Part of #391.

Description

Expose enough request and lifecycle evidence for operators without returning internal details or trusting client-supplied forwarding headers. This is native runtime integration, not a replacement observability framework.

Current implementation on audited main

Prerequisite PR and scope correction

PR #275 is the prerequisite for error and response-lifecycle integration. It already introduces category-only public error rendering and response-owned exactly-once completion/observer paths. See error.rs at the inspected PR head and response.rs at the inspected PR head. The legacy source-string wire errors described in the main baseline must not be repaired in a second independent implementation.

Reuse its terminal outcome/completion contracts for native diagnostics, including aborted admissions and streams that fail after headers. Trusted incoming forwarding-header policy and operator-visible log/metrics integration remain separately scoped work.

PR #389 is additional related timing work, not the primary HTTP prerequisite. Check compatibility between its generic collector and #275's request-owned monotonic clock and egress lifecycle before using them together. Neither should introduce another clock origin or expose sensitive payloads.

Proposed solution

  1. Reuse PR feat: implement portable outbound HTTP and request lifecycle contracts #275's category-only error rendering. Add leakage tests only for uncovered production boundaries; do not implement another sanitizer or return provider source strings.
  2. Wire log/tracing facades into its admission, response completion, upstream failure, stream failure, and process lifecycle hooks. Record correlated method/route/status/duration/terminal outcome without bodies, secret values, or sensitive raw queries.
  3. Integrate shared observability work and, when available and compatible, PR feat(core): add shared request timing collector and middleware #389's timing collector with feat: implement portable outbound HTTP and request lifecycle contracts #275's monotonic clock. Use the established exactly-once completion/observer boundaries rather than another finalization framework.
  4. Define explicit trusted incoming proxies, defaulting to no trust. Normalize or ignore untrusted Forwarded/X-Forwarded-* input before effective host/scheme/client metadata is consumed; document direct adapter bypass.
  5. Preserve custom app logging ownership and document native metrics/diagnostics integration. Operator TLS termination and application authorization stay separate.

Done when

  • Re-audit the accepted PR feat: implement portable outbound HTTP and request lifecycle contracts #275 implementation and record which checks are already covered. Add only residual changes and missing production acceptance evidence.
  • A sentinel secret in conversion, router, or upstream errors never appears in the client response or ordinary request logs; internal diagnostics retain only explicitly permitted detail.
  • Request completion and failure records correlate status, duration, and request identity, including stream failures after headers are sent.
  • Custom logging ownership does not install a competing logger/subscriber.
  • Spoofed Forwarded/X-Forwarded-* headers from untrusted peers cannot change effective production metadata; accepted trusted-proxy cases still work.
  • Tests cover trusted proxy matching, header precedence, malformed forwarding values, and direct-peer fallback.
  • Operators have a documented way to observe active work, rejections, and shutdown, using the existing observability contracts.
  • No secret values or body payloads become metric labels or diagnostic fields.

Related work and boundaries

Coordinate with #92 and #118 through #122; keep those issues under their existing parents. PR #389 provides the shared request-timing collector and should be checked for compatibility with PR #275 before telemetry integration. Application authentication, rate limiting, and a new tracing vendor integration are not included.

Affected areas

Adapter -- Axum; Core (routing, extractors, middleware); Documentation.

Dependencies

Primary prerequisite: PR #275. Build residual production work on its accepted contracts, not the superseded main implementation.

Coordinate with #393 and #392 so rejection and draining events use the same request/lifecycle state.

Acceptance sequence and deployment guidance

Integrate diagnostics after the accepted PR #275 error/completion contract and coordinate #389 timing separately. Define the native logging/proxy-trust evidence consumed by #398/#399 and #400; keep app-specific privacy/exposure policy downstream.

#392 lifecycle and #396 configuration/storage are the first production runtime integration priorities. Docker/architecture preparation does not wait for a new general metrics framework. Record verified error redaction, forwarding trust, and custom logging ownership rather than claiming production readiness from a passing HTTP probe.

Implementation and relationship tracking

Implementation: draft PR #409 on issue-397-diagnostics-proxy-trust-spec, based on PR #407 above #406 and #275. The open-stack implementation exception does not satisfy accepted-stack or final-image acceptance, bypass #394/#395 experiments, or make optional PR #389 a prerequisite. GitHub Development now links this issue to PR #409, whose description includes Closes #397. The issue remains open now; the PR remains draft pending its integration and acceptance gates.

Metadata

Metadata

Labels

No labels
No labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions