Repository navigation
Conversation
There was a problem hiding this comment.
🔵 Needs a closer look
It introduces broad observability infrastructure (runtime module-patching instrumentation, a frontend process.exit monkey-patch, OTLP/gRPC bundling into the Next.js standalone server, and new Helm/Docker/CI wiring) whose runtime behavior cannot be fully verified without building and deploying, so it warrants final human review.
0 open findings
What changed in this PR
This PR replaces the wallet's legacy Prometheus-based metrics (@prometheus-io/client) with a full OpenTelemetry (OTel) implementation. It introduces a new shared workspace package @shared/telemetry that centralizes OTel SDK setup, env-based configuration/validation, URL redaction, and a root-span-dropping sampler. The wallet backend (Express/NestJS) and frontend (Next.js) server are wired to emit traces and metrics over OTLP/gRPC to a collector. A local, opt-in observability stack (otel-collector → Prometheus + Tempo → Grafana) is added behind a Docker Compose observability profile, and the testnet-wallet Helm chart gains an in-release OTel collector (Deployment/Service/ConfigMap/ServiceMonitor) plus telemetry env wiring.
Changes:
- New
@shared/telemetrypackage: opt-in (TELEMETRY_ENABLEDdefaults false) OTel SDK bootstrap with strict zod env validation, URL redaction span processor, and aParentBased(DropRootSpan)sampler; backend and frontend consume it with their own instrumentation lists. - Local observability:
local/observability.yaml+local/config/*and newpnpm dev:observability/local:up:observability/local:down:observabilityscripts, documented inREADME.md. - Helm:
testnet-walletchart addsotelCollectortemplates,telemetryData/telemetryEndpointhelpers, telemetry config keys, and a ServiceMonitor; removes the old metrics port/service.
| File | Description |
|---|---|
README.md |
Documents opt-in observability stack, Grafana access, and new dev scripts |
pnpm-lock.yaml |
Drops @prometheus-io/client; adds OTel API/SDK/instrumentation + @grpc/grpc-js/protobuf deps and the new @shared/telemetry importer |
packages/shared/telemetry/* |
New package: config parsing, SDK/provider setup, redaction, sampler, secret paths, tests |
packages/wallet/backend/src/index.ts, telemetry.ts, app.ts, config/env.ts |
Telemetry-first import, graceful shutdown w/ flush, backend instrumentation list, telemetry env schema; metrics server removed |
packages/wallet/frontend/src/instrumentation*.ts, next.config.js |
Next.js instrumentation hook + frontend telemetry bootstrap with process.exit flush workaround |
packages/shared/backend/src/middleware/* |
Removes metrics.ts and its export |
local/observability.yaml, local/config/* |
otel-collector, Prometheus, Tempo, Grafana containers + configs/dashboards |
helm/testnet-wallet/** |
OTel collector templates, _helpers.tpl telemetry helpers, configMap/secrets/values wiring, ServiceMonitor, tests, README |
root package.json, .github/labeler.yml |
New observability scripts; labeler entry for the new package |
Files not reviewed (1)
- pnpm-lock.yaml: Generated file
🧠 Review effort: Balanced
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
Proposed changes
This PR introduces the capturing of metrics and traces using the Open Telemetry SDK in both he wallet backend and frontend (SSR). It also modifies the helm charts so that releases will start a deployment for a open telemetry collector which will in turn expose the teletry to prometheus and Grafana.
Context
In previous iterations of testnet we had endless problems tracking the root causes for performance and functional issues. This motivated me to ensure that we add metrics and traces.
Previously we added metrics by adding the prometheus sdk but after some consideration and experiments it has now become clear that we should be using the open telemetry collector pattern instead so that we can do traces in a consistent way as well.
Screenshot
Example of a trace showing the a frontend trace

Example showing a backend trace
