Observability: OpenTelemetry (OTEL) logging + tracing + alerting #17

Closed
opened 2026-07-02 17:44:38 +00:00 by JMR-dev · 0 comments
JMR-dev commented 2026-07-02 17:44:38 +00:00 (Migrated from github.com)

Context

Cross-cutting operational ticket. Without this, a failed weekly publish run (#14/#15) or a broken ingest endpoint (#7) could go unnoticed.

Updated: use OpenTelemetry (OTEL) for both logging and tracing. The OTLP submission/export endpoint (collector/backend) is TBD — keep it configurable and decide the destination later.

Scope

  • Instrument the Worker with OpenTelemetry:
    • Traces: spans for each ingest request and for the weekly publish run (a run-level span + per-report child spans), capturing outcome (accepted / rejected / rate-limited / error; per-report success/failure).
    • Structured logs: OTEL log records for the same events (ingest accepted/rejected/rate-limited; publish per-report success/failure), correlated with trace/span IDs.
    • (Metrics optional — e.g. ingest error rate / publish counts — if cheap to add.)
  • Export via OTLP. The exporter endpoint/backend is TBD — make it configurable via a Worker secret/env (e.g. OTEL_EXPORTER_OTLP_ENDPOINT plus any auth headers as a secret), never hardcoded. Selecting the actual collector/backend is a separate follow-up decision.
  • TinyGo/Wasm constraint to evaluate: the full OpenTelemetry-Go SDK may be too heavy or incompatible with the TinyGo→Wasm Worker build. Assess this; if it doesn't fit, use a minimal OTLP/HTTP exporter (lightweight or hand-rolled, OTLP-JSON or -protobuf over HTTP) that runs under TinyGo. Document the choice and rationale.
  • Alerting: driven from the chosen OTEL backend (e.g. alert rules on publish-run-failure spans or elevated ingest error rate) once the endpoint is decided; wire and document it then.

Acceptance criteria

  • Ingest (#7) and the weekly publish run (#14/#15) emit OTEL traces + structured logs over OTLP to the configured endpoint.
  • A failed publish run or a spike in ingest errors is visible in the telemetry and can trigger the chosen alert path in the backend.
  • The OTLP endpoint + credentials are configuration/secrets, not committed.
  • Unit tests cover the instrumentation wiring (e.g. spans/log records emitted for accepted/rejected/error paths) using an in-memory/test exporter.

Dependencies

Best added once #7 and #14 exist to instrument, but OTEL setup + logging/tracing conventions can be scaffolded earlier. OTLP endpoint/backend selection is a separate decision (TBD).

## Context Cross-cutting operational ticket. Without this, a failed weekly publish run (#14/#15) or a broken ingest endpoint (#7) could go unnoticed. **Updated:** use **OpenTelemetry (OTEL)** for both logging and tracing. The OTLP submission/export endpoint (collector/backend) is **TBD** — keep it configurable and decide the destination later. ## Scope - Instrument the Worker with **OpenTelemetry**: - **Traces:** spans for each ingest request and for the weekly publish run (a run-level span + per-report child spans), capturing outcome (accepted / rejected / rate-limited / error; per-report success/failure). - **Structured logs:** OTEL log records for the same events (ingest accepted/rejected/rate-limited; publish per-report success/failure), correlated with trace/span IDs. - (Metrics optional — e.g. ingest error rate / publish counts — if cheap to add.) - **Export via OTLP.** The exporter **endpoint/backend is TBD** — make it configurable via a Worker secret/env (e.g. `OTEL_EXPORTER_OTLP_ENDPOINT` plus any auth headers as a secret), never hardcoded. Selecting the actual collector/backend is a separate follow-up decision. - **TinyGo/Wasm constraint to evaluate:** the full OpenTelemetry-Go SDK may be too heavy or incompatible with the TinyGo→Wasm Worker build. Assess this; if it doesn't fit, use a minimal OTLP/HTTP exporter (lightweight or hand-rolled, OTLP-JSON or -protobuf over HTTP) that runs under TinyGo. Document the choice and rationale. - **Alerting:** driven from the chosen OTEL backend (e.g. alert rules on publish-run-failure spans or elevated ingest error rate) once the endpoint is decided; wire and document it then. ## Acceptance criteria - Ingest (#7) and the weekly publish run (#14/#15) emit OTEL **traces + structured logs** over OTLP to the configured endpoint. - A failed publish run or a spike in ingest errors is visible in the telemetry and can trigger the chosen alert path in the backend. - The OTLP endpoint + credentials are configuration/secrets, not committed. - Unit tests cover the instrumentation wiring (e.g. spans/log records emitted for accepted/rejected/error paths) using an in-memory/test exporter. ## Dependencies Best added once #7 and #14 exist to instrument, but OTEL setup + logging/tracing conventions can be scaffolded earlier. OTLP endpoint/backend selection is a separate decision (**TBD**).
Sign in to join this conversation.