Boundlayer
OpenTelemetryObservabilityDistributed TracingSREDevOps

OpenTelemetry in Production: Observability Architecture for Traces, Metrics, Logs, and Cost Control

A practical guide to production OpenTelemetry architecture, semantic conventions, context propagation, sampling, cardinality, Collector reliability, SLOs, security, and cost control.

20 min read

By BoundLayer Engineering Team

BoundLayer is a senior engineering partner for SaaS, fintech, AI automation, cloud infrastructure, legacy modernization, Web3, IoT, GPU computing, and data systems.

Many engineering teams have monitoring but still cannot answer a simple production question: why did this customer workflow fail?

They may have infrastructure graphs, application logs, error tracking, and a cloud dashboard. During an incident, however, engineers still search across unrelated tools, guess which service handled the request, and discover that the important context was never recorded.

OpenTelemetry provides a vendor-neutral foundation for generating, correlating, processing, and exporting traces, metrics, and logs. It solves an important interoperability problem, but installing an SDK does not make a system observable.

A production observability platform also needs a telemetry contract, context propagation, controlled cardinality, sampling, redaction, service ownership, SLOs, useful alerts, reliable collectors, and a cost model.

At BoundLayer, we design observability around production decisions: detect customer impact, locate the responsible component, understand the cause, verify recovery, and measure whether the system meets its commitments.

This guide explains how to build that architecture with OpenTelemetry without creating an expensive stream of data nobody can use.


Observability Is an Outcome, Not a Data Volume

Telemetry is the data emitted by a system. Observability is the ability to use that data to understand the system's behavior, including failures that were not predicted in advance.

A useful platform should answer questions such as:

  • Which customer journeys are failing?
  • Did the problem begin after a deployment or configuration change?
  • Which dependency added latency?
  • Is the issue isolated to one tenant, region, endpoint, or version?
  • Are retries hiding a downstream failure?
  • Did a queue delay work or did a worker process it slowly?
  • Is database time spent waiting for a connection, a lock, or the query itself?
  • Did the recovery restore the business workflow, not only CPU usage?

Collecting every possible log and span does not guarantee these answers. It can instead increase ingestion cost, slow investigation, and leak sensitive information.

Start with the operational questions and service objectives, then design the signals needed to answer them.


What OpenTelemetry Provides

OpenTelemetry is a set of APIs, SDKs, semantic conventions, protocols, instrumentation libraries, and Collector components. It is not a storage or visualization backend by itself.

The project supports several telemetry signals:

  • traces describe the path and timing of a request or workflow;
  • metrics measure behavior over time using counters, gauges, and distributions;
  • logs record discrete events and diagnostic detail;
  • baggage propagates selected context between components;
  • profiles describe code-level resource consumption as that signal matures.

OpenTelemetry Protocol, usually called OTLP, gives applications and collectors a standard way to exchange telemetry. The Collector can receive, process, filter, enrich, sample, route, and export data to one or more backends.

This separates instrumentation from backend selection. A team can change vendors or route different environments differently without rewriting every application integration.

Vendor neutrality is valuable, but it is not zero effort. Semantic consistency and operating the telemetry pipeline remain the organization's responsibility.


Begin With Service-Level Objectives

Infrastructure metrics are necessary, but customers do not experience CPU utilization. They experience successful or failed workflows and acceptable or unacceptable latency.

Define service-level indicators for critical capabilities:

  • successful checkout rate;
  • payment authorization latency;
  • document-processing completion time;
  • API availability by endpoint class;
  • job completion before a business deadline;
  • fresh telemetry or data-pipeline lag;
  • percentage of devices reporting within the expected interval.

Then define service-level objectives, such as:

99.9% of eligible API requests succeed over 30 days
99% of interactive searches complete within 800 ms
99.5% of accepted jobs finish within 10 minutes

An error budget translates the SLO into a limited amount of tolerated failure. It gives teams a rational way to balance release speed and reliability.

Alert on rapid or sustained error-budget consumption rather than every small metric fluctuation. A page should indicate likely customer impact and require timely human action. Lower-priority conditions can create tickets, dashboards, or capacity reports.

Without SLOs, observability becomes a collection of charts with no shared definition of healthy.


A Production OpenTelemetry Architecture

A scalable design normally includes application instrumentation, local or node-level collection, a gateway tier, and one or more storage backends.

Applications, workers, databases, proxies, infrastructure
                        |
              OpenTelemetry SDKs / agents
                        |
           Local or node Collector layer
             - batching and buffering
             - host and platform context
             - local enrichment
                        |
              Collector gateway tier
             - authentication
             - redaction and filtering
             - sampling
             - routing and export
                        |
       Traces | Metrics | Logs | Archive / secondary backend
                        |
          Dashboards, SLOs, alerts, investigations

The exact topology depends on scale and runtime.

Agent pattern

A Collector runs next to the workload as a sidecar, host process, or Kubernetes DaemonSet. This reduces application-to-collector network distance and can add node or pod metadata.

Gateway pattern

A shared Collector deployment receives telemetry from many workloads. It centralizes policy, sampling, routing, and export credentials, and can scale independently.

Many environments combine both patterns. The OpenTelemetry documentation describes agent deployment as straightforward for local collection but less flexible alone for complex or evolving deployments.

Applications should usually export to a nearby Collector rather than directly to a vendor endpoint. This provides a control point for retries, batching, policy changes, and backend migration.


Instrument Business Boundaries, Not Every Function

Automatic instrumentation can capture HTTP, database, messaging, and common framework operations with little code. It is a strong baseline, especially for legacy systems.

It cannot understand the important business operations automatically.

Add manual instrumentation around boundaries such as:

  • order submitted;
  • payment authorization attempted;
  • report generation completed;
  • account provisioning failed;
  • workflow entered human review;
  • firmware rollout reached a cohort;
  • data pipeline checkpoint advanced.

Useful spans describe a meaningful operation and its outcome. They should not trace every helper function or loop iteration.

For each critical workflow, record:

  • stable operation name;
  • service and version;
  • environment and region;
  • outcome status and error type;
  • dependency name and operation;
  • retry count;
  • queue or workflow identifier where safe;
  • bounded business category;
  • relevant timing phases.

Do not use raw SQL, complete URLs, request bodies, prompts, documents, or customer records as default attributes. Useful context must remain bounded and safe.


Establish a Telemetry Contract

Without shared conventions, one service reports service, another reports app_name, and a third omits ownership entirely. Cross-service queries and alerts become fragile.

Define required resource attributes for every workload:

service.name
service.version
deployment.environment.name
cloud.region
service.namespace
team.name

Use official OpenTelemetry semantic conventions for standard technologies such as HTTP, databases, RPC, messaging, cloud, and generative AI. Add organization-specific attributes only when there is a clear query or control that depends on them.

A telemetry contract should specify:

  • attribute names, types, and allowed values;
  • ownership and service catalog identity;
  • span naming rules;
  • status and error classification;
  • metric units and histogram boundaries;
  • log severity mapping;
  • required propagation formats;
  • sensitive-data rules;
  • cardinality budgets;
  • schema version and deprecation process.

Check the contract in CI and deployment validation. It is cheaper to reject an unbounded label before production than to discover it in a storage bill.


Propagate Context Across Every Boundary

Distributed tracing works only when trace context crosses service boundaries.

For synchronous HTTP or RPC, standard instrumentation usually injects and extracts W3C Trace Context headers. Asynchronous systems require more deliberate work.

When a service publishes a message or starts a job:

  1. Capture the current trace context.
  2. Add supported context to message metadata.
  3. Create a producer span.
  4. Extract context in the consumer.
  5. Create a consumer or processing span with the correct relationship.
  6. Record queue waiting time separately from processing time.

Long-running workflows may use links instead of a strict parent-child relationship when one job aggregates several causes or processing occurs much later.

OpenTelemetry's context propagation guidance also highlights security concerns. Sanitize untrusted inbound trace headers and avoid propagating internal context to external systems unnecessarily.

Baggage is especially risky because it travels across boundaries and may be copied into spans or logs. Never place credentials, tokens, personal data, or unrestricted customer input in baggage.


Use Traces, Metrics, and Logs Together

The signals are complementary.

Metrics detect

Metrics are efficient for alerting, SLO calculation, capacity trends, and aggregate behavior. They answer "how often" and "how much."

Traces localize

Traces show where time was spent and how one workflow crossed services, queues, databases, and providers. They answer "where in this request."

Logs explain

Logs capture structured events and diagnostic details that do not belong on every span. They answer "what exactly happened here."

Correlate them using trace and span IDs. OpenTelemetry's logging model supports carrying trace context into existing logging systems, allowing an engineer to move from an alert, to an example trace, to the relevant logs without manually matching timestamps.

A practical incident flow is:

SLO alert
  -> metric dimension identifies service and operation
  -> exemplar or query opens relevant traces
  -> slow/error span identifies component
  -> correlated logs provide controlled diagnostic detail
  -> deployment metadata identifies recent change

Do not duplicate every log line as a span event. Design each signal for its strength.


Control Metric Cardinality Before Production

Metrics systems create a separate time series for each unique combination of label values. An apparently harmless label can multiply cost and memory usage dramatically.

Never use unbounded values such as these as metric labels:

  • user or account ID;
  • request or trace ID;
  • complete URL including identifiers;
  • email address;
  • SQL statement;
  • error message text;
  • session ID;
  • file or document name.

Use bounded dimensions:

  • route template such as /orders/{id};
  • HTTP method and status class;
  • service and version;
  • region and environment;
  • known error category;
  • small tenant tier or product plan set.

High-cardinality context can belong in sampled traces or structured logs with appropriate access and retention. It does not belong in a metric label merely because a dashboard tool allows it.

Track active series, label value count, ingestion rate, and cost by team. Enforce limits in instrumentation reviews and Collector processing.


Design Sampling Around Investigation Value

Keeping every trace may be affordable at low volume. At scale, traces require a deliberate retention and sampling policy.

Head sampling

The decision is made when the trace begins. It is simple and reduces processing early, but the sampler does not yet know whether the request will become slow or fail.

Tail sampling

The decision is made after spans have arrived. It can retain errors, slow traces, rare routes, selected tenants, or unusual attributes while dropping routine success traffic.

Tail sampling requires the Collector tier to temporarily hold trace data. All spans for a trace must reach a compatible sampling decision point, which affects routing, memory, scaling, and failure behavior.

A practical policy may retain:

  • all unexpected errors;
  • all traces above a latency threshold;
  • all critical workflow failures;
  • a baseline probability sample of normal traffic;
  • a higher sample for new releases or canaries;
  • a capped sample from noisy repeated errors;
  • selected synthetic tests.

The official OpenTelemetry sampling guidance notes that sampling has direct compute cost, engineering cost, and the opportunity cost of missing useful data. Treat sampling as an operating policy that needs review, not a one-time percentage.

Metrics used for SLOs should not be derived only from sampled traces unless the statistical implications are understood. Generate reliable aggregate metrics independently or through correctly designed span-to-metric processing.


Make the Collector Reliable

The Collector is part of the production data path for observability. If it drops telemetry during the incident you most need to investigate, the platform has failed.

Configure and monitor:

  • memory limits and backpressure;
  • batching;
  • bounded sending queues;
  • retry behavior;
  • exporter timeout and failure rate;
  • dropped spans, metrics, and logs;
  • receiver throughput;
  • processing latency;
  • queue utilization;
  • process CPU and memory;
  • configuration and component versions.

Scale gateways by measured throughput and failure behavior. Tail sampling, transformations, and high-cardinality attributes can make resource use different from raw event counts.

Avoid an unbounded local disk queue that silently consumes the host. Decide whether temporary loss, local persistence, or reduced sampling is the correct behavior when the backend is unavailable.

Use health checks and readiness appropriately. A Collector process can be alive while its exporter is permanently rejected by the backend.

Keep a minimal independent path for Collector health so a failure in the main telemetry pipeline can still be detected.


Redact Sensitive Data Before Export

Observability systems often have broad internal access, long retention, vendor replication, and powerful search. Treat them as sensitive data systems.

Potentially dangerous telemetry includes:

  • authorization headers and cookies;
  • API keys and database credentials;
  • request and response bodies;
  • personal and financial information;
  • document content;
  • AI prompts and retrieved context;
  • SQL parameters;
  • internal tokens in URLs;
  • secrets embedded in exception messages.

Apply data minimization in application instrumentation and enforce central policies in the Collector. Central redaction is useful, but it cannot guarantee safety if raw payloads are exported through a path that bypasses it.

Use allowlists for sensitive domains rather than trying to identify every secret after collection. Hashing is not automatically anonymization when the input domain is small or can be correlated.

Define access controls, audit logs, retention, data residency, and incident procedures for telemetry backends. Production debugging convenience does not override privacy or contractual obligations.


Control Observability Cost as an Engineering Metric

Observability cost grows from event volume, payload size, active metric series, indexing, retention, query frequency, network transfer, and backend pricing.

Attribute cost by service, team, environment, and signal. Useful unit metrics include:

  • telemetry cost per million requests;
  • bytes per request or job;
  • spans per transaction;
  • log volume per customer workflow;
  • active metric series per service;
  • cost as a percentage of infrastructure spend.

Reduce cost in this order:

  1. Stop collecting data with no operational use.
  2. Remove sensitive or verbose payloads.
  3. Fix unbounded cardinality.
  4. Aggregate repetitive metrics and logs.
  5. Apply service-specific sampling.
  6. Tier retention by value.
  7. Route debug data to lower-cost storage when justified.

Do not solve cost only by dropping all traces. A small volume of well-selected, well-instrumented traces is more useful than a large volume of incomplete ones.

Managed and self-hosted backends have different cost structures. Include engineering, on-call, upgrades, storage operations, and disaster recovery when comparing them.


Build Actionable Alerts

An alert should include enough context for a responder to begin work:

  • affected service and operation;
  • customer or business impact;
  • current and expected value;
  • start time and trend;
  • recent deployment or configuration change;
  • links to SLO, traces, logs, and runbook;
  • ownership and escalation path.

Good paging signals include:

  • fast or slow error-budget burn;
  • sustained critical workflow failure;
  • queue age exceeding a business deadline;
  • severe data freshness violation;
  • fleet or region-wide device loss;
  • inability to process or settle financial operations.

CPU or memory alerts may be useful, but only when they predict failure or require a defined action. A warning that resolves itself before a human can respond should not page someone repeatedly.

Review alerts after incidents. Remove noisy conditions, add missing context, and verify that runbooks match the real recovery process.


Tie Telemetry to Deployments

Every trace, metric, and log should identify the deployed service version. Record deployment markers and configuration changes alongside service health.

During a rollout, compare:

  • error rate by old and new version;
  • latency distribution;
  • dependency behavior;
  • resource use;
  • critical business outcomes;
  • new error categories;
  • telemetry volume and cardinality.

Use those signals for canary analysis and automated rollback only when the metrics are reliable and the thresholds are tested.

This is particularly important during database changes. Our zero-downtime PostgreSQL migration playbook shows how application versions, lock behavior, backfills, and data reconciliation need one correlated operational view.


Observability for Queues and Workflows

Request traces are not enough for asynchronous systems.

Measure each stage:

  • message accepted;
  • time waiting in the queue;
  • delivery attempts;
  • processing duration;
  • external dependency time;
  • retry delay;
  • dead-letter or terminal failure;
  • total time to business completion.

A worker can have low processing latency while customers wait an hour in a backlog. Alert on oldest-item age and completion SLO, not only consumer CPU.

Use a stable workflow or correlation ID in controlled logs and traces. Do not use it as an unbounded metric label.

For event and data pipelines, include source position, consumer lag, checkpoint age, schema failure, and reconciliation status. Our guide to Change Data Capture and real-time data pipelines explains the end-to-end correctness signals required beyond connector health.


A Practical Implementation Plan

Phase 1: Identify critical journeys

Select a small number of customer and operational workflows. Define owners, SLIs, SLOs, and incident questions.

Phase 2: Establish the contract

Define service identity, versions, semantic conventions, propagation, cardinality limits, sensitive-data rules, and minimum signals.

Phase 3: Instrument one vertical slice

Trace a workflow through API, database, queue, worker, and external dependency. Add RED metrics: rate, errors, and duration. Correlate structured logs.

Phase 4: Deploy Collectors

Introduce local collection where useful and a gateway for central policy, routing, sampling, and backend credentials. Monitor the pipeline itself.

Phase 5: Build SLOs and alerts

Create customer-impacting indicators, error-budget alerts, ownership, and runbooks. Test alerts with controlled failures.

Phase 6: Govern cost and quality

Track signal volume, cardinality, dropped data, sampling outcomes, cost by service, and contract compliance.

Phase 7: Expand deliberately

Add services and workflows by value. Retire redundant agents, dashboards, and duplicate log paths only after the replacement is proven.

This approach works well during legacy software modernization, where telemetry creates the baseline needed to prioritize changes and verify that each modernization step improved the system.


Common Failure Modes

Sending applications directly to one vendor

Backend credentials, routing, and policy become scattered across services. A Collector layer provides a controlled abstraction.

Instrumenting every function

Trace volume grows while investigations become noisier. Instrument service, dependency, and business boundaries.

High-cardinality metric labels

User IDs, URLs, and error text create explosive series growth. Enforce bounded dimensions before deployment.

Sampling only at the trace head

Rare errors and slow requests disappear with normal traffic. Evaluate tail sampling or targeted always-on traces for critical workflows.

Tail sampling without trace-aware routing

Different gateway instances see partial traces and make incomplete decisions. Route all spans for a trace consistently.

Full payload logging

Costs and security exposure increase quickly. Capture controlled metadata and retrieve source records through authorized systems when needed.

Monitoring infrastructure but not outcomes

All servers appear healthy while checkout or settlement fails. Define business SLIs.

Ignoring Collector health

Dashboards go quiet and are misread as recovery. Alert on telemetry pipeline loss and export failure.

Creating dashboards without ownership

Panels become stale and alerts have no responder. Bind services and SLOs to teams and runbooks.


How BoundLayer Can Help

BoundLayer designs and implements observability for backend, cloud, data, IoT, fintech, and AI systems.

We can help you:

  • audit existing metrics, logs, traces, dashboards, alerts, and cost;
  • define service-level indicators and objectives around business workflows;
  • instrument applications with OpenTelemetry SDKs and auto-instrumentation;
  • design Collector agent and gateway deployments;
  • implement trace context across HTTP, messaging, jobs, and workflows;
  • establish semantic conventions and telemetry governance;
  • control metric cardinality, sampling, retention, and vendor spend;
  • redact sensitive data and secure telemetry access;
  • build actionable alerts, runbooks, and incident workflows;
  • migrate between observability backends with less application coupling.

Observability is also a core part of our AWS infrastructure audit and optimization, where we review reliability, deployment, security, cost, backups, and operational readiness together.

For AI systems, the same foundation extends into model, tool, evaluation, and cost signals. Our AgentOps guide covers those additional concerns.


Final Takeaway

OpenTelemetry standardizes how telemetry moves through a system. The engineering value comes from what the organization builds on that foundation.

Define customer-facing objectives. Instrument meaningful boundaries. Propagate context across synchronous and asynchronous work. Correlate traces, metrics, and logs. Control cardinality, sampling, sensitive data, and cost. Monitor the Collector as production infrastructure.

The result should not be more dashboards. It should be faster detection, shorter investigations, safer releases, and a clear understanding of whether the software is delivering the outcome customers expect.

Need observability that helps engineers resolve real incidents?

We audit and build OpenTelemetry platforms, SLOs, tracing, metrics, logging, alerting, and cost controls around your production workflows.

Free consultation

Get a Free 30-Minute Technical Consultation

Share a few details about your project and we'll get back to you within 48 hours with a clear next step.

  • No sales pressure — a senior engineer, not a sales rep
  • Clear next step within 48 hours
  • We can sign an NDA before we talk

By submitting, you agree to be contacted about your request. We respect your privacy and can sign an NDA on request.