The Practical OpenTelemetry Tracing Guide (with Code Examples)
When a user clicks "Add to Cart" on an e-commerce site, that single action can fan out across a dozen services: a frontend gateway, an auth service, a cart service, an inventory check, or a database write. If any step slows down or fails outright, the user sees a spinner or an error.
You, the product owner, often see nothing because each service logs in isolation. You know something in your application system misbehaved, but not the source of the problem nor its cause.
Distributed tracing solves this by recording the full journey of a request as it moves through your system. Every service the request touches produces a timestamped record of the work it performed: how long it took, whether it succeeded, and what it called next. These records are stitched together into a single trace: an end-to-end map of everything that happened to fulfill that request, across every service boundary.
OpenTelemetry (OTel) is the open-source, vendor-neutral standard for generating and collecting this trace data. A Cloud Native Computing Foundation graduated project, it provides the APIs, SDKs, and tooling to instrument applications in virtually any language without coupling your code to a specific observability backend.
This guide covers how OpenTelemetry's tracing signal works from the inside, explaining:
- the OpenTelemetry span data model,
- how OpenTelemetry traces are propagated across contexts,
- how OpenTelemetry APIs and SDKs work,
- what a production-grade tracing pipeline looks like, and
- some common pitfalls to be aware of.
How OpenTelemetry Tracing Works
Traces and Spans
A trace represents the complete path of a single request through a distributed system. It is identified by a globally unique TraceID (a 16-byte identifier) that every service involved in handling that request shares.
A trace is composed of spans. Each span represents a single, named, timed operation, such as an HTTP request being handled, a database query executing, a message being published to a queue. A span captures when the operation started, when it ended, whether it succeeded, and the metadata that makes debugging possible. Every span carries its own unique SpanID (an 8-byte identifier).
Together, traces and spans give you something that logs alone cannot: a causal, time-ordered view of everything that happened to a request, across every process it touched.
The Span Data Model
The real diagnostic power of a span lies in the structured data it carries. Each span records the following fields:
-
Name: A human-readable string that identifies the class of operation:
POST /api/checkout,SELECT orders,payment.authorize. Backends use span names to group and aggregate similar operations, so the name must be low-cardinality. -
Kind: Defines the span's role in a distributed interaction. There are five kinds:
SERVER: handling an inbound request (e.g., HTTP server, gRPC handler). Duration measures server-side processing time.CLIENT: making an outbound request (e.g., HTTP client call, database query). Duration measures the full round-trip including network latency.PRODUCER: sending an asynchronous message (e.g., Kafka publish). In high-throughput environments, these spans often end before the consumer begins processing the message.CONSUMER: processing an asynchronous message. The span covers the time from receiving the message to completing its handling.INTERNAL: an in-process operation that does not cross a service boundary. Use it for instrumenting important business logic.
-
Status: Marks the outcome of the operation. There are three status types:
Unset: means no status was explicitly recorded. This is the default status.Ok: explicitly confirms success.Error: marks a failure. It is what backends use to calculate error rates and trigger alerts.
-
Attributes: Key-value pairs that add queryable, structured metadata to a span. While the span name should be generic, attributes should carry the specifics:
http.request.method = POST,http.response.status_code = 500,user.id = abc-123.

OpenTelemetry defines Semantic Conventions, which are standardized attribute names for common operations. OpenTelemetry backends utilize these standards to better understand and correlate events, and to provide richer out-of-the-box analysis.
-
Events: Timestamped log messages attached to a span, marking specific moments during execution. A common use case is recording exceptions; when an error occurs, you capture the exception type, message, and optionally the stack trace as an event on the active span.
-
Links: References to spans in other traces. Unlike parent-child relationships, links describe causal connections that don't fit into a single trace tree. For example, a batch job that processes messages from multiple producers.
Trace Structure: Parents, Children, and Root Spans
Spans don't exist in isolation. Rather, they form a tree that reflects the execution flow of the request. When one operation calls another, the calling span becomes the parent and the called span becomes the child. The child stores the parent's SpanID as its ParentSpanID, creating an explicit link.
This hierarchy is what allows tracing backends to render the waterfall (or Gantt chart) view that makes distributed tracing visually powerful. You can see at a glance which operations triggered which others, how they overlap in time, and where latency accumulated.
The first span in a trace, the one with no parent, is called the root span. It represents the entry point of the request (typically an incoming HTTP call at the edge) and its duration is what the end user actually experienced.

Context Propagation
A trace that stops at a service boundary is useless. The entire point is continuity across processes, and that requires context propagation.
When a service creates a span for an outgoing request, the OpenTelemetry SDK serializes the current trace context (the TraceID, the active SpanID, and trace flags like sampling status) into the request's headers. The downstream service extracts this context, and any spans it creates automatically become children of the calling span. This is how a single logical trace spans multiple processes without any manual wiring.
The standard format for this is the W3C Trace Context specification, which defines two headers:
traceparentcarries the version, trace ID, parent span ID, and trace flags. Example:00-5b8efff798038103d269b633813fc60c-eee19b7ec3c1b174-01tracestatecarries optional vendor-specific or application-specific data.
Within a single process, context propagation is handled differently depending on the language. The Python SDK uses context variables (contextvars), which also propagate automatically across asyncio tasks. Other languages handle this based on their own capabilities and design patterns.
Regardless of the mechanism, the result is the same: when you create a new span, the SDK knows which span is currently active and sets the parent-child relationship accordingly.
For most common protocols (HTTP, gRPC, messaging systems), OpenTelemetry's instrumentation libraries handle injection and extraction automatically. You don't write propagation code by hand unless you are working with a custom transport, or utilizing OpenTelemetry in languages like Rust.
The API and SDK Split
OpenTelemetry splits its tracing surface into two layers: the Tracing API, which your application code calls to create spans and propagate context, and the Tracing SDK, which does the actual work of recording, processing, and exporting those spans.
The separation exists for a specific reason. Without an SDK present, every API call quietly returns a valid but empty result. This means that the code doesn't emit any spans, doesn't export any data and introduces no additional overhead.
This makes the API safe to depend on in shared libraries and frameworks, as the library can instrument itself without forcing its users to run any particular observability stack.
Whether tracing is actually active becomes a decision for the application owner, made at deployment time by installing and configuring the SDK.
We won't dive any deeper into the comparison here. If you're interested in learning more, we have a detailed breakdown for OpenTelemetry API vs SDK.
OTLP: The Wire Format
Once spans are recorded by the SDK, they need to be transmitted to an observability backend. The native protocol for this is OTLP (OpenTelemetry Protocol): a high-performance, vendor-neutral format designed for the three primary telemetry signals (traces, metrics, and logs).
OTLP organizes data in a three-tiered hierarchy of Resource → Scope → Data that eliminates metadata redundancy across batches. It supports transmission over both gRPC (the default, on port 4317) and HTTP (on port 4318), with Protobuf encoding for compact, schema-safe payloads.
Because OTLP is an open standard, any compatible observability backend can ingest your traces without translation or data loss. This is what makes OpenTelemetry genuinely vendor-neutral — you instrument once, and the choice of backend is a configuration change.
Generating Traces from your Applications
Understanding the data model is step one. The next step is generating trace data from your applications. OpenTelemetry provides two complementary approaches: auto-instrumentation, which covers common libraries and frameworks with minimal setup, and manual instrumentation, which lets you trace the business logic that only you know about.
Auto-Instrumentation
For web services in most popular programming languages, auto-instrumentation is the fastest path to useful traces. It works by wrapping business-critical libraries (Flask, Django, Requests, SQLAlchemy, Redis, among others) so that they automatically produce spans for every HTTP request, database query, and outbound call your application handles.
There are two ways to enable it:
Option 1: Zero-code with the opentelemetry-instrument CLI
Install the base packages and let the bootstrap command detect your installed libraries:
pip install opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap --action=installThen run your application through the opentelemetry-instrument wrapper:
OTEL_SERVICE_NAME=cart-service \
OTEL_TRACES_EXPORTER=otlp \
OTEL_EXPORTER_OTLP_ENDPOINT=https://ingest.<region>.signoz.cloud:443 \
OTEL_EXPORTER_OTLP_HEADERS="signoz-ingestion-key=<your-ingestion-key>" \
opentelemetry-instrument python app.pyWith this setup, all major API operations produce spans automatically. No application code changes are required.
Option 2: Programmatic setup
If you need finer control over which libraries are instrumented, you can configure them explicitly in code:
from opentelemetry.instrumentation.flask import FlaskInstrumentor
from opentelemetry.instrumentation.requests import RequestsInstrumentor
from opentelemetry.instrumentation.sqlalchemy import SQLAlchemyInstrumentor
FlaskInstrumentor().instrument() # All Flask routes → SERVER spans
RequestsInstrumentor().instrument() # All outbound HTTP calls → CLIENT spans
SQLAlchemyInstrumentor().instrument() # All SQL queries → CLIENT spansEvery Python instrumentation object has an uninstrument method that you can use to disable spans for specific libraries to reduce noise or potentially manage costs.
from opentelemetry.instrumentation.sqlalchemy import SQLAlchemyInstrumentor
SQLAlchemyInstrumentor().uninstrument() # Performs the uninstrumentation flowBoth approaches produce the same result: spans that capture timing, status codes, and relevant attributes for every instrumented operation. The CLI approach is simpler for getting started, while the programmatic approach gives you control over instrumentation ordering and configuration.
Manual Instrumentation
Auto-instrumentation handles the infrastructure layer (HTTP, databases, messaging) but it knows nothing about your business logic. If you need to trace domain-specific operations like payment processing, order validation, or inventory reservation, you instrument those manually using the Trace API.
The foundation is a Tracer, which you obtain once from the global OpenTelemetry instance:
from opentelemetry import trace
tracer = trace.get_tracer("cart-service")Creating a span
The simplest way to create a span in Python is with start_as_current_span, used as a context manager. The span starts when the block is entered and ends automatically when the block exits, even if an exception is raised:
def process_checkout(order_id: str, user_id: str):
with tracer.start_as_current_span("checkout.process") as span:
span.set_attribute("order.id", order_id)
span.set_attribute("user.id", user_id)
validate_order(order_id)
charge_payment(order_id)
confirm_order(order_id)Adding attributes
Attributes are how you attach the specifics that make a span useful for debugging. Use them for identifiers, status codes, and business context that you will want to filter and query later:
span.set_attribute("order.id", order_id)
span.set_attribute("payment.provider", "stripe")
span.set_attribute("payment.amount_cents", 4999)
span.set_attribute("customer.tier", "premium")Where possible, follow OpenTelemetry's Semantic Conventions for standardized attribute names. For custom business attributes, use a consistent namespace (like order.* or payment.*) to keep things organized.
Recording events
Events mark specific moments within a span's lifetime. They are useful for logging state transitions or capturing details that does not warrant a separate span:
with tracer.start_as_current_span("checkout.process") as span:
validate_order(order_id)
span.add_event("order.validated", {"validation.result": "passed"})
charge_payment(order_id)
span.add_event("payment.charged", {
"payment.provider": "stripe",
"payment.amount_cents": 4999,
})Nested spans and automatic parenting
When you create a span inside the scope of another span, the SDK automatically sets up the parent-child relationship. You do not need to pass span references between functions:
def process_checkout(order_id: str):
with tracer.start_as_current_span("checkout.process"):
validate_order(order_id) # creates a child span
charge_payment(order_id) # creates another child span
def validate_order(order_id: str):
with tracer.start_as_current_span("checkout.validate") as span:
span.set_attribute("order.id", order_id)
# validation logic...
def charge_payment(order_id: str):
with tracer.start_as_current_span("payment.charge") as span:
span.set_attribute("order.id", order_id)
span.set_attribute("payment.provider", "stripe")
# payment logic...The resulting trace tree looks like this:
checkout.process (350ms)
├── checkout.validate (45ms)
└── payment.charge (280ms)Recording errors
When an operation fails, the span should carry both the exception detail and an error status. start_as_current_span does most of this for you: if an exception propagates out of the with block, the SDK automatically records it as an exception event and sets the span status to Error. For the common case, where you let the error bubble up, you don't write any error-handling code at all:
def charge_payment(order_id: str):
with tracer.start_as_current_span("payment.charge") as span:
span.set_attribute("order.id", order_id)
# If charge() raises, the span is automatically marked Error
# and the exception is recorded as an event on the span.
return payment_gateway.charge(order_id)You only record the exception yourself when you catch it without re-raising, for example when you recover from the failure but still want it on the trace. In that case the exception never escapes the with block, so the SDK doesn't see it and you capture it manually:
from opentelemetry.trace import StatusCode
def charge_payment(order_id: str):
with tracer.start_as_current_span("payment.charge") as span:
span.set_attribute("order.id", order_id)
try:
return payment_gateway.charge(order_id)
except PaymentError as e:
span.record_exception(e) # adds the event
span.set_status(StatusCode.ERROR, str(e)) # record_exception alone won't set status
return None # handled, not re-raisedDon't do both: calling record_exception manually and letting the exception escape the with block records it twice, producing duplicate exception events on the same span.
Choosing What to Instrument
Not every function needs a span. Over-instrumentation clutters your traces with noise and adds measurable overhead in high-throughput services. Focus manual instrumentation on operations where timing and metadata provide actionable insight:
- Service entry points: incoming HTTP handlers, gRPC methods, message consumers. Auto-instrumentation typically covers these already, so verify before duplicating.
- Outbound calls: HTTP requests to other services, database queries, cache lookups. Also often covered by auto-instrumentation.
- Business-critical operations: payment processing, order validation, inventory checks, permission evaluations. These are the operations where failures and latency directly impact users, and they are invisible to auto-instrumentation.
- Long-running or expensive operations: batch jobs, report generation, file processing. Spans here help you understand where time is actually being spent.
Skip internal utility functions, data transformations, and in-memory computations unless you have a specific latency concern.
A good rule of thumb to keep in mind is that a trace with 10 meaningful spans tells you more than one with 200 trivial ones.
The Production Tracing Pipeline
The previous section showed how to generate spans from your application code. In production, those spans need to travel from your services to an observability backend reliably, efficiently, and without impacting application performance. This is where a production-grade observability pipeline becomes necessary.
The OpenTelemetry Collector
While it is possible to export spans directly from your application to a backend, most production deployments place an OpenTelemetry Collector between the two. The Collector is a standalone process that receives, processes, and exports telemetry data.
Why not export directly? A few reasons:
- Decoupling: Your application does not need to know about your backend's endpoint, authentication, or retry logic. It sends spans to a local Collector, and the Collector handles the rest.
- Processing: You can filter, transform, enrich, and sample spans before they leave your infrastructure.
- Reliability: The Collector can buffer spans during backend outages and retry failed exports, preventing data loss without blocking your application.
- Fan-out: A single Collector can forward spans to multiple backends simultaneously.
The Collector processes data through a three-stage pipeline:
-
Receivers accept incoming telemetry. The most common is the OTLP receiver, which listens on ports 4317 (gRPC) and 4318 (HTTP). You can also configure receivers for legacy formats like Jaeger or Zipkin.
-
Processors transform data in flight. Common processing steps include batching spans for efficient export, adding or removing attributes, filtering out noisy spans (like health check endpoints), and applying sampling policies.
-
Exporters send processed data to its final destination. The OTLP exporter is the standard choice for OpenTelemetry-native backends like SigNoz. You can configure multiple exporters to send the same data to different backends simultaneously.
Sampling: Controlling Volume and Cost
A high-traffic service can generate millions of spans per minute. Storing every single one is expensive and rarely necessary, since the vast majority represent normal, successful requests. Sampling lets you keep the traces that matter while controlling costs.
There are two fundamental approaches:
Head-based sampling makes the decision at the root of the trace, before any child spans are created. With a parent-based sampler (the default in the SDKs), downstream services inherit this decision through the propagated context, so a trace is either fully captured or entirely dropped. It is simple and stateless, but it is also blind: it cannot know at the time of the decision whether the trace will end up containing an error or being unusually slow.
Tail-based sampling defers the decision until the complete trace has been assembled. This means you can write rules like "keep all traces with errors" or "keep all traces slower than 2 seconds" while sampling everything else. The tradeoff is complexity: it requires a stateful component (typically the Collector) to buffer complete traces before deciding, which adds memory overhead and operational surface area.
For most teams, starting with head-based sampling at the SDK level and graduating to tail-based sampling at the Collector as volume grows is a practical path. For a detailed breakdown and configuration examples, see our guide on sampling strategies.
Configuring the SDK for Export
While the opentelemetry-instrument CLI handles most configuration through environment variables, you can also configure the SDK programmatically when you need more control:
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.resources import Resource
resource = Resource.create({
"service.name": "cart-service",
"service.version": "1.4.2",
"deployment.environment.name": "production",
})
provider = TracerProvider(resource=resource)
provider.add_span_processor(
BatchSpanProcessor(OTLPSpanExporter(
endpoint="https://ingest.<region>.signoz.cloud:443",
headers={"signoz-ingestion-key": "<your-ingestion-key>"},
))
)
trace.set_tracer_provider(provider)The BatchSpanProcessor is critical here: it queues completed spans in memory and sends them in batches rather than making a network call for every individual span. This keeps export overhead low and predictable, and is the recommended choice in production.
Environment Variables
OpenTelemetry SDKs are designed to be configured through environment variables so that you can change tracing behavior without modifying code or redeploying. The most commonly used variables include:
| Variable | Purpose | Example |
|---|---|---|
OTEL_SERVICE_NAME | Identifies your service in traces | cart-service |
OTEL_EXPORTER_OTLP_ENDPOINT | Where to send telemetry | https://ingest.<region>.signoz.cloud:443 |
OTEL_TRACES_SAMPLER | Which sampler to use | parentbased_traceidratio |
OTEL_TRACES_SAMPLER_ARG | Sampler configuration | 0.1 (10% sampling) |
OTEL_EXPORTER_OTLP_HEADERS | Auth headers for the backend (unneeded if using a local Collector) | signoz-ingestion-key=<key> |
OTEL_RESOURCE_ATTRIBUTES | Additional resource metadata | deployment.environment.name=production |
This convention means the same application image can run in development (with 100% sampling to a local Collector) and production (with 10% sampling to a remote backend) by changing environment variables at deploy time, with no code changes.
Common Pitfalls
Tracing is straightforward to set up but easy to get wrong in subtle ways. The issues below account for the majority of support questions in production tracing deployments.
Broken or Disconnected Traces
The most common problem is traces that appear fragmented: spans exist, but they show up as separate root spans instead of forming a connected tree. The usual culprit here is initializing the OpenTelemetry SDK too late.
If the SDK is configured after your application has already imported and initialized its HTTP or database libraries, the instrumentation hooks miss their window. The libraries start handling requests before the SDK can wrap them, so spans are created without trace context. The fix is to always initialize the SDK before importing your application code and frameworks.
Missing Spans
Sometimes traces are connected but incomplete. Expected spans simply do not appear.
Instrumentation library not installed. Auto-instrumentation only works for libraries that have a corresponding instrumentation package. If you use psycopg2 for PostgreSQL but only installed the SQLAlchemy instrumentor, your raw psycopg2 queries will not produce spans. Run opentelemetry-bootstrap --action=install to detect and install all matching packages.
Aggressive sampling. If your sampler is set to a low ratio (e.g., 1%), most traces are dropped before any spans are exported. During debugging, temporarily set OTEL_TRACES_SAMPLER=always_on to rule out sampling as the cause.
Export failures. Spans are generated but never reach the backend due to network issues, incorrect endpoints, or authentication problems. Check your exporter logs. The BatchSpanProcessor logs warnings when exports fail, but these are easy to miss among noisy application logs.
High-Cardinality Span Names
This was covered in the span data model section, but it is worth repeating because it is one of the most expensive mistakes in production. Span names that include variable data (user IDs, request parameters, timestamps) create an unbounded number of unique span groups in your backend. This degrades query performance, inflates storage costs, and breaks any dashboard or alert built on span name aggregations.
The fix is always the same: keep the span name generic and move variable data into attributes.
Performance Overhead
Tracing should not choke your application's performance. If the opposite is the case, then the cause is usually one of these:
Synchronous exporters. If spans are exported inline with request processing rather than batched in the background, every span adds network latency to the critical path. Always use BatchSpanProcessor in production, never SimpleSpanProcessor.
Over-instrumentation. Wrapping every internal function in a span generates noise without diagnostic value. Each span carries overhead for creation, context management, and export. Focus instrumentation on the operations listed in the "Choosing What to Instrument" section.
Unbounded attribute sizes. Attaching large payloads as span attributes (full request bodies, SQL result sets, serialized objects) inflates span size and can cause export failures. Keep attributes small and bounded.
Analyzing Traces with SigNoz
Once your application is instrumented and spans are flowing, you need an observability backend to store, query, and visualize the data. SigNoz is built natively on OpenTelemetry, which means it ingests OTLP data directly and preserves the full trace data model without translation or flattening.
Trace Explorer
SigNoz's Trace Explorer lets you run aggregated queries across your trace data, filtering by attributes like service.name, http.method, status.code, and any custom attributes you have added. This is where you go to answer questions like "what is the p99 latency for the checkout service this week?" or "which endpoints have the highest error rate?"

Flamegraphs and Gantt Charts
For individual trace inspection, SigNoz renders each trace as a flamegraph and Gantt chart. You can see the full breakdown of a request: which spans executed, in what order, how long each took, and where latency accumulated. This is the view that makes the parent-child span hierarchy practically useful for debugging.

SigNoz Cloud is the easiest way to run SigNoz. Sign up for a free account and get 30 days of unlimited access to all features.
You can also install and self-host SigNoz yourself since it is open-source. With 24,000+ GitHub stars, open-source SigNoz is loved by developers. Find the instructions to self-host SigNoz.