Processing Pipelines
Configuration and Processing Pipelines¶
The OpenTelemetry Collector processes telemetry data through a modular pipeline composed of receivers, processors, and exporters. This architecture allows you to filter, transform, and route data before exporting it to observability platforms. Understanding how to configure these components is essential for optimizing data flow, reducing noise, and ensuring compliance with organizational requirements.
Pipeline Structure Overview¶
A pipeline is defined in the config.yaml file and consists of:
1. Receivers: Capture telemetry data (e.g., from agents, APIs, or protocols like OTLP).
2. Processors: Modify, filter, or enrich data (e.g., dropping traces, adding metadata, or converting formats).
3. Exporters: Send data to backend systems (e.g., Prometheus, Jaeger, or cloud platforms).
Each pipeline is a sequence of stages, and you can define multiple pipelines for different use cases. For example, one pipeline might handle logs, while another focuses on traces.
# Example pipeline configuration
service:
pipelines:
traces:
receivers: [otlp]
processors: [batch, filter]
exporters: [otlpexporter]
metrics:
receivers: [prometheus]
exporters: [prometheusexporter]
Key Components and Use Cases¶
1. Receivers¶
Receivers are the entry point for telemetry data. Common receivers include:
- otlp: For OTLP protocol (used by OpenTelemetry SDKs).
- prometheus: For scraping metrics from Prometheus endpoints.
- jaeger: For ingesting Jaeger traces.
Example:
2. Processors¶
Processors perform transformations, filtering, or enrichment. Key use cases:
- Filtering: Drop traces/logs that don’t meet criteria (e.g., service.name != "production").
- Batching: Reduce overhead by grouping events (e.g., batch processor with max_queue_size).
- Log/Trace Enrichment: Add metadata like environment or user IDs.
Example:
processors:
filter:
attributes:
- key: "service.name"
value: "production"
action: "drop"
batch:
timeout: 5s
max_queue_size: 1000
3. Exporters¶
Exporters send data to observability tools. Examples:
- otlpexporter: For sending traces/metrics to an OTLP endpoint.
- prometheusexporter: For exposing metrics via Prometheus.
- logging: For debugging (logs data to stdout).
Example:
Advanced Processing: Routing and Conditional Logic¶
Use the router processor to direct data to different exporters based on attributes. For example, route traces from a specific service to a dedicated backend:
processors:
router:
routes:
- match:
service.name: "db-cluster"
exporter: "prometheusexporter"
- exporter: "otlpexporter"
This ensures metrics from the db-cluster service are exported to Prometheus, while all other traces go to the OTLP endpoint.
Note: Routes are evaluated in the order they are defined, and later routes override earlier ones if there's overlap.
Best Practices¶
- Modularize pipelines: Separate traces, metrics, and logs into distinct pipelines for clarity.
- Test configurations: Use the
otelcol-contribCLI to validate configs before deployment: - Monitor processor performance: Batching and filtering can impact latency; tune parameters like
timeoutormax_queue_sizebased on workload.
Key takeaways¶
- Pipeline stages (receivers → processors → exporters) enable flexible telemetry processing.
- Processors are critical for filtering noise, batching data, and enriching signals.
- Routing allows conditional data distribution based on attributes like
service.name. - Modular configuration and testing ensure reliability and scalability in production.