Don't stuff everything into metrics: log dedupe can cut costs a hundredfold
Of the three signals — logs, metrics and traces — logs are the easiest to pile up meaninglessly; using log drain for templating and then pairing it with log dedupe can cut log volume a hundredfold without losing fidelity.
The video won't play here. Listen to the audio instead:
The argument · timestamps estimated from transcript position
Duplicate logs aren't fidelity, they're a bill
Mike Goldsmith puts the starting point of cost control on ‘redundant information’: the same seemingly identical record sent a hundred thousand times adds no value — you only need to know how often it happens, then extrapolate. But whether the data lands in a backend analytics system or a local system inside your infrastructure, you pay to process, store and analyze it. So his goal isn't to cut fidelity, but to keep ‘what matters and why it matters’ while no longer keeping an exact copy of every occurrence. Martin Thwaites frames this as an iron triangle: you want good telemetry, you want it fast, you want it cheap, and you're constrained by memory and CPU at the same time. Mike adds a counterintuitive note: if cost control is done too crudely, the cost of running that cost control itself can be about the same as just sending all the data.
— Mike GoldsmithThe biggest trap is using the wrong signal
Ken Rimple asks what the most common pain point is, and Mike says ‘they're using the wrong telemetry signal’. His observation is that teams are overly metrics-driven: metrics can tell you something happened and how many times, but because of the way they're collected, they can't give you the context of why it happened. Logs and traces are the signals that tell the why. Martin adds a distinction: little M metrics are the graphs you draw and the visualizations you stare at, and you always need those; big M metrics are pre-conceived time-series aggregations that serve ‘known problems’. If you have a dashboard that needs to be produced in real time, metrics are the right thing to use; but to see what happened inside a single user request, or what different requests have in common, you need logs and traces.
— Mike Goldsmithlog drain splits logs into a template plus variables
log drain comes from an old Python library and a research paper, was ported to Go and made it into the Collector. It analyzes each incoming log: which parts are variables and which are static, and after a learning period it generates a template. For example, ‘user Mike logged in at IP address 192.168...’ — the template turns Mike and the IP into two placeholders, with the rest static text. After that, when Martin logs in or Ken logs in, the template doesn't need to change, only the two variables differ. Once the variables are extracted into attributes, you can make informed trade-offs: I don't care about this user, I don't care about this internal IP range. Martin suggests it should be renamed ‘AI log templating’ so it sells better, and maybe even raises VC money.
— Mike GoldsmithTemplate plus dedupe divides log volume by a hundred
log dedupe used on its own has a precondition problem: if the log body is mixed with variables like Mike, Ken and Martin, they all look different, and it's hard to tell which ones are the same. Pair the template generated by log drain with log dedupe, and you can say ‘this log shape happened a hundred times’, so you keep one copy with a count attached, and log volume drops to one percent with no loss of fidelity. Martin immediately raises a counterexample: failed logins are usually compliance logs, must all be kept, and can't be deduped. Mike's answer is that the template already tells you what's in the placeholders, so you can filter and route between the two, or add attributes like ‘login status’ into the dedupe logic and retain at different rates.
— Mike GoldsmithDon't send evidence you must keep into the dedupe pipeline
Martin follows up on whether reroute means you can split to different destinations, and Mike confirms: with log template plus a routing connector, you make decisions based on the attributes extracted from the template — successful logins go through log dedupe and are kept at one in a thousand, because you don't care; failed logins bypass dedupe entirely and go straight to S3 or another store. His principle is blunt: log dedupe is a tool for actively reducing volume, and if you know something must be kept, don't let it go through that pipeline. From this Martin derives the key structure of cost control: hot storage holds only the operational analytics data needed to answer questions quickly, while another pipeline keeps one hundred percent of the data for forensic analysis. He also stresses that using a managed backend doesn't mean SREs and platform teams don't have to touch the Collector — quite the opposite, cleaning up telemetry and implementing log dedupe and log drain is exactly their new job.
— Mike GoldsmithLogs that should have been metrics: turn them into metrics
The signal to metrics connector allows cross-signal conversion inside the Collector: if a log is describing something that would be better described as a metric, convert it into a metric, which has a much lower collection interval, then filter out the original log. Mike says this may be a capability many people don't know about — the same information and context can move between signals. For metrics themselves, his warning is not to over-reduce cardinality — that damages the fidelity metrics can give you, and ‘aggregate on aggregate, and you're probably going to have a bad time’. The interval processor, by contrast, is safe: it aggregates data scraped every second or every five seconds into thirty- or sixty-second windows, with the same total and fidelity, just at a lower frequency. Martin's position is that if you want second-level granularity, you should use sampled traces and logs, not metrics.
— Mike GoldsmithHead sampling eats rare but important paths
OpenTelemetry currently has three sampling strategies. Head sampling is usually done at the SDK layer, taking a fixed percentage and not looking at what happens inside the trace at all: great for cost control, but the important yet rare paths in your system get eaten whole by the sampler. The tail sampling processor is an improvement — it gets the full trace context and decides after the trace ends — but it doesn't have enough knobs, especially for high-cardinality attributes: values that change frequently like user ID, session ID, wallet ID, where you can't guarantee every variation is represented. That's exactly the problem the dynamic sampler is meant to solve.
— Mike GoldsmithBucket by fingerprint, keep at least one of every variation
The dynamic sampler uses Refinery's algorithm: you tell it the attributes you care about — service name, user ID, wallet ID, session ID — and whenever a combination of those attributes appears, it gets its own bucket and its own sampling, with a guarantee that each bucket keeps at least one trace. As the volume in a bucket grows to ten, a hundred, a thousand, the threshold automatically rises, meaning fewer are kept. You can go by fixed percentage (say about ten percent) or by throughput (how many spans per minute). Martin gives the example of a checkout flow: ten different loyalty or discount algorithms, one of whose paths is only taken one time in twenty — if you sample at ten percent or even one percent because volume is too high, that path is essentially impossible to catch; put it into fingerprint buckets and every algorithm keeps at least one trace.
— Mike GoldsmithWhat sampling really trades is confidence in your answers
Mike frames the core tension of sampling strategy as: how confident are you in the answers you give based on this telemetry. Head sampling is basically hopeless with low-probability events and a high sampling threshold, unless you go hunting specifically; tail sampling is theoretically workable but has a very high bar — you have to fully organize every policy and understand how each variation is applied. The dynamic sampler lowers complexity and raises confidence, because you know every combination will keep at least one trace, which gives you a foothold to start answering questions from. On release status, it's already available in the Honeycomb Collector distribution, and the official OpenTelemetry Contrib distribution is still a few weeks away for final polish. Martin notes it's alpha right now, with breaking changes coming in the next few weeks, and suggests waiting for beta or general availability.
— Mike GoldsmithIn their own words · checked verbatim
So when you've got redundant information, you're having something that looks identical, and you'll send it 100,000 times, and it doesn't have a lot of value.
Mike Goldsmith0:00
If you do some cost controls just really heavy-handedly, the cost of running that cost control could not be that much different from just sending all of that data.
Mike Goldsmith0:00
It doesn't give you context of why something happened. It tells you that something happened and how many times that thing happened.
Mike Goldsmith6:30
I can have one copy of it and tell you that it happened 100 times. And those two, in that scenario there, it's very simplistic, but you've reduced the log volume by 100 times and you haven't lost any fidelity.
Mike Goldsmith16:00
Log Dedupe is a tool that is used to intentionally reduce volume. If you know that you need to keep something, don't pass your data through that.
Mike Goldsmith20:30
If you aggregate and aggregate, you're probably in for a bad time.
Mike Goldsmith24:30
With the dynamic sampler, it reduces the complexity and increases your confidence because you know that you're going to get variations of everything.
Mike Goldsmith31:30
I think a lot of people have this almost misconception that instrumentation and a telemetry pipeline is something that you can just build once and forget. And that's not the case.
Mike Goldsmith35:30
Figures
| Number of OpenTelemetry sampling strategies | 3 (head sampling, tail sampling processor, dynamic sampler) | 29:00 |
| Log volume reduction from log dedupe paired with templates | 100x reduction | 16:00 |
| interval processor aggregation window | 30 seconds or 60 seconds | 24:30 |
| When the dynamic sampler becomes available in the official Contrib distribution | expected in the coming weeks | 35:30 |
Glossary
- log drain
- A Collector component that analyzes logs, distinguishes static from variable parts, generates templates and extracts the variables into attributes.
- log dedupe
- A processor that merges logs of the same shape into one entry with an occurrence count, used to actively reduce log volume.
- tail sampling
- A sampling approach that waits until the trace ends and decides whether to keep it based on the full trace context.
- head sampling
- Deciding whether to keep a trace at its start by a fixed percentage, without looking at what happens inside the trace.
- cardinality guardian
- A community-contributed Collector component used to reduce metric cardinality.
- Refinery
- Honeycomb's long-standing proprietary sampling tool, whose algorithm the dynamic sampler ports.
How to listen
SREs, platform engineers and observability leads already running the OpenTelemetry Collector in production and squeezed by telemetry bills or log volume.
The opening small talk about ‘enjoying chewing on barbed wire’ can be skipped.