How to Build for Observability Without Turning Everything Into a Metric

I’ve watched teams treat observability like a scavenger hunt. They instrument every function, log every variable, and ship metrics for things nobody will ever look at. The end result? A system that’s technically observable but practically a mess. You get dashboards that look slick in a demo but leave you squinting at a screen at 2 a.m. with no real answers. The skill isn’t hoovering up more data—it’s knowing what to ignore.

Observability isn’t really about metrics. It’s about questions. When something breaks, you need to poke your system with arbitrary questions you didn’t think of beforehand. That’s the definition I lean on, lifted from the control theory folks who first kicked the term around. Metrics are just one way to answer those questions, and honestly, they’re often the wrong tool for the job. Let’s walk through how to build systems that are genuinely observable without drowning in a sea of numbers.

Start With the Questions, Not the Data

Most teams do this backwards. They grab a monitoring tool, skim the docs, and start blasting out metrics for every knob the library exposes. Request counts, error rates, latency percentiles—the usual greatest hits. Then they plaster dashboards over those metrics and pat themselves on the back. The snag is, this approach pretends you already know what’s going to break. You don’t.

Instead, sit with your team and list the questions you’ve actually asked during incidents over the last six months. Not hypothetical ones—real, sweaty-palms questions. Stuff like: “Which users got hit by the payment timeout?” or “What was the exact call chain that caused that deadlock?” or “Why did the cache miss rate go nuts for this one key pattern?” Those questions tell you what data you need, and just as importantly, what data you can ditch.

I once worked on a system where we logged every database query with full parameters. Felt smart at the time—until we noticed we were sitting on terabytes of logs nobody ever searched. The only question we kept asking was “Which queries are slow right now?” We tore out the full-parameter logging, kept the query templates and durations, and our log volume tanked by 80% without losing any real debugging muscle. That’s the kind of tradeoff you make when you lead with questions.

Team collaborating around a whiteboard with system architecture diagrams

Structured Events Beat Metric Aggregates

Metrics are summaries. They take a stream of events and squash them into a single number: a counter, a gauge, a histogram bucket. Handy for spotting trends, sure—but they shred the context you need when you’re actually debugging. If your error rate jumps from 0.1% to 5%, a metric tells you that it happened. It doesn’t whisper why.

Structured events keep the context alive. An event is just a record of something that happened, with fields that describe it: timestamp, user ID, request ID, service name, error type, whatever actually matters. You can spit them out as log lines, but the trick is they’re structured—JSON or something close—so you can query them later. When that error rate spikes, you can filter events by error type, user segment, or upstream dependency and spot the pattern in minutes instead of hours.

I’m not saying toss metrics in the bin. They’re fine for dashboards and alerts. But treat them as derived views, not the source of truth. Emit structured events from your code, then roll them up into metrics downstream if you need to. That way you keep the raw material for investigation while still getting the at-a-glance stuff. Honeycomb has pushed this approach hard, and their observability engineering manifesto is worth a read if you want the full argument.

What Makes a Good Structured Event?

A good event is wide, not dense. Lots of fields, but each field is simple. One event might carry 20 keys, but every value is a string, number, or boolean—never a nested blob you have to unpick later. This keeps querying fast and predictable. Don’t be shy about high-cardinality fields like user IDs and request IDs; modern observability backends are built to chew on those. The fields you pick should map straight to the questions you listed earlier. If a field doesn’t help answer any of those questions, axe it.

One trap I see constantly is teams adding fields “just in case.” They log the whole request body, all headers, environment variables, stack traces for every call. That’s not observability—that’s digital hoarding. It burns cash on storage, slows your queries to a crawl, and makes the signal harder to find. Be brutal. If you haven’t needed a field in three months, stop emitting it.

Close-up of a developer typing code on a laptop with multiple screens showing logs

Traces Are Your Best Friend (If You Use Them Right)

Distributed tracing gets a lot of hype, but most setups I see are surface-level. Teams drop in an auto-instrumentation library, admire the pretty waterfall charts, and figure they’re done. That’s the bare minimum. The real juice in tracing comes from custom spans and attributes that map to your actual business logic.

Auto-instrumentation hands you the skeleton: HTTP calls, database queries, queue operations. Fine for spotting slow services. But when a user reports a bug, you need to trace their specific path through your system. That means adding spans for things like “validate_coupon_code” or “calculate_shipping_tax” and attaching attributes like the coupon code itself or the user’s loyalty tier. Now you can search traces for every user who touched a specific coupon and see exactly where it fell over.

Don’t trace everything. Pick the paths that matter—the ones that make you money or lose you customers when they break. For an e-commerce site, that’s checkout, payment, and order fulfillment. For a SaaS product, it’s login, feature access, and data export. Instrument those paths deeply with custom spans. For internal batch jobs that run once a day, a few log lines might be plenty. You’re not chasing 100% tracing coverage; you’re making the important bits transparent.

Sampling Without Guilt

If you trace every request, you’ll bankrupt yourself on telemetry storage. Sampling isn’t a compromise—it’s a deliberate design choice. The trick is to sample with a brain. Keep 100% of traces for errors and high-latency requests, but only 10% of successful fast requests. That way you hang onto the interesting data while keeping costs from going haywire. Tail-based sampling, where the keep/drop call happens after the request finishes, is the ideal because you can decide based on the outcome. If your tracing backend doesn’t support it, push for it or find one that does.

I’ve seen teams freak out about sampling because they’re scared of missing a rare bug. But if a bug is so rare it only shows up in 0.01% of traces, you won’t catch it by staring at dashboards anyway. You’ll catch it when a user complains, and then you can search by their request ID. That’s why you log request IDs in your events and traces—so you can always pull the full trace for a specific incident, even if it wasn’t sampled at first.

Server racks with blinking lights in a data center

Dashboards Are a Crutch, Not a Solution

Walk into any ops center and you’ll see walls plastered with dashboards. Most of them are dead weight. They show metrics that never budge, or they’re so cluttered you can’t pick out an anomaly if it slapped you. A dashboard only earns its keep if it helps you answer a specific question faster than digging through raw data. If you’re building a dashboard for “general health,” you’re building wall art.

I cap my teams at three dashboards per service: one for business metrics (signups, purchases, whatever the product folks care about), one for SLOs (error budgets, latency against targets), and one for debugging (request rates by endpoint, error breakdowns, dependency health). That’s the lot. If someone wants a fourth, they have to explain which question it answers and why the existing three don’t cover it. Usually, they realize they just need a saved query, not a whole new dashboard.

The debugging dashboard is the one that tends to balloon. People pile on charts for every metric they can dream up. Push back on that. Instead, make the debugging dashboard a launchpad for exploration. Include links to log queries, trace searches, and the underlying events. Teach your team to start with the dashboard to spot anomalies, then jump into the raw data to investigate. The dashboard is the map, not the territory.

Alert on Symptoms, Not Causes

This is an old SRE principle that teams still fumble. An alert should fire when users are feeling pain—high error rates, sluggish responses, dropped requests. It should not fire because CPU is above 80% or memory is running low. Those are possible causes, not symptoms. If your CPU is pegged but users are humming along fine, nobody needs to be woken up. Fix it during office hours.

Symptom-based alerting forces you to think about what actually matters. It also cuts down alert fatigue, which is the silent killer of operational chops. When every alert is a real user-impacting problem, people pay attention. When half your alerts are false alarms, people start tuning out all of them—including the real ones.

To pull this off, you need clear service-level objectives (SLOs). Define what “good enough” looks like for each service: 99.9% of requests finish in under 200ms, or 99.5% of checkouts succeed. Then alert when you’re chewing through your error budget too fast. This ties observability straight to user experience and gives you a sane basis for deciding when to dig in.

Log Levels Are a Poor Man’s Observability

I’ve grown pretty sour on traditional log levels: DEBUG, INFO, WARN, ERROR. They’re too blunt and too subjective. One developer’s WARN is another’s DEBUG. They also nudge people toward the bad habit of flipping log levels at runtime to “get more detail,” which means you’re shipping data you normally suppress—data that might be exactly what you need during an incident but isn’t there because you turned it off.

Instead, emit all your structured events at a single level and use fields to separate severity. Add a field like event_type with values request, error, warning, debug. Then you can filter on that field when querying, but you never lose data. Storage is cheap enough that the cost of keeping debug events is usually trivial next to the cost of not having them when you’re in a pinch. If volume is a worry, sample the debug events aggressively but keep 100% of errors.

This approach also simplifies your logging config. No more per-class log level fiddling, no more arguments about whether something is WARN-worthy. Just emit the event with the right type and move on. Your future self, knee-deep in an incident at midnight, will be grateful you kept everything.

Testing Your Observability Setup

You wouldn’t ship code without tests. Don’t ship observability without testing it either. The best way is through chaos engineering, but you don’t need to go full Netflix. Start with game days: block out a two-hour session where you inject failures into a staging environment and have the on-call engineer debug them using only the observability tools. No SSH, no direct database access—just dashboards, logs, and traces.

If they can’t find the root cause within 30 minutes, your observability is broken. Note what data they asked for but couldn’t get, and add it. Note what data they had but never touched, and think about removing it. Run these game days monthly, and you’ll keep refining your setup based on real use, not guesswork.

One team I coached discovered during a game day that their trace sampling was dropping all traces for a specific error path because the error happened after the sampling decision. They switched to tail-based sampling and suddenly had full visibility into that failure mode. They would never have spotted that by just staring at dashboards.

Frequently Asked Questions

How do I convince my team to stop over-instrumenting?

Show them the bill. Calculate how much you’re spending on telemetry storage and query capacity, then map that to the percentage of data that’s actually queried. Most teams find that 90% of their telemetry data is never accessed. That’s a pretty compelling argument for trimming back. Also, run a game day and highlight how the noise slows down debugging. People change habits when they feel the pain firsthand.

What’s the minimum observability I need for a new service?

Structured request logs with request ID, user ID, service name, operation, latency, and error details. Distributed traces for all inbound and outbound calls, with custom spans for critical business logic. A single SLO-based dashboard with error rate and latency against targets. That’s enough to debug most issues and know when things are broken. Add more only when a real incident exposes a gap.

Should I use OpenTelemetry or stick with vendor-specific agents?

Use OpenTelemetry. It’s becoming the standard for a reason: it decouples your instrumentation from your backend. You can switch from one observability vendor to another without rewriting your code. The auto-instrumentation is solid for most languages, and the API for custom spans and events is clean. Vendor agents lock you in and often lag behind on features. The only exception is if you’re all-in on a vendor that provides significant value beyond telemetry collection, but even then, I’d push for OTel as the collection layer.

Observability is a practice, not a product. It’s about making your system’s behavior explainable without precomputing all the explanations. Metrics have their place, but they’re just the tip of the iceberg. The bulk of your observability investment should go into structured events, thoughtful tracing, and the discipline to ask questions before you instrument. Do that, and you’ll spend less time staring at dashboards and more time fixing real problems.