The short answer
Quick answer: Logs, metrics and traces are three kinds of data that a running system emits about itself. Metrics are numbers measured over time, such as requests per second or error rate. They are cheap, good for dashboards and alerts, and tell you that something is wrong. Logs are timestamped records of individual events. They carry detail and tell you what happened. Traces follow a single request as it passes through every service it touches, showing where the time went or where it failed. Used together, linked by shared identifiers, they let you move from "the site is slow" to the exact cause.
Why you cannot just attach a debugger
On your own machine you can pause the program and inspect it. In production:
- There are many servers, and you do not know which one handled the failing request.
- A single user action may pass through a dozen services.
- The problem may be intermittent and gone before you look.
- You cannot stop a live system to examine it.
So the system has to report on itself continuously. That reporting is called telemetry, and a system you can understand from its telemetry is said to be observable.
Monitoring asks questions you prepared in advance: is the error rate above 1%? Observability is the ability to ask new questions about problems you did not predict, without shipping new code to answer them.
The OpenTelemetry project calls these data types signals.
Logs
A log is a record of a discrete event, with a timestamp.
Unstructured, the traditional kind:
2026-10-04 14:32:07 ERROR Payment failed for order 8812: card declined
Structured, usually JSON:
{
"timestamp": "2026-10-04T14:32:07.412Z",
"level": "error",
"service": "payments",
"message": "payment failed",
"order_id": 8812,
"reason": "card_declined",
"trace_id": "4bf92f3577b34da6a3ce929d0e0e4736"
}
Structured logs can be searched and filtered by field: "all errors for order 8812", or "all log lines for this trace". Prefer them.
Levels indicate severity: debug, info, warning, error. Production systems usually record info and above.
Strengths: rich detail about one specific event.
Weaknesses: high volume, expensive to store and search, and hard to summarise.
Guidelines:
- Log events that matter, with context, not every line executed.
- Send logs to a central system, since individual servers come and go.
- Never log secrets or sensitive personal data: passwords, tokens, card numbers.
- Set retention periods; logs are costly.
Metrics
A metric is a numeric measurement recorded at intervals. Each data point is a number, a timestamp and a few labels.
| Type | Meaning | Example |
|---|---|---|
| Counter | A value that only increases | Total requests served |
| Gauge | A value that goes up and down | Memory in use, queue length |
| Histogram | The distribution of values | Request duration |
Strengths: compact and cheap, because data is aggregated. Fast to query. Ideal for dashboards, trends and alerts. Metrics are also what drives auto scaling.
Weaknesses: no detail about individual requests.
Why percentiles, not averages
An average hides the experience of the unluckiest users. If 99 requests take 100 ms and one takes 10 seconds, the average is about 200 ms, which looks fine. But one user in a hundred waited 10 seconds. Percentiles capture that tail: the 99th percentile (p99) is the time within which 99% of requests finished. Watch p50, p95 and p99.
Cardinality
Each unique combination of label values creates a separate time series. Labels with many possible values, such as user ID or request ID, cause the number of series to explode and can overwhelm the metrics system. Keep labels to values with a small, bounded set: service, endpoint, status code, region.
What to measure
Google's SRE book chapter on monitoring distributed systems recommends four golden signals for any user-facing service:
| Signal | Question |
|---|---|
| Latency | How long do requests take? (Track successes and failures separately) |
| Traffic | How much demand is there? |
| Errors | What fraction of requests fail? |
| Saturation | How full is the service: CPU, memory, connections, queues? |
Two similar shorthand methods: RED (rate, errors, duration) for services, and USE (utilisation, saturation, errors) for resources such as CPUs and disks.
Traces
In a system made of many services, a slow request might be slow anywhere. See monolith vs microservices. A distributed trace records the whole journey of one request.
- A trace represents the complete request.
- It is made of spans. Each span is one unit of work: an HTTP call, a database query, a function. A span has a name, a start time, a duration and attributes.
- Spans nest: a parent span contains the child operations it triggered.
GET /checkout 820 ms
├── auth-service: verify token 15 ms
├── cart-service: get cart 40 ms
│ └── redis: GET cart:42 2 ms
├── payment-service: charge 710 ms
│ ├── postgres: SELECT customer 8 ms
│ └── card-provider: POST /charges 690 ms <- the slow part
└── email-service: enqueue receipt 12 ms
The picture answers at once what would otherwise take hours of reading logs: the delay is in the external card provider.
How it works
When a request enters the system, it is given a unique trace ID. Every service passes that ID to the next one, in an HTTP header (the W3C standard is traceparent). This is context propagation. Each service records its spans tagged with the trace ID, and the tracing system assembles them.
Sampling
Recording every request in full is expensive, so systems sample: keep a percentage, or decide after the fact to keep all traces that were slow or failed.
Strengths: shows causality and where time is spent across services.
Weaknesses: needs every service to be instrumented and to pass the context along; sampled data may miss a rare case.
How they work together
| Metrics | Logs | Traces | |
|---|---|---|---|
| Tells you | Something is wrong, and how much | What happened in detail | Where in the request path |
| Data | Aggregated numbers | Individual events | Per-request timelines |
| Cost | Low | High | Medium, with sampling |
| Best for | Alerts and dashboards | Root-cause detail | Latency and dependencies |
A typical investigation:
- An alert fires: the checkout error rate is 5%. (Metric)
- A dashboard shows it started at 14:30, just after a deployment, and only in one region. (Metrics)
- You open a trace of a failed request and see the payment service timing out when calling the database. (Trace)
- You pull the logs for that trace ID and find "connection pool exhausted". (Logs)
The link between them is correlation: put the trace ID in every log line, and attach example trace IDs to metrics, so you can jump from one to the next.
OpenTelemetry
For years, every monitoring vendor had its own agents and libraries. OpenTelemetry is an open, vendor-neutral standard for producing all three signals. You instrument your code once, often automatically through libraries for common frameworks, and send the data to whichever back end you choose. Popular open-source back ends include Prometheus for metrics, Grafana for dashboards, and Jaeger or Tempo for traces.
Continuous profiling, which shows which lines of code consume CPU and memory in production, is increasingly treated as a fourth signal.
Alerting well
- Alert on symptoms users feel, such as error rate and latency, not on every internal cause.
- Every alert should need a human to do something. If it does not, it is noise.
- Avoid alert fatigue. Too many alerts teach people to ignore them.
- Define service level objectives (SLOs): a target such as "99.9% of requests succeed", and alert when you are using up the allowed failures too fast.
- Link each alert to a runbook describing what to check.
Common mistakes
- Logging everything, which costs a great deal and buries the useful lines.
- Logging secrets or personal data.
- Unstructured logs that cannot be queried.
- Watching averages.
- High-cardinality metric labels.
- No correlation IDs, so logs from one request cannot be connected.
- Dashboards nobody looks at.
- Alerts with no clear action.
Frequently asked questions
What are the three pillars of observability?
Logs, metrics and traces: three complementary types of data that together let you understand what a system is doing.
What is the difference between monitoring and observability?
Monitoring checks for known problems with predefined dashboards and alerts. Observability is the ability to investigate unknown problems using the system's telemetry.
What is distributed tracing?
A technique that follows one request across all the services it touches, recording how long each step took.
What is OpenTelemetry?
An open standard and set of tools for generating and exporting logs, metrics and traces in a vendor-neutral way.
Conclusion
Metrics tell you something is wrong, traces show you where, and logs explain why. None is enough alone, and their value multiplies when they share identifiers. Instrument your services from the start, watch the golden signals, and make sure every alert leads to a clear action. The time to set this up is before the incident, not during it.
Related articles
- Why Distributed Systems Are So Hard
- Monolith vs Microservices: Why Companies Switch (Both Ways)
- How Blue-Green and Canary Deployments Reduce Risk
- How Auto-Scaling Handles Sudden Traffic Spikes
