The short answer
Quick answer: A message queue sits between services. A producer writes a message to the queue and moves on; a consumer reads and processes it later. Neither needs to know about the other, be running at the same moment, or work at the same speed. That decoupling absorbs traffic spikes, survives downstream failures, and lets many systems react to the same event. Traditional brokers such as RabbitMQ deliver each message to a worker and then delete it. Log-based systems such as Kafka keep an ordered, durable log that many consumers read at their own pace and can replay.
The problem with direct calls
Suppose placing an order must charge the card, send a confirmation email, update inventory and notify analytics. With direct synchronous calls:
- The user waits for all four to finish.
- If the email service is down, the whole order fails.
- If analytics is slow, checkout is slow.
- Adding a fifth action means changing the checkout code.
- A sudden burst of orders hits every downstream service at once.
With a queue, the checkout service publishes one "order placed" message and returns. Each downstream service consumes it independently.
What a queue gives you
- Asynchrony. The caller does not wait for slow work.
- Buffering. If producers briefly outpace consumers, messages wait in the queue instead of overwhelming anything.
- Resilience. If a consumer is down, messages accumulate and are processed when it returns.
- Independent scaling. Add consumers to work through a backlog faster.
- Fan-out. New subscribers can be added without touching the producer.
Two models
Work queue (point to point)
Each message goes to one consumer. Several workers pull from the same queue and share the load. This is ideal for background jobs: resize this image, send that email.
Publish/subscribe
Each message goes to every subscriber. The order service publishes "order placed"; inventory, email and analytics each receive their own copy.
Most systems support both.
Traditional brokers
RabbitMQ, Amazon SQS and ActiveMQ follow the classic model:
- A producer sends a message to the broker.
- The broker stores it in a queue.
- A consumer receives it and processes it.
- The consumer acknowledges success.
- The broker deletes the message.
If the consumer crashes before acknowledging, the broker redelivers the message to another consumer.
RabbitMQ adds flexible routing. Producers publish to an exchange, which routes to queues according to rules: by exact key, by pattern, or to all bound queues. The AMQP concepts guide explains the model.
Log-based systems: Kafka
Kafka, along with systems like Amazon Kinesis and Apache Pulsar, works differently. It is a distributed, append-only log, described in the Kafka documentation.
- A topic is a named stream of messages.
- Each topic is split into partitions. A partition is an ordered log; new messages are appended to the end.
- Every message in a partition has a sequential offset.
- Messages are not deleted when read. They are kept for a configured time or size, regardless of consumption.
- Each consumer group tracks its own offset per partition, meaning how far it has read.
This design has important consequences:
- Many independent readers. Several consumer groups read the same topic without affecting each other.
- Replay. A consumer can reset its offset and reprocess history, which is invaluable after fixing a bug or adding a new downstream system.
- Scaling. Within a group, each partition is assigned to one consumer. More partitions allow more consumers in parallel.
- Throughput. Appending to a log and reading it sequentially is extremely fast. It is the same principle as a write-ahead log.
- Durability. Each partition is replicated across several brokers.
Ordering
Kafka guarantees order within a partition, not across a topic. Producers choose a partition by hashing a key, so all messages with the same key (say, one customer's ID) land in the same partition and stay in order.
| Traditional queue (RabbitMQ, SQS) | Log (Kafka) | |
|---|---|---|
| After consumption | Message is deleted | Message is retained |
| Replay | No | Yes |
| Multiple independent consumers | Needs a queue each | Built in via consumer groups |
| Ordering | Per queue, weaker with several consumers | Strict per partition |
| Routing | Rich | Simple: topic and key |
| Typical use | Task queues, request routing | Event streams, pipelines, high volume |
Delivery guarantees
Networks fail and processes crash, so messaging systems offer one of three guarantees:
| Guarantee | Meaning | Risk |
|---|---|---|
| At most once | Send and never retry | Messages can be lost |
| At least once | Retry until acknowledged | Messages can be duplicated |
| Exactly once | Each message takes effect once | Hard; achieved only within limits |
At least once is the practical default. A consumer might process a message, crash before acknowledging, and receive it again. So consumers must be idempotent: processing the same message twice must have the same effect as processing it once. Typical techniques are deduplicating by message ID or using naturally idempotent writes. See how payment systems avoid charging you twice.
"Exactly once" in systems like Kafka means exactly-once effect within Kafka's own transactions. Once you call an external system, such as an email provider, you are back to at-least-once plus idempotency. The reason is fundamental; see the Two Generals' Problem.
When things go wrong
- Poison messages. A message that always fails would be retried forever and block the queue. After a few attempts it is moved to a dead-letter queue for inspection.
- Backlog. If consumers are too slow, the queue grows. Monitor consumer lag (how far behind consumers are) and scale consumers or apply backpressure.
- Retries with backoff. Wait longer between attempts so a struggling dependency can recover.
- The dual-write problem. Saving to your database and publishing a message are two separate actions; one can succeed while the other fails. The outbox pattern fixes it: write the message to an "outbox" table in the same database transaction as the business data, and have a separate process publish from that table.
What queues are used for
- Background jobs: emails, image processing, report generation.
- Smoothing spikes: ticket sales, flash sales, bulk imports.
- Event-driven microservices: services react to each other's events instead of calling directly. See monolith vs microservices.
- Data pipelines: streaming changes from databases to warehouses, search indexes and caches.
- Activity and log collection at high volume.
- Notifications. See how notification systems work.
When not to use one
A queue adds a component to operate and makes behaviour asynchronous, which is harder to debug and reason about.
- If the caller needs the result immediately, use a direct call.
- If the system is small, a database table used as a job queue may be enough.
- Do not introduce Kafka for a few hundred messages a day.
Frequently asked questions
What is the difference between Kafka and RabbitMQ?
RabbitMQ is a message broker that routes messages to queues and deletes them once consumed. Kafka is a distributed log that retains messages so many consumers can read and replay them.
What is a consumer group?
A set of consumers that share the work of reading a topic. Each partition is read by one member of the group, and the group tracks its position with offsets.
Does a message queue guarantee order?
Only within limits. Kafka guarantees order within a partition. With several competing consumers on one queue, processing order is not guaranteed.
What is a dead-letter queue?
A holding queue for messages that repeatedly failed to process, so they do not block others and can be examined later.
Conclusion
A message queue replaces "do this now and wait" with "this happened; deal with it when you can". That one change decouples services in time, in speed and in failure. Choose a broker for task distribution and routing, a log for event streams and replay, assume at-least-once delivery, and make every consumer idempotent.
Related articles
- Monolith vs Microservices: Why Companies Switch (Both Ways)
- How Payment Systems Avoid Charging You Twice (Idempotency)
- How Write-Ahead Logs Prevent Data Loss During Crashes
- How Notification Systems Send Millions of Push Alerts
