The short answer
Quick answer: "Serverless" does not mean there are no servers. It means you do not manage them. You upload code, usually as small functions, and the cloud provider runs it whenever an event occurs: an HTTP request, a file upload, a message in a queue, a timer. The provider handles provisioning, scaling, patching and availability. It scales automatically from zero to thousands of simultaneous executions, and you pay only for the time your code actually runs. The trade-offs are delays when a function starts from idle (cold starts), limits on running time and state, less control, and closer ties to one provider.
A ladder of abstraction
Each step hands more responsibility to the provider.
| Model | You manage | Provider manages |
|---|---|---|
| Your own servers | Everything | Nothing |
| Virtual machines | Operating system, runtime, scaling, your code | Hardware |
| Containers on an orchestrator | Images, cluster configuration, scaling rules | Hardware, perhaps the control plane |
| Platform as a service | Your application | Servers, runtime, some scaling |
| Serverless functions | Your functions | Everything else |
Serverless is the far end: you think about code and events, not machines.
Two things the word covers
Mike Roberts's article Serverless Architectures distinguishes two meanings.
Functions as a service (FaaS). You deploy individual functions that the platform runs on demand. Examples: AWS Lambda, Google Cloud Functions, Azure Functions, Cloudflare Workers.
Backend as a service. Fully managed services that you use through an API with no capacity to manage: object storage, managed databases, authentication, message queues.
A serverless application typically combines both: functions as glue between managed services. There are also serverless containers (such as Google Cloud Run and AWS Fargate), which run a container image on the same pay-per-use, scale-to-zero basis.
How a function runs
Using AWS Lambda as the model:
- You write a handler and upload it.
def handler(event, context):
name = event.get("queryStringParameters", {}).get("name", "world")
return {"statusCode": 200, "body": f"Hello, {name}!"}
- You connect it to a trigger.
- When an event arrives, the platform finds or creates an execution environment: a small, isolated sandbox containing your code and its runtime. Providers use lightweight virtual machines or isolates to keep customers apart; see how Docker containers work for the underlying ideas.
- Your handler runs and returns a result.
- The environment is kept warm for a while in case another event comes. If none does, it is discarded.
Each environment handles one request at a time on most platforms. If a hundred events arrive together, the platform runs a hundred environments.
Common triggers:
| Trigger | Typical use |
|---|---|
| HTTP request | APIs and webhooks |
| File uploaded to storage | Resize an image, scan a document |
| Message on a queue or stream | Background processing; see how message queues work |
| Schedule | Nightly reports, clean-up jobs |
| Database change | Update a search index, send a notification |
Why people choose it
- No server management. No operating systems to patch, no capacity to plan.
- Automatic scaling. From nothing to a large burst and back, with no configuration. See how auto-scaling works.
- Pay per use. You are billed for requests and running time, in fine increments. An idle function costs nothing.
- Fast to ship. Small teams can deliver without infrastructure expertise.
- High availability built in. The provider spreads execution across data centres.
For workloads that are spiky or quiet most of the time, the cost difference from always-on servers can be large.
The trade-offs
Cold starts
When no warm environment is available, the platform must create one: start the sandbox, load the runtime, load your code, run its initialisation. That cold start adds delay, from tens of milliseconds to several seconds, depending on the language, package size and what the code does at start-up.
It matters for user-facing requests. Mitigations:
- Keep deployment packages small and initialisation light.
- Choose a runtime that starts quickly.
- Pay to keep a number of environments permanently warm.
- Use a platform built on lightweight isolates, which start in milliseconds.
Statelessness
An environment may be destroyed at any moment, and consecutive requests may hit different ones. Nothing important can live in memory or on local disk. All state goes to external services: a database, a cache, object storage.
Limits
- Maximum running time. Lambda functions, for example, are limited to 15 minutes.
- Memory and package size caps.
- No long-lived connections in ordinary functions, which makes things like WebSockets need a dedicated managed service.
Databases
A burst of traffic creates many environments, each opening its own database connection. A traditional database can be overwhelmed by the sheer number. The usual answers are a connection proxy or a database designed for this pattern. See why your database needs connection pooling.
Cost at steady high volume
Pay-per-use is cheap when usage is low or uneven. For a service that is busy all day, every day, the per-request price can exceed the cost of running your own always-on containers or virtual machines.
Lock-in
Triggers, permissions, configuration and surrounding services are specific to each provider. The function code may be portable; the architecture around it usually is not.
Harder to test and observe
A system made of dozens of functions and managed services is difficult to run locally and to trace. Good logging, metrics and tracing are essential.
When it fits
| Good fit | Poor fit |
|---|---|
| Event-driven processing: uploads, queue messages | Long-running computations |
| Traffic that is spiky or unpredictable | Constant, heavy traffic |
| Scheduled jobs | Workloads needing very low, consistent latency |
| APIs with light to moderate traffic | Stateful, long-lived connections |
| Prototypes and small teams | Applications needing specialised hardware or system-level control |
| Glue between cloud services | Workloads that must be portable between providers |
Serverless and containers
They are not opposites, and many systems use both.
| Serverless functions | Containers on Kubernetes | |
|---|---|---|
| Unit | A function | A long-running service |
| Scaling | Automatic, per request, to zero | Configured; usually keeps a minimum running |
| Billing | Per invocation and duration | Per provisioned capacity |
| Start-up | Possible cold start | Already running |
| Control | Little | Extensive |
| Operational effort | Low | High |
See what Kubernetes does for the other side. Serverless container platforms sit in between: you supply any container, and it scales to zero.
At the edge
A newer variation runs functions in data centres close to users, often on a CDN's network. See how CDNs work. These edge functions typically use lightweight isolates instead of containers, so they start almost instantly, with tighter limits on what the code can do.
Good practice
- Keep functions small and focused.
- Make handlers idempotent. Events can be delivered more than once; see idempotency.
- Initialise expensive resources outside the handler so warm environments reuse them.
- Set timeouts and concurrency limits, to protect downstream systems and your bill.
- Give each function the minimum permissions it needs.
- Define everything with infrastructure as code.
Frequently asked questions
Does serverless mean there are no servers?
No. Servers run your code, but the provider owns and operates them. You never provision or maintain one.
What is a cold start?
The extra delay when a function is invoked and no ready environment exists, so one must be created first.
Is serverless cheaper?
For low, irregular or bursty traffic, usually yes. For constant high traffic, dedicated capacity is often cheaper.
What is the difference between serverless and containers?
Containers package an application to run as a continuously running service you manage. Serverless functions run on demand, scale automatically to zero, and are billed per use.
Conclusion
Serverless is an operating model, not an absence of hardware: you write functions, the provider runs them when events happen, and you pay for what you use. It removes a great deal of operational work and scales effortlessly. In return, you accept cold starts, stateless design, platform limits and dependence on your provider. Use it where traffic is uneven and the work is event-shaped.
Related articles
- How Auto-Scaling Handles Sudden Traffic Spikes
- How Docker Containers Work (and How They Differ From VMs)
- Why Your Database Needs Connection Pooling
- What Kubernetes Does and Why Companies Use It
