← Back to blog

2026-10-01 · 8 min read

AWS Lambda Performance: Cold Starts, Memory, and Cost

How Lambda cold starts actually work, why memory is the only performance dial that matters, and how to tune functions for speed and cost at the same time.

#aws#serverless#lambda#performance#cost-optimization
AWS Lambda Performance: Cold Starts, Memory, and Cost

AWS Lambda Performance: Cold Starts, Memory, and Cost

Lambda promises "just run code, forget servers," and mostly delivers. But two things trip up almost everyone tuning it for production: cold starts and the counter-intuitive way memory drives both speed and cost. Get these two right and most Lambda performance problems disappear.

What a cold start actually is

When Lambda needs a new execution environment, first invoke, scale-up, or after an idle period, it has to:

  1. Download your code/image,
  2. Start the runtime,
  3. Run your initialization code (everything outside the handler),
  4. Then run the handler.

Steps 1-3 are the cold start. Subsequent invokes reuse that warm environment and skip straight to the handler. The cold-start penalty varies a lot by runtime and package size.

Cold start by runtime

  • Interpreted (Python, Node.js): fast cold starts, typically tens to a couple hundred ms. Great default for latency-sensitive functions.
  • JVM/.NET: slower cold starts (JIT, larger runtime): can be hundreds of ms to seconds. Fine for steady workloads, painful for spiky user-facing ones.
  • Go/Rust (compiled, small binary): very fast.
  • Container images: can be slower to start than zip packages if the image is large: keep images lean.

If cold-start latency on the user path matters, runtime choice is your first lever.

Reducing cold starts

  • Shrink the package. Less to download = faster start. Strip unused dependencies; for Node, bundle and tree-shake; use Lambda layers judiciously.
  • Minimize init code. Everything outside the handler runs on every cold start. Lazy-load what you don't always need. But, initialize SDK clients and DB connections outside the handler so they're reused across warm invokes (init once, reuse many times).
  • Provisioned Concurrency keeps N environments warm and ready: eliminates cold starts for that N, at a cost. Use it for latency-critical, predictable-traffic functions (checkout, auth).
  • SnapStart (Java) snapshots an initialized environment and restores it, cutting JVM cold starts dramatically, and it's free.

Memory is the real performance dial

This is the thing people miss: Lambda allocates CPU proportionally to memory. You don't set CPU directly, you set memory, and CPU (and network) scale with it. So a function that's slow isn't necessarily memory-bound; it might just be CPU-starved because you gave it too little memory.

The counter-intuitive result: bumping memory often makes a function both faster and cheaper.

You pay for GB-seconds = memory × duration. If doubling memory more than halves the duration (common for CPU-bound work, because you also got more CPU), your bill goes down while latency improves.

Tune it with Power Tuning, not guesses

Don't eyeball memory settings. Use AWS Lambda Power Tuning (an open-source Step Functions state machine). It runs your function across a range of memory settings and charts cost vs speed, so you can pick the optimum, minimum cost, minimum time, or the balanced point.

The typical finding is that the default 128 MB is wrong, it's often both slower and more expensive than a higher setting for any non-trivial work.

Other cost/perf levers

  • Graviton (arm64) for Lambda: ~20% better price/performance for free, if your deps are ARM-safe (see Graviton & Savings Plans).
  • Set sane timeouts. A function hung waiting on a dependency burns money until timeout. Match the timeout to realistic worst-case, not the 15-minute max.
  • Tune concurrency. Reserved concurrency protects downstream systems from being overwhelmed during a spike; it also caps blast radius.
  • Watch the duration distribution, not the average. p99 is what your users feel and what cold starts show up in.

When Lambda isn't the answer

Performance tuning includes knowing when to stop. If a function runs constantly at high concurrency, Fargate or EKS may be cheaper and more predictable than Lambda: serverless economics favor spiky or low-volume workloads (see EKS vs ECS vs Fargate). And for steady high throughput, Provisioned Concurrency starts to look like "paying for servers anyway."

The short version

  • Cold starts = environment setup + your init code; reduce package size and init work.
  • Use Provisioned Concurrency / SnapStart for latency-critical, predictable functions.
  • Memory sets CPU: more memory often means faster and cheaper; never leave it at 128 MB blindly.
  • Tune with Lambda Power Tuning, not guesswork.
  • Consider Graviton, sane timeouts, and whether a container would be cheaper at high steady load.

Building cost-efficient serverless on AWS is part of my consulting work, reach out. See also my serverless API reference.

Share:LinkedInXWhatsApp

Related articles

Reactions & comments