top of page

What Is a Cold Start in Serverless Computing, What Causes It, and Which Fix Best Balances Latency and Cost?

11 minutes ago
29 min read
Serverless cold start balancing latency and cost.

Most requests to a serverless function are fast. The ones people remember are the few that are not: the first call after a quiet period, the burst that arrives when a campaign goes out, or the request that lands on a freshly deployed version. Those slow requests are cold starts, and the usual advice for them, from keep-warm pings to paying for always-on capacity, ranges from useful to wasteful depending on the workload. This guide explains what a cold start is, what causes it, how to measure it, and which fix best balances latency and cost on AWS Lambda, Azure Functions, Google Cloud Run, Vercel, and Cloudflare Workers, using current provider documentation as of October 2026.


TL;DR


  • Definition: A cold start is the extra latency a request pays when a serverless platform has no ready execution environment and must create and initialize one before your code can run.

  • Causes: First use after scale-to-zero, bursts that need more concurrent environments, new deployments, and environment recycling. Steady traffic does not guarantee zero cold starts.

  • Scale of the problem: AWS reports that cold starts typically affect under 1% of Lambda invocations in production workloads. That is an AWS-specific figure, and the slow tail (p95 and p99) is where users feel it.

  • Verdict: Measure first, remove avoidable initialization work, use provider-native acceleration such as Lambda SnapStart or Cloud Run startup CPU boost where it fits, then buy the smallest warm baseline your latency objective justifies.

  • Not a default: Provisioned Concurrency, Azure Always Ready, and Cloud Run minimum instances solve real problems but bill for idle capacity. Scheduled keep-warm pings are a limited workaround, not a best practice.


Quick answer


A serverless cold start is the extra startup time when a platform has no ready execution environment and must create and initialize one before running your code. Causes include scale from zero, traffic bursts and new deployments. The best latency and cost balance usually comes from trimming initialization first, then using snapshots or minimal warm capacity.


Table of contents



What Is a Cold Start in Serverless Computing?


A cold start is the additional latency a serverless request experiences when the platform has no ready execution environment for that function and must create and initialize one first. The time is spent before your handler begins useful work: allocating compute, starting the language runtime, loading code and dependencies, and running initialization code. Once an environment exists and has served a request, later requests can reuse it, which is why the same function can respond slowly once and quickly the next time.


Serverless computing is a cloud model in which the provider manages servers and scaling and bills for usage rather than for provisioned machines. Function as a Service (FaaS) is the form where you deploy individual functions that run in response to events or HTTP requests. Many FaaS platforms support scale to zero, meaning they run no capacity for an idle function and charge nothing for idle time. That economy is the root cause of cold starts: if nothing is running, something has to start.


Providers describe the idea in different words. AWS documents Lambda cold starts as the initialization steps that happen when a function is invoked after inactivity or during rapid scale-up. Google Cloud says a request may have to wait for a new container instance to start, most commonly when a service scales from zero. Microsoft notes that Azure Functions apps that scale to zero can see added latency at startup. Cloudflare and Vercel use the term too, but their execution models differ enough that the comparison needs care, as the platform sections later in this article explain.


A compact cold-start lifecycle


  1. A request arrives and no idle, initialized environment is available for the function.

  2. The platform allocates or activates an execution environment. Depending on the provider, that is a container, microVM, process, or isolate.

  3. The runtime starts and your code and dependencies are loaded.

  4. Your initialization code runs: imports, framework setup, client creation.

  5. The handler runs and returns the response.

  6. The environment stays available for some time, so a later request may reuse it. Providers do not guarantee how long.


Everything before step five is cold-start overhead. In this article, "cold start" means that extra startup work, "warm start" means reuse of an already-initialized environment, and "latency" always names which part of the request path is being discussed.


Cold Start vs. Warm Start: What Actually Changes?


A warm start reuses an execution environment that has already been initialized. On AWS Lambda, objects declared outside the handler stay initialized when the function is invoked again, the contents of the /tmp directory persist while the environment is frozen, and a database connection opened during initialization can be reused. AWS also states that Lambda terminates execution environments every few hours for runtime updates and maintenance, even for functions invoked continuously, so reuse is an optimization, not a promise.


Aspect

Cold start

Warm start

Environment

New environment created or restored

Existing initialized environment reused

Runtime and code

Loaded on this request's path

Already loaded

Global or static initialization

Runs once, before the handler

Already done

Clients and connections

Created or re-established

Often reused; verify before use

Effect on the request

Added startup latency

No startup latency

What triggers it

Idle gaps, scale-out, deployments, recycling

Steady reuse of existing environments

Who can influence it

Platform and developer

Developer, through reuse-safe code


Warm does not mean permanent, and cold and warm executions coexist. If a function has two warm environments and five requests arrive at once on a one-request-per-environment model, two requests can reuse them while the other three need new environments. Platforms that run several requests per instance, such as Cloud Run, Azure Flex Consumption, and Vercel's fluid compute, soften this effect because one instance can serve several requests at once.


What Happens During a Serverless Cold Start?


A cold start is a stack of steps, and only some belong to you. The layers below follow the order in which they happen.


  1. Environment allocation (mostly the platform): compute is reserved and sized. AWS describes this as container provisioning based on the function's configured memory.

  2. Runtime, process, or isolate startup (platform and runtime choice): the language runtime or engine starts.

  3. Artifact or image loading (shared): Lambda downloads and unpacks code, Cloud Run starts a container using image streaming, and Cloudflare fetches and compiles the script.

  4. Dependency loading (developer): libraries are imported or class-loaded.

  5. Application, static, and global initialization (developer): framework setup, dependency injection, configuration, and module-level code.

  6. Network, client, and database setup (developer and architecture): VPC attachment, DNS, TLS, secrets retrieval, and connection establishment.

  7. Handler execution (developer): the actual work. It is not cold-start overhead, but it sits in the same latency number.


AWS states that the largest contributor to latency before function execution is initialization code, and that its cost depends on package size, the amount of initialization work, and how quickly libraries and services set up connections. Cloudflare's engineers describe the same shape for Workers: fetching the script, compiling it, running top-level code, and then serving the first invocation. The practical consequence is that you cannot tune the provider's allocation time, but you often own the largest slice of the delay.


A scenario makes the order concrete. A request arrives after the platform scaled the function to zero. No reusable environment exists, so the platform provisions one, starts the runtime, loads the code, and runs your imports and client setup. Only then does the handler run. A later request may reuse that environment and skip every step before the handler.


What Causes Cold Starts?


Cold starts happen whenever demand needs an environment that is not already initialized and idle. The table summarizes the main causes and the controls that address them.


Cause

Why it happens

What addresses it

First invocation or scale from zero

No environment exists yet

Minimum or provisioned capacity; snapshots

Traffic burst or scale-out

Concurrent requests exceed existing environments

Higher per-instance concurrency; capacity buffer; faster startup

Idle retirement

Providers reclaim unused environments; timing is not guaranteed

Baseline capacity, not assumptions about retention

New deployment or version

New code needs environments built from it

Pre-initialize the new version; staged rollouts

Configuration change

Environments are rebuilt from the new definition

Plan changes; pre-warm where supported

Platform recycling

AWS says Lambda terminates environments every few hours

Tolerate it; baseline capacity helps

Crash or failure

Environment is reset and re-initialized

Handle errors; avoid process exits


Frequent traffic does not guarantee zero cold starts, because concurrency drives environment count. If a one-request-per-environment function normally serves two concurrent requests and a spike brings twenty, eighteen requests need new environments no matter how recently the function was called. AWS's guidance on Lambda cold starts makes the same point: they occur after inactivity or during rapid scale-up.


Traffic shape matters too. A USENIX ATC 2020 study of the full Azure Functions production workload found that most functions are invoked very infrequently, with an eight-order-of-magnitude range in invocation frequency. That result comes from one provider's workload in one period, but it explains why rarely used functions meet idle-gap cold starts so often. Cloudflare also reports that when requests spread across many servers, each server can see a given Worker rarely, which makes cold starts more likely until traffic is coalesced. On distributed and edge platforms, cold starts are therefore per location and per server, not per function.


How Much Latency Does a Cold Start Add?


There is no universal number, and any article that gives one is generalizing from a single platform. Cold-start latency depends on the platform architecture, the runtime, package size, framework weight, CPU and memory allocation, network attachment, database connection setup, JIT compilation and class loading, whether snapshots are used, and how traffic arrives.


Provider-published figures show the range. AWS's Lambda documentation says cold starts typically occur in under 1% of invocations and that their duration varies from under 100 ms to over 1 second. That is an AWS Lambda statement based on AWS's analysis of production workloads, not a measurement of the whole industry. The AWS SnapStart documentation adds that one-time initialization, such as loading modules or frameworks, can take several seconds for some functions, and that SnapStart can reduce this to as low as sub-second in optimal scenarios. Google reported in 2022 that startup CPU boost cut startup time roughly in half for some workloads it measured, with Java applications benefiting most. Cloudflare's 2025 engineering post notes that an early figure of 5 milliseconds was accurate at the time, then explains that larger Workers and higher startup limits made cold starts longer for complex applications. Each number describes one platform, one period, and one workload shape.


Why p50 hides the problem and p99 shows it


Percentiles describe the latency distribution. The p50 is the median, p95 means 95% of requests were faster, and p99 means 99% were faster. Cold starts are rare events with large penalties, so they barely move averages and medians but dominate the tail. As illustrative arithmetic, not a measurement: if a request normally takes 50 ms, and 1% of requests add 1,000 ms of cold-start time, the mean rises only to about 60 ms and the median stays near 50 ms, while the slowest 1% of requests take over a second. Users who hit that tail see a very different product than the dashboard average suggests.


Keep these quantities separate when you diagnose slow requests: cold-start frequency, cold-start duration, normal handler duration, downstream latency, queue latency, and scale-out latency. On Cloud Run, for example, a request waiting for a new instance is held in a queue for up to 3.5 times the average container startup time or 10 seconds, whichever is greater, according to Google's documentation. That wait is startup-related queueing, not handler time.


Which Runtimes and Frameworks Are Most Affected?


Runtime choice matters, but application initialization often matters as much. AWS's cold-start guidance says interpreted languages such as Python and Node.js typically initialize faster, compiled languages such as Java and .NET may take longer because of steps like class loading, and custom or OS-only runtimes running compiled binaries commonly offer the fastest cold starts. Those are tendencies, not laws, and AWS also recommends keeping runtimes current because newer versions often improve startup.


  • Java: startup includes JVM start, class loading, and often a heavy framework. Snapshot and CPU-acceleration features target this case: Lambda SnapStart supports Java 11 and later, and Google reported large Cloud Run startup gains from CPU boost for Java.

  • .NET: just-in-time compilation and library initialization add first-call cost. AWS documents an environment variable, AWS_LAMBDA_DOTNET_PREJIT, that controls ahead-of-time JIT compilation on the .NET 8 runtime, and SnapStart supports .NET 8 and later.

  • Node.js: module loading and bundle size dominate. Bundling, tree shaking, and importing individual SDK clients reduce work; Vercel applies bytecode caching to Node.js 20 and later in production.

  • Python: import time is the usual cost, and large scientific or machine-learning libraries can make it significant. SnapStart supports Python 3.12 and later on Lambda.

  • Go, Rust, and other compiled runtimes: small binaries and fast process start help, yet initialization such as configuration loading, secrets fetches, and connection setup still counts.


Frameworks deserve separate attention. Dependency-injection containers, classpath or module scanning, ORM initialization, and large middleware stacks run before the first request and can outweigh the language runtime. The sensible conclusion is not "switch languages". Profile initialization, then decide whether the cost is the runtime, the framework, or your own code. Switching languages is expensive and may not move the number you care about.


Does Package or Container Size Cause Cold Starts?


Sometimes, but size is a proxy rather than a cause. Three costs are easy to confuse: transfer and extraction of the artifact, runtime loading of code, and application initialization. A small package with heavy initialization can start slower than a larger, lean one.


AWS's cold-start guidance says larger Lambda packages can add latency through S3 download time, ZIP extraction, and layer mounting, and that each added dependency increases what Lambda must download, unpack, and initialize. ZIP deployments allow up to 50 MB uploaded directly or 250 MB unzipped through S3. For container images, Lambda pulls the image from Amazon ECR, and AWS notes that pulling large images can contribute to cold start latency, so image size should be kept minimal.


Platforms differ here. Google Cloud's documentation says that because of Cloud Run's container image streaming, image size does not affect container startup time. That statement is specific to Cloud Run; it should not be applied to other platforms. On Cloudflare Workers, Cloudflare's engineers explain that larger scripts increase the data fetched from script storage and the time to compile. Microsoft's Flex Consumption documentation suggests mounting Azure Files shares so that large binaries stay out of the deployment package, keeping deployments small.


  • Remove unused dependencies and use tree shaking where your toolchain supports it.

  • Exclude tests, documentation, and other files that the function never reads.

  • Prefer lightweight libraries, and import individual SDK clients instead of whole SDKs.

  • Watch dynamic imports and framework discovery, which can hide expensive work at startup.

  • Treat native libraries and bundled models as large artifacts; consider mounted storage or on-demand loading where the platform supports it.


Why Autoscaling Creates a Latency–Cost Tradeoff


Serverless platforms save money by running nothing when nothing is happening. Fast responses require capacity that is ready before the request arrives. Ready capacity costs money, so every cold-start mitigation is an economic decision about how much idle capacity to pay for and where.


A conceptual relationship helps structure that decision. It is a way of thinking, not a provider billing formula:


Expected user impact ≈ cold-start probability × incremental cold-start latency × importance of the affected request


Each term has its own lever. Probability falls with minimum or provisioned capacity, higher per-instance concurrency, and avoiding scale-to-zero on critical paths. Incremental latency falls with leaner initialization, snapshots, and startup CPU acceleration. Importance is a business judgment: a cold start on a background job rarely matters, while one on a login or checkout path can. The cheapest improvement usually reduces latency or targets importance, because those levers do not require paying for idle capacity around the clock.


Compare the cost of mitigation with the cost of the latency, not with zero. Include engineering time, operational complexity, the cost of overprovisioning, and the cost of slow or failed requests. Do not assume the lowest infrastructure bill is the best business outcome, and do not assume the fastest option is worth its price. The decision sections later in this article apply this reasoning workload by workload.


How Do You Measure Cold Starts Correctly?


Measure cold starts in production telemetry and in controlled tests. A single manual request after a quiet period tells you almost nothing: it is one sample, it may or may not land on a warm environment, and it cannot reveal scale-out, deployment, or tail behavior.


On AWS Lambda, the REPORT log line for a cold invocation includes an Init Duration. Lambda always emits an INIT_REPORT log for functions using provisioned concurrency or SnapStart, and SnapStart functions also report a Restore Duration. Failures during initialization appear in INIT_REPORT with a status. AWS documents an environment variable, AWS_LAMBDA_INITIALIZATION_TYPE, whose value is provisioned-concurrency or on-demand, so you can tell how an environment was created. AWS also notes suppressed inits: after an invoke failure Lambda re-initializes the environment on the next invocation without a separate INIT line, so the reported duration can include extra initialization time. On other platforms, rely on the provider's request logs and tracing, and on application-level logging that marks the first request an instance serves.


  1. Mark every request as cold or warm. Log a flag from module-level state, or use the provider's initialization indicator, and attach it to traces.

  2. Compute cold-start rate. Cold invocations divided by total invocations, per function, version, and region.

  3. Report p50, p95, and p99 separately for cold and warm requests, so the penalty is visible.

  4. Correlate with concurrency and deployments. Cold starts that cluster at deploy time or during spikes point to different fixes than idle-gap cold starts.

  5. Trace downstream calls. Database, auth, and third-party latency can look like cold-start latency if you only watch the total.

  6. Load test with realistic traffic shape, including an idle period followed by a single request, a gradual ramp, and a sudden burst.

  7. Repeat after every runtime, dependency, or platform change, because provider behavior and defaults evolve.


Vercel's Observability view reports Active CPU, Provisioned Memory, and invocations for functions, and Cloud Run exposes a billable instance time metric that is useful when you tune concurrency and minimum instances for cost. Use these cost signals next to the latency signals, because a fix that improves p99 and doubles the bill needs a business case.


How to Reduce Cold Starts Without Paying for Always-On Capacity


These techniques shrink or remove work on the cold path. They cost engineering time rather than idle compute, which is why they come first in the decision hierarchy.


  • Trim dependencies and artifacts, as described above, and keep runtimes current.

  • Right-size memory and CPU. On Lambda, more memory also means more CPU, which can shorten initialization. AWS recommends its Lambda Power Tuning tool to compare speed against the extra cost, because higher memory raises the per-millisecond price even when it shortens the run.

  • Tune concurrency. Higher per-instance concurrency means fewer instances for the same load, so fewer cold starts. Google notes that a lower Cloud Run concurrency setting generally lowers per-request latency but needs more instances, and that the best setting is found by load testing.

  • Attach to a VPC only when needed. AWS notes that VPC attachment involves creating network interfaces and can add latency.

  • Place compute near its data to reduce connection and query time inside initialization and handlers.

  • Use startup CPU boost on Cloud Run, which adds CPU during instance startup.

  • Use provider-native snapshot or caching features, such as Lambda SnapStart and Vercel's bytecode caching.


Eager versus lazy initialization


Neither extreme is right. AWS says static initialization is often the best place to open database connections so they are reused across invocations, and also says to lazily load objects that only some code paths use. Google Cloud warns that global variables always initialize at startup, which lengthens startup, and recommends lazy initialization for infrequently used objects, while cautioning that lazy initialization adds latency to the first requests on a new instance and can cause overscaling and dropped requests when you deploy a new revision under load. AWS also notes that with provisioned concurrency, initialization runs ahead of time, so moving more work outside the handler is beneficial there. The principle: initialize reusable essentials once, defer nonessential work, and measure both phases.


Example from the AWS Lambda documentation: imports, logger setup, and the S3 client are created during initialization, before the handler runs.


import os
import json
import cv2
import logging
import boto3

s3 = boto3.client('s3')
logger = logging.getLogger()
logger.setLevel(logging.INFO)

def lambda_handler(event, context):
  # Handler logic...

Snapshot and Restore: When Is It the Sweet Spot?


Snapshot-based startup runs initialization once, ahead of time, saves the initialized state, and restores from that state instead of initializing from scratch. It targets the same costs as other fixes but without keeping a running environment waiting.


AWS Lambda SnapStart is the clearest current example. When you publish a function version with SnapStart enabled, Lambda initializes it, takes an encrypted Firecracker microVM snapshot of memory and disk state, and caches it. On first invocation and as invocations scale up, Lambda resumes new environments from that snapshot. AWS says this can cut startup latency from several seconds to as low as sub-second, works best with invocations at scale, and may help less for infrequently invoked functions.


Per AWS documentation at the time of writing, SnapStart supports Java 11 and later, Python 3.12 and later, and .NET 8 and later, in ZIP and container image formats on AWS base images. It does not support other managed runtimes such as Node.js 24 or Ruby 4.0, OS-only runtimes, provisioned concurrency, Amazon EFS, Amazon S3 Files, or ephemeral storage above 512 MB. It works only on published versions and aliases, not $LATEST, and it is unavailable in the Asia Pacific (New Zealand) and Asia Pacific (Taipei) Regions. Check the page before you plan around it, because this list changes.


Snapshots create correctness obligations. State captured during initialization is shared across environments, so unique IDs, secrets, and random seeds must be generated after restore. Network connections made during initialization may not survive, so validate and re-establish them. Temporary credentials and cached timestamps should be refreshed in the handler. Runtime hooks let you run code before the checkpoint and after the restore.


Pricing has a different shape from always-on capacity. AWS states there is no additional SnapStart cost for Java managed runtimes. For other supported runtimes, you pay for snapshot caching, with a three-hour minimum for each published version while it stays active, and a restore charge each time an environment is restored; duration charges also cover initialization code and runtime hooks. As illustrative arithmetic using the sample US East (N. Virginia) rates on AWS's pricing page, caching one 1 GB version for 30 days comes to roughly $3.90, and 100,000 restores of 1 GB add roughly $14. Rates change by region and over time, so recalculate from the current pricing page. AWS's pricing page also lists Lambda MicroVMs, which it describes as snapshot-based images that start quickly and can suspend while idle; evaluate that separately.


SnapStart is not a replacement for provisioned concurrency in every case. AWS says provisioned concurrency keeps functions initialized and ready to respond in double-digit milliseconds, and recommends it when strict cold-start requirements cannot be met by SnapStart. For compatible runtimes with heavy initialization, SnapStart often delivers most of the benefit without a standing baseline, which is why it is the middle ground in this article's hierarchy.


Provisioned, Minimum, and Always-Ready Capacity


These features keep a baseline of initialized capacity so that traffic within the baseline does not meet a cold start. They share a concept, and they are not identical.


  • AWS Provisioned Concurrency: pre-initialized execution environments for a published version or alias. Billing runs from enablement until disabled, rounded up to the nearest five minutes, plus request and duration charges when used, and the free tier does not apply. Traffic beyond the configured number uses on-demand environments, which can cold start.

  • Azure Flex Consumption Always Ready: one or more instances always running for an HTTP, Blob, or Durable group or an individual function. The default is zero. You pay for baseline memory, active execution time, and executions, with no free grants. Scale beyond them uses on-demand instances.

  • Azure Premium plan: always-ready instances keep workers perpetually warm, and Microsoft says at least one instance per plan must always be kept warm.

  • Cloud Run minimum instances: instances kept running so the service can take its configured concurrency without starting a new one. Google calls this a best-effort target, and minimum instances can be restarted. With request-based billing, idle minimum instances bill at a lower idle rate; with instance-based billing they bill at the default rate for their whole lifetime.


The common lesson is sizing. Capacity sized for normal load still sees on-demand scale-out during spikes. AWS gives a sizing formula, concurrency equals average requests per second multiplied by average request duration in seconds, and suggests a 10% buffer. Application Auto Scaling can adjust provisioned concurrency on a schedule, which suits predictable peaks, or by target tracking, which AWS says needs a burst to persist for at least three minutes and may react poorly to very short bursts unless you use the Maximum statistic. As illustrative arithmetic at the sample US East (N. Virginia) rate on AWS's pricing page, ten 1 GB provisioned environments held for 30 days come to roughly $108 before request and duration charges. Compare that figure with the revenue or experience at stake on the endpoint.


Example from the AWS Lambda documentation: configure provisioned concurrency for an alias.


aws lambda put-provisioned-concurrency-config --function-name my-function \
  --qualifier BLUE \
  --provisioned-concurrent-executions 100

Provisioned capacity does not remove slow downstream dependencies, handler-time latency, or cold starts beyond the baseline, and it does not make deployment-time starts disappear unless you pre-initialize the new version.


Do Keep-Warm Pings Actually Work?


Sometimes, within narrow limits. A scheduled request every few minutes can keep one environment of a low-traffic function in use and reduce idle-gap cold starts. It is a workaround, not a best practice.


  • A ping warms one environment. Concurrent real requests still need additional environments that were never warmed.

  • Providers do not guarantee retention. AWS says Lambda terminates environments every few hours, and Cloud Run notes that minimum instances can be restarted.

  • Deployments and new versions still start fresh environments.

  • Pings add invocations, cost, and noise in metrics, and you must keep the schedule and function code compatible.

  • Native controls, such as provisioned concurrency, always-ready instances, minimum instances, and snapshots, express the actual intent and cover scale-out better.


If a single low-traffic endpoint tolerates an occasional slow request, even a ping is probably unnecessary. If it does not, use the provider's native control.


AWS Lambda: Which Cold-Start Fixes Work Best?


Based on current AWS documentation and pricing, a sensible Lambda hierarchy is:


  1. Measure Init Duration and cold-start rate from REPORT and INIT_REPORT logs.

  2. Cut initialization: trim dependencies, import only what you use, and keep runtimes current.

  3. Tune memory with a tool such as AWS Lambda Power Tuning, since memory also buys CPU.

  4. Enable SnapStart if the runtime is Java 11+, Python 3.12+, or .NET 8+ and initialization is heavy. Make code snapshot-safe first.

  5. Add Provisioned Concurrency for the endpoints with strict latency requirements, sized from measured concurrency and scaled on a schedule or by target tracking.


The two accelerators are alternatives rather than layers: AWS states that SnapStart does not support provisioned concurrency. Choose one per function version based on the tail-latency target. For steady, high-volume workloads, AWS also offers Lambda Managed Instances, which run Lambda functions on EC2 capacity in your account with multi-concurrency and a management fee on top of EC2 pricing. That is a different cost model, not a cold-start toggle.


Azure Functions: Which Cold-Start Fixes Work Best?


Microsoft's hosting documentation lists Flex Consumption as the recommended serverless plan, the Premium plan for always-warm workers, and Consumption as legacy. It labels Linux Consumption as retired for new use and states that hosting on Linux Consumption retires on 30 September 2028, so new serverless apps should target Flex Consumption.


On Flex Consumption, the main levers are always-ready instances, per-instance concurrency, and instance size. Always-ready instances can be assigned to HTTP, Blob, or Durable groups or to individual functions, and are not subject to the on-demand scale-out rate or the maximum instance count. Raising concurrency lets each instance do more work, so fewer new instances start. The 512 MB, 2,048 MB, and 4,096 MB sizes map to 0.25, 1, and 2 cores, and Microsoft suggests 2,048 MB as the default. Flex Consumption supports code-only deployments on Linux, with no deployment slots and no in-place migration from other plans. Keep deployment packages lean, and consider Azure Files mounts for large binaries. The cost model is on-demand execution plus, when enabled, the always-ready baseline.


Premium keeps at least one instance warm by design and offers the most predictable billing, which suits nearly continuous workloads. For container control, Azure Functions can also run on Azure Container Apps, where a minimum replica count of one or more removes cold starts and zero keeps scale-to-zero.


Google Cloud Run and Cloud Run Functions: Which Fixes Work Best?


Google Cloud's Cloud Run guidance names four levers: fast container startup, startup CPU boost, minimum instances, and tuned concurrency.


Startup CPU boost temporarily raises CPU during instance startup and for ten seconds after it. It is billed: Google's example says a container with 2 CPU is charged for 4 CPU during startup and that extra window. Enable it with one flag:


gcloud run services update SERVICE --cpu-boost

Minimum instances remove the zero-to-one cold start and let the service accept its configured concurrency without starting another instance, but scale-out beyond that still starts new instances. Avoid process exits that shut instances down. Global initialization always runs at startup, so use lazy initialization for rarely used objects, with the deploy-time caveat described earlier. Cloud Run's pricing offers request-based billing, where idle minimum instances bill at a lower idle rate, and instance-based billing, where instances bill for their whole lifetime. Google's 2022 announcement said Cloud Functions 2nd gen used startup CPU boost by default; confirm current behavior in the docs. Cold starts for functions follow the same container-startup rules as services.


Vercel and Edge/Isolate Platforms: Does “Cold Start” Mean the Same Thing?


Not exactly. Containers, microVMs, processes, snapshots, and isolates start differently, so the same term covers different costs.


Vercel's fluid compute lets multiple invocations share one function instance, for Node.js and Python, and prefers existing idle resources before allocating new ones. Vercel says it is enabled by default for projects created since April 23, 2025, adds function pre-warming on production deployments, and applies bytecode caching to Node.js 20+ in production only; the first request is not cached, and later cold starts benefit. Vercel contrasts this with traditional serverless, which uses a microVM per function. These are Vercel's claims, and this research found no independent measurement to confirm them. Pricing uses Active CPU, Provisioned Memory, and invocations. Provisioned memory is billed until the last in-flight request finishes, Active CPU is billed only while code runs, and between requests the instance is paused with no charge.


Cloudflare Workers use V8 isolates, lightweight contexts that Cloudflare says are designed to start very quickly and that eliminate the cold starts of the virtual machine model. Cloudflare also notes isolates are not necessarily long-lived and may be evicted. In a 2025 post, Cloudflare said "eliminate" is a strong word, that cold starts for complex Workers had begun to exceed the TLS handshake used for pre-warming, and that routing changes raised the warm request rate for enterprise traffic from 99.9% to 99.99%. Those are Cloudflare-reported figures for Cloudflare traffic. Isolates reduce startup cost, but code size, startup CPU time, and eviction still shape latency, so treat "zero cold starts" as a scoped vendor claim.


Cold-Start Mitigation Compared: Latency, Cost, and Complexity


The ratings below are qualitative judgments drawn from the provider documentation cited in this article, not benchmarks. Portability reflects how transferable the technique is across platforms.


Technique

Cold-start reduction

Idle or baseline cost

Engineering effort

Predictability

Burst behavior

Portability

Best for

Main downside

Init and dependency optimization

Medium to high

None

Medium

Medium

Helps every new instance

High

Every workload

Gains plateau

Memory and CPU sizing

Low to medium

None; rate rises

Low

Medium

Helps new instances

High

CPU-bound init

Can raise bill

Concurrency tuning

Medium

None

Low to medium

Medium

Fewer scale-outs

Medium

I/O-bound APIs

Needs thread-safe code

Snapshot and restore

High where supported

Low; runtime-dependent

Medium

High

Restores on scale-out

Low

Heavy init on Java, Python, .NET

Runtime limits; uniqueness

Startup CPU boost

Medium

None; boosted startup CPU billed

Low

Medium

Speeds every cold start

Low

CPU-bound startup

Billed; not elimination

Minimum or always-ready instances

High within baseline

Yes

Low

High

Cold starts beyond baseline

Medium

Latency-sensitive APIs

Idle cost

Provisioned concurrency

Very high within baseline

Yes

Low to medium

Very high

Spillover can cold start

Low

Strict p99 on Lambda

Idle cost; no SnapStart

Scheduled capacity

High at known peaks

Only when scheduled

Medium

High if predictable

Misses surprises

Medium

Predictable peaks

Forecast errors

Keep-warm pings

Low

Low

Low

Low

Does not help

High

Rare single-instance cases

Misses concurrency and deploys

Move to isolate or edge execution

Varies

Varies

High

Varies

Platform-specific

Low

Short, lightweight handlers

Rewrite; runtime limits


Serverless Platform Comparison


Platform

Startup model

Main mitigation controls

Baseline capacity option

Cost posture

Key caveat

AWS Lambda

Execution environments with Init, Invoke, Restore

Init tuning, memory, SnapStart, Provisioned Concurrency

Provisioned Concurrency

On-demand duration; baseline billed per GB-second

SnapStart and PC cannot combine

Azure Functions

Instances per function group (Flex Consumption)

Always ready, concurrency, instance size

Always ready; Premium

Execution plus always-ready baseline

Flex is Linux, code-only

Google Cloud Run and functions

Container instances with queued requests

Startup CPU boost, minimum instances, concurrency

Minimum instances

Request-based or instance-based billing

Minimum instances are best-effort

Vercel Functions

Shared instances under fluid compute

Optimized concurrency, bytecode caching, pre-warming

No documented fixed baseline found

Active CPU, provisioned memory, invocations

Vendor-reported claims

Cloudflare Workers

V8 isolates

Pre-warming, sharded routing, small scripts

Not applicable

Request and CPU based

Isolates can be evicted


Which Cold-Start Fix Best Balances Latency and Cost?


Optimize initialization first, then use provider-native acceleration, and pay for warm capacity only on the endpoints and concurrency levels whose p95 or p99 objectives justify it. The current documentation supports this order: the first steps cost engineering time rather than idle compute, while paid capacity buys predictability rather than speed. Starting with always-on capacity usually hides initialization waste instead of removing it, and every warm instance repeats that waste.


On AWS, the middle ground is SnapStart for compatible runtimes with heavy initialization. On Azure, it is Flex Consumption with a small always-ready count and tuned concurrency. On Google Cloud, it is startup CPU boost plus one minimum instance and tuned concurrency where needed. On Vercel, it is fluid compute with the platform defaults. Use the matrix to match the approach to the workload.


Workload

Recommended approach

Why

Cost posture

Do not pay for

Low-traffic or hobby APIs

Accept cold starts; trim init

Occasional delay is tolerable

Near zero

Warm capacity or pings

Typical production APIs

Init optimization, sizing, SnapStart or boost

Cuts duration without baseline

Low

Blanket provisioned capacity

User-facing SaaS where p95 matters

Above plus 1 to few warm instances

Removes idle-gap starts

Low to moderate

Peak-sized baseline

Login, auth, or checkout

Warm baseline on those functions only

High value per request

Moderate

Warm capacity on all functions

Strict p99 or SLO workloads

Provisioned or always-ready, sized with buffer and scaling

Predictability beats savings

Highest

Unmeasured oversizing

Java, .NET, or framework-heavy

SnapStart or boost first; then baseline

Heavy init benefits most

Low to moderate

Rewrites before measuring

Highly bursty event processing

Concurrency tuning, buffering, queues

Baseline cannot cover spikes

Low

Large idle baseline

Predictable peaks

Scheduled provisioned or minimum capacity

Capacity only when needed

Moderate

All-day baseline

Background or async jobs

Do nothing; tolerate cold starts

Latency rarely matters

Lowest

Any warm capacity


Do not treat the lowest bill as the goal. A slow login or checkout can cost more than the baseline that prevents it, and a warm baseline on a low-value path is waste. Use measured cold-start rate, tail latency, and the importance of the request, as in the earlier relationship, to decide where paid capacity belongs.


When Should You Stop Using Serverless Because of Cold Starts?


Cold starts alone are rarely enough reason to leave serverless, because most have a mitigation and migration has its own cost. Other capacity can fit better in four cases: utilization is permanently high, so reserved capacity costs less than per-use billing (AWS positions Lambda Managed Instances for steady, high-volume workloads); latency objectives are stricter than any available mitigation delivers; the workload depends on specialized, stateful, or warm resources such as large in-memory models; or the mitigations you need are unsupported for your runtime or configuration, for example SnapStart on Node.js. Even then, test a warm baseline before replatforming.


Common Serverless Cold-Start Myths


  • "Every request is a cold start." AWS says cold starts typically affect under 1% of Lambda invocations in production workloads.

  • "High traffic means no cold starts." Rising concurrency needs new environments.

  • "A cron ping eliminates them." It warms at most one environment.

  • "It is purely the provider's problem." AWS says initialization code is the largest contributor to pre-execution latency.

  • "More memory always costs more." More CPU can shorten duration, so the total can fall; Power Tuning tests this.

  • "Containers remove cold starts." Image pulls and initialization still count, except where a platform streams images.

  • "Provisioned Concurrency is always best." It bills while idle and spillover can still cold start.

  • "Cold starts always take seconds." AWS reports a range from under 100 ms to over 1 second.

  • "Switching languages always fixes it." Framework and dependency initialization often dominate.


Practical Cold-Start Optimization Playbook


  1. Define the latency objective and which endpoints it covers.

  2. Measure cold-start rate and p50, p95, and p99 for cold and warm requests.

  3. Locate controllable initialization time.

  4. Optimize artifacts, dependencies, runtime version, and initialization.

  5. Test memory, CPU, and concurrency settings.

  6. Apply provider-native acceleration where compatible.

  7. Add minimum or provisioned capacity only where the objective requires it.

  8. Load test bursts, deployments, and idle returns.

  9. Compare monthly cost with the measured latency gain.

  10. Repeat after major runtime, dependency, or platform changes.


FAQ


What is a serverless cold start?


A serverless cold start is the extra latency a request pays when the platform has no ready execution environment and must create and initialize one first. The time covers allocating compute, starting the runtime, loading code, and running initialization. Later requests can reuse the initialized environment and skip that work.


What causes Lambda cold starts?


Lambda creates a new execution environment on first use after inactivity, during scale-out when concurrency exceeds existing environments, after deployments, and when AWS recycles environments, which it says happens every few hours. The delay is the Init phase: extensions, runtime bootstrap, and your static code.


How long can a cold start take?


AWS says Lambda cold starts typically occur in under 1% of invocations and last from under 100 ms to over 1 second. That is an AWS-specific statement. AWS's SnapStart documentation says heavy framework initialization can take several seconds. Other platforms differ, so measure your own p95 and p99.


Can cold starts be eliminated completely?


No platform promises zero cold starts without qualification. Provisioned or minimum capacity removes them for traffic within the baseline, but spillover, deployments, and recycling can still create fresh environments. Treat claims of zero cold starts as scoped to a specific mechanism and capacity.


What is the difference between SnapStart and Provisioned Concurrency?


SnapStart restores environments from a cached snapshot and suits heavy initialization on Java, Python 3.12+, and .NET 8+ without a standing baseline. Provisioned Concurrency keeps pre-initialized environments ready and bills continuously. AWS says SnapStart does not support Provisioned Concurrency, so choose one per function version.


Does increasing Lambda memory reduce cold starts?


Often, for CPU-bound initialization. AWS says Lambda allocates more CPU with more memory, which can shorten initialization, but the per-millisecond price rises. Use a tool such as AWS Lambda Power Tuning to compare speed and cost, because faster execution can offset the higher rate.


Do keep-warm pings work?


They can reduce idle-gap cold starts for one low-traffic environment. They do not warm additional environments for concurrent requests, survive deployments, or defeat recycling, and they add cost and noise. Native controls such as provisioned concurrency, always-ready instances, or minimum instances express the intent better.


Which runtime has the fastest cold start?


No runtime wins universally. AWS says interpreted runtimes such as Python and Node.js typically initialize faster than Java and .NET, and compiled custom runtimes are commonly fastest, but framework and dependency initialization often dominate. Measure Init Duration in your own function before changing languages.


Does Cloud Run have cold starts?


Yes. Cloud Run starts a container when a request needs a new instance, commonly when scaling from zero. Minimum instances, startup CPU boost, tuned concurrency, and lean initialization reduce the effect. Google notes that container image size does not affect startup because of image streaming.


Do Vercel and edge platforms still have cold starts?


Both reduce them through different architectures. Vercel says fluid compute shares instances, caches bytecode for Node.js 20+ in production, and pre-warms production deployments. Cloudflare says Workers' isolates eliminate the virtual machine model's cold starts and reports a 0.01% cold start rate for enterprise traffic. Both are vendor claims.


When is provisioned capacity worth paying for?


When a latency objective on a valuable endpoint cannot be met by initialization work, SnapStart, or CPU boost. Size the baseline from measured concurrency, scale it on a schedule for predictable peaks, and keep it off low-value paths. Minimum and always-ready instances bill while idle at provider-specific rates.


Are cold starts a reason to avoid serverless?


Rarely by themselves. Mitigations exist, and migration has its own costs. Consider other capacity when utilization is steadily high, latency objectives exceed what mitigations deliver, or required features are unavailable. Otherwise, fix initialization and add targeted warm capacity.


Key Takeaways


  • A cold start is initialization overhead on a request that finds no ready environment.

  • Concurrency, deployments, and recycling cause cold starts even under steady traffic.

  • Judge impact at p95 and p99, not the average.

  • Initialization work is usually the largest slice you control.

  • SnapStart and startup CPU boost cut startup without a standing baseline where supported.

  • Provisioned, always-ready, and minimum capacity buy predictability and bill while idle.

  • Pings are a limited workaround, and "zero cold starts" claims are always scoped.


Actionable Next Steps


  1. Add a cold-start flag and Init Duration to your logs and traces.

  2. Chart cold-start rate and p99 for your top five endpoints.

  3. Remove unused dependencies and trim heavy imports.

  4. Run a memory and concurrency test.

  5. Check SnapStart or startup CPU boost compatibility.

  6. Pick one latency-critical endpoint for a small warm baseline.

  7. Load test a burst and a deployment.

  8. Review cost against latency monthly.


Glossary


  • Cold start: extra startup latency when no ready environment exists.

  • Warm start: a request served by an already-initialized environment.

  • Serverless computing: a cloud model where the provider manages servers and bills by use.

  • FaaS: deploying individual functions that run on events or requests.

  • Execution environment: the isolated sandbox that runs your function.

  • Initialization (INIT): startup work before the handler runs.

  • Runtime: the language engine that executes your code.

  • Scale to zero: running no capacity while a function is idle.

  • Concurrency: requests being handled at the same time.

  • Tail latency: the slowest requests, such as p99.

  • p50: the median; half of requests are faster.

  • p95: 95% of requests are faster.

  • p99: 99% of requests are faster.

  • Provisioned Concurrency: an AWS Lambda setting that keeps pre-initialized environments ready.

  • Minimum instances: a Cloud Run setting that keeps instances running.

  • Always Ready: an Azure Flex Consumption setting for always-running instances.

  • Snapshot/restore: saving initialized state and resuming from it.

  • Lambda SnapStart: a Lambda feature that restores from a cached snapshot.

  • Startup CPU boost: extra CPU during Cloud Run instance startup.

  • Container image: packaged code, runtime, and dependencies.

  • Dependency: a library your code needs.

  • JIT compilation: compiling code to machine code at run time.

  • SLO: an internal target for service performance.

  • SLA: a contractual service commitment.


Sources & References


bottom of page