
Cloud Run or GKE? How I Decide in Real Projects
Almost every GCP project I join ends up having the same discussion at some point. A new service is ready to ship, and someone asks whether it should run on Cloud Run or go onto the GKE cluster that's already there. Both run containers, so it can look like a question of taste. In practice the two platforms are built on quite different assumptions about who operates the infrastructure and who pays when nothing is happening.
For my own projects the answer has sometimes been neither. When I started Lieferhand, a marketplace connecting restaurants with delivery drivers, the whole MVP ran on one Hetzner server for €4.50 a month. Spring Boot on Java 21, PostgreSQL, Caddy for HTTPS and Stripe for payments, all in a single Docker Compose file. I decided against Kubernetes on purpose, because an MVP simply doesn't need it.
That changes once a service has to scale and I'm on Google Cloud anyway.
Cloud Run Today
You give Cloud Run a container image and it runs it. Google handles the servers and the scaling, TLS certificates included. With request-based billing you only pay while an instance is actually processing requests, rounded up to 100 ms, so a service nobody calls costs you nothing. There are no node pools to patch, and nobody has to debug an ingress controller at 2 AM.
Deploying is one command:
gcloud run deploy my-api \
--image europe-west1-docker.pkg.dev/my-project/services/my-api:1.4.2 \
--region europe-west1 \
--allow-unauthenticatedThe image path in that snippet points to Artifact Registry on purpose. Container Registry (gcr.io) is deprecated, even though plenty of older tutorials still use it. I'd also keep --allow-unauthenticated only for endpoints that are really meant to be public. Anything internal should require IAM authentication.
Concurrency gets surprisingly little attention in most comparisons. By default a single Cloud Run instance processes up to 80 requests at once, and you can raise that limit to 1,000. AWS Lambda works differently, there every invocation runs in its own environment. For an I/O-bound Spring Boot or Node API you need far fewer instances than the raw request numbers would suggest, and you see that directly on the bill.
The myths left over from 2019
A lot of engineers still judge Cloud Run by the version that launched in 2019, and back then it really was limited. Since then WebSockets and gRPC streaming have arrived. You can run sidecars next to your app, an OpenTelemetry collector for example, and GPUs are there if you need inference. Batch work and scheduled tasks go to Cloud Run Jobs. With Direct VPC egress your service reaches Cloud SQL or internal VMs over private IPs, without a connector in between.
Background work is the objection I hear most often, and it's outdated as well. With instance-based billing the CPU stays allocated between requests. For consumers that never receive an HTTP request at all, say a Pub/Sub or Kafka worker, Google added worker pools. It's worth checking whether they're out of preview before you build on them.
Cold starts are the real trade-off
The catch with scale to zero is the first request after a quiet phase. Cloud Run has to start a fresh instance, and whoever sent that request waits for it. For a small Go service that's usually a few hundred milliseconds, while an untuned Spring Boot app can easily need several seconds.
That one hits close to home, since most of my backends are Spring Boot, the Lieferhand API included. On the Hetzner box startup time never mattered because the app just keeps running. If I moved it to Cloud Run with scale to zero, it's the first thing I'd look at.
The common fix is a minimum instance count. Keeping one instance warm removes most cold starts, but you pay for it around the clock, at a reduced rate for idle time. Depending on CPU and memory that ends up at a few tens of euros a month, which is still little compared to running a cluster.
There are two more levers that don't cost anything extra. Startup CPU boost gives the container additional CPU while it boots. And the app itself can start faster, for example with Spring AOT or a GraalVM native image.
GKE gives you everything, including the work
With GKE, Google runs the Kubernetes control plane and leaves the rest to you. That means node pools, requests and limits, ingress and autoscaling, often with a service mesh and a monitoring stack on top. Some workloads need exactly that. Databases or message brokers with persistent volumes run well on a cluster, and custom operators or raw TCP and UDP traffic have no place on Cloud Run anyway.
I spent a good part of my time at Deutsche Bahn on this kind of setup. In passenger information we designed and built a message broker component on Kafka and RabbitMQ, packaged the services as Helm charts and deployed them to Kubernetes through GitLab CI pipelines. That cluster ran on AWS EKS rather than GKE, but the reasoning carries over directly. A landscape of services talking to each other through brokers is pretty much the textbook case for Kubernetes.
GKE Autopilot is a middle ground. You still work with Kubernetes objects, but Google manages the nodes and bills you per pod. For teams that need the Kubernetes API and don't want to look after node pools, it's often the sensible choice, even if Kubernetes itself stays just as complex.
Operating it feels different too. Getting a fresh GKE cluster up takes several minutes before the first pod runs, and afterwards someone has to tune the HPA and possibly the VPA. That someone should have real Kubernetes experience, which not every team has.
What it costs in practice
For a rough comparison I took a typical REST API in europe-west1, 100,000 requests a day with an average response time of 200 ms. On Cloud Run with 1 vCPU, 512 MB and scale to zero, list prices put that at around €15 to €25 a month. Good concurrency usually pushes it lower, and part of it falls into the free tier anyway.
The same API on GKE Autopilot with two replicas for availability comes to roughly €50 to €70 including the load balancer. That only works out if the free tier credit covers your cluster fee, which it does for one Autopilot or zonal cluster per billing account. Each further cluster adds about €70 a month in management fees. With GKE Standard and three e2-medium nodes running around the clock, I end up at €85 to €100.
These are estimates based on list prices, not numbers from an actual bill. Your own setup can look quite different, so it's worth running it through the Google Cloud pricing calculator before you make a decision based on cost.
The Terraform difference
On Cloud Run, the whole service fits in about twenty lines:
resource "google_cloud_run_v2_service" "api" {
name = "my-api"
location = "europe-west1"
template {
containers {
image = "europe-west1-docker.pkg.dev/my-project/services/my-api:1.4.2"
resources {
limits = {
cpu = "1000m"
memory = "512Mi"
}
}
}
scaling {
min_instance_count = 0
max_instance_count = 10
}
}
}The image is pinned to a fixed version on purpose. With :latest Terraform has no way of seeing that anything changed, and at some point nobody knows for sure what's running in production.
To get the same service running on GKE you need a VPC, the cluster and either a node pool or an Autopilot config. On top of that come a Deployment, a Service, an Ingress and an HPA. You're easily past 150 lines before the first request gets through, and every one of them has to be maintained.
How I decide
I start with state. If the service needs persistent storage or holds state itself, it goes to GKE. Otherwise I check whether it needs Kubernetes as a platform, for example because of custom operators or protocols beyond HTTP and gRPC. If neither applies, it goes on Cloud Run.
I rarely switch just because of latency. When cold starts become a problem, a minimum instance usually fixes it for a fraction of what a cluster costs.
Plenty of setups use both anyway, with the APIs on Cloud Run and anything stateful on GKE in the same VPC.
If a service outgrows Cloud Run, moving it to GKE is manageable. Cloud Run follows the Knative serving model, so the container stays exactly the same and mostly the deployment config changes. The opposite direction tends to be harder, because over the years a lot grows into a cluster that you'd first have to untangle. For new stateless services I therefore start on Cloud Run and only move them when there's a concrete reason to.


