Skip to main content

Scale and availability

Four different questions share one screen in the console, and it helps to know which is which.

The questionThe name
How many copies of the application exist?Instances
How much can each one use?Instance size — see Resources and limits
Does the number follow the load?Autoscaling
Does it survive losing a machine?High availability

Instances​

The number of copies of your application running at the same time. More instances serve more concurrent requests and consume proportionally more resources.

The console shows how many are up and how many were requested — "3 of 4 up". The difference matters: showing only the requested number would make the screen agree with whoever asked and disagree with reality.

Autoscaling​

Instead of a fixed number, you declare a range and a CPU usage target. The platform adds instances when usage goes above the target and removes them when it stays below for a few minutes.

FieldWhat it is
MinimumThe floor. Never fewer than this, even with no traffic
MaximumThe ceiling. Scale never goes past it
CPU usage to scaleThe target, as a percentage of each instance's reserved CPU

The target is measured against the reservation, not the maximum. That is why every service needs declared reserved CPU: without it autoscaling has nothing to compare against and simply does not happen.

70% is a reasonable starting point: it leaves room for the peak between two evaluations. A 90% target keeps the application near its ceiling all the time, and scaling only reacts after latency has already risen.

Scaling up is fast; scaling down is slow​

The platform adds capacity without waiting — a spike needs instances now. To scale down, it waits five minutes and uses the highest recommendation in that window.

The asymmetry is deliberate. Scaling down quickly produces sawtooth behaviour: load drops, instances disappear, traffic returns to the survivors, and they scale up again.

Minimum of one, never zero​

An online application with zero instances is an application that is down without anyone having said so. The accepted minimum is 1.

High availability​

A service's instances are spread across different machines, automatically. It is not a setting to turn on: it is the behaviour of every published application.

Without it, four instances can land on the same machine — four processes with a single fate. With spreading, losing a machine means reduced capacity, not an application that is down.

Maintenance respects this too: when the platform needs to drain a machine, it removes one instance at a time and waits for the replacement to be ready before continuing.

The exact reach of the protection​

Spreading protects against the failure of one machine. It does not protect against losing the availability zone — the machines in the current installation live in a single zone. The same notice appears in the console: the product does not promise what the infrastructure does not deliver.

When capacity runs out​

If you ask for more instances than fit in the pool, the ones that do not fit wait for room. The platform raises an alert when that happens — there is no automatic machine growth in this installation, so increasing capacity is a decision for whoever operates it.

Adjusting​

  1. Open the project and go to Scale and availability.
  2. Pick the service and click Adjust.
  3. Choose Fixed number of instances or Autoscaling.
  4. For autoscaling, set the minimum, maximum and CPU target.
  5. Save.

Like every configuration change, it takes effect once applied, not once saved: the platform republishes the service with the same image and follows it to the end.

When scaling does not happen​

The console explains instead of staying silent:

What the screen saysWhat to check
Not receiving the usage measurementThe service has no reserved CPU declared
At the configured limitThe range hit its maximum — or its minimum
Instance with nowhere to runThe pool ran out of capacity; talk to whoever operates the platform