> ## Documentation Index
> Fetch the complete documentation index at: https://help.the-meridian.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Autoscaling

Meridian runs your app on Cloud Run and scales it with demand. You set the bounds; the platform adds and removes instances between them as traffic rises and falls.

## Scaling controls

Each environment has:

* **Minimum instances**: how many instances stay warm at all times. Every environment runs at 0, so it scales to zero when idle. A reserved instance is available on Enterprise as the [always-on server](/always-on-server), which Meridian turns on for the environment you name.
* **Maximum instances**: the ceiling autoscaling won't exceed, protecting you from runaway cost under a traffic spike. Your plan sets the most you can choose: up to 1 on Lite, 2 on Pro, 4 with Medium and 8 with Large. A production environment that sets no maximum runs with that ceiling. Staging and development always run on one instance.
* **Concurrency**: how many simultaneous requests a single instance handles before another is added.

CPU and memory per instance are configured alongside these on the environment.

## Choosing bounds

Scaling to zero suits most apps: Meridian keeps production environments warm on a best-effort basis (see below), and a cold start on the first request is acceptable for low-traffic or custom apps. An app that must answer instantly at all times needs the [always-on server](/always-on-server), which is part of Enterprise. See [hosting upgrades](/add-ons) for the instance limits each size includes.

## Warm instances without always-on

Meridian sends a small request to the production environment of every deployed app every four to five minutes. Staging and development environments are not kept warm: they scale to zero when idle, and the first request after a quiet spell pays a cold start. Cloud Run keeps an instance around for a while after each request, so in practice your app is usually warm when a merchant arrives, and the first request takes a fraction of a second rather than a full cold start. These requests are made by Meridian, are not counted in your bandwidth or uptime, and cost you nothing.

This is a best effort, not a guarantee. Cloud Run can still replace an instance on its own schedule, in which case the next request pays one cold start. If your app must never cold start, the [always-on server](/always-on-server) keeps a dedicated instance warm at all times. It is part of Enterprise only.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.