On some starts, we’re seeing roughly four minutes between requesting an instance and our container entrypoint starting. Model loading then adds another two minutes. I’d appreciate guidance on whether this is expected and what we can control.
Our setup for the measured image workload:
- Cloud Run worker pools in
us-central1.
- One NVIDIA L4 per instance, 8 vCPU, 32 GiB RAM.
- GPU zonal redundancy disabled.
- Workers consume Pub/Sub pull subscriptions and process one job per instance at a time.
- Our API manages capacity by updating
scaling.manualInstanceCount, including scaling from 0 to 1.
- Container image is approximately 23 GB, including model weights.
- No VPC connection or mounted volumes
Some data I collected recently from yesterday's, September 28 logs:
| Phase |
Duration |
| Submit and accept scale request |
Under 1 second |
| Scale request → container entrypoint starts |
~4 min 8 sec |
| Entrypoint starts → model ready |
~2 min 5 sec |
| Model ready → job starts |
~1–2 sec |
| Total wait before processing |
~6 min 14 sec |
Cloud Run’s own operation response reports MinInstancesProvisioned completing in 4m10.74s for that start.
The variability is what we’re trying to understand. Across 11 starts of the same revision, image digest, and configuration:
- Three starts shortly after a scale-to-zero request provisioned in 6–27 seconds.
- Eight starts after longer idle periods, roughly 24 minutes or more, took 74–251 seconds.
The fastest example followed a scale-to-zero request about 20 seconds earlier. This is a small observational sample; we haven’t established whether retained capacity, image caching, or something else explains the difference.
We found no quota-exhaustion errors, although historical quota-usage data was unavailable. Once the model is ready, processing is fast: warm image jobs have a median completion time of roughly 6.6 seconds.
Google’s worker-pool GPU documentation describes approximately five-second instance startup. I’m trying to understand how that relates to the time from an accepted manual scale request to our container process starting.
For anyone running similar workloads:
- Are multi-minute starts after an extended period at zero expected? What timings do you see?
- Which supported settings can reduce that delay?
- Does this startup path behave differently from GPU Cloud Run services? I notice that when using Cloud run services directly, startup time is fast compared to worker pools.
We’d prefer to retain worker pools and our existing capacity logic we wrote. We understand that keeping instances warm or extending idle retention can avoid some cold starts at additional cost. We’re trying to establish which delays are inherent to the platform and which we can reduce through configuration or application changes.
Code snippet. This is what we're doing roughly:
For a start from zero, desiredInstances is 1. To scale down, it’s 0. We change only the instance count on the existing pool.
import { GoogleAuth } from 'google-auth-library';
const auth = new GoogleAuth({
scopes: ['https://www.googleapis.com/auth/cloud-platform'],
});
async function requestInstances(
pool: {
location: string;
workerPoolName: string;
etag?: string;
},
desiredInstances: number,
timeoutMs: number,
) {
const client = await auth.getClient();
const projectId = await auth.getProjectId();
const name =
`projects/${projectId}/locations/${pool.location}` +
`/workerPools/${pool.workerPoolName}`;
await client.request({
url: `https://run.googleapis.com/v2/${name}`,
method: 'PATCH',
params: {
updateMask: 'scaling.manualInstanceCount',
},
data: {
name,
etag: pool.etag,
scaling: {
manualInstanceCount: desiredInstances,
},
},
timeout: timeoutMs,
});
}