# Replicate status

> Is Replicate down? Live status for the Replicate API, uptime history, and per-region response times from our own probes on three continents.

Published: 2026-08-30 | Canonical: https://yoping.me/status/replicate

## How we check Replicate

We watch api.replicate.com from all five of our probe regions, every 2 minutes. A region that sees a failure does not make this page say "down" on its own: a second region has to agree first, the same confirmation rule every YoPingMe monitor uses.

We watch the API host your application creates predictions against and pin the response an unauthenticated request receives rather than assuming a 2xx. Model execution happens on separate infrastructure that can queue or slow independently, so this measures whether the API is reachable, not how long a given model takes to run.

The live board - the current verdict, 24-hour, 7-day, and 30-day uptime, and per-region response times - is on the HTML page at https://yoping.me/status/replicate. Those numbers change too fast to repeat honestly in a static mirror, so where an answer below says "the board above", it means that page.

## What is Replicate?

Replicate runs machine learning models as an API. Rather than
provisioning hardware, installing dependencies, and keeping a model
loaded, you make a request naming a model and receive a prediction, with
the platform handling packaging, scheduling, and scaling. It lowers the
cost of using a model from an infrastructure project to an HTTP call,
which is why it turns up in products that would never justify running
inference hardware of their own.

The asynchronous model is the first thing to understand, because it
changes what failure looks like. A request creates a prediction, and that
prediction moves through states while the work is scheduled and
executed. There is no single moment where a request either succeeds or
fails; there is an object whose state you poll or receive a webhook
about. An application that treats prediction creation as the whole
interaction will report success for work that later failed, and one that
relies purely on webhooks will lose track of runs whose notification was
not delivered. Reconciling against the API is the only reliable way to
know what happened.

Cold starts are the second property worth planning around. Models are
loaded onto hardware when needed, and loading a large model takes
substantially longer than running it. A model called continuously stays
warm and feels fast; a model called occasionally pays that cost on nearly
every request. From the outside this looks like wildly inconsistent
performance, and it is a predictable consequence of running many models
on shared capacity rather than a sign of instability. Applications that
need consistent latency either keep a model warm deliberately or design
the interface around the wait.

Queueing follows from the same design. When capacity for a model is
scarce, predictions wait, and a queue is not an error. This is where an
outside availability signal has real limits, and it is worth being honest
about that: our probes can tell you the API is answering, which does not
tell you that a particular model has capacity right now. The prediction's
own state is the authority on that, and no external monitor can
substitute for it.

The product design lesson is the same one that applies to any inference
provider. Work that runs asynchronously should be presented
asynchronously, with a state your users can see and a path that does not
end in a spinner. Products that hide a queue behind a synchronous request
break in a particularly unpleasant way the first time the queue gets
long.

## Frequently asked questions

### Is Replicate down right now?

Check the board above; it reflects our own probes against Replicate's API from all five of our regions, refreshed every couple of minutes. This measures whether the API responds, which is narrower than whether a particular model is running quickly.

### My prediction is stuck in the queue. Is Replicate down?

Usually not. A prediction sits in the queue while capacity for that model is found, and a cold model has to be loaded onto hardware before it can run, which takes noticeably longer than a warm one. The prediction object reports its own state, and a status of starting rather than failed means the platform is working through it rather than dropping it.

### Why is the first request to a model so much slower than the rest?

Because the model has to be loaded before it can run, and that cold start dominates the first request. A model used constantly stays warm and responds quickly; one called a few times a day pays the loading cost almost every time. This is a property of running many models on shared hardware rather than a fault, and it will never appear on a status page.

### Does an outage lose predictions that were already running?

Predictions are tracked as objects with their own state, so the right move after any incident is to check their final status rather than assume. A prediction can end as succeeded, failed, or canceled, and a webhook that was not delivered during the incident does not change what actually happened to the run. Reconciling from the API rather than from your webhook log is the reliable path.

### How is this different from Replicate's own status page?

Theirs reports what Replicate's team has confirmed and published from inside the platform. Ours comes from outside it, on a fixed schedule, independent of whether anyone has written an update yet - narrower in detail, but it does not wait on a human to post it.
