# Groq status

> Is Groq down? Live status for the Groq API, uptime history, and per-region response times measured by our own probes every 2 minutes.

Published: 2026-08-30 | Canonical: https://yoping.me/status/groq

## How we check Groq

We watch api.groq.com from all five of our probe regions, every 2 minutes. A region that sees a failure does not make this page say "down" on its own: a second region has to agree first, the same confirmation rule every YoPingMe monitor uses.

We watch the API host applications send inference requests to and pin the response an unauthenticated request receives, rather than expecting a 2xx. The console and marketing site are served separately and can look entirely healthy while the inference API is refusing or slowing traffic, which is the failure that reaches your users.

The live board - the current verdict, 24-hour, 7-day, and 30-day uptime, and per-region response times - is on the HTML page at https://yoping.me/status/groq. Those numbers change too fast to repeat honestly in a static mirror, so where an answer below says "the board above", it means that page.

## What is Groq?

Groq runs inference on custom hardware designed for the job, and speed is
the product. Where general-purpose accelerators are optimised for
training and adapted for serving, Groq's architecture targets the
generation of tokens one after another, which is what makes the
difference noticeable in an interactive application. Teams choose it when
the response time of a model is a user-facing feature rather than an
implementation detail.

That focus shapes what an incident looks like. For an application built
around fast responses, a degradation that leaves every request succeeding
but doubles the time to produce them is a genuine problem, even though
nothing has failed and no status page will mark it. This is the reason
the per-region response times above sit next to the up-or-down verdict
rather than behind it. A monitor that only asks whether a request
succeeded would report a perfect day through exactly the incident most
likely to affect a product whose value is speed.

Rate limiting is the most common cause of errors that get mistaken for
outages. Limits are enforced per model and per account, across both
request counts and token throughput, and a batch job or a traffic spike
can consume the allowance quickly. The response is a 429 with headers
describing what remains, which is the API behaving correctly and telling
you so. Applications that treat every non-200 as an outage will report
one; applications that read the headers will back off and continue.

Model availability is the second thing to separate from platform health.
Hosted inference providers add and retire models on their own schedule,
and a request naming a model that has been deprecated fails while
everything else works. The error is specific and names the model, but
integrations that log a generic failure hide that detail, which turns a
one-line configuration change into an outage investigation.

Finally, it is worth thinking about what your application should do when
inference is unavailable. Products that call a model in the request path
have made that model a hard dependency of the feature, and often of the
page. A fallback that degrades to a simpler behaviour, a queued retry, or
even an honest message beats a spinner that never resolves, and the
decision is much easier to make in advance than during the first
incident.

## Frequently asked questions

### Is Groq down right now?

Check the board above; it reflects our own probes against Groq's API from all five of our regions, refreshed every couple of minutes and confirmed across two regions before this page would call it down.

### I am getting 429 errors from Groq. Is that an outage?

No, a 429 is rate limiting and means the API answered you. Limits apply per model and per account across requests and tokens, and a burst of traffic or a batch job can exhaust them while the service is entirely healthy. The response headers report your remaining allowance, which tells you immediately whether you are being limited or something is actually wrong.

### My requests are slower than usual. Does that count as downtime?

Not as downtime, but it is worth watching, which is why the response times above are shown alongside the up or down state. Inference latency varies with model, prompt length, and current load, and a degradation that never fails a request can still break a product built around fast responses. A page that only tracked availability would call that a perfect day.

### Does a model being unavailable mean Groq is down?

No, and this is a common source of confusion. Models are added, deprecated, and occasionally taken out of service individually, so a request naming a model that no longer exists returns an error while every other model responds normally. The error names the model, and the current model list in Groq's documentation is the place to confirm.

### How is this page different from Groq's own status page?

Theirs reports what Groq's engineers have confirmed and published, including per-component detail. This page reports what our probes see from the outside, on a fixed schedule, independent of when or whether an update gets posted. During an incident, the two together are more useful than either alone.
