Skip to content

Incident detail

Data Center Failure Impacting Capacity

Resolved incidentMinor

Timeline window

to

Outage alerts

Get alerted the next time Groq breaks

Free email alerts for up to 5 providers. No card, live in about a minute. Paid plans add Slack, Teams, Discord, and webhook delivery across your whole stack, plus higher API quotas.

Timeline

Incident updates

Every update Groq posted, oldest to newest, exactly as it appeared on their official status page.

  1. Investigating

    We are currently investigating a potential issue at one of our US Central data centers that is impacting system capacity. Users may experience higher latencies due to reduced resources. Our engineering team is working to identify the root cause and remediate the problem.

  2. Identified

    We have identified a cooling system failure at one of our US Central data centers that is causing reduced capacity. The team is working on restoring capacity. Users continue to see elevated latencies for specific models.

  3. Identified

    A power loss issue at approximately 22:20pm UTC led to a subsequent cooling system failure at one of our US Central data centers that is causing reduced capacity, primarily for the `openai/gpt-oss-20b` model. The team is working on restoring capacity.

  4. Monitoring

    We have redistributed and allocated more capacity to production models in the US Central region. Users should see performance return to normal. We’re now monitoring the infrastructure to ensure stability.

  5. Resolved

    The capacity issue has been resolved. All services are operating at normal capacity. The incident was caused by a power loss issue that led to a cooling system failure at one of our US Central data centers. We apologize for any disruption and appreciate your patience.

Keep exploring

More from Groq

Neighboring incidents on Groq's timeline and the rest of their record on OutageDeck.