Groq Status: Real-Time Visibility Into GroqCloud's System Health and Status
Data Center Failure Impacting Capacity

Started

Resolved

Duration

1 hr 31 min

Resolvedminor impact

Update timeline

Resolved

The **capacity issue has been resolved**. All services are operating at normal capacity. The incident was caused by a **power loss issue that led to a cooling system failure at one of our US Central data centers**. We apologize for any disruption and **appreciate your patience**.

Monitoring

We have **redistributed and allocated more capacity to production models** in **the US Central region.** Users should **see performance return to normal**. We’re now **monitoring** the infrastructure to ensure stability.

Identified

A power loss issue at approximately 22:20pm UTC led to a subsequent **cooling system failure** at one of our **US Central** data centers that is causing **reduced capacity**, primarily for the `openai/gpt-oss-20b` model. The team is **working on restoring capacity**.

Identified

We have **identified a cooling system failure** at one of our **US Central** data centers that is causing **reduced capacity**. The team is **working on restoring capacity**. Users **continue to see elevated latencies for specific models**.

Investigating

We are currently investigating a potential issue at one of our US Central data centers that is impacting system capacity. Users **may experience higher latencies** due to reduced resources. Our engineering team is working to identify the root cause and remediate the problem.

Data Center Failure Impacting Capacity — Groq Status: Real-Time Visibility Into GroqCloud's System Health and Incident Timeline & Status — DevHelm