Started
Resolved
Duration
1 hr 4 min
Update timeline
This incident is now resolved.
Our upstream GPU provider has applied a mitigation and gemma-4 error rates have returned to normal. We have alternate routing paths staged and ready to enable if the issue recurs. We are monitoring before marking this resolved.
We've confirmed the issue originates with our GPU provider, who has declared a networking incident on their side. We're testing mitigations to route gemma-4 traffic elsewhere. Conversations using other LLMs are unaffected. If you need stability now, you can point your persona at a different LLM.
The issue is isolated to our default LLM (gemma-4), served by a GPU provider that began returning errors at 17:40 UTC. As a temporary mitigation we are routing gemma-4 traffic to an alternate model. Customers on gemma-4 may notice a change in response style while this is in effect; we will revert once the provider recovers.
This is affecting gemma-4; we recommend our customers switch to a different LLM for immediate needs.
Starting at 17:40 UTC, conversations using our default LLM began experiencing long delays and unresponsive replicas. A GPU provider is returning errors and our automatic failover is recovering most turns, but with significant added latency. We are routing traffic to a backup provider.
It appears this issue is coming from one of our downstream providers, Crusoe. We're continuing to investigate full scope of conversations affected.
We are aware of an issue currently impacting tavus conversations, we are investigating.