Kustomer Status
Platform Events Service Disruption (PROD1)

Started

Resolved

Duration

3 hrs 21 min

Postmortemminor impactVendor link

Update timeline

Resolved

Kustomer has resolved an event affecting platform events in PROD1 orgs that caused delays on event-based data. After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at support@kustomer.com if you have additional questions or concerns.

Monitoring

Kustomer has implemented an update to address an event affecting platform events in PROD1 orgs that caused delays on event-based data. Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer Support at support@kustomer.com if you have additional questions or concerns.

Identified

Kustomer continues to work on the issue affecting platform events in PROD1 orgs that may cause delays on event-based data. Our team is actively working to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer Support at support@kustomer.com for any further questions or updates.

Identified

Kustomer has identified an event affecting platform events in PROD1 orgs that may cause delays on event-based data. Our team is currently actively working to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer Support at support@kustomer.com for any further questions or updates.

Identified

Kustomer has identified an event affecting platform events in PROD1 orgs that may cause delays on event-based data. Our team is currently working to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer Support at support@kustomer.com for any further questions or updates.

Postmortem

## Summary On July 25, 2026, some customers experienced elevated platform latency affecting messaging, conversation updates, routing, voice, and other customer-service workflows. Customers may have seen messages remain in a sending state, delayed inbound or outbound messages, slower page and API responses, delayed routing or assignment, and intermittent voice-call delays. The incident was caused by an unusually large burst of background data-processing activity. This created a sudden increase in demand on shared platform infrastructure. Automatic scaling added capacity, but it could not absorb the burst quickly enough to prevent degradation in dependent workflows. We restored service by increasing available capacity, reducing the rate of the initiating workload, and carefully processing delayed work while monitoring platform health. ## Impact * **Customer effect:** Intermittent platform latency; delayed inbound and outbound messaging; slower conversation and API updates; delayed routing or assignment; and intermittent voice-call delays * **Duration:** Customer-visible degradation began at approximately 5:22 PM ET. Service initially recovered at approximately 7:31 PM ET, but degradation later recurred. Broad platform performance was restored and the incident was resolved at 10:30 PM ET. * **Scope:** The incident affected a subset of customers within one production environment. We did not identify corresponding customer-facing degradation in other production environments. ## Timeline * **Approximately 5:22 PM ET:** Processing latency and message backlogs began increasing. * **6:51 PM ET:** The incident response team began coordinated investigation and recovery. * **Approximately 7:00–7:23 PM ET:** We identified affected processing paths and began carefully recovering delayed work at controlled rates. * **7:31 PM ET:** Additional capacity had come online, customer-facing performance had materially improved, and the incident was initially resolved while monitoring continued. * **8:28 PM ET:** The incident was reopened after renewed reports of intermittent degradation. * **Approximately 8:38 PM ET:** Investigation confirmed that messaging, routing, channel, and voice symptoms shared the same underlying platform dependency. * **Approximately 9:52 PM ET:** We reduced the rate of the initiating background workload to protect customer-facing traffic. * **10:08 PM ET:** Monitoring confirmed that workload pressure had fallen substantially and platform health continued improving. * **10:30 PM ET:** After continued monitoring confirmed recovery, the incident was resolved. * **After resolution:** Recovery of remaining delayed work continued at controlled rates under monitoring. ## Root cause A burst of high-volume background data-processing activity generated more downstream work than shared platform infrastructure could safely absorb over a short period. Automatic scaling responded and added capacity, but capacity was added incrementally and did not come online quickly enough for the size and speed of the burst. While the platform was catching up, increased request latency and timeouts affected customer-facing workflows that depended on the same infrastructure. ## Resolution We restored platform performance by: * Increasing available capacity and improving the rate at which additional capacity could be added. * Reducing the rate of the initiating background workload. * Recovering delayed work at controlled rates to avoid creating another traffic spike. * Continuing monitoring after customer-facing performance returned to expected levels. ## Preventative actions We are taking the following actions to reduce the likelihood and impact of recurrence: * Add stronger limits and backpressure controls for high-volume background operations. * Improve the speed at which shared infrastructure scales during sudden traffic increases. * Improve isolation between background processing and latency-sensitive customer workflows. * Expand alerting for rapid backlog growth, resource pressure, and unusual workload patterns. * Strengthen controlled recovery procedures for delayed work. * Expand burst-load and failure-recovery testing for shared platform services. ## Current status Platform performance returned to expected levels, and the initiating workload remained controlled. We continued monitoring service health and the recovery of delayed work after resolution.