Started
Resolved
Duration
less than a minute
Update timeline
Between 7:00 AM and 7:20 AM EST, we observed an increase in 5xx error rates, which may have resulted in intermittent errors or failed requests for some customers. The error rate has since returned to normal, and we are investigating the root cause to ensure we have fully resolved the issue. We apologize for any disruption and appreciate your patience.
**Root Cause Analysis** **Incident:** Intermittent elevated 5xx errors across services **Date Range:** September 9, 2026, 7:00 AM – 9:20 AM EDT \(approximately 2 hours 20 minutes of intermittent impact\) **Impact:** Some requests submitted to the EmailAuthScore API, and related services, intermittently returned error responses. Impact was intermittent rather than continuous and was limited to defined windows described in the timeline below. **Status:** Resolved ### **1. Summary** On September 9, 2026, a planned increase in backend processing capacity led to elevated connection load on an internal infrastructure component used to support downstream processing for several services. This component reached an internal capacity threshold, which introduced processing delays and resulted in intermittent error responses on the EmailAuthScore API and related services for a subset of requests. The issue was identified through a combination of internal monitoring and reports during the impact window. Socure engineering correlated the errors with the recent capacity change and reverted it, which stopped the initial increase in errors. A second, related increase in errors then occurred when a separate capacity change, made earlier for a different internal service, placed additional load on the same shared component; engineering applied further mitigations, after which all affected services returned to normal, stable operation. Section 5 outlines the corrective and preventive actions Socure is committing to, with priority on improving visibility into infrastructure capacity and preventing delays in any single internal component from affecting customer-facing request processing. ### **2. Timeline** | **Date / Time \(EDT\)** | **Event** | | --- | --- | | **September 9, 2026 07:00 AM** | Following a planned increase in backend processing capacity, the EmailAuthScore API began intermittently returning error responses. | | **September 9, 2026 07:20 AM** | The recent capacity change was reverted; error rates returned to normal. | | **September 9, 2026 08:45 AM** | A separate capacity change made earlier for a different internal service placed additional load on the same shared internal component, and the EmailAuthScore API again began intermittently returning error responses. | | **September 9, 2026 09:20 AM** | Additional mitigations were applied; error rates returned to normal and remained stable. | | **September 9, 2026 01:13 PM** | As a precaution, Socure temporarily rerouted a portion of internal processing away from the affected internal component to safely complete the remaining capacity increase. No further errors were observed after this change. | ### **3. Root Cause** * **Primary Root Cause:** As part of planned work to increase system capacity ahead of anticipated demand, Socure increased the processing capacity of backend services. This change increased the number of connections to an internal messaging component used to support downstream processing for these services. That component reached an internal connection capacity limit, which resulted in connection and processing delays. These delays introduced additional latency and, in some cases, backpressure into request processing, which contributed to intermittent elevated error rates on the EmailAuthScore API and related services for a subset of requests during the impact windows described above.. * **Contributing Factors:** * A sudden increase in infrastructure capacity demand exceeded the capacity that was immediately not available within the internal messaging component. ### **4. Resolution** Upon identifying that the elevated error rates were linked to the recent capacity change, Socure engineering reverted the change, which stopped the initial increase in errors. A second, related increase in errors then occurred when a separate capacity change, made earlier for a different internal service, placed additional load on the same shared component. Engineering applied additional mitigations, including temporarily rerouting a portion of internal processing away from the affected component for the impacted services. Following these actions, all affected services returned to normal, stable operation, which Socure confirmed through continued monitoring. ### **5. Corrective and Preventive Actions** | **Action** | **Description** | **ETA / Status** | | --- | --- | --- | | Optimize infrastructure capacity utilization | Adjusted backend service capacity and connection usage to ensure processing remains within supported infrastructure limits and to reduce the risk of similar capacity-related issues. | Done | | Increase messaging infrastructure capacity | Increased the capacity of the underlying messaging infrastructure to provide additional headroom and support future increases in traffic volume. | Done | | Enhance capacity monitoring and alerting | Add proactive monitoring and alerts to identify potential infrastructure capacity constraints before they affect service performance. | In progress. ETA 30 Sep 2026. | ### **6. Lessons Learned** * Proactive monitoring of infrastructure health needs to be in place ahead of capacity changes, not only application-level monitoring. ### **7. Next Steps & Ongoing Commitment** We are committed to: * Completing the corrective and preventive actions listed in Section 5. * Improving visibility into the health and capacity of critical internal infrastructure ahead of future scaling activity. * Providing a follow-up update once all corrective actions are complete.