Flexera System Status Dashboard Status
Spot - Service Degradation

Started

Resolved

Duration

2 hrs 13 min

Postmortemmajor impactVendor link

Update timeline

Resolved

Our investigation determined that an internal service responsible for monitoring and resource state management became unresponsive, which contributed to the issues reported during the incident, including instances remaining in a Resuming state longer than expected, resource state discrepancies, and under-capacity alerts. Our technical teams performed recovery actions, including restarting the affected service, and have since observed stable service operation. Following validation and monitoring, we have confirmed that affected services have returned to normal operation. This incident is now considered resolved. A post-incident review is ongoing, and we will continue to evaluate opportunities for improvement.

Monitoring

We are continuing to observe stable service behavior across the affected areas, and the previously reported symptoms are no longer being observed. Our technical teams are closely monitoring the environment and validating service stability across affected customer workloads.We will continue to monitor and provide a further update once monitoring activities are complete or if additional information becomes available.

Investigating

Our technical teams are observing early signs of recovery and are actively monitoring service behavior while validating the impact across affected customer workloads. Investigation and recovery efforts remain ongoing. We will continue to provide updates as we learn more and confirm service stability.

Investigating

Incident Description:
 We are investigating an issue affecting portions of the Spot platform. Customers may experience instances remaining in a Resuming state longer than expected, under-capacity alerts, and discrepancies between resource status information displayed in Spot and AWS. In some cases, customers may observe differences between the status shown in the Spot console and the actual status reported by AWS for affected resources. Our technical teams are actively investigating the scope and impact of the issue. Priority: P2 Restoration Activity:
 Our technical teams are actively investigating the issue and assessing its impact across affected services and customer workloads. Current efforts are focused on identifying the underlying cause, validating the full scope of customer impact, and restoring normal service behavior. Further updates will be provided as more information becomes available.

Postmortem

**Description:** Spot Ocean - Resource State Discrepancies and Under-Capacity Alerts **Timeframe:** August 25, 2026, 6:15 PM PDT – August 25, 2026, 10:22 PM PDT **Incident Summary** On Tuesday, August 25, 2026, at 6:15 PM PDT, Ocean customers began experiencing service degradation affecting resource state visibility and management operations. During the incident, customers may have observed instances remaining in a Resuming state longer than expected, under-capacity alerts, and discrepancies between resource status information displayed in Spot and AWS. In some cases, resource state information displayed within Spot did not accurately reflect the actual state of resources in AWS. Technical teams investigated the issue and identified degradation within a central Ocean service responsible for processing resource monitoring and management requests. As a result, monitoring and management operations slowed or stopped, resulting in broader service degradation across Ocean. The affected service was restarted, restoring normal request processing and service functionality. Following validation and monitoring activities, service stability was confirmed and the incident was resolved. **Root Cause** Investigation determined that a service responsible for processing resource monitoring and management requests experienced prolonged delays while communicating with a dependent service. As those delays accumulated, the service became unable to effectively process new requests. Although the service continued appearing operational, it was no longer able to process requests normally. This resulted in delayed resource state updates and contributed to the customer-facing symptoms observed during the incident, including resource state discrepancies, under-capacity alerts, and instances remaining in a Resuming state longer than expected. **Remediation Actions** The following actions were taken during the incident response: 1. **Incident Investigation Initiated:** Technical teams responded to production alerts and customer reports to assess the scope and impact of the issue. 2. **Service Degradation Identified: I**nvestigation determined that the affected service was no longer able to effectively process new monitoring and resource management requests. 3. **Service Recovery Performed:** The affected service components were restarted, restoring normal request processing and service functionality. 4. **Post-Recovery Validation Completed:** Technical teams monitored service behavior and validated stability before declaring the incident resolved. **Future Preventative Measures** Technical teams have initiated follow-up work to reduce the likelihood of similar issues occurring in the future. 1. **Monitor Service Reliability Improvements:** Technical teams have identified an existing service behavior that contributed to this incident and are implementing improvements to enhance overall service reliability. 2. **Service Design and Configuration Review:** Technical teams are reviewing the service design, configuration, and request processing behavior involved in this incident to identify additional opportunities to improve resiliency. 3. **Additional Investigation and Corrective Actions:** Technical teams are continuing to assess the findings identified during the investigation and will implement any additional corrective actions determined to be relevant to this incident.