Redox Status
Log visibility is delayed

Started

Resolved

Duration

5 hrs 56 min

Postmortemminor impactVendor link

Update timeline

Resolved

This incident has been resolved. Log visibility is restored. If you have any questions, please reach out to support@redoxengine.com

Monitoring

A fix has been implemented and visibility for the logs in the backlog is returning. We are actively monitoring the situation and ensuring that all affected logs are visible again.

Identified

We have pushed a fix that will make new logs visible, however logs received between approximately 5:45PM and 8:40AM will still not be viewable. The backlogs of logs will continue to process through the queue while it catches up.

Identified

We have identified the cause for the log visibility issues and are actively working on a fix. We will update the status of this incident when the fix has been deployed.

Investigating

We are aware of an issue where logs are not visible in the dashboard. While logs are still processing properly, the dashboard is not displaying them as expected. If you have questions about specific logs or your connection and whether it has been affected, please contact us at support@redoxengine.com

Postmortem

## Summary Following an emergency failover carried out in response to a third-party outage on September 8, a configuration change made to stabilize that failover infrastructure reduced capacity on the system powering dashboard and log visibility. Message processing itself was not affected — data continued to be sent and received normally throughout this incident. Customers began experiencing delayed visibility into message logs and payload details overnight; we identified the issue by 7:45 AM Central on September 9 and fully resolved it by 1:51 PM Central. No data was lost at any point. ## What Happened As part of stabilizing the backup infrastructure used to restore service during a separate third-party outage earlier that day, our team identified an opportunity to make its scaling configuration more resilient to future failover deployments, and moved quickly to put that improvement in place. That change was deployed around 4:34 PM Central on September 8. It did not have the intended effect, and capacity on the backup system supporting dashboard and log visibility was reduced as a result. Message processing itself was not affected. Monitoring coverage on that temporary infrastructure was still being extended at the time, so the reduced capacity was not immediately flagged, and a visibility backlog accumulated overnight. ## Impact Customers experienced delayed visibility into message logs and payload details in the Customer Dashboard, beginning overnight and continuing into the morning of September 9. Message processing was not affected — data continued to be sent and received normally. The impact was limited to the dashboard and log visibility layer. ## How We Resolved It Customer reports and internal alerts led our team to begin investigating at approximately 7:30 AM Central on September 9. Within 15 minutes, we identified that the failover infrastructure was underscaled and began remediation. Between 8:00 AM and 12:00 PM Central, we: * Increased infrastructure capacity to add processing headroom * Scaled the services powering dashboard and log visibility * Distributed traffic across infrastructure to accelerate recovery Visibility into message logs returned to fully caught-up status by 12:51 PM Central. We returned to normal operating capacity by 1:51 PM Central and confirmed resolution on our status page. ## What We're Doing About This * Re-evaluate failover playbooks, so scaling, monitoring, and cleanup steps are rehearsed and routine * Require monitoring and alerting to be fully wired up on any temporary or backup infrastructure before a team stands down from an incident * Add automated deployment tests to catch configuration resets before they reach production, including on temporary or failover infrastructure * Consider implementing absolute minimum scale limits so services cannot drop below production-safe thresholds, while preserving a standard, well-documented method for intentionally adjusting capacity when needed We appreciate your patience during this incident. If you have questions or concerns, please reach out to your Redox account team. ‌ _This issue arose while our team was stabilizing infrastructure used in an emergency failover for a separate third-party outage. See \[Incident 1: Third-Party Service Outage & Processing Delay\] for details._

Log visibility is delayed — Redox Incident Timeline & Status — DevHelm