Started
Resolved
Duration
12 hrs 31 min
Update timeline
This incident has been resolved.
Fix has been implemented in Prod 1 and we are monitoring for any further issues.
Fix has been implemented in Prod Eu 1 and we are monitoring for any further issues.
A fix has been implemented and we are monitoring the results.
The issue has been identified and a fix is being implemented.
We are currently investigating this issue.
# Root Cause Analysis Report ## IaCM Module Registry Timeouts ## 1. Summary On Monday, September 15, 2026, some customers using Infrastructure as Code Management \(IaCM\) Module Registry on Prod EU1, and later Prod1, experienced timeouts and failed requests. Affected customers saw Module Registry operations hang or fail in the UI, and IaCM pipelines that look up modules or module-registry connectors during Initialize or Plan could stall and then expire. This was not a platform-wide outage. Other Harness modules were not broadly impacted. We have no report of customer data loss. Workspace plan/apply that did not need Module Registry during the failing calls continued as normal. # 2. Impact * Product area: IaCM Module Registry \(UI and API\) and IaCM pipelines that call Module Registry during Initialize or Plan. * Customer experience: timeouts and failed Initialize/Plan steps; Module Registry UI/API delays or failures. In at least one case, the Module Registry CLI continued to work while the plan path timed out. * Environments: Prod EU1, then Prod1. This was not a full multi-cluster outage. * Severity: S-2. * Data: no data loss reported. # 4. Root Cause A sudden traffic surge caused backend connections to hit capacity limits, leading to client-side timeouts. Increasing capacity and restarting the IaCM service cleared remediated the issue # 5. Remediation **Immediate:** restarted the IaCM service to release stale connections and restore Module Registry and related pipeline steps. **Capacity:** increased the backend capacity for Module Registry database connection **Permanent:** Tune and optimize the backend for latency and resiliency so that we can absorb traffic spikes without impact. # Next Steps **To prevent such issues happening again** _We apologize for the disruption to Module Registry and to pipelines that depend on it. Please reach out to your Harness support contact if you have any further questions about this incident._