Recurring EU1 Service Disruptions (August–September 2026)

Uptime Impact: 3 hours and 44 minutes
Resolved
Resolved

Root Cause Analysis — Recurring EU1 Service Disruptions (August–September 2026)

Affected environment: EU1 Incidents covered:

  • August 13, 2026 — 11:57–12:34 CET (37 minutes)
  • August 31, 2026 — 10:11–10:24 CET (13 minutes)
  • September 1, 2026 — 10:33–13:37 CET (~3 hours 4 minutes)

Summary

Between August 13 and September 1, 2026, the EU1 environment experienced three separate service disruptions affecting our RabbitMQ messaging cluster. In each case, a brief initial interruption led to extended instability that required manual intervention from our engineering team to resolve — ranging from a short manual restart to several hours of controlled recovery involving active management of client reconnections.

Timeline

  • August 13, 2026, 11:57 CET — RabbitMQ interruption detected on EU1. Manual intervention was performed; stability was restored by 12:34 CET.
  • August 31, 2026, 10:11 CET — RabbitMQ interruption detected on EU1. Manual intervention was performed; stability was restored by 10:24 CET.
  • September 1, 2026, 10:33 CET — RabbitMQ interruption detected on EU1. Multiple restart attempts were required, and engineers applied a manual connection gate, allowing only a fraction of clients to reconnect at a time. Full stability was restored by 13:37 CET.

Root Cause

All three disruptions share the same underlying cause. Following a brief interruption to the RabbitMQ cluster — in some cases lasting only seconds — client installations running outdated versions of our MRC client began reconnecting aggressively and continuously. This reconnection behavior generated a volume of traffic against the cluster comparable to a denial-of-service attack, preventing RabbitMQ from stabilizing on its own and requiring manual intervention in every case to restore service. In the most severe incident (September 1), the volume of reconnection attempts was high enough that engineers had to manually gate client access, allowing only a fraction of clients to reconnect at a time to give the cluster room to stabilize.

This behavior was first identified in early 2026, and a fix preventing this aggressive reconnection pattern was released in MRC shortly after. All three disruptions described here were caused by client installations that had not yet been upgraded to a version containing this fix.

Remediation

  • Each incident was resolved through manual restart and stabilization of the affected RabbitMQ nodes by our engineering team.
  • During the September 1 incident, engineers additionally applied a temporary manual connection gate to control the rate of client reconnections and allow the cluster to recover safely.

Prevention

  • A fix for the underlying reconnection behavior has been available in MRC since early 2026. We strongly recommend all customers confirm they are running the latest MRC version, as this is the only way to be fully protected from this issue going forward.

We apologize for the disruption these incidents caused and are committed to working with affected customers to complete their MRC upgrade as quickly as possible.

Avatar for
Began at:

Affected components