EU2 RabbitMq Server Outage

Uptime Impact: 1 hour and 17 minutes
Resolved
Resolved

Postmortem Incident Report

Incident Summary

• Date: July 22, 2026
• Incident window: 10:49 – 21:15 CET
• Total service downtime: ~1 hour 17 minutes (across five separate outage periods)
• Impact:
  • Repeated AMQP service disruptions on the EU2 environment following a RabbitMQ node failure.
  • The cluster failed to stabilize after several restart attempts, resulting in multiple short outages throughout the day.
  • Full stability was restored following a planned complete cluster restart in the evening.

We sincerely apologize for the disruption this caused. Reliable messaging is a critical component of our platform, and we deeply regret the inconvenience experienced by our users and teams.


Resolution

Following the initial node failure, our team performed several targeted restart attempts to restore cluster stability. As instability persisted, we scheduled a complete cluster restart outside business hours, which fully resolved the issue.

Timeline:

• 10:49 CET – RabbitMQ node failure detected; first restart completed within ~3 minutes but did not fully stabilize the cluster.
• 12:35 – 12:47 CET – Second outage.
• 12:57 – 13:14 CET – Third outage.
• 13:37 – 13:57 CET – Fourth outage; service remained unstable afterward.
• 20:50 – 21:15 CET – Planned complete cluster restart performed; system has been stable since.

Avatar for
Began at:

Affected components