Summary
On the evening of 17th and 18th of October we experienced issues with an underlying service providers autoscaling feature causing issues with our ordering instances, resulting in a subset of orders failing between the following time-slots:
2025-10-17:
17:04 - 17:11
2025-10-18:
17:02 - 17:11
17:35 - 17:42
17:46 - 17:49
18:05 - 18:14
The issue was resolved automatically each time by the underlying service providers auto heal feature.
Root cause was related with an underlying service provider and its autoscaling feature. With multiple steps failing during the autoscaling procedure causing a chain reaction, adding complexity to finding and applying the correct mitigation. The issue was fully resolved after de-activating the autoscaling and replacing it with a baseline set of hardware.
Scope
Kiosk, app & online.
Severity
Guest's were unable to complete orders via kiosk, app & online.
Timeline
All timestamps and dates are CET.
- 2025-10-17 - 17:04 Ordering Instance went down.
- 17:13 - Ordering instance recovered.
- 17:19 - FO Support notified of ongoing incident.
- 17:22 - Issue is escalated internally.
- 17:25 - Development team confirms previous partial outage.
- 17:35 - Customers informed via tickets of previous partial outage.
- 2025-10-18 -
- 17:02 - Ordering instance went down.
- 17:11 - Ordering instance recovered.
- 17:17 - Internal escalation to development teams and investigation instigated after customer incident report.
- 17:26 - Development team confirms previous partial outage. Additional investigation continues.
- 17:34 - Customers informed.
- 17:35 - Ordering instance went down.
- 17:42 - Ordering instance recovered.
- 17:43 - Additional Internal escalation to development teams and investigation instigated after new customer incident reports.
- 17:46 - Ordering instance went down.
- 17:49 - Ordering Instance recovered.
- 17:56 - Development team confirms previous partial outage. Additional investigation continues.
- 18:05 - Ordering instance went down.
- 18:13 - Development team applied long term mitigation of de-activating autoscaling feature and replacing it with a baseline set of hardware.
- 18:14 - Ordering instance recovered.
- 18:45 - Customers informed of initial root cause analysis and long term mitigation.
Cause
- Underlying service providers autoscaling feature.
Planned improvements
- Added separate system warmup procedures to prevent auto load balancing issues. (Deployed and confirmed working.)
- Improvement to incident management process to improve communication.