Summary
On the evening of 25th of October we had an re-occurring issue with our underlying service providers autoscaling feature causing issues with our ordering instances, resulting in a subset of orders failing between 17:00 - 17:11.
Root cause was related with an underlying service provider and its autoscaling feature. With multiple steps failing during the autoscaling procedure causing a chain reaction. The issue was resolved after de-activating the autoscaling and replacing it with a baseline set of hardware.
One of the planned improvements implemented after the Autoscaling incident on the 17th & 18th of October did not fully work as expected, resulting in the issue re-appearing when the autoscaling failed.
Planned improvements moving forward: we are going to implement additional separate warmup functionality. Until this in place, static and proactive scaling will be used.
Scope
Kiosk, app & online.
Severity
Guest's were unable to complete orders via kiosk, app & online.
Timeline
All timestamps and dates are CET.
- 17:00 - Ordering Instance went down.
- 17:11 - Development team applies mitigation of de-activating autoscaling feature and replacing it with a baseline set of hardware. And observe recovery of the service immediately.
Cause
- Underlying service providers autoscaling feature.
Planned improvements
- Adding additional separate system warmup procedures to prevent auto load balancing issues.
- Manual and proactive scaling until long term solution is in place.