Increased NVMe Storage Latency in Zone RMA1

Minor incident Region RMA (Rümlang, ZH, Switzerland) Linux Cloud Servers (RMA1)
2026-07-01 12:41 CEST · 1 day, 23 hours, 49 minutes

Updates

Post-mortem

Timeline
cloudscale is planning to replace and extend the hardware for the NVMe-SSD and Bulk volumes. As a preparatory step, we needed to perform some rebalancing in the NVMe-SSD cluster in our RMA1 zone, i.e. move existing customer data to different storage hosts and/or disks.

On 2026-07-01 at 12:41 CEST, we started the first rebalancing process. While we were aware that rebalancing would lead to higher I/O latency for our customers to a limited degree, prior tests in our lab environment had not shown a latency increase big enough to make us expect issues for our customers’ workloads. However, during the process, we got some customer feedback stating that they had received short-term alerts from monitoring systems and were noticing the higher latency, causing some research effort for the root cause on their end but no sustained issues with their workloads.

On 2026-07-02 at 13:45 CEST (and with communication to the customers who had contacted us the day before), we started the next step of the necessary rebalancing. While we still did not expect a problematic impact on customer workloads, we soon received customer feedback about even higher latency and actual impact on their workloads, such as slowed down applications or increased memory consumption.

From 14:30 CEST, cloudscale engineers joined an emergency call, closely watching the situation and evaluating the best path forward. From 15:00 CEST, our engineers applied a number of tuning settings in order to slow down the rebalancing process and free up capacity to serve I/O requests from customers more quickly. However, the tuning only had a limited effect. But as the rebalancing progressed, the impact on I/O latency gradually diminished.

Between 2026-07-03 05:55 and 12:30 CEST, we attempted the remaining rebalancing in small partial steps. However, latency measurements indicated a similar impact as before, although for a much shorter duration in each individual step.

At 12:30 CEST, considering that neither the tuning settings nor performing the rebalancing in small partial steps reduced its impact enough, we postponed the remaining rebalancing tasks. In our lab environment, we completely reverted the previously performed migration and started over, comparing in-depth metrics between lab and production at the very beginning of the initial migration. We wanted to fully understand why this process had much more impact on our customers than it used to have in the past, and how we could best mitigate the effect.

Our investigation revealed that Ceph, which our storage clusters are based on, introduced new defaults and tuning options (e.g. for processes like rebalancing and recovery after hardware failures) in recent versions while transitioning to “mClock” as the default scheduler. This explained why some of our documented, proven mitigations turned out ineffective in production this time. On the other hand, mClock provides new options for fine-grained QoS. Once we grew confident with a new set of configuration settings in our lab, we were able to resume the rebalancing process in our productive RMA1 NVMe-SSD cluster at 17:54 CEST, with no further impact on customer workloads.

Next Steps
We will make the new, optimized configuration settings for mClock the default across our Ceph-based storage clusters so in similar circumstances, no impact on customer workloads should occur in the first place.

We will also extend our monitoring and establish additional triggers to relevant metrics in order to better represent the storage performance as it is perceived by our customers’ workloads. For monitoring alerts based on the new triggers, we will prepare appropriate action plans to provide guidance to our on-call engineers and support them in quickly and effectively resolving similar incidents in the future.

We sincerely apologize for the inconvenience this issue may have caused you and your customers.

July 6, 2026 · 16:05 CEST
Retroactive

Between 2026-07-01 12:41 and 2026-07-03 12:30 CEST, a planned maintenance on the NVMe-SSD storage cluster in the RMA1 zone, where we had not expected any relevant impact on our customers, has led to multiple phases of increased storage latency with varying levels of impact to some customers’ workloads.

July 6, 2026 · 16:00 CEST

← Back