Incident of Object Storage and Bulk Volumes in Region RMA

Major incident Region RMA (Rümlang, ZH, Switzerland) Linux Cloud Servers (RMA1) Object Storage Service (RMA)
2026-07-21 10:42 CEST · 27 minutes

Updates

Post-mortem

Management Summary
Between 2026-07-21 10:40 and 11:09 CEST, access to Bulk volumes and object storage in the RMA region was blocked. This was caused by a follow-up task to a prior maintenance. The follow-up task was not supposed to make actual changes to the affected running storage hosts.

Timeline
During the operating system upgrade on our storage hosts in the RMA (Rümlang) region, a typo had been found in network-related config files. While this typo had been there for years with no effect under the old operating system, it became relevant under the new operating system. For a number of storage hosts mainly serving our object storage and Bulk volumes, the running network configuration was fixed back then while the typo was left untouched in the on-disk config file for the time being.

At 2026-07-21 10:39 CEST, as a follow-up task to the operating system upgrade, we intended to persist the correct network config to the on-disk config files on the relevant storage hosts. As the running network configuration was already correct, no actual change was supposed to be made there.

At 10:40 CEST, the roll-out playbook failed, leaving the network connections of the affected storage hosts down. This not only blocked access to Bulk volumes and object storage in the RMA region for customers, but also management access to the hosts for our engineers.

At 10:43 CEST, after a quick assessment of the situation at hand, more engineers were involved. Following our prepared emergency procedures, we first re-established management access to the affected storage hosts via remote KVM, and re-enabled networking for actual storage contents shortly thereafter.

At 11:09 CEST, we confirmed that access to Bulk volumes and object storage in the RMA region for customers had been fully restored.

For some time after restoring network connectivity of the affected storage hosts, recovery and backfill processes have been running as designed in a Ceph cluster, ensuring correct distribution of data between storage hosts and rebuilding triple replication of data where necessary. These background processes may have had a minor impact on the storage performance until completion.

Next Steps
We will review the failed roll-out playbook, and evaluate additional safeguards for both the tools in use as well as our operating procedures. Independent of this incident, an upgrade to our lab environment has already been initiated, eliminating some of the discrepancies between lab and production environments, and thus enabling better testing of changes like this one.

We sincerely apologize for the inconvenience this issue may have caused you and your customers.

July 21, 2026 · 17:00 CEST
Resolved

The issue has been fully resolved and the storage cluster is back to normal.

We will follow up with a detailed incident report later.

Please accept our apologies for the inconvenience this issue may have caused you and your customers.

July 21, 2026 · 14:32 CEST
Monitoring

Access to Bulk volumes and object storage in region RMA has been restored. However, storage performance may be lower than usual for some time due to recovery processes running in the background.

We keep watching the state and will update this incident ticket if necessary.

Please accept our apologies for the inconvenience this issue may have caused you and your customers.

July 21, 2026 · 11:09 CEST
Issue

Our engineers are currently investigating an incident with our storage clusters in region RMA.

Bulk volumes and the object storage may be inaccessible for the time being.

We will keep you posted.

July 21, 2026 · 10:45 CEST

← Back