Started September 8, 2026 at 11:43 PM UTC · Ongoing
IssuesMinor incident
Affected components
ProvisioningDatabase as a Service (DBaaS)NetworkComputeStorage
monitoring
We are currently investigating network connectivity alerts related to our TXL datacenter. While connectivity appears to be restored already, we are actively monitoring the situation to ensure stability. We will provide updates as new information becomes available.
identified
We see indications of a recurrence of the issue and are setting the status page to active. Our Network Team is investigating and we will share updates here as they become available.
The issue is currently negatively affecting Compute, Provisioning and Network Components.
identified
We are continuing to work on a fix for this issue.
identified
Our Network Team has identified an issue related to a recent router replacement and applied a mitigation. Compute and Storage systems are showing signs of recovery. We are actively monitoring the environment until all services are fully restored
identified
We are currently investigating potential residual impact on DBaaS Services.
Compute and Block Storage have reported full recovery.
monitoring
DBaaS service has recovered, as well. We are setting the incident into Monitoring status, while the teams are ensuring that all affected services operate normally.
Cloud Support: Limited Phone Support Availability
Started September 8, 2026 at 2:15 PM UTC · 7h 42m
IssuesMinor incident
Affected components
Cloud Support
identified
While phone support will generally be available, our Support capacity is reduced due to multiple cases of sick leave.
We want to inform you that our phone support will be limited during the following time slots, leading to increased response times on the phone channel.
In these cases, please reach out per email. Thank you.
resolved
We have people available now, Support is reachable as usual.
Cloud Support: Limited Phone Support Availability
Started September 4, 2026 at 4:53 PM UTC · 3d 13h
IssuesMinor incident
Affected components
Cloud Support
investigating
While phone support will generally be available, our capacity is reduced due to multiple cases of sick leave.
We want to inform you that our phone support will be limited during the following time slots, leading to increased response times on the phone channel.
In these cases, please reach out per email. Thank you.
identified
We still run short on coverage of service, and we are currently doing changes on our ticketing system.
So we kindly ask for your understanding if Support responses are not as timely as you are accustomed to.
Thank you.
resolved
We managed to fully re-establish our service, Support is back.
thank you for your patience.
Partner Subcontracts can not access new location de/fra/1
Started September 1, 2026 at 1:36 PM UTC · 2h 37m
IssuesMinor incident
Affected components
Data Center Designer (DCD)Cloud API
identified
For our partners with subcontracts:
We found that subcontracts cannot access the new Frankfurt location "de/fra/1".
We are currently working on a fix to make this available,
we will keep you informed when it's done.
monitoring
Wir konnten das Problem beheben, sodass die neue Region nun für alle Verträge verfügbar ist.
Wir monitoren, ob es dazu noch weitere Rückmeldungen gibt.
resolved
This incident has been resolved.
DCD is Currently Unavailable
Started August 31, 2026 at 3:35 PM UTC · 1h 28m
OutageMajor incident
Affected components
Data Center Designer (DCD)
investigating
The DCD is currently unavailable, displaying an endless "Loading DCD..." page.
We are currently investigating this issue, and will keep you updated.
monitoring
The DCD is now available again, and we are monitoring the situation.
resolved
We are marking this incident as resolved.
Our network team is investigating negative impact on the IAM Services caused by a change rollout. We will update the Status Page once the root cause has been established.
postmortem
# **Preliminary Root Cause Analysis**
This Root Cause Analysis is preliminary as research is still being conducted to determine the technical root cause of the incident.
## **What happened?**
On August 31, 2026, between 15:00 UTC and 15:43 ITC, and again between 16:32 UTC and 16:35 UTC, customers were unable to reach IONOS Cloud's Identity and Access Management \(IAM\) service and the Data Center Designer \(DCD\), which is dependent on this service. The disruption affected direct logins, partner and reseller portal access, and associated management consoles. The total customer-facing impact lasted approximately 48 minutes across both intervals.
Cloud APIs \([api.ionos.com](http://api.ionos.com)\) remained fully operational throughout the incident. The underlying application services were healthy at all times - the failure was confined to the network edge layer.
## **How was this possible? \(Root Cause\)**
During a scheduled maintenance window on August 31, 2026, a network configuration update was applied to the edge network infrastructure with the intent of optimizing routing filters. The configuration was verified as correct prior to and during application. Following the rollout, a routing propagation anomaly emerged on one edge network switch, causing asymmetric routing behavior: incoming TCP connection requests from clients were silently dropped \(black-holed\) at the network edge before reaching the application cluster.
Because the application services themselves remained up and healthy, this failure mode was not immediately visible through internal health checks - the services were isolated from receiving inbound public internet traffic rather than failing.
The root cause of why this specific switch exhibited asymmetric routing behavior following an otherwise valid configuration change remains under active investigation. Whether this was triggered by a switch platform behavior or a software version-specific bug is being determined through staging environment reproduction.
## **What are we doing to prevent recurrence?**
### **Immediate Actions**
* **Configuration Rollback:** Upon identifying the routing anomaly, a full rollback of the network configuration was executed across all affected edge switches. Routing announcements and TCP reachability of the public production IP addresses were verified following the rollback. All affected services - DCD, Partner Portal, Reseller Portal, and IAM - were confirmed fully operational by 16:35 UTC. \(DONE\)
### **Short-term**
* **Staging Environment Replication:** Detailed test cases are being executed in the staging environment to reproduce the exact routing propagation behavior under the same switch configuration conditions. The goal is to determine whether the anomaly is attributable to a specific software version bug or a switch platform behavior, so that the precise failure condition can be isolated and addressed before any future rollout. ETA: Within two weeks
### **Mid-term**
* **Revised Rollout Strategy:** Based on the findings from staging, a revised rollout approach will be designed to ensure that any future application of these routing filter optimizations can be performed with greater stability guarantees - including more granular validation checkpoints between switch-level changes. ETA: October 2026
## **Closing remarks**
We recognize that loss of access to IAM and DCD carries real operational impact. The fact that the underlying services were healthy throughout is not a mitigation of that impact.
We are committed to ensuring that the root cause is fully understood before any re-attempt of the original change, and that the revised rollout strategy addresses the conditions that led to the asymmetric routing behavior.
We thank you for your patience while we conclude the investigation.
Object Storage - Increased latency in eu-central-1
Started August 26, 2026 at 9:53 AM UTC · 8d 3h
IssuesMinor incident
Affected components
Object Storage
investigating
We are currently investigating increased latency affecting S3 Object Storage in the eu-central-1 region. Some customers may experience slow response times for read and write operations. Our engineering team is actively working on resolving the issue. We will provide updates as more information becomes available.
identified
The issue has been identified and a fix is being implemented.
identified
We see recurring latency spikes affecting our S3 service in eu-central-1. The Object Storage and Network teams are jointly investigating. Although a technical root cause remains undetermined at this time, our highest priority is implementing measures to mitigate the frequency and amplitude of the spikes. We appreciate your patience and will keep you informed.
identified
Our engineering teams have identified a path to remediation. Initial measures have been applied, improving the situation for some services. Customers using Object Storage in eu-central-1 may continue to experience elevated latency, which may vary in severity. Work continues on a comprehensive fix. We will provide further updates
monitoring
Response times for Object Storage in eu-central-1 have improved significantly and continue to stabilize. We are closely monitoring system performance. We will provide further updates as remediation progresses.
resolved
The increased latency affecting S3 Object Storage in eu-central-1 has been resolved. Response times have returned to normal levels. We will continue to monitor the service.
postmortem
# Root Cause Analysis
## What happened?
Starting approximately 19:00 UTC on 24 August 2026, customers accessing S3 Object Storage in the Frankfurt \(FRA4\) data center experienced elevated latency across all operations - uploads, downloads, metadata requests, and deletions. Intermittent HTTP 503 Service Unavailable and 404 Not Found errors were observed on object read requests. The impact was measurable for all customer which had buckets in the both affected datacenters of the region, with some customers experiencing severe degradation depending on their bucket configuration and access patterns.
The incident remained in progress with high priority from 25 August through 3 September 2026. The latency issue was mitigated at approximately 13:40 UTC on 3 September 2026.
## How was this possible? \(Root Cause\)
**Primary cause - Software bug in Quality of Service \(QoS\) subsystem**
IONOS S3 Object Storage in the Frankfurt region is using a distributed object storage system. A QoS feature of this service - implemented via a redis-qos service - applies rate limiting to S3 requests at the cluster level. A bug in the S3 service leads to not well distributed queries to the Redis-QOS service \(which is served by multiple server for high availability\) this caused the Redis-QOS process to reach and sustain 100% CPU utilization, progressively slowing down all S3 request processing at the cluster level. This affected every request passing through the affected nodes, irrespective of the operation type or the specific bucket accessed.
This is an internal defect within the software used. The bug caused the S3 service to consume all available resources before load reached request levels that would normally trigger throttling, meaning the degradation occurred continuously rather than only under peak conditions.
In close cooperation with the software vendor, we disabled the QoS rate-limiting function on 3 September, which fully resolved the latency issue. S3 is currently operating without some QoS features while a permanent fix is prepared by the vendor.
## What are we doing to prevent recurrence?
**Already completed:**
* QoS disabled in the S3 service - fully resolved the latency issue. \(DONE\)
* Requesting permanent solution from the software vendor. \(INPROGRESS\)
**Short-term - ETA: within 2 weeks:**
* Permanent fix for the Cloudian QoS bug: IONOS Cloud is in active coordination with the vendor to obtain and deploy a fix for the redis-qos defect. Once the fix is validated, lost QoS features will be re-enabled.
* Database partition monitoring: We are implementing monitoring that alerts on partition size growth before any individual partition approaches a problematic threshold. This will allow our team to identify and address bucket layout issues proactively.
**Mid-term - ETA: 1 to 3 months:**
* QoS architecture review: Following the permanent QoS fix, we will review the architectural isolation of the QoS service together with the vendor to ensure that a future resource contention event in the rate-limiting layer cannot propagate to the request path at the same scale.
* Monitoring and alerting improvements: We are extending cluster-level monitoring to surface redis-qos CPU saturation and database compaction backlog as first-class incident signals, with automated escalation before customer-visible latency develops.
## Closing remarks
An incident of this duration in a core infrastructure service is not acceptable. The high-latency period persisted for nine days, during which customer workloads depending on S3 in the Frankfurt region were degraded. Multiple optimisations and mitigation strategies were implemented during the course of the incident, but could only improve the situation for individual buckets and only to a certain extent. Detecting the underlying QoS bug and developing a mitigation required coordination with the vendor’s engineering team.
While the latency issue is mitigated, we remain in close contact with the vendor. The engineering work to deliver a permanent QoS fix, reduce database partition pressure, and prevent recurrence is in progress. We are also working closely with our technology partner to understand delays in the analysis of the root cause of this incident. We will conduct a joint post mortem to identify areas where collaboration during incidents can be improved.
We recognise the impact this incident caused to your operations. We believe that the listed measures will help us prevent similar error patterns and speed up analysis and recovery for software related issues in the future.
We thank you for your patience during the incident.
Object Storage Service Restrictions
Started August 23, 2026 at 12:47 PM UTC · 5h 0m
IssuesMinor incident
Affected components
Data Center Designer (DCD)Object StorageObject StorageObject StorageObject Storage
investigating
We are currently investigating an issue where Buckets and Object Storage Keys are not displayed in the Data Center Designer.
It is currently not possible to access, alter, create and delete Buckets and Keys via the Data Center Designer.
Currently it's not possible to reserve or handle IP blocks, neither in DCD, nor via API.
Our teams are investigating the issue.
We will keep you updated.
identified
The issue has been identified and a fix is being implemented.
resolved
This incident has been resolved.
Cloud Support: Telephone Support reduced capacity
Started August 14, 2026 at 2:50 PM UTC · 5d 17h
IssuesMinor incident
Affected components
Cloud Support
identified
There may be periods of increased waiting times when calling IONOS Cloud Support by phone. We ask Customers and Partners to contact Cloud Support via the DCD Form or Email, instead.
identified
The situation is still unchanged, so best would be to please reach out via email.
resolved
We managed to get back to normal capacity, so Support is back to normal availability.
Thank you for the patience.
we are currently investigating increased error rate for our provisioning service.
Limited access to provisioning services
Started August 11, 2026 at 4:05 PM UTC · 19h 47m
IssuesMinor incident
Affected components
ProvisioningProvisioningData Center Designer (DCD)ProvisioningProvisioningProvisioningProvisioningCloud APIProvisioningProvisioningProvisioning
investigating
Currently, there is an increased processing time for provisioning orders that are initiated via Data Center Designer or API.
Occasionally, connections may be lost in the direction of the Data Center Designer.
Availability and accessibility of your virtual data center resources will remain unaffected.
We will inform you as soon as the functionality has been restored.
identified
We have identified a likely culprit. The Provisioning Team has implemented a mitigation. We see the performance of the service stabilizing.
identified
We are still seeing residual 500 errors from the Cloud API and are working towards resolving the remaining service degradation.
monitoring
We confirm that the provisioning service has recovered and is operating normally. Our teams will continue to monitor the environment to ensure its stability and continued operation.
resolved
This incident has been resolved.
Managed Kubernetes - Intermittent Control Plane Unavailability
We are aware of intermittent control plane unavailability affecting a subset of Managed Kubernetes customers. Affected customers may experience API call failures, deployment timeouts, and temporary disruption of cluster management operations.
Our engineering team is actively working on both immediate mitigations and longer-term architectural improvements.
Several mitigations have already been deployed, including maintenance schedule optimization, compaction regression fixes, and storage performance improvements. Additional measures - including infrastructure migration, dedicated event etcd clusters, improved load balancing, and horizontal scaling - are in progress.
We are providing regular updates on this page. Customers experiencing issues are encouraged to subscribe to this incident for timely notifications.
identified
During a service rollout today, a subset of control planes experienced temporary restarts. Affected customers may notice brief API unavailability while these control planes recover. The team is monitoring the recovery.
Separately, work on improving infrastructure capacity and load distribution continues as described in our initial update.
We will post another update once the affected control planes have fully stabilized.
identified
We are currently rolling out memory scaling measures to address recurring stability issues during compaction operations on the affected etcd clusters. Additionally, we are planning to roll out further horizontal scaling for the affected clusters today.
We will provide another update once these measures have been applied and we can assess their impact.
identified
Our plans for further horizontal scaling of the control plane are progressing. We expect to be able to do a dry run and further testing within the next hours, before we proceed with migrations.
We aim to finish work on horizontal scaling until EOD.
We are rolling out memory configuration improvements in parallel.
monitoring
Memory adjustments have been rolled out and show positive effects.
We are starting the migrations planned to further improve control plane performance for all our customers.
We estimate that the migration will be completed in the next hours.
Control Plane performance is expected to improve already during the migration.
We are setting this incident into Monitoring status and will provide an update once the migration is completed.
identified
Migration of the first batches has been completed. The Kubernetes Team has identified a remaining issue preventing further migration. We are setting this incident back to active until the issue is resolved and the migration completed.
identified
Migration has been picked up again. We already see encouraging results after completion of first batches. In the next hours we focus on completing the migration and horizontal scaling.
During the migration, single etcds can be temporarily unavailable for time periods lasting around 30 seconds.
We expect further performance and stability improvements for all customers during and after the migration.
identified
Memory limit adjustments and migration have had positive effects on the first control plane cluster. Customer situated in the first control plane cluster should already see substantial improvements in performance and stability.
We have started to roll out memory adjustments in the remaining control plane cluster, as well. Rolling out the memory limits can lead to temporary unavailability of affected control planes. These interruptions should be brief and will not affect running workloads.
After memory limit adjustments are fully rolled out on the second control plane cluster, horizontal scaling and migrations will resume on both control plane clusters for the next hours until workload is distributed optimally.
monitoring
Memory limit adjustments have been fully rolled out across all control plane clusters and are showing positive effects. We have significantly expanded the underlying infrastructure capacity and migration of customer workloads is progressing well.
We are observing substantial improvements in control plane stability and performance. The recurring disruption patterns described in earlier updates are no longer present.
Migration work will continue throughout the day. We will provide an update once migrations are completed or if any changes in status occur.
monitoring
Migration of customer workloads on the first control plane cluster has been completed ahead of schedule. Customer clusters have been redistributed across expanded infrastructure and all migrated workloads are running without issues. We continue to observe stable control plane performance with no new disruptions reported.
resolved
We are marking this incident as resolved.
The measures already put into place have had the desired effect on the service and have improved performance and stability.
While this incident is marked as resolved, we are continuing executing on our action plan to improve the performance and reliability of our Managed Kubernetes Control Planes. Our current focus:
- Further load balancing on our Control Plane Clusters
- Migration of control planes to improved infrastructure
Cloud Support: Telephone Line Availability Degraded
Started August 4, 2026 at 8:20 AM UTC · 1h 42m
IssuesMinor incident
Affected components
Cloud Support
identified
IONOS Cloud Support is temporarily not always available via phone.
We ask Customers and Partners to contact Cloud Support via the DCD Form or Email, instead.
resolved
We were able to assign additional personnel and will set this status page to resolved.
Connectivity Issues with DBaaS (MongoDB)
Started July 28, 2026 at 1:46 PM UTC · 2h 50m
OutageMajor incident
Affected components
Database as a Service (DBaaS)
investigating
We are currently investigating a connectivity issue impacting our DBaaS (MongoDB) services. Our team is working to resolve the problem, and we will update you as soon as full functionality has been restored.
investigating
We are continuing to investigate this issue.
identified
We have identified an issue with DNS. Our DBaaS Team is currently analyzing the service. A potential culprit has been identified.
We will provide another update here latest 14:15
monitoring
The DBaaS Team has performed a rollback of a change deployed prior that updated DNS services. We are seeing services recovering. We are monitoring the situation closely.
resolved
We are marking this incident as resolved as no further abnormalities could be detected. We will share an RCA as soon as it is compiled.
LAS: Storage Loss of Redundancy
Started July 25, 2026 at 6:27 AM UTC · 1h 22m
IssuesMinor incident
Affected components
StorageNetwork
investigating
We are investigating alerts related to loss of redundancy on storage servers. There is currently no customer impact, our storage team is investigating and working to restore redundancy. We are suspecting a faulty networking component.
monitoring
We have found a likely root cause. A management component was causing excessive memory consumption introducing instability on the affected storage servers. This has been mitigated. We are currently monitoring the environment and will close the incident if no further anomalies are observed.
resolved
No more anomalies were detected.
The Root Cause of the loss of redundancy was identified as a memory leak in a management component. A durable mitigation was put in place to avoid recurrence. The memory leak will be addressed in an upcoming update to the management component.
Cloud Support: Limited Phone Support Availability
Started July 24, 2026 at 6:51 AM UTC · 2d 22h
IssuesMinor incident
Affected components
Cloud Support
identified
While phone support will generally be available, our capacity is reduced due to multiple cases of sick leave.
We want to inform you that our phone support will be limited during the following time slots, leading to increased response times on the phone channel.
In these cases, please reach out per email. Thank you.
resolved
Cloud Support availability is back.
Thank you for your understanding.
AI Model Hub - Service Degradations
Started July 23, 2026 at 10:44 AM UTC · 4d 20h
OutageMajor incident
Affected components
AI Model Hub
investigating
We are currently investigating increased AI Model Hub error rates and latencies. More details will be shared as they become available.
Affected services: AI Model Hub
Location: Global Services
monitoring
A fix has been implemented, which has reduced the number of 4xx and 5xx responses to nominal levels. We will continue to monitoring the results.
resolved
We are marking this incident as resolved. Our AI Modelhub Team has addressed the issue, which was caused by acute resource constraints. This was resolved by scaling out the bottlenecks.
MK8s - Connectivity Issue
Started July 18, 2026 at 1:33 PM UTC · 16d 20h
IssuesMinor incident
Affected components
Managed Kubernetes
investigating
We are currently investigating a suspected network connectivity issue influencing the MK8s service. We will keep the status page updated with new information as they become available.
investigating
Initial investigation makes a network connectivity issue unlikely. The team is focusing investigation on Managed Kubernetes side.
identified
We have identified a spike in our provisioning engine queue that is a likely root cause for the issues observed. Our provisioning team is informed and has joined the response.
identified
We have narrowed down the issue and believe we have the culprit identified. We are currently confirming the finding.
identified
Our storage team has confirmed the suspected culprit. We are currently mitigating the issue and will monitor provisioning job execution afterwards.
identified
Storage issue has successfully been resolved, however, job provisioning is currently not progressing successfully. Our provisioning team is investigating.
identified
While provisioning block could be resolved, we are investigating an increased number of storage related errors. We are directing attention to these remaining issues.
Customers could still see issues with attaching storages on Kubernetes.
identified
We have found an issue with one storage server belonging to a redundant storage server pair. Our team is implementing a mitigation.
monitoring
The incident should now be mitigated. The affected storage server is currently getting restored. Once this is completed, redundancy will be fully restored. We do not see any remaining provisioning issues any longer. We are monitoring the situation and the restoration and will then set the incident to resolved.
monitoring
Due to the ongoing restoration effort on the storage server customers might still see residual impact when attaching/detaching storage until redundancy is fully restored. We are resetting the impact of the service back to Degraded Performance.
monitoring
The team has encountered a complication during the recovery of the second storage server in the pair. Hardware needs to be replaced. Our datacenter team is working on this task. For customers still affected we are working on implementing a mitigation to unblock storage operations in parallel.
monitoring
Hardware replacement is being conducted. The team continues to mitigate acute storage provisioning issues until hardware replacement and recovery is completed to minimize customer impact.
monitoring
Hardware replacement was completed. Restoration of redundancy is continuing.
monitoring
The first hardware replacement was unsuccessful. Another one is attempted. In the meanwhile a data migration is running to migrate data to another storage target. Due to the amount of data to be migrated the migration is estimated to take several hours.
monitoring
The migration of storages has been completed, which should prevent further issues with attaching storage. Migration of snapshot data is currently ongoing, so snapshot related activities might still fail to complete successfully.
resolved
we are marking this incident as resolved as the underlying storage issue has been resolved. We will create a follow up incident for Managed Kubernetes performance and stability issues to avoid confusion.
Ionos Cloud outage history and incident timeline | Uptimus