DigitalOcean está experimentando problemas con su panel de API y control y puede encontrarse con errores y mayor latencia mientras persiste el problema
resolved
Este incidente ha sido resuelto.
Traducido automáticamente desde la actualización oficial del incidente.
Servidores inalcanzables - Azure West US
Comenzó 23 de julio de 2026 a las 15:33 UTC · 6h 15m
OutageIncidente mayor
Componentes afectados
Microsoft AzureDedicated servers
investigating
Hemos recibido varias alertas notificándonos acerca de un par de servidores desplegados en Azure West US que no son accesibles. Aparentemente esto parece estar relacionado específicamente con esta región impactando el 25% de los servidores.
monitoring
Microsoft Azure confirmó que estaban teniendo problemas de red en la región de Estados Unidos Occidental. Por ahora, la mayoría de los servidores están operativos. Seguiremos monitoreando las siguientes horas.
resolved
Este incidente ha sido resuelto.
Traducido automáticamente desde la actualización oficial del incidente.
False positive server down messages sent
Comenzó 25 de junio de 2026 a las 14:10 UTC · 1d 0h
IssuesIncidente menor
Componentes afectados
BackendDedicated servers
investigating
We are currently investigating this issue.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem
**Summary**
Between June 25 and June 26, our monitoring system incorrectly reported some dedicated servers as unreachable and sent "server down" notifications for servers that were in fact operating normally. This was a monitoring-side issue only — no customer servers or services were actually down, and no data or message delivery was affected.
**Impact**
Affected customers received one or more alerts indicating their dedicated server was down. These alerts were false positives. The servers themselves remained healthy and fully available throughout the incident.
**Root cause**
Before connecting to a server, our monitoring system performs a network authorization step. A defect in how one component handled a transient network hiccup could leave a monitoring process unable to complete that step, after which it failed to connect to the servers it was responsible for checking. The servers themselves stayed healthy and available throughout — the monitoring process had simply lost its ability to reach them, and reported them as down.
**Resolution**
We identified the affected monitoring processes, confirmed the root cause in production, and deployed a fix that makes this authorization step resilient to such transient failures and able to recover automatically. After deploying, we monitored the system to confirm that notifications returned to normal.
We apologize for the confusion these false notifications may have caused.
Monitoring Systems — False Alarms
Comenzó 30 de mayo de 2026 a las 22:25 UTC · 13h 4m
Pending
monitoring
Between 17:30 and 20:00 UTC on 2026-05-30, our server monitoring system generated a large number of false "server_unreachable" alerts. Clusters were operational and healthy throughout this period — no customer data or message delivery was affected.
The alerts were caused by a rolling restart of our monitoring service, during which individual monitoring workers went offline mid-cycle and lost state. This caused the workers to report connection timeouts for servers they could no longer reach — even though those servers were fully healthy.
We apologize for any concern this may have caused.
resolved
This incident has been resolved.
Connectivity issues in Azure West Europe
Comenzó 23 de mayo de 2026 a las 16:21 UTC · 1h 1m
IssuesIncidente menor
Componentes afectados
Microsoft Azure
investigating
We are investigating connectivity issues affecting some CloudAMQP instances in Azure West Europe. Affected instances may see connection failures or unexpected restarts.
The root cause is an ongoing Microsoft Azure platform incident in the West Europe region (started 14:31 UTC, 23 May 2026). Microsoft is investigating.
resolved
Resolved — Microsoft has confirmed that the Azure West Europe Virtual Machines incident is resolved. The impact window was 14:09–14:13 UTC on 23 May 2026, during which a limited number of customers may have experienced connection failures or unexpected Virtual Machine restarts. The Azure environment self-healed and the service has been confirmed restored.
Affected CloudAMQP instances should now be operating normally.
System metrics not forwarded to legacy integration
Comenzó 22 de mayo de 2026 a las 3:16 UTC · 1h 34m
IssuesIncidente menor
Componentes afectados
Metrics
investigating
Since around 8pm UTC last night, some metrics are not being forwarded to legacy integration. Prometheus integrations not affected. We are investingating.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem
# Legacy host metrics integrations degraded
Incident window: 2026-05-21 18:37 UTC – 2026-05-22 04:05 UTC \(9h 28m\)
Affected service: Legacy metric integrations \(CloudWatch, Datadog, Librato, New Relic, Splunk, Stackdriver\). Host metrics affected, not broker metrics, nor was internal monitoring or console graphs.
## Summary
For ~9.5 hours, the service that collects host-level metrics \(CPU, memory, disk, network\) for legacy third-party integrations entered a crash loop and stopped shipping data.
## Root Cause
An internal credential-signing service was migrated to a new container runtime the afternoon before. Due to a bug with environment variables the collector could not renew it's credentials. The collector's existing credential remained valid for several hours, so the failure only surfaced when it tried to renew.
## Resolution
On-call staff detected the failure early morning, developers helped restore the service and collector recovered at 04:05 UTC.
## Prevention
* All services, both signing and collector report failed authentication attempts earlier and with higher severity
* Metrics pipeline alarm thresholds was tightened so a similar drop in data triggers alarms faster.
Backend slow
Comenzó 19 de mayo de 2026 a las 13:21 UTC · 28m
IssuesIncidente menor
Componentes afectados
Backend
investigating
We are currently experiencing issues with our backend services handling account and server creation.
This does not affect any running customer servers but might delay provisioning new servers.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Backend slow
Comenzó 5 de mayo de 2026 a las 10:23 UTC · 3h 0m
IssuesIncidente menor
Componentes afectados
Backend
investigating
We are currently experiencing issues with our backend services handling account and server creation.
This does not affect any running customer servers but might delay provisioning new servers.
monitoring
A fix has been implemented and we are monitoring the results.
monitoring
Still seeing slow requests, doing another investigation.
monitoring
Looks like we're back at full performance now.
resolved
This incident has been resolved.
Connecting issues in Azure West Europe
Comenzó 3 de mayo de 2026 a las 19:11 UTC · 6h 34m
IssuesIncidente menor
Componentes afectados
Microsoft Azure
investigating
We are currently investigating this issue.
identified
Servers in Azure West Europe can experience connection issues due to underlying issues at provider.
Latest update from Azure
"Current Status: We continue to investigate an availability issue caused by a subset of unhealthy backend storage infrastructure supporting Azure Virtual Machines and Azure Direct Drive in West Europe."
resolved
This incident has been resolved.
Backend slow/unresponsible
Comenzó 28 de abril de 2026 a las 10:27 UTC · 38m
OutageIncidente mayor
Componentes afectados
Backend
investigating
We are currently experiencing issues with our backend services handling account and server creation.
This does not affect any running customer servers but might delay provisioning new servers.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
We had an abnormal increase in traffic causing timeouts and slow responses for customer.cloudamqp.com, we increased capacity and are back to normal. This did not affect any running clusters, just provisioning and configuration updates.
Connection issues in Azure region westus
Comenzó 24 de abril de 2026 a las 9:36 UTC · 29m
OutageIncidente mayor
Componentes afectados
Microsoft Azure
investigating
We are getting indications of instances being down in Azure region westus. We are investigating the issue.
resolved
This incident has been resolved.
Issues with Azure West US 3 region
Comenzó 23 de abril de 2026 a las 10:10 UTC · 45m
IssuesIncidente menor
Componentes afectados
Microsoft Azure
investigating
Servers in Azure West US 3 can experience connection issues due to underlying issues at provider. We monitor and will update when we know more.
resolved
This incident has been resolved.
Metrics integration issues
Comenzó 23 de abril de 2026 a las 3:30 UTC · 2h 30m
Pending
Componentes afectados
Metrics
investigating
There appears to be an issue regarding sending metrics to all Legacy metrics integrations. Alarms based on broker metrics also got affected.
We are currently investigating the issue
Prometheus based metrics (Datadog v3 etc) not affected.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem
Broker metrics and metrics integrations degradation
Window: 2026-04-23 03:25 – 05:35 UTC \(2h 10m\)
## Impact
Broker metrics \(connections, channels, queues, consumers, message rates, node and netsplit state\) were degraded, progressively dropping until recovery at 05:35 UTC. Downstream effects:
* Queue, Consumer, Connection, Channel, Connection flow, and Netsplit alarms: threshold breaches during the window may not have triggered, or triggered late. Alarms already tripped before 03:25 UTC kept
their state.
* Metrics Integrations \(Datadog, CloudWatch, New Relic, Splunk, Dynatrace, etc.\): delivery of broker metrics was reduced for the duration of the window. Alarms configured on the receiving side may
likewise have missed or lagged.
Not affected: server metrics \(CPU, memory, disk\) and their corresponding alarms, Notice alarms, broker availability, and the metric graphs shown in the CloudAMQP console \(these use a separate data path
and continued to update normally\).
## Timeline \(UTC\)
* 03:25 — broker-metrics collection begins degrading
* 03:35 — ~30% of broker-metric samples missing
* 03:50 — ~50% plateau
* 05:25 — collector workers cycled and re-initialised
* 05:35 — full throughput restored
## Root cause
Broker metrics are polled from each cluster's management HTTP API by a pool of collector workers and published to an internal message bus that both the Alarms service and Metrics Integrations consume
from. At 03:25 UTC this service became unresponsive without raising an error. Affected workers silently stopped polling while remaining alive to the platform, so automatic restart did not trigger. On-call
staff began investigating around 05:00 UTC; force-restarting the service restored metric flow.
The exact trigger has not been identified. Our focus is on ensuring the condition is detected and handled promptly if it recurs.
## What we are changing
* Added metrics and alarms specific to this service, with high-urgency paging for on-call.
* Reliability improvements in the service itself to detect and restart stalled work.
## What you can do
CloudAMQP offers a new generation of metrics integrations based on Prometheus, with a re-engineered pipeline that does not rely on these centralised services — each server forwards data directly to your
endpoint. These have been running in production for some time and have proved very reliable.
* [https://www.cloudamqp.com/blog/prometheus-metrics-integrations.html](https://www.cloudamqp.com/blog/prometheus-metrics-integrations.html)
* [https://www.cloudamqp.com/docs/monitoring\_metrics\_datadog\_v3.html](https://www.cloudamqp.com/docs/monitoring_metrics_datadog_v3.html)
* Background: [https://www.cloudamqp.com/blog/decentralized-observability-with-open-telemetry-part-1.html](https://www.cloudamqp.com/blog/decentralized-observability-with-open-telemetry-part-1.html)
Monitoring Systems Short Outage and "Connection issues" Alarms
Comenzó 15 de abril de 2026 a las 14:00 UTC · 0m
Pending
resolved
One of our backend components suffered a small outage due to an unexpected full services restart. Connectivity with the servers were broken and service triggered alarms for many clusters.
We apologize for the inconvenience. All clusters were operational and healthy within the incident/alarms timespan.
Management Interface Unavailable - 502 Bad Gateway Response
Comenzó 2 de abril de 2026 a las 14:41 UTC · 31m
IssuesIncidente menor
Componentes afectados
Shared serversDedicated servers
identified
Customers are experiencing Management UI being unavailable with 502 Bad Gateway.
Timeline:
- 14:16 UTC -> The issue was identified
- 14:30 UTC -> Recent code changes were reverted. No relationship with incident noted.
identified
The Management UI and brokers are serving traffic normally. You can still use the Management UI by using the Broker URL and properly typing User and Password.
The impact is related to our backend service logs into Management UI via CloudAMQP SSO.
identified
We are continuing to work on a fix for this issue.
resolved
The Management UI SSO shortcut is back to service.
Impacted Clusters:
All customers with clusters created more than 6 months ago
Cause:
Some code changes unexpectedly changed the communication schema between our backend services and the SSO component on the clusters. The proxy configuration conflicted with those changes causing the 502 Bad Gateway error to show up.
Impact:
Only the SSO via CloudAMQP Console was impacted given it required communication with our backend components. The AMQP Clusters were healthy and servicing traffic via both AMQP and HTTP on Management UI and API.
Logging into the Management UI via Credentials and all Authentications Backends was functioning normally.
Metrics delivery delay for 12% of the fleet
Comenzó 25 de marzo de 2026 a las 17:45 UTC · 17m
Pending
Componentes afectados
Metrics
identified
We're currently experiencing issues with metrics delivery, this affects 12% of the API based v1/v2 integrations. We've added additional capacity and expect have handled the backlog within 20 minutes.
resolved
This incident has been resolved.
Queue metrics for legacy integrations delayed
Comenzó 18 de marzo de 2026 a las 21:25 UTC · 6h 24m
IssuesIncidente menor
Componentes afectados
Metrics
identified
We noticed an issue with our queue metrics being sent to metrics integrations. there is currently a delay in processing, but we have identified an issue and expect the delay to be fixed shortly.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Coralogix log delivery performance issues
Comenzó 3 de marzo de 2026 a las 21:11 UTC · 17h 3m
Pending
investigating
We're currently experiencing issues with Coralogix log delivery.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Metrics delivery outage for v3 Prometheus based metrics.
Comenzó 3 de marzo de 2026 a las 18:28 UTC · 2h 23m
Pending
Componentes afectados
Metrics
investigating
We're currently experiencing issues with metrics delivery, this affects v3 (Prometheus based) metrics.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem
### Summary
On March 3, an internal monitoring metric stopped being reported for approximately 14 hours. No customer metrics delivery was affected, all integrations continued operating normally. The missing metric caused our dashboards to display false failure rates, which was
incorrectly reported as a metrics delivery outage.
### Impact
None. Metrics continued to be delivered to all customer endpoints throughout the incident.
### Resolution
The monitoring metric was restored and we have updated our alerting queries to be resilient to missing data, preventing false positives in the future.
Historial de interrupciones de Cloudamqp | Uptimus