共享服务器“ Seal” 上的问题
- investigating
我们目前正在调查这一问题.
- identified
这一问题已经确定,一个解决办法正在实施之中.
- monitoring
一项措施已经执行,我们正在监测结果.
- resolved
这一事件已经得到解决.
自动翻译自官方事件更新。
54 Cloudamqp incidents · 2024年12月 — official updates, affected components, duration and resolution details.
我们目前正在调查这一问题.
这一问题已经确定,一个解决办法正在实施之中.
一项措施已经执行,我们正在监测结果.
这一事件已经得到解决.
自动翻译自官方事件更新。
DigitalOcean的API和控制面板目前出现问题,在问题持续存在时,您可能会遇到出错和更长时间的问题
这一事件已经得到解决.
自动翻译自官方事件更新。
我们收到各种警报 通知我们部署在Azure West US上的几台服务器 无法到达 显然,这似乎与这个区域有关,影响到25%的服务器.
微软Azure证实他们在美国西部地区有网络问题。 目前,大多数服务器已投入使用。 我们将持续监控 接下来的几个小时.
这一事件已经得到解决.
自动翻译自官方事件更新。
We are currently investigating this issue.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
**Summary** Between June 25 and June 26, our monitoring system incorrectly reported some dedicated servers as unreachable and sent "server down" notifications for servers that were in fact operating normally. This was a monitoring-side issue only — no customer servers or services were actually down, and no data or message delivery was affected. **Impact** Affected customers received one or more alerts indicating their dedicated server was down. These alerts were false positives. The servers themselves remained healthy and fully available throughout the incident. **Root cause** Before connecting to a server, our monitoring system performs a network authorization step. A defect in how one component handled a transient network hiccup could leave a monitoring process unable to complete that step, after which it failed to connect to the servers it was responsible for checking. The servers themselves stayed healthy and available throughout — the monitoring process had simply lost its ability to reach them, and reported them as down. **Resolution** We identified the affected monitoring processes, confirmed the root cause in production, and deployed a fix that makes this authorization step resilient to such transient failures and able to recover automatically. After deploying, we monitored the system to confirm that notifications returned to normal. We apologize for the confusion these false notifications may have caused.
Between 17:30 and 20:00 UTC on 2026-05-30, our server monitoring system generated a large number of false "server_unreachable" alerts. Clusters were operational and healthy throughout this period — no customer data or message delivery was affected. The alerts were caused by a rolling restart of our monitoring service, during which individual monitoring workers went offline mid-cycle and lost state. This caused the workers to report connection timeouts for servers they could no longer reach — even though those servers were fully healthy. We apologize for any concern this may have caused.
This incident has been resolved.
We are investigating connectivity issues affecting some CloudAMQP instances in Azure West Europe. Affected instances may see connection failures or unexpected restarts. The root cause is an ongoing Microsoft Azure platform incident in the West Europe region (started 14:31 UTC, 23 May 2026). Microsoft is investigating.
Resolved — Microsoft has confirmed that the Azure West Europe Virtual Machines incident is resolved. The impact window was 14:09–14:13 UTC on 23 May 2026, during which a limited number of customers may have experienced connection failures or unexpected Virtual Machine restarts. The Azure environment self-healed and the service has been confirmed restored. Affected CloudAMQP instances should now be operating normally.
Since around 8pm UTC last night, some metrics are not being forwarded to legacy integration. Prometheus integrations not affected. We are investingating.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
# Legacy host metrics integrations degraded Incident window: 2026-05-21 18:37 UTC – 2026-05-22 04:05 UTC \(9h 28m\) Affected service: Legacy metric integrations \(CloudWatch, Datadog, Librato, New Relic, Splunk, Stackdriver\). Host metrics affected, not broker metrics, nor was internal monitoring or console graphs. ## Summary For ~9.5 hours, the service that collects host-level metrics \(CPU, memory, disk, network\) for legacy third-party integrations entered a crash loop and stopped shipping data. ## Root Cause An internal credential-signing service was migrated to a new container runtime the afternoon before. Due to a bug with environment variables the collector could not renew it's credentials. The collector's existing credential remained valid for several hours, so the failure only surfaced when it tried to renew. ## Resolution On-call staff detected the failure early morning, developers helped restore the service and collector recovered at 04:05 UTC. ## Prevention * All services, both signing and collector report failed authentication attempts earlier and with higher severity * Metrics pipeline alarm thresholds was tightened so a similar drop in data triggers alarms faster.
We are currently experiencing issues with our backend services handling account and server creation. This does not affect any running customer servers but might delay provisioning new servers.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently experiencing issues with our backend services handling account and server creation. This does not affect any running customer servers but might delay provisioning new servers.
A fix has been implemented and we are monitoring the results.
Still seeing slow requests, doing another investigation.
Looks like we're back at full performance now.
This incident has been resolved.
We are currently investigating this issue.
Servers in Azure West Europe can experience connection issues due to underlying issues at provider. Latest update from Azure "Current Status: We continue to investigate an availability issue caused by a subset of unhealthy backend storage infrastructure supporting Azure Virtual Machines and Azure Direct Drive in West Europe."
This incident has been resolved.
We are currently experiencing issues with our backend services handling account and server creation. This does not affect any running customer servers but might delay provisioning new servers.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results.
We had an abnormal increase in traffic causing timeouts and slow responses for customer.cloudamqp.com, we increased capacity and are back to normal. This did not affect any running clusters, just provisioning and configuration updates.
We are getting indications of instances being down in Azure region westus. We are investigating the issue.
This incident has been resolved.
Servers in Azure West US 3 can experience connection issues due to underlying issues at provider. We monitor and will update when we know more.
This incident has been resolved.
There appears to be an issue regarding sending metrics to all Legacy metrics integrations. Alarms based on broker metrics also got affected. We are currently investigating the issue Prometheus based metrics (Datadog v3 etc) not affected.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
Broker metrics and metrics integrations degradation Window: 2026-04-23 03:25 – 05:35 UTC \(2h 10m\) ## Impact Broker metrics \(connections, channels, queues, consumers, message rates, node and netsplit state\) were degraded, progressively dropping until recovery at 05:35 UTC. Downstream effects: * Queue, Consumer, Connection, Channel, Connection flow, and Netsplit alarms: threshold breaches during the window may not have triggered, or triggered late. Alarms already tripped before 03:25 UTC kept their state. * Metrics Integrations \(Datadog, CloudWatch, New Relic, Splunk, Dynatrace, etc.\): delivery of broker metrics was reduced for the duration of the window. Alarms configured on the receiving side may likewise have missed or lagged. Not affected: server metrics \(CPU, memory, disk\) and their corresponding alarms, Notice alarms, broker availability, and the metric graphs shown in the CloudAMQP console \(these use a separate data path and continued to update normally\). ## Timeline \(UTC\) * 03:25 — broker-metrics collection begins degrading * 03:35 — ~30% of broker-metric samples missing * 03:50 — ~50% plateau * 05:25 — collector workers cycled and re-initialised * 05:35 — full throughput restored ## Root cause Broker metrics are polled from each cluster's management HTTP API by a pool of collector workers and published to an internal message bus that both the Alarms service and Metrics Integrations consume from. At 03:25 UTC this service became unresponsive without raising an error. Affected workers silently stopped polling while remaining alive to the platform, so automatic restart did not trigger. On-call staff began investigating around 05:00 UTC; force-restarting the service restored metric flow. The exact trigger has not been identified. Our focus is on ensuring the condition is detected and handled promptly if it recurs. ## What we are changing * Added metrics and alarms specific to this service, with high-urgency paging for on-call. * Reliability improvements in the service itself to detect and restart stalled work. ## What you can do CloudAMQP offers a new generation of metrics integrations based on Prometheus, with a re-engineered pipeline that does not rely on these centralised services — each server forwards data directly to your endpoint. These have been running in production for some time and have proved very reliable. * [https://www.cloudamqp.com/blog/prometheus-metrics-integrations.html](https://www.cloudamqp.com/blog/prometheus-metrics-integrations.html) * [https://www.cloudamqp.com/docs/monitoring\_metrics\_datadog\_v3.html](https://www.cloudamqp.com/docs/monitoring_metrics_datadog_v3.html) * Background: [https://www.cloudamqp.com/blog/decentralized-observability-with-open-telemetry-part-1.html](https://www.cloudamqp.com/blog/decentralized-observability-with-open-telemetry-part-1.html)
One of our backend components suffered a small outage due to an unexpected full services restart. Connectivity with the servers were broken and service triggered alarms for many clusters. We apologize for the inconvenience. All clusters were operational and healthy within the incident/alarms timespan.
Customers are experiencing Management UI being unavailable with 502 Bad Gateway. Timeline: - 14:16 UTC -> The issue was identified - 14:30 UTC -> Recent code changes were reverted. No relationship with incident noted.
The Management UI and brokers are serving traffic normally. You can still use the Management UI by using the Broker URL and properly typing User and Password. The impact is related to our backend service logs into Management UI via CloudAMQP SSO.
We are continuing to work on a fix for this issue.
The Management UI SSO shortcut is back to service. Impacted Clusters: All customers with clusters created more than 6 months ago Cause: Some code changes unexpectedly changed the communication schema between our backend services and the SSO component on the clusters. The proxy configuration conflicted with those changes causing the 502 Bad Gateway error to show up. Impact: Only the SSO via CloudAMQP Console was impacted given it required communication with our backend components. The AMQP Clusters were healthy and servicing traffic via both AMQP and HTTP on Management UI and API. Logging into the Management UI via Credentials and all Authentications Backends was functioning normally.
We're currently experiencing issues with metrics delivery, this affects 12% of the API based v1/v2 integrations. We've added additional capacity and expect have handled the backlog within 20 minutes.
This incident has been resolved.
We noticed an issue with our queue metrics being sent to metrics integrations. there is currently a delay in processing, but we have identified an issue and expect the delay to be fixed shortly.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We're currently experiencing issues with Coralogix log delivery.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We're currently experiencing issues with metrics delivery, this affects v3 (Prometheus based) metrics.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
### Summary On March 3, an internal monitoring metric stopped being reported for approximately 14 hours. No customer metrics delivery was affected, all integrations continued operating normally. The missing metric caused our dashboards to display false failure rates, which was incorrectly reported as a metrics delivery outage. ### Impact None. Metrics continued to be delivered to all customer endpoints throughout the incident. ### Resolution The monitoring metric was restored and we have updated our alerting queries to be resilient to missing data, preventing false positives in the future.