我们的错误升高2
- investigating
我们正看到我们两个组的异常流量正在影响着一部分流量。 影响仅限于某些请求,而大多数交通继续正常运行. 我们正在积极调查原因并努力减轻影响.
- investigating
我们看到错误的短暂而频繁地激增,影响到我们两个组群的一小部分交通。 这种影响是断断续续的,并因客户而异,绝大多数客户没有受到影响。 我们的系统已自动扩大能力,以应对不寻常的交通暴动,我们正在继续密切监测局势.
- resolved
这一事件已经得到解决.
自动翻译自官方事件更新。
52 Pusher incidents · 2023年11月 — official updates, affected components, duration and resolution details.
我们正看到我们两个组的异常流量正在影响着一部分流量。 影响仅限于某些请求,而大多数交通继续正常运行. 我们正在积极调查原因并努力减轻影响.
我们看到错误的短暂而频繁地激增,影响到我们两个组群的一小部分交通。 这种影响是断断续续的,并因客户而异,绝大多数客户没有受到影响。 我们的系统已自动扩大能力,以应对不寻常的交通暴动,我们正在继续密切监测局势.
这一事件已经得到解决.
自动翻译自官方事件更新。
我们正在调查影响我们2组的更高误差率。 所有服务仍在运作,但一些客户可能会遇到间歇性问题.
一项固定措施已经执行,我们正在监测结果。 所有出版请求现已成功.
这一事件已经得到解决.
自动翻译自官方事件更新。
We are currently investigating an issue with intermittent errors for the Channels API in US2 cluster.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
Customers on the SA1 cluster may have experienced delayed or missed webhook deliveries, including cache channel webhooks.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently investigating this issue.
The fix has been implemented and we are monitoring the results. The issue affected only part of the cluster traffic. All services are currently operational
This incident has been resolved.
We are currently investigating the issue with increased error rate for the Channels API in cluster us3
A fix has been implemented and we are monitoring the situation
This incident has been resolved.
We are currently investigating this issue.
This incident has been resolved.
We’ve been made aware of an issue caused by our billing provider, Maxio, where a limited number of our customer accounts were incorrectly downgraded due to an incident on their side. We are actively working with Maxio to restore all affected accounts to their correct subscription levels.
All affected accounts have been restored to their correct subscription levels. We sent an email yesterday to all impacted customers. If you notice any issues with your subscription, please contact us.
We are currently investigating this issue.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently investigating this issue.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently experiencing disruptions with socket connections in AP1 cluster, which began at 10:00 UTC. Clients may experience reconnections. Our team is actively investigating the issue and working to restore full connectivity as quickly as possible.
A fix is implemented. We are seeing improvements and continue to actively monitor the situation to ensure full stability. During the incident, many clients successfully connected using fallback protocols.
We observed unusual activity on the cluster, likely caused by at least one customer experiencing a cyber attack. Our team is mitigating the issue while simultaneously increasing cluster resources to minimize impact on other customers. The cluster availability is returning to normal. We continue to monitor closely.
All services are now fully operational. We continue to monitor the cluster.
This incident has been resolved.
We're experiencing an elevated level of API errors and latency in AP4 and are currently looking into the issue.
The team has identified the issue with one of our backend caching servers. The team is working to restore the connections to the server.
The team continues to make progress restoring the cache services. We anticipate full resolution in the next 10-15 minutes.
A fix has been implemented and we are monitoring. The cluster will take a few more minutes to fully stabilize across all nodes.
We are currently investigating an issue affecting the Channels API. A large number of customers may be unable to publish new messages through the API in the AP4 cluster.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
## **Root Cause Analysis: Elevated API Errors and Outage in AP4 Cluster** **Incident Date:** October 20, 2025 **Status:** Resolved ### **Summary** Between **October 20 and October 21, 2025**, customers using the **AP4 cluster** experienced elevated API errors, latency, and message publishing failures. The issue primarily affected the **Channels API**, preventing customers from publishing new messages and leading to degraded real-time functionality for end-users. During system recovery and while implementing mitigations from a previous incident on Oct 18th, a **misconfigured Redis container** in the AP4 cluster failed to start correctly, preventing caching operations needed for API requests. This misconfiguration went undetected by proactive monitoring, delaying full recovery until October 21 at 17:43 UTC. ### **Impact** Throughout the incident period, customers in the **AP4 cluster** experienced: * **High API error rates** when attempting to publish messages through the Channels API * **Failed or delayed message delivery** for connected clients * **Temporary downtime** for end-customer applications relying on real-time messages Other clusters remained operational, though some minor latency was observed in isolated regions due to dependencies on shared services. ### **Root Cause** This incident resulted from **a chain of events** involving both external and internal factors: 1. \*\*Major AWS Outage \(October 20\)\*\*A large-scale **AWS outage in the US-East region** disrupted multiple dependent systems, impacting several Pusher clusters. 2. **Misconfigured Redis Container \(October 21\)** As systems in the AP4 cluster attempted to scale during recovery, one of the backend **Redis cache containers** failed to start due to a **misconfigured environment variable**. This prevented Redis from initializing properly, resulting in API operations failing or timing out. 3. **Monitoring Gap** Existing monitoring did not capture the **Redis startup failure** because the specific failure mode occurred after initialization checks had passed. This delayed internal detection until API error rates increased and customer impact was observed. 4. **Delayed Customer Communication** Initial updates to customers were delayed while the team triaged the issue and verified the failure pattern, prolonging the time before external notification. ### **Detection and Response** The issue was detected through a combination of **monitoring alerts** showing elevated error rates and **customer reports** of publishing failures. **Timeline of Events:** * **October 20** – AWS outage began, affecting multiple Pusher clusters leading to increased delays and errors. **October 20, evening UTC** – Pusher clusters began recovery as AWS services were restored. * **October 21, 15:06 UTC** – Internal monitoring detected elevated API errors in AP4; engineers began investigation. Incident was unrelated to prior AWS outage. * **October 21, 15:09 UTC** – Root cause identified as a failed Redis caching container. * **October 21, 15:27 UTC** – Restoration of Redis connections underway. * **October 21, 15:29 UTC** – Fix implemented; cluster began gradual recovery. * **October 21, 17:39 UTC** – Full stabilization confirmed across AP4 nodes. * **October 21, 17:43 UTC** – Incident marked resolved after sustained recovery. ### **Resolution** To restore full functionality, the engineering team: * Corrected the **Redis container configuration** preventing startup * Restarted and validated cache services across all AP4 nodes * Confirmed API endpoints were fully operational and message publishing resumed * Monitored latency and error metrics to confirm sustained stability ### **Preventative Actions** To reduce recurrence risk and improve detection and response, Pusher is implementing the following: * **Enhanced Redis Monitoring:** Extending monitoring coverage to detect Redis startup and post-init failures. * **Customer Communication Enhancements:** Improving internal escalation and communication processes to ensure faster external updates. ### **Next Steps and Commitment** We recognize the importance of reliable API performance for our customers. Our teams are conducting a full review of caching dependencies and configuration management across all clusters to prevent similar incidents. We sincerely apologize for the disruption caused by this event and appreciate your patience as we worked through a complex multi-day recovery scenario. Pusher remains committed to transparency, reliability, and continuous improvement in service resilience.
We're seeing an increased latency and a number of 503 errors on us3 cluster. The team is investigating.
We are continuing to investigate this issue.
We are continuing to investigate this issue.
We are currently experiencing a major outage in US3 cluster. Our engineering team is actively working to restore full functionality as quickly as possible. In the meantime, we recommend customers temporarily switch to an alternative cluster to minimize service disruption.
Our team has identified the root cause of the issue and implemented a fix. The affected cluster is now in the process of recovering, and performance is gradually returning to normal. We will continue to closely monitor the situation until full recovery is confirmed.
All services are now fully operational. Our team continues to monitor the system closely to ensure ongoing stability and performance.
This incident has been resolved.
## Root Cause Analysis: Redis Cluster Startup Failures – US3 Cluster **Incident Date:** October 20, 2025 **Duration:** 18:36 UTC – 20:41 UTC **Status:** Resolved ## **Summary** On October 20, 2025, customers experienced a significant outage affecting the US3 cluster. The disruption began with increased latency and 503 errors before escalating into full service downtime as Redis clusters in the US3 region failed to start successfully during an infrastructure update. Service was fully restored at 20:41 UTC after engineers identified and resolved the underlying startup issue, confirming stability across all Redis clusters. ## **Impact** Between **18:36 UTC and 20:41 UTC**, customers experienced: * Major outage in the **US3** cluster, impacting message delivery and connection reliability * Elevated error rates and timeouts across dependent APIs * Temporary need for customers to switch to alternate clusters to maintain service continuity No customer data was lost. However, applications relying solely on the affected cluster experienced full downtime for a large portion of the incident. ## **Root Cause** The outage was caused by a failure of Redis clusters to start correctly following an infrastructure update due to a missing configuration flag upon startup. An unexpected upgrade to the Docker runtime running our Redis cluster introduced a breaking change that prevented container startup for certain Redis deployments. When replacement Redis instances in the US3 redis-main cluster were launched, they failed initialization checks and repeatedly restarted, rendering the cluster unavailable. The incompatibility remained undetected until a routine node replacement in the US3 cluster introduced new Redis instances to the cluster. ## **Detection and Response** Monitoring systems first detected increased error rates and latency at **18:36 UTC**, followed by a rise in 503 responses from the affected APIs. **Timeline of events:** * **18:36 UTC** – Increased latency and 503 errors observed in US3 cluster * **19:01 UTC** – Engineering began investigation into Redis startup failures * **19:04 UTC** – Incident declared a major outage; mitigation efforts initiated * **19:59 UTC** – Root cause identified and configuration fix applied * **20:20 UTC** – Services operational; monitoring for recovery stability * **20:41 UTC** – All Redis nodes confirmed healthy; incident resolved ## **Resolution** The engineering team: * Implemented a temporary fix to the Redis environments ensure Redis instances could initialize successfully * Blocked further automated replacements in other clusters until validated * Verified recovery and stability across all Redis clusters After deployment of the fix, Redis instances in the US3 redis-main cluster started correctly and full service was restored. ## **Preventative Actions** To prevent recurrence, the team has: * Rolled out a permanent fix across all Redis clusters used in all regions * Planned a long-term remediation to modernize Redis image packaging for compatibility with current and future Docker releases ## **Next Steps and Commitment** We are conducting a broader review of infrastructure upgrade processes to better detect runtime incompatibilities before they impact production. We apologize for the disruption this incident caused. Ensuring reliability and transparency remains our highest priority, and we continue to strengthen our processes to maintain consistent, predictable service for all Pusher customers.
Our cloud provider is currently experiencing an incident that's causing increased latency on Pusher Channels main cluster `mt1` . This issue is external to our systems and is not related to our infrastructure. Our team is actively investigating to determine if there are any additional impacts and working to minimize downstream effects while our provider implements a fix.
This incident has been resolved.
Our cloud provider is currently experiencing an incident that's affecting Pusher Beams operations. This issue is external to our systems and is not related to our infrastructure. Our team is actively investigating to determine if there are any additional impacts and working to minimize downstream effects while our provider implements a fix.
This incident has been resolved.
Our cloud provider is currently experiencing an incident affecting webhook deliveries in Beams. This issue is external to our infrastructure and is limited to webhooks. Other Beams functionality remains operational. Our team is actively investigating to identify any additional impacts while monitoring the provider's resolution progress.
The impact on our services is decreasing and we have not observed any additional issues with other services. We continue to monitor the situation closely.
Our cloud provider is reporting recovery across most affected services. On our side, we haven’t noticed any issues in the past hour, and all services are operating normally. We continue to monitor the situation closely and will provide updates if anything changes.
Our cloud provider is currently experiencing an incident that's causing increased latency and webhook delivery failures. This issue is external to our systems and is not related to our infrastructure. Our team is actively investigating to determine if there are any additional impacts and working to minimize downstream effects while our provider implements a fix.
The impact on our services is decreasing and we have not observed any additional issues with other services. We continue to monitor the situation closely.
Our cloud provider is reporting recovery across most affected services. On our side, we haven’t noticed any issues in the past hour, and all services are operating normally. We continue to monitor the situation closely and will provide updates if anything changes.
This incident has been resolved.
Increased latency and degraded performance on clusters main, us2 and us3.
The issue has been identified and we are applying mitigations. We are seeing improved latency in main and us3 clusters. us2 still has elevated latency.
We have scaled all three clusters out to handle the increased traffic. We anticipate the latency to slowly improve.
Latency has remained stable over the past three hours. Our investigation identified multiple contributing factors to the earlier latency increase, including network saturation and capacity limits at our cloud providers affecting the US2 cluster. The US3 cluster issue was resolved quickly through autoscaling, while the US2 cluster experienced scaling delays that required manual intervention. We are implementing corrective measures to prevent recurrence and improve system resilience.
## **Root Cause Analysis: Increased Latency and Message Delivery Failures – US2 Cluster** **Incident Date:** October 19, 2025 **Duration:** 00:44 UTC – 08:17 UTC **Status:** Resolved ### **Summary** On October 19, 2025, customers using Pusher experienced increased latency and message delivery failures. These issues primarily affected the US2 cluster, with intermittent impact also observed in the MT1 and US3 clusters. The incident resulted in delayed or undelivered messages for many applications. Latency stabilized at 08:17 UTC after mitigation actions were completed. ### **Impact** Between 00:44 UTC and 08:17 UTC, multiple customers experienced: * **Delayed or failed message delivery** across affected clusters * **Degraded performance** in connection establishment and publishing The most significant and prolonged impact occurred in the **US2 cluster**, while **MT1** and **US3** clusters saw elevated latency for a shorter period before stabilizing. ### **Root Cause** The primary cause of the incident was IP address saturation within the subnet assigned to the public Pusher clusters. When traffic levels increased, the **US2 cluster** was unable to scale out further because the available IP addresses in its subnet were fully utilized. This IP scaling limitation prevented the creation of additional instances needed to handle the load. Secondary factors included temporary **network saturation** and **capacity limits** at our cloud provider, which amplified the latency in the early stages of the incident. ### **Detection and Response** The issue was first detected through a combination of **customer reports** and **internal monitoring alerts** showing elevated response times and connection errors. The timeline of actions was as follows: * **00:44 UTC** – Monitoring alerted the team to increased latency across MT1, US2, and US3 clusters. * **03:11 UTC** – Engineers identified subnet capacity as a contributing factor; mitigations began. * **04:03 UTC** – All clusters were scaled out to distribute traffic; latency began to improve in MT1 and US3 clusters. * **08:17 UTC** – Manual intervention allowed US2 to scale successfully, restoring normal latency. ### **Resolution** To restore service, the engineering team: * Scaled out **MT1** and **US3** clusters to handle increased traffic loads * Monitored all clusters to confirm sustained stability Once additional capacity was provisioned and high loads normalized, latency levels returned to normal and remained stable. ### **Preventative Actions** To prevent recurrence, Pusher has initiated the following actions: * **Rate Limits:** We will re-evaluate how rate limits are implemented to better mitigate content from neighboring customers on the shared clusters. * **Subnet Expansion:** Re-evaluating and increasing the size of subnets assigned to shared clusters to ensure sufficient IP availability for future scaling events. * **Load Balancer Enhancements:** We will implement load balance sharding in order to better distribute connections. * **Capacity Planning Improvements:** Enhancing internal monitoring and alerting for subnet and IP utilization thresholds. ### **Next Steps and Commitment** We recognize that message latency and delivery reliability are critical to our customers’ applications. Our team is continuing a full review of cluster capacity management and provider configuration to improve resilience under high traffic conditions. We apologize for the disruption this incident caused and appreciate your patience while we worked to resolve it. Ensuring reliability and transparency remains our highest priority.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently investigating this issue.
This incident has been resolved.