Stiamo vedendo traffico insolito sul cluster us2 che sta influenzando un sottoinsieme di traffico. L'impatto è limitato a determinate richieste, mentre la maggior parte del traffico continua a funzionare normalmente.
Stiamo attivamente indagando sulla causa e lavorando per mitigare l'impatto.
investigating
Stiamo vedendo brevi ma frequenti punte in errori che interessano un piccolo sottoinsieme di traffico sul cluster us2. L'impatto è intermittente e varia tra i clienti, con la maggior parte dei clienti che non hanno alcun impatto.
I nostri sistemi hanno scalato automaticamente la capacità in risposta agli insoliti scoppi di traffico, e stiamo continuando a monitorare la situazione da vicino.
resolved
Questo incidente è stato risolto.
Tradotto automaticamente dall'aggiornamento ufficiale dell'incidente.
Errori elevati sul cluster us2
Inizio 22 agosto 2026 alle ore 03:17 UTC · 1h 31m
Pending
Componenti interessati
Channels presence channelsChannels WebSocket client APIChannels REST API
investigating
Stiamo indagando su un aumento del tasso di errore che interessa il cluster us2.
Tutti i servizi rimangono operativi, ma un sottoinsieme di clienti può sperimentare problemi intermittenti.
monitoring
Una soluzione è stata implementata e stiamo monitorando i risultati. Tutte le richieste di pubblicazione sono ora in successo.
resolved
Questo incidente è stato risolto.
Tradotto automaticamente dall'aggiornamento ufficiale dell'incidente.
Intermittent API errors in US2 cluster
Inizio 6 aprile 2026 alle ore 01:52 UTC · 1h 4m
IssuesIncidente minore
Componenti interessati
Channels WebSocket client API
investigating
We are currently investigating an issue with intermittent errors for the Channels API in US2 cluster.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Cluster SA1: Delayed Channels webhook delivery on SA1 cluster
Inizio 31 marzo 2026 alle ore 12:33 UTC · 11m
IssuesIncidente minore
Componenti interessati
Channels Webhooks
identified
Customers on the SA1 cluster may have experienced delayed or missed webhook deliveries, including cache channel webhooks.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Increased error rate on the US2 cluster
Inizio 15 febbraio 2026 alle ore 20:33 UTC · 1h 41m
Pending
investigating
We are currently investigating this issue.
monitoring
The fix has been implemented and we are monitoring the results. The issue affected only part of the cluster traffic. All services are currently operational
resolved
This incident has been resolved.
Errors increase for API in cluster us3
Inizio 29 gennaio 2026 alle ore 10:38 UTC · 21m
Pending
Componenti interessati
Channels REST API
investigating
We are currently investigating the issue with increased error rate for the Channels API in cluster us3
monitoring
A fix has been implemented and we are monitoring the situation
resolved
This incident has been resolved.
Increased error rate on the US2 cluster impacting presence events
Inizio 26 gennaio 2026 alle ore 22:29 UTC · 14m
Pending
Componenti interessati
Channels presence channelsChannels Webhooks
investigating
We are currently investigating this issue.
resolved
This incident has been resolved.
Incorrect subscription changes affecting some accounts
Inizio 22 gennaio 2026 alle ore 16:38 UTC · 19h 20m
Pending
Componenti interessati
Channels REST APIChannels Webhooks
identified
We’ve been made aware of an issue caused by our billing provider, Maxio, where a limited number of our customer accounts were incorrectly downgraded due to an incident on their side. We are actively working with Maxio to restore all affected accounts to their correct subscription levels.
resolved
All affected accounts have been restored to their correct subscription levels. We sent an email yesterday to all impacted customers. If you notice any issues with your subscription, please contact us.
Intermittent errors with message delivery on cluster us2
Inizio 16 gennaio 2026 alle ore 08:48 UTC · 3h 52m
IssuesIncidente minore
Componenti interessati
Channels presence channels
investigating
We are currently investigating this issue.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Intermittent errors with message delivery
Inizio 16 gennaio 2026 alle ore 01:07 UTC · 31m
IssuesIncidente minore
Componenti interessati
Channels presence channelsChannels WebSocket client API
investigating
We are currently investigating this issue.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
AP1 cluster - Socket connection failures
Inizio 27 ottobre 2025 alle ore 11:50 UTC · 1h 45m
IssuesIncidente minore
Componenti interessati
Channels presence channelsChannels WebSocket client API
investigating
We are currently experiencing disruptions with socket connections in AP1 cluster, which began at 10:00 UTC. Clients may experience reconnections. Our team is actively investigating the issue and working to restore full connectivity as quickly as possible.
identified
A fix is implemented. We are seeing improvements and continue to actively monitor the situation to ensure full stability. During the incident, many clients successfully connected using fallback protocols.
monitoring
We observed unusual activity on the cluster, likely caused by at least one customer experiencing a cyber attack. Our team is mitigating the issue while simultaneously increasing cluster resources to minimize impact on other customers. The cluster availability is returning to normal. We continue to monitor closely.
monitoring
All services are now fully operational. We continue to monitor the cluster.
resolved
This incident has been resolved.
Elevated API Errors in AP4 Cluster
Inizio 21 ottobre 2025 alle ore 15:06 UTC · 2h 37m
OutageIncidente maggiore
Componenti interessati
Channels WebSocket client APIChannels Stats IntegrationsChannels REST API
investigating
We're experiencing an elevated level of API errors and latency in AP4 and are currently looking into the issue.
identified
The team has identified the issue with one of our backend caching servers. The team is working to restore the connections to the server.
identified
The team continues to make progress restoring the cache services. We anticipate full resolution in the next 10-15 minutes.
monitoring
A fix has been implemented and we are monitoring. The cluster will take a few more minutes to fully stabilize across all nodes.
identified
We are currently investigating an issue affecting the Channels API. A large number of customers may be unable to publish new messages through the API in the AP4 cluster.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem
## **Root Cause Analysis: Elevated API Errors and Outage in AP4 Cluster**
**Incident Date:** October 20, 2025
**Status:** Resolved
### **Summary**
Between **October 20 and October 21, 2025**, customers using the **AP4 cluster** experienced elevated API errors, latency, and message publishing failures. The issue primarily affected the **Channels API**, preventing customers from publishing new messages and leading to degraded real-time functionality for end-users.
During system recovery and while implementing mitigations from a previous incident on Oct 18th, a **misconfigured Redis container** in the AP4 cluster failed to start correctly, preventing caching operations needed for API requests. This misconfiguration went undetected by proactive monitoring, delaying full recovery until October 21 at 17:43 UTC.
### **Impact**
Throughout the incident period, customers in the **AP4 cluster** experienced:
* **High API error rates** when attempting to publish messages through the Channels API
* **Failed or delayed message delivery** for connected clients
* **Temporary downtime** for end-customer applications relying on real-time messages
Other clusters remained operational, though some minor latency was observed in isolated regions due to dependencies on shared services.
### **Root Cause**
This incident resulted from **a chain of events** involving both external and internal factors:
1. \*\*Major AWS Outage \(October 20\)\*\*A large-scale **AWS outage in the US-East region** disrupted multiple dependent systems, impacting several Pusher clusters.
2. **Misconfigured Redis Container \(October 21\)** As systems in the AP4 cluster attempted to scale during recovery, one of the backend **Redis cache containers** failed to start due to a **misconfigured environment variable**. This prevented Redis from initializing properly, resulting in API operations failing or timing out.
3. **Monitoring Gap** Existing monitoring did not capture the **Redis startup failure** because the specific failure mode occurred after initialization checks had passed. This delayed internal detection until API error rates increased and customer impact was observed.
4. **Delayed Customer Communication** Initial updates to customers were delayed while the team triaged the issue and verified the failure pattern, prolonging the time before external notification.
### **Detection and Response**
The issue was detected through a combination of **monitoring alerts** showing elevated error rates and **customer reports** of publishing failures.
**Timeline of Events:**
* **October 20** – AWS outage began, affecting multiple Pusher clusters leading to increased delays and errors. **October 20, evening UTC** – Pusher clusters began recovery as AWS services were restored.
* **October 21, 15:06 UTC** – Internal monitoring detected elevated API errors in AP4; engineers began investigation. Incident was unrelated to prior AWS outage.
* **October 21, 15:09 UTC** – Root cause identified as a failed Redis caching container.
* **October 21, 15:27 UTC** – Restoration of Redis connections underway.
* **October 21, 15:29 UTC** – Fix implemented; cluster began gradual recovery.
* **October 21, 17:39 UTC** – Full stabilization confirmed across AP4 nodes.
* **October 21, 17:43 UTC** – Incident marked resolved after sustained recovery.
### **Resolution**
To restore full functionality, the engineering team:
* Corrected the **Redis container configuration** preventing startup
* Restarted and validated cache services across all AP4 nodes
* Confirmed API endpoints were fully operational and message publishing resumed
* Monitored latency and error metrics to confirm sustained stability
### **Preventative Actions**
To reduce recurrence risk and improve detection and response, Pusher is implementing the following:
* **Enhanced Redis Monitoring:** Extending monitoring coverage to detect Redis startup and post-init failures.
* **Customer Communication Enhancements:** Improving internal escalation and communication processes to ensure faster external updates.
### **Next Steps and Commitment**
We recognize the importance of reliable API performance for our customers. Our teams are conducting a full review of caching dependencies and configuration management across all clusters to prevent similar incidents.
We sincerely apologize for the disruption caused by this event and appreciate your patience as we worked through a complex multi-day recovery scenario. Pusher remains committed to transparency, reliability, and continuous improvement in service resilience.
US3 cluster - major outage
Inizio 20 ottobre 2025 alle ore 18:36 UTC · 2h 6m
OutageIncidente maggiore
Componenti interessati
Channels presence channelsChannels WebSocket client APIChannels REST API
investigating
We're seeing an increased latency and a number of 503 errors on us3 cluster. The team is investigating.
investigating
We are continuing to investigate this issue.
identified
We are continuing to investigate this issue.
identified
We are currently experiencing a major outage in US3 cluster. Our engineering team is actively working to restore full functionality as quickly as possible. In the meantime, we recommend customers temporarily switch to an alternative cluster to minimize service disruption.
identified
Our team has identified the root cause of the issue and implemented a fix. The affected cluster is now in the process of recovering, and performance is gradually returning to normal. We will continue to closely monitor the situation until full recovery is confirmed.
monitoring
All services are now fully operational. Our team continues to monitor the system closely to ensure ongoing stability and performance.
resolved
This incident has been resolved.
postmortem
## Root Cause Analysis: Redis Cluster Startup Failures – US3 Cluster
**Incident Date:** October 20, 2025
**Duration:** 18:36 UTC – 20:41 UTC
**Status:** Resolved
## **Summary**
On October 20, 2025, customers experienced a significant outage affecting the US3 cluster.
The disruption began with increased latency and 503 errors before escalating into full service downtime as Redis clusters in the US3 region failed to start successfully during an infrastructure update.
Service was fully restored at 20:41 UTC after engineers identified and resolved the underlying startup issue, confirming stability across all Redis clusters.
## **Impact**
Between **18:36 UTC and 20:41 UTC**, customers experienced:
* Major outage in the **US3** cluster, impacting message delivery and connection reliability
* Elevated error rates and timeouts across dependent APIs
* Temporary need for customers to switch to alternate clusters to maintain service continuity
No customer data was lost. However, applications relying solely on the affected cluster experienced full downtime for a large portion of the incident.
## **Root Cause**
The outage was caused by a failure of Redis clusters to start correctly following an infrastructure update due to a missing configuration flag upon startup.
An unexpected upgrade to the Docker runtime running our Redis cluster introduced a breaking change that prevented container startup for certain Redis deployments. When replacement Redis instances in the US3 redis-main cluster were launched, they failed initialization checks and repeatedly restarted, rendering the cluster unavailable.
The incompatibility remained undetected until a routine node replacement in the US3 cluster introduced new Redis instances to the cluster.
## **Detection and Response**
Monitoring systems first detected increased error rates and latency at **18:36 UTC**, followed by a rise in 503 responses from the affected APIs.
**Timeline of events:**
* **18:36 UTC** – Increased latency and 503 errors observed in US3 cluster
* **19:01 UTC** – Engineering began investigation into Redis startup failures
* **19:04 UTC** – Incident declared a major outage; mitigation efforts initiated
* **19:59 UTC** – Root cause identified and configuration fix applied
* **20:20 UTC** – Services operational; monitoring for recovery stability
* **20:41 UTC** – All Redis nodes confirmed healthy; incident resolved
## **Resolution**
The engineering team:
* Implemented a temporary fix to the Redis environments ensure Redis instances could initialize successfully
* Blocked further automated replacements in other clusters until validated
* Verified recovery and stability across all Redis clusters
After deployment of the fix, Redis instances in the US3 redis-main cluster started correctly and full service was restored.
## **Preventative Actions**
To prevent recurrence, the team has:
* Rolled out a permanent fix across all Redis clusters used in all regions
* Planned a long-term remediation to modernize Redis image packaging for compatibility with current and future Docker releases
## **Next Steps and Commitment**
We are conducting a broader review of infrastructure upgrade processes to better detect runtime incompatibilities before they impact production.
We apologize for the disruption this incident caused. Ensuring reliability and transparency remains our highest priority, and we continue to strengthen our processes to maintain consistent, predictable service for all Pusher customers.
MT1 cluster - increased latency
Inizio 20 ottobre 2025 alle ore 17:26 UTC · 13h 49m
OutageIncidente maggiore
Componenti interessati
Channels presence channelsChannels WebSocket client APIChannels DashboardChannels REST API
investigating
Our cloud provider is currently experiencing an incident that's causing increased latency on Pusher Channels main cluster `mt1` . This issue is external to our systems and is not related to our infrastructure. Our team is actively investigating to determine if there are any additional impacts and working to minimize downstream effects while our provider implements a fix.
resolved
This incident has been resolved.
Pusher Beams degradation
Inizio 20 ottobre 2025 alle ore 17:21 UTC · 13h 52m
IssuesIncidente minore
Componenti interessati
Beams
investigating
Our cloud provider is currently experiencing an incident that's affecting Pusher Beams operations. This issue is external to our systems and is not related to our infrastructure. Our team is actively investigating to determine if there are any additional impacts and working to minimize downstream effects while our provider implements a fix.
resolved
This incident has been resolved.
Beams - webhook delivery issues
Inizio 20 ottobre 2025 alle ore 08:19 UTC · 1h 59m
Pending
Componenti interessati
Beams
investigating
Our cloud provider is currently experiencing an incident affecting webhook deliveries in Beams. This issue is external to our infrastructure and is limited to webhooks. Other Beams functionality remains operational. Our team is actively investigating to identify any additional impacts while monitoring the provider's resolution progress.
monitoring
The impact on our services is decreasing and we have not observed any additional issues with other services. We continue to monitor the situation closely.
resolved
Our cloud provider is reporting recovery across most affected services. On our side, we haven’t noticed any issues in the past hour, and all services are operating normally. We continue to monitor the situation closely and will provide updates if anything changes.
MT1 cluster - increased latency and webhook delivery failures
Inizio 20 ottobre 2025 alle ore 08:10 UTC · 4h 46m
IssuesIncidente minore
Componenti interessati
Channels Webhooks
investigating
Our cloud provider is currently experiencing an incident that's causing increased latency and webhook delivery failures. This issue is external to our systems and is not related to our infrastructure. Our team is actively investigating to determine if there are any additional impacts and working to minimize downstream effects while our provider implements a fix.
monitoring
The impact on our services is decreasing and we have not observed any additional issues with other services. We continue to monitor the situation closely.
monitoring
Our cloud provider is reporting recovery across most affected services. On our side, we haven’t noticed any issues in the past hour, and all services are operating normally. We continue to monitor the situation closely and will provide updates if anything changes.
resolved
This incident has been resolved.
Increased latency - US2 cluster
Inizio 19 ottobre 2025 alle ore 00:44 UTC · 7h 32m
IssuesIncidente minore
Componenti interessati
Channels presence channelsChannels WebSocket client API
investigating
Increased latency and degraded performance on clusters main, us2 and us3.
identified
The issue has been identified and we are applying mitigations. We are seeing improved latency in main and us3 clusters. us2 still has elevated latency.
identified
We have scaled all three clusters out to handle the increased traffic. We anticipate the latency to slowly improve.
resolved
Latency has remained stable over the past three hours. Our investigation identified multiple contributing factors to the earlier latency increase, including network saturation and capacity limits at our cloud providers affecting the US2 cluster. The US3 cluster issue was resolved quickly through autoscaling, while the US2 cluster experienced scaling delays that required manual intervention. We are implementing corrective measures to prevent recurrence and improve system resilience.
postmortem
## **Root Cause Analysis: Increased Latency and Message Delivery Failures – US2 Cluster**
**Incident Date:** October 19, 2025
**Duration:** 00:44 UTC – 08:17 UTC
**Status:** Resolved
### **Summary**
On October 19, 2025, customers using Pusher experienced increased latency and message delivery failures. These issues primarily affected the US2 cluster, with intermittent impact also observed in the MT1 and US3 clusters. The incident resulted in delayed or undelivered messages for many applications.
Latency stabilized at 08:17 UTC after mitigation actions were completed.
### **Impact**
Between 00:44 UTC and 08:17 UTC, multiple customers experienced:
* **Delayed or failed message delivery** across affected clusters
* **Degraded performance** in connection establishment and publishing
The most significant and prolonged impact occurred in the **US2 cluster**, while **MT1** and **US3** clusters saw elevated latency for a shorter period before stabilizing.
### **Root Cause**
The primary cause of the incident was IP address saturation within the subnet assigned to the public Pusher clusters.
When traffic levels increased, the **US2 cluster** was unable to scale out further because the available IP addresses in its subnet were fully utilized. This IP scaling limitation prevented the creation of additional instances needed to handle the load.
Secondary factors included temporary **network saturation** and **capacity limits** at our cloud provider, which amplified the latency in the early stages of the incident.
### **Detection and Response**
The issue was first detected through a combination of **customer reports** and **internal monitoring alerts** showing elevated response times and connection errors.
The timeline of actions was as follows:
* **00:44 UTC** – Monitoring alerted the team to increased latency across MT1, US2, and US3 clusters.
* **03:11 UTC** – Engineers identified subnet capacity as a contributing factor; mitigations began.
* **04:03 UTC** – All clusters were scaled out to distribute traffic; latency began to improve in MT1 and US3 clusters.
* **08:17 UTC** – Manual intervention allowed US2 to scale successfully, restoring normal latency.
### **Resolution**
To restore service, the engineering team:
* Scaled out **MT1** and **US3** clusters to handle increased traffic loads
* Monitored all clusters to confirm sustained stability
Once additional capacity was provisioned and high loads normalized, latency levels returned to normal and remained stable.
### **Preventative Actions**
To prevent recurrence, Pusher has initiated the following actions:
* **Rate Limits:** We will re-evaluate how rate limits are implemented to better mitigate content from neighboring customers on the shared clusters.
* **Subnet Expansion:** Re-evaluating and increasing the size of subnets assigned to shared clusters to ensure sufficient IP availability for future scaling events.
* **Load Balancer Enhancements:** We will implement load balance sharding in order to better distribute connections.
* **Capacity Planning Improvements:** Enhancing internal monitoring and alerting for subnet and IP utilization thresholds.
### **Next Steps and Commitment**
We recognize that message latency and delivery reliability are critical to our customers’ applications. Our team is continuing a full review of cluster capacity management and provider configuration to improve resilience under high traffic conditions.
We apologize for the disruption this incident caused and appreciate your patience while we worked to resolve it. Ensuring reliability and transparency remains our highest priority.
Higher latency on our AP1 cluster for some customers
Inizio 19 settembre 2025 alle ore 21:15 UTC · 1d 10h