Global Errors and Failures with Publish and Functions, Delays with Events & Actions
开始时间 2026年8月14日 UTC 14:29 · 0m
Pending
受影响的组件
North America Points of PresencePublish/Subscribe ServiceAsia Pacific Points of PresenceSouthern Asia Points of PresenceEuropean Points of PresenceFunctions Service
resolved
From 13:55 to 14:15 UTC, users globally may have experienced errors with failure on Publish, and published messages may not have been received by subscribers. We also experienced errors and failures with PubNub Functions, and delays with Events & Actions during the period. Services have returned to normal and we are actively monitoring.
A root cause analysis will be published in the coming days. If you believe you were impacted and would like to speak with us, please report impact to [email protected].
postmortem
### **Problem Description, Impact, and Resolution**
At approximately **13:55 UTC on Aug 14, 2026**, we observed elevated publish errors and message replication failures in our publish/subscribe service, which also caused latency in other PubNub services globally. Customers may have experienced increased publish error rates, delayed or missed message delivery, delayed message persistence, and increased latency for Functions and Events & Actions workflows.
The root cause of the incident was an unusually large concentration of global publish traffic that was not limited by our throttling layers. That traffic created resource pressure in the publish and replication layers, increased load on storage systems, and caused downstream processing delays in dependent services. We mitigated the issue by adjusting targeted traffic controls, increasing capacity for affected publish and replication components, and isolating the high-volume traffic pattern to reduce broader platform impact. The issue was resolved at approximately **14:15 UTC on Aug 14, 2026**.
This issue occurred because of an issue with our automated controls specific to this exceptional traffic pattern. As a result, the increased load affected multiple services before a mitigation could be fully applied.
### **Mitigation Steps and Recommended Future Preventative Measures**
To prevent a similar issue from occurring in the future, we have isolated the identified high-volume traffic pattern onto dedicated infrastructure and also increased baseline capacity for the affected components across our PoPs.
We are also further strengthening our traffic detection and management processes for this exceptional load pattern so it will be identified and contained earlier without cascading impact across dependent services.
Replication failures
开始时间 2026年6月10日 UTC 20:14 · 31m
Pending
受影响的组件
North America Points of PresencePublish/Subscribe Service
investigating
Starting at 17:00 UTC on June 10, a subset of publishes originating from the North America POP failed to replicate to subscribers globally. PubNub Technical Staff is investigating, and more information will be posted as it becomes available.
monitoring
The PubNub Technical Staff identified that failures are limited to publishes from US-East and has identified, have applied a fix, and are now monitoring the results.
If you are experiencing issues that you believe to be related to this incident, please report the details to PubNub Support ([email protected]).
resolved
Services have returned to normal. A root cause analysis will be published in the coming days. If you believe you were impacted and would like to speak with us, please report impact to [email protected].
postmortem
## Problem Description, Impact, and Resolution
At 19:50 UTC on June 10, 2026, we observed a small fraction of publishes originating from US-EAST-1 failing to replicate to subscribers globally. We removed the degraded publisher pod from service and the issue was resolved at 21:21 UTC on June 10, 2026. The root cause of the incident was triggered by a single process that fell into a degraded state where it continued receiving inbound traffic and passing health checks, but traffic sent outbound from the process was failing at an abnormally high rate. Our automated health check/recovery system did not auto-detect and replace the degraded process because its health check API reported itself as healthy.
## Mitigation Steps and Recommended Future Preventative Measures
To prevent a similar issue from occurring in the future, we are improving the data within our health check APIs to return more complete performance metrics over a rolling time window. We are enhancing the issue detection logic to detect more patterns that infer process failure, even if the process itself is reporting as healthy.
Global Errors and Failures with Publish and Functions, Delays with Events & Actions
开始时间 2026年6月9日 UTC 14:00 · 0m
Pending
resolved
From 13:20 to 13:48 UTC, users globally may have experienced errors with failure on Publish, and published messages may not have been received by subscribers. We also experienced errors and failures with PubNub Functions, and delays with Events & Actions during the period. Services have returned to normal and we are actively monitoring.
A root cause analysis will be published in the coming days. If you believe you were impacted and would like to speak with us, please report impact to [email protected].
postmortem
### **Problem Description, Impact, and Resolution**
At approximately **13:20 UTC on June 9, 2026**, we observed elevated publish errors and message replication failures in our publish/subscribe service, which also caused latency in other PubNub services globally. Customers may have experienced increased publish error rates, delayed or missed message delivery, delayed message persistence, and increased latency for Functions and Events & Actions workflows.
The root cause of the incident was an unusually large concentration of global publish traffic that was not limited by our throttling layers. That traffic created resource pressure in the publish and replication layers, increased load on storage systems, and caused downstream processing delays in dependent services. We mitigated the issue by adjusting targeted traffic controls, increasing capacity for affected publish and replication components, and isolating the high-volume traffic pattern to reduce broader platform impact. The issue was resolved at approximately **13:48 UTC on June 9, 2026**.
This issue occurred because we did not have sufficient automated controls and isolation processes in place to protect shared infrastructure from this type of exceptional traffic pattern. As a result, the increased load affected multiple services before mitigation could be fully applied.
### **Mitigation Steps and Recommended Future Preventative Measures**
To prevent a similar issue from occurring in the future, we have isolated the identified high-volume traffic pattern onto dedicated infrastructure and also increased baseline capacity for the affected components across our PoPs.
We are also further strengthening our traffic detection and management processes for exceptional load patterns so they can be identified and contained earlier without cascading impact across dependent services.
Connectivity Issues Affecting a Subset of Subscriptions
开始时间 2026年3月24日 UTC 20:21 · 52m
Pending
受影响的组件
Publish/Subscribe Service
investigating
As of 19:27 UTC, we are investigating reports of connectivity issues affecting a subset of subscriptions across several global regions. While the majority of the PubNub network is operating normally, users on the affected segment may experience:
- Connection delays and increased latency.
- Intermittent errors when subscribing to channels.
Our engineering team has identified the cause and is actively working to restore stability. We will provide updates as soon as more information is available.
resolved
The incident has been resolved. If you believe you were impacted by the incident and wish to discuss it with our team, please contact us by email at [email protected]
postmortem
### **Problem Description, Impact, and Resolution**
On March 24, 2026, at 19:27 UTC, one network shard experienced intermittent connectivity affecting a subset of customers. The affected users may have experienced elevated latency and temporary error responses related to their subscription requests. The instability was caused by an atypical surge in message volume within a shared processing environment that had improperly configured resource limits. This led to high resource utilization and triggered automated system restarts. PubNub Engineering resolved the issue by implementing the proper limits after expanding infrastructure capacity to accommodate the increased load. Service was fully stabilized once the environment was tuned to the new traffic profile.
### **Mitigation Steps and Recommended Future Preventative Measures**
**Infrastructure Tuning:** Adjusted automated scaling parameters to provide greater headroom for rapid traffic fluctuations.
**Enhanced Traffic Management:** Deployed refined monitoring heuristics to better isolate and manage high-volume traffic patterns without impacting shared resources.
**Dynamic Resource Allocation:** Accelerating the rollout of enhanced vertical scaling technology to allow individual processing nodes to adapt more fluidly to demand spikes.
**Operational Coordination:** Strengthening internal protocols for high-capacity events to ensure large-scale traffic shifts are proactively transitioned to dedicated environments.
Delay in Publishing Messages to Storage Globally
开始时间 2026年1月1日 UTC 00:25 · 35m
Issues轻微事件
受影响的组件
North America Points of PresenceStorage and Playback ServiceAsia Pacific Points of PresenceSouthern Asia Points of PresenceEuropean Points of Presence
investigating
Starting at 12:13AM GMT, we began seeing delays in published messages being written to history storage. Users may be unable to access recently stored messages through History API calls. The messages will eventually be published.
Our engineers are actively working to address the issue and we will provide updates here. If you feel your service has been impacted and you would like to discuss, please email [email protected].
monitoring
A fix has been applied and we are monitoring the results. Published messages waiting on queue to be written to history have now been successfully written. We will continue to monitor for the next 15 minutes.
resolved
The incident has been resolved. Published messages continue to be written to history storage normally. No messages were lost during the incident. If you believe you were impacted by the incident and wish to discuss it with our team, please contact us by email at [email protected]
postmortem
### **Problem Description, Impact, and Resolution**
On January 1, 2026 at 00:00 UTC, we observed elevated latency in our History service across multiple regions. Customers may have experienced delays in message persistence and history availability during this period.
The issue was caused by a mismatch in newly created persistence tables. Specifically, required columns for message metadata were missing from the new tables, resulting in failed write operations and backed-up queues. This created downstream pressure on our storage systems, leading to higher latency in history processing.
We mitigated the issue by manually applying the correct updates across all affected persistence spaces. After the updates were applied, message processing returned to normal and queue latency cleared.
This issue occurred because we did not have proper controls in place to ensure schema consistency for newly generated monthly persistence tables.
### **Mitigation Steps and Recommended Future Preventative Measures**
To resolve the issue, we manually applied the required schema updates globally. In the coming days, we will update our change management processes to ensure schema changes are correctly applied to all future monthly tables. We are also auditing our schema tracking and automating validation to prevent inconsistencies across environments.
These improvements will ensure that future table generation includes all necessary columns and reduce the risk of similar issues impacting History service performance.
Increased errors observed and resolved
开始时间 2025年11月20日 UTC 16:30 · 0m
Issues轻微事件
resolved
Increased errors (5XX) observed
On Nov. 20, 2025 between 16:36 UTC and 16:48 UTC, publish and replication errors (5xx) were observed in North American and Asia Pacific regions, and immediately resolved.
Please contact [email protected] if you believe you were impacted by the issue and would like to discuss it with our team.
postmortem
**Problem Description, Impact, and Resolution**
Starting at 16:46 UTC on Nov. 20, 2025, we noticed a small number of errors with the publish API in the North American and Asia Pacific regions. The system automatically recovered with all functionality fully restored by 16:50 UTC on Nov. 20, 2025.
Increased latency and errors observed in US-West
开始时间 2025年10月20日 UTC 13:37 · 10h 11m
Issues轻微事件
受影响的组件
North America Points of PresencePublish/Subscribe ServicePresence ServiceAccess Manager Service
investigating
Beginning at 12:40 UTC, we observed increased latency and errors in our US-West region for multiple services. PubNub Technical Staff is currently investigating, and more updates will follow once available.
investigating
We are continuing to investigate this issue.
investigating
Apologies, but we are continuing to investigate this issue.
identified
We are still working with our infrastructure provider to mitigate the effects of the ongoing issues. We have been monitoring and handling traffic to minimize disruptions. While points of presence in the US may see slightly higher latencies, we are stable. Please reach out to PubNub support to report issues.
identified
We are still stable, but seeing minor issues. We are still working with our infrastructure provider to address the remaining issues. Please reach out to PubNub support with any issues.
monitoring
PubNub remains stable and fully operational. Our infrastructure provider is continuing to apply mitigation steps to restore normal health across their network and related services. We are closely monitoring the situation and coordinating with them to ensure continued stability. Please reach out to PubNub Support if you experience any service issues.
identified
The issue has been identified and a fix is being implemented.
identified
PubNub continues to remains stable, with only minor issues being observed. Our infrastructure provider is making progress in restoring their services, with early signs of recovery in several regions. We continue to work closely with them as they complete mitigations across the remaining areas. Please reach out to PubNub Support if you experience any impact.
identified
PubNub remains stable, with only minor issues being observed. Our infrastructure provider is making progress in restoring their services, with early signs of recovery in several regions. We continue to work closely with them as they complete mitigations across the remaining areas. Please reach out to PubNub Support if you experience any impact.
identified
Our status is the same: stable, and working with our infrastructure provider to restore services to their normal state. Our normal policy is to update every 30 minutes, but we will stop issuing this status; we will update this page when a change to the status occurs. As always, reach out to PubNub support with any issues.
monitoring
PubNub remains stable, and our infrastructure provider continues to make steady progress in restoring services. System performance and reliability are improving, and functionality across affected components is returning to normal levels. We continue to monitor closely and will share additional updates as recovery progresses. We will provide another update when there are any changes in status which we expect within the hour. Please contact PubNub Support if you experience any impact.
monitoring
PubNub remains stable, and our infrastructure provider has restored normal service capacity. Systems are now operating at pre-event levels, and most dependent services are performing as expected. We continue to monitor closely and will share further updates as needed. Please contact PubNub Support if you experience any impact.
resolved
This incident has been fully resolved. Our infrastructure provider has restored normal operations, and all systems are functioning as expected. Service performance and reliability have returned to pre-event levels. PubNub continues to monitor systems closely to ensure ongoing stability and we will provide an RCA in in the next couple of day.
Thank you for your patience and understanding throughout this event. If you experience any further issues, please reach out to PubNub Support.
postmortem
### **Problem Description, Impact, and Resolution**
On October 20th, 2025 at 07:06 UTC, our monitoring systems alerted us to elevated error levels across multiple PubNub services in the IAD region \(US-East\). Some customers may have experienced increased error rates and latency, as well as intermittent issues with Presence service availability across IAD \(US-East\), SJC \(US-West\), and HND \(AP-Northeast\).
We quickly determined the issue was caused by a broader infrastructure outage affecting our cloud provider \(AWS\) in the IAD region. We initiated regional failover procedures and re-routed new connections to alternate regions. However, due to undefined steps in some of our failover processes and delays accessing some tools due to the provider issue, existing connections for some services remained degraded for longer than expected.
To restore full service, we manually reset established connections, re-routed Presence traffic to Frankfurt \(EU-Central\), and brought on additional infrastructure in other regions to absorb traffic. Errors were mitigated by 09:20 UTC. Later in the day, additional regional load in US-West triggered a new wave of service degradation. We responded by isolating the US-East region again and scaling up balancer capacity in US-West. PubNub services were stabilized by 13:20 UTC, and remained in a monitoring state while our infrastructure provider worked to fully resolve the underlying issue.
By 22:35 UTC, our provider reported full restoration of service. After validating stability in US-East, we completed rebalancing traffic by 23:48 UTC, and declared the incident resolved.
### **Mitigation Steps and Recommended Future Preventative Measures**
While this incident was caused by an external infrastructure outage, we’ve identified several opportunities to strengthen our internal readiness and response procedures.
We are consolidating and centralizing our regional failover procedures to ensure they are immediately accessible and complete for all production services. Any gaps in our process documentation for newer services will be addressed to ensure readiness before they are fully adopted into production. Additionally, we are reviewing and resolving issues with internal tooling, including inventory and DNS resolution problems, which made mitigation more difficult during the incident.
These improvements will ensure faster and more consistent responses to future infrastructure-level disruptions, and reduce potential impact on customer traffic across regions.
Elevated latencies and errors for multiple services in US-west and US East
开始时间 2025年10月20日 UTC 07:32 · 1h 47m
Issues轻微事件
受影响的组件
North America Points of PresencePublish/Subscribe ServiceDNS ServiceStorage and Playback ServicePresence ServiceAccess Manager ServiceApp Context ServiceMobile Push GatewayStream Controller Service
investigating
At approximately 07:02 UTC, PubNub services began experiencing elevated latencies and server errors in the US-West and US-East regions. PubNub Technical Staff is currently investigating, and more updates will follow once available.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
With no further issues observed for the past 30 minutes, the incident has been resolved. We will follow up soon with a root cause analysis. If you believe you experienced an impact related to this incident, please report it to PubNub Support at [email protected].
postmortem
### **Problem Description, Impact, and Resolution**
On October 20th, 2025 at 07:06 UTC, our monitoring systems alerted us to elevated error levels across multiple PubNub services in the IAD region \(US-East\). Some customers may have experienced increased error rates and latency, as well as intermittent issues with Presence service availability across IAD \(US-East\), SJC \(US-West\), and HND \(AP-Northeast\).
We quickly determined the issue was caused by a broader infrastructure outage affecting our cloud provider \(AWS\) in the IAD region. We initiated regional failover procedures and re-routed new connections to alternate regions. However, due to undefined steps in some of our failover processes and delays accessing some tools due to the provider issue, existing connections for some services remained degraded for longer than expected.
To restore full service, we manually reset established connections, re-routed Presence traffic to Frankfurt \(EU-Central\), and brought on additional infrastructure in other regions to absorb traffic. Errors were mitigated by 09:20 UTC. Later in the day, additional regional load in US-West triggered a new wave of service degradation. We responded by isolating the US-East region again and scaling up balancer capacity in US-West. PubNub services were stabilized by 13:20 UTC, and remained in a monitoring state while our infrastructure provider worked to fully resolve the underlying issue.
By 22:35 UTC, our provider reported full restoration of service. After validating stability in US-East, we completed rebalancing traffic by 23:48 UTC, and declared the incident resolved.
### **Mitigation Steps and Recommended Future Preventative Measures**
While this incident was caused by an external infrastructure outage, we’ve identified several opportunities to strengthen our internal readiness and response procedures.
We are consolidating and centralizing our regional failover procedures to ensure they are immediately accessible and complete for all production services. Any gaps in our process documentation for newer services will be addressed to ensure readiness before they are fully adopted into production. Additionally, we are reviewing and resolving issues with internal tooling, including inventory and DNS resolution problems, which made mitigation more difficult during the incident.
These improvements will ensure faster and more consistent responses to future infrastructure-level disruptions, and reduce potential impact on customer traffic across regions.
Potential for some missed messages for subscribers in IAD
开始时间 2025年10月17日 UTC 06:42 · 1h 24m
Outage重大事件
受影响的组件
North America Points of PresencePublish/Subscribe Service
investigating
We are currently investigating an incident that could lead to some missed messages by subscribers in the IAD region. All messages are being received and persisted, and can be retrieved from the Storage service.
This incident started around 22:07 UTC (03:07 PDT) on 16th Oct 2025. We suspect a moderate impact. Please report any impact related to this incident to [email protected] with any details that you can provide.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
There have been no further issues for the past 45 minutes. We are resolving this issue, and we will follow up with a post-mortem soon.
postmortem
### **Problem Description, Impact, and Resolution**
On October 17, 2025 at 04:51 UTC, some customers may have experienced elevated latency and error rates with the Pub/Sub service in the IAD region \(US-East\). Our engineering teams began immediate investigation and identified a spike in errors related to a recent update to the Pub/Sub service.
We began formal incident response and initiated rollback of the service deployment shortly thereafter. The issue was fully resolved by 06:50 UTC, and rollback across all regions was completed by 08:00 UTC.
The issue occurred because a misconfiguration in the release caused incorrect behavior in the channel cleanup logic. Additionally, our alerting configuration did not include coverage for the synthetic test failures that would have surfaced this issue sooner, delaying detection.
### **Mitigation Steps and Recommended Future Preventative Measures**
To prevent a similar issue from occurring in the future, our engineering teams have written a simpler and more reliable replacement for the faulty logic. That code is currently undergoing rigorous testing before being reintroduced in a future release.
We are also addressing the lack of proper alerting that contributed to a delayed response. Synthetic tests have been reviewed, and appropriate alerting will be implemented to ensure similar regressions are detected earlier. In parallel, we are updating our development and testing processes to catch such issues before code reaches production. Lastly, we are conducting a refresher training on our incident response process to ensure faster execution and coordination in the future.
Elevated Event & Action Error In FRA Region
开始时间 2025年10月6日 UTC 10:13 · 47m
Issues轻微事件
受影响的组件
European Points of Presence
investigating
At about 08:40 UTC, the Event & Action publish operation began to experience elevated error rates. Our Technical Staff is actively investigating, and more information will be posted as it becomes available.
If you are experiencing issues that you believe are related to this incident, please report the details to PubNub Support ([email protected]).
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
With no further issues observed, the incident has been resolved. We will follow up soon with a root cause analysis.
If you believe you experienced an impact related to this incident, please report it to PubNub Support at [email protected].
postmortem
### **Problem Description, Impact, and Resolution**
At 08:40 UTC on October 6, 2025, we observed elevated error rates in the Events & Actions service in our EU-Central \(FRA\) region, which led to delays in processing publish-triggered events. Some customers may have experienced slower-than-expected execution of their event workflows during this time.
We identified a malformed payload that was causing backend consumers to fail when attempting to process the queue. We deployed an updated build with improved parsing logic, which cleared the blockage and restored normal service. The issue was fully resolved by 11:00 UTC on October 6, 2025.
This issue occurred because our event processing service did not correctly handle a malformed message format, which caused the processing queue to stall. Additionally, the alerting system in place was not configured to detect this failure mode promptly, delaying our response.
### **Mitigation Steps and Recommended Future Preventative Measures**
To reduce the risk of similar delays in the future, we are refining our alert thresholds and naming conventions to improve early detection and clarity during response. We are also reviewing validation logic to ensure malformed messages are consistently isolated before reaching backend queues.
Increased Error Rate and Latency for Presence
开始时间 2025年9月7日 UTC 18:45 · 17m
Pending
受影响的组件
Presence Service
monitoring
We experienced increased latencies and error rates in our Tokyo, San Jose, and Virginia points-of-presence from 18:14 - 18:17 UTC. A fix was deployed and we are monitoring the situation.
resolved
The issues affected the Presence service from 18:14 - 18:17 UTC. We will follow up with a post-mortem soon.
We apologize for the impact this may have had on your service. Please reach out to us by contacting PubNub Support ([email protected]) if you wish to discuss the impact on your service.
postmortem
### **Problem Description, Impact, and Resolution**
At 18:14 UTC on September 7, 2025 we observed increased error rates and latency for our Presence service in our San Jose, Virginia, and Tokyo regions. We increased capacity in those regions and the issue was resolved at 18:17 UTC. This issue was a recurrence of the issue [experienced on September 2, 2025](https://status.pubnub.com/incidents/1n8xk6w5y9lk), where a bug in one of our APIs allowed a request to execute an operation that exceeded assumed limits in extreme cases, causing out-of-memory conditions for the Presence service.
### **Mitigation Steps and Recommended Future Preventative Measures**
In the previous instance of this issue, we placed restrictions on the API in question; those changes were not restrictive enough, which allowed for this recurrence. We have corrected that oversight, as well as increased memory capacities in this area of our system as an additional safeguard.
Presence is experiencing elevated latencies and error rates
开始时间 2025年9月2日 UTC 18:34 · 16m
Pending
受影响的组件
North America Points of PresencePresence ServiceAsia Pacific Points of Presence
monitoring
At about 11:09, the Presence began to experience elevated latencies and error rates. PubNub Technical Staff is investigating and more information will be posted as it becomes available.
If you are experiencing issues that you believe to be related to this incident, please report the details to PubNub Support ([email protected]).
resolved
The issues affected the Presence service from 18:08 to 18:16 UTC. We will follow up with a post-mortem soon.
We apologize for the impact this may have had on your service. Please reach out to us by contacting PubNub Support ([email protected]) if you wish to discuss the impact on your service.
postmortem
### **Problem Description, Impact, and Resolution**
At 18:09 UTC on September 2, 2025 we observed increased error rates and latency for our Presence service in our San Jose, Virginia, and Tokyo regions. We increased capacity in those regions and the issue was resolved at 18:16 UTC. This issue occurred because a bug in one of our APIs allowed a request to execute an operation that exceeded assumed limits in extreme cases. In this case, a large number of such requests were executed that resulted in out-of-memory conditions for the Presence service.
### **Mitigation Steps and Recommended Future Preventative Measures**
To prevent a similar issue from occurring in the future we have enforced the intended limit on the API in question. We have also added additional testing and monitoring in this area.
Presence service errors in multiple regions
开始时间 2025年7月31日 UTC 22:14 · 59m
Issues轻微事件
受影响的组件
North America Points of PresencePresence ServiceAsia Pacific Points of Presence
investigating
We have detected elevated error levels with the Presence service in multiple regions. Our Engineers are actively working to mitigate the issues and return service to normal levels. We will provide updates here.
If you believe you have been impacted by the issue, please report impact to [email protected].
identified
The issue has been identified and a fix is being implemented. We will provide updates here as progress is made.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
With no further issues observed, the incident has been resolved. We will follow up soon with a root cause analysis.
If you believe you experienced an impact related to this incident, please report it to PubNub Support at [email protected].
postmortem
### **Problem Description, Impact, and Resolution**
On August 1, 2025 at 22:15 UTC, we observed elevated 5xx errors across the Presence service in multiple regions. Customers may have experienced intermittent failures when attempting to receive or update presence messages.
We identified a subset of channels exhibiting highly concentrated activity patterns and applied a targeted configuration change to rebalance traffic across the cluster. The issue was resolved by 23:15 UTC on August 1, 2025.
This issue occurred because our infrastructure lacked proactive safeguards to evenly distribute presence traffic across nodes in scenarios where a small number of channels receive a disproportionately high number of presence updates. This resulted in resource saturation on some nodes without triggering early mitigation.
### **Mitigation Steps and Recommended Future Preventative Measures**
To prevent a similar issue from occurring in the future, we have applied a sharding configuration to affected channel patterns, which redistributes load more evenly across infrastructure components. This approach reduces the risk of overload caused by concentrated traffic.
In the coming days we will be:
* Reviewing long-term suitability of the sharding configuration applied during this incident.
* Investigating automation options for dynamically applying sharding logic based on real-time usage patterns.
* Enhancing internal tooling and monitoring to better detect and respond to load imbalance scenarios before they cause service degradation.
Elevated Presence Latency and Errors in FRA
开始时间 2025年7月24日 UTC 15:30 · 0m
Pending
resolved
Between 15:33-15:59 UTC we observed elevated levels of latency and errors with our presence service in our FRA region. The issue has been resolved and remains stable.
We apologize for the impact this may have had on your service. Please reach out to us by contacting PubNub Support ([email protected]) if you wish to discuss the impact on your service.
postmortem
### **Problem Description, Impact, and Resolution**
At 14:43 UTC on July 22, 2025, we observed elevated latency and service degradation in our Presence system in the FRA region, impacting customers' ability to receive timely presence events. We implemented changes in our infrastructure to mitigate the pressure on the system, and the incident ended at 15:06 UTC on July 22, 2025. Unfortunately, before we were able to conclusively determine and address the root cause, the issue re-occurred from 15:33 to 15:59 UTC on July 24, 2025.
The failures started with an out-of-memory condition in some Presence service nodes. The root cause of the memory condition was a caching system—operated by a third-party vendor—that we rely heavily on for this service began responding slowly at first, and later timing out. While nodes that run into critical issues are generally taken out of service and replaced automatically, this issue began affecting our system faster than new capacity could be brought online to mitigate the incident.
Furthermore, a bug in a downstream internal system caused it to also experience failures under the pressure due to a misconfiguration in how it retried failed requests caused by the Presence system’s issues.
On the 22nd, we took manual action to break the retry cycle, scale capacity, and allow the Presence service to recover. On the 24th, we identified and addressed the root cause at the caching layer.
### **Mitigation Steps and Recommended Future Preventative Measures**
In the short term, we have scaled the capacity of the Presence system and we implemented changes in the affected caching layer to mitigate the issue of the vendor-provided cache system. Longer term, we are working with the vendor to fix their underlying issue, while also exploring alternatives. We also changed the retry configuration to prevent this kind of retry storm from happening again. We implemented monitoring and alerts for this failure mode, to allow us to more quickly identify this kind of issue.
Intermittent Presence Latency Spikes in FRA
开始时间 2025年7月22日 UTC 18:18 · 0m
Issues轻微事件
resolved
Between 14:43-15:06 UTC we observed intermittent spikes in errors and latency with our presence service in our FRA region. The issue has been resolved and remains stable.
postmortem
### **Problem Description, Impact, and Resolution**
At 14:43 UTC on July 22, 2025, we observed elevated latency and service degradation in our Presence system in the FRA region, impacting customers' ability to receive timely presence events. We implemented changes in our infrastructure to mitigate the pressure on the system, and the incident ended at 15:06 UTC on July 22, 2025. Unfortunately, before we were able to conclusively determine and address the root cause, the issue re-occurred from 15:33 to 15:59 UTC on July 24, 2025.
The failures started with an out-of-memory condition in some Presence service nodes. The root cause of the memory condition was a caching system—operated by a third-party vendor—that we rely heavily on for this service began responding slowly at first, and later timing out. While nodes that run into critical issues are generally taken out of service and replaced automatically, this issue began affecting our system faster than new capacity could be brought online to mitigate the incident.
Furthermore, a bug in a downstream internal system caused it to also experience failures under the pressure due to a misconfiguration in how it retried failed requests caused by the Presence system’s issues.
On the 22nd, we took manual action to break the retry cycle, scale capacity, and allow the Presence service to recover. On the 24th, we identified and addressed the root cause at the caching layer.
### **Mitigation Steps and Recommended Future Preventative Measures**
In the short term, we have scaled the capacity of the Presence system and we implemented changes in the affected caching layer to mitigate the issue of the vendor-provided cache system. Longer term, we are working with the vendor to fix their underlying issue, while also exploring alternatives. We also changed the retry configuration to prevent this kind of retry storm from happening again. We implemented monitoring and alerts for this failure mode, to allow us to more quickly identify this kind of issue.
Increased Latency and Errors in Presence in the US (East and West) and AP-North regions
开始时间 2025年7月8日 UTC 17:36 · 40m
Issues轻微事件
受影响的组件
North America Points of PresencePresence ServiceAsia Pacific Points of Presence
investigating
At approximately 17:18 UTC, PubNub Presence services began experiencing elevated latencies and server errors in the Asia Pacific-North and US regions. PubNub Technical Staff is currently investigating, and more updates will follow once available.
investigating
We are continuing to investigate this issue.
monitoring
Services have returned to normal and we will continue monitoring for the next 30 minutes.
resolved
With no further issues observed for the past 30 minutes, the incident has been resolved. We will follow up soon with a root cause analysis. If you believe you experienced an impact related to this incident, please report it to PubNub Support at [email protected].
postmortem
### **Problem Description, Impact, and Resolution**
On July 8, 2025 at 17:15 UTC on, we observed elevated errors and increased latency for customers using our Presence service in the US East, US West, and Tokyo regions. We identified an abnormal concentration of traffic and applied a configuration change to redistribute load across our infrastructure. The issue was resolved by 17:45 UTC the same day.
### **Mitigation Steps and Recommended Future Preventative Measures**
To prevent a similar issue from occurring in the future, we have implemented targeted balancing for Presence traffic patterns that exhibit such a concentrated load. This allows us to distribute traffic more evenly across infrastructure components. In the coming days, we will work to identify additional patterns that may require similar configuration changes to ensure even load balancing.
Increased errors and latency in FRA region
开始时间 2025年6月30日 UTC 13:39 · 1h 22m
Issues轻微事件
受影响的组件
North America Points of PresenceEuropean Points of Presence
investigating
At approximately 13:11 UTC, PubNub services began experiencing elevated latencies and server errors in the Europe & North US region. PubNub Technical Staff is currently investigating, and more updates will follow once available.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
With no further issues observed, the incident has been resolved. We will follow up soon with a root cause analysis.
If you believe you experienced an impact related to this incident, please report it to PubNub Support at [email protected].
postmortem
Beginning on Friday, June 27, 2025 at 08:15 UTC, there were occasional, intermittent increases in latency and errors in three of our services: Pub/Sub, History, and Presence. The root cause discussed in this analysis was identified and corrected on Monday, June 30.
### **Problem Description, Impact, and Resolution**
Recently, to ensure PubNub had access to more cloud server capacity across our many regions, we introduced new instance types to our system to provide a more heterogeneous set of instance types on which PubNub’s services run. Over time, PubNub has created many OS/kernel-level configurations to optimize the performance of each server. However, with the more heterogeneous instance types, an underlying setting that we were explicitly specifying, which controls limits on network connectivity, was being silently overridden by our upstream load balancers. When we introduced the new instance types, they would reach connectivity limits. Unfortunately, the errors we initially encountered pointed us in incorrect directions, causing the investigation to take longer than we normally strive for.
The issue was mitigated once we identified this issue and configured the affected services to run on other instance types and launched more capacity.
### **Mitigation Steps and Recommended Future Preventative Measures**
To prevent recurrence, we modified the new instance types to emit metrics related to these OS thresholds and limits, enabling us to detect when these limits are approached or exceeded, regardless of instance type. This change allows us to scale proactively and properly route traffic based on instance type, ensuring we are more dynamic in heterogenous instance type deployment configuration.
Again, we apologize for the incidents outlined above and are committed to maintaining transparency when issues affect our customers. Should you have any questions regarding this analysis, please reach out to our support team at [[email protected]](mailto:[email protected]).
Increased errors and latency in FRA region
开始时间 2025年6月30日 UTC 08:12 · 1h 19m
Issues轻微事件
受影响的组件
European Points of Presence
investigating
At approximately 07:30 UTC, PubNub services began experiencing elevated latencies and server errors in the Europe region. PubNub Technical Staff is currently investigating, and more updates will follow once available.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
With no further issues observed, the incident has been resolved. We will follow up soon with a root cause analysis.
If you believe you experienced an impact related to this incident, please report it to PubNub Support at [email protected].
postmortem
Beginning on Friday, June 27, 2025 at 08:15 UTC, there were occasional, intermittent increases in latency and errors in three of our services: Pub/Sub, History, and Presence. The root cause discussed in this analysis was identified and corrected on Monday, June 30.
### **Problem Description, Impact, and Resolution**
Recently, to ensure PubNub had access to more cloud server capacity across our many regions, we introduced new instance types to our system to provide a more heterogeneous set of instance types on which PubNub’s services run. Over time, PubNub has created many OS/kernel-level configurations to optimize the performance of each server. However, with the more heterogeneous instance types, an underlying setting that we were explicitly specifying, which controls limits on network connectivity, was being silently overridden by our upstream load balancers. When we introduced the new instance types, they would reach connectivity limits. Unfortunately, the errors we initially encountered pointed us in incorrect directions, causing the investigation to take longer than we normally strive for.
The issue was mitigated once we identified this issue and configured the affected services to run on other instance types and launched more capacity.
### **Mitigation Steps and Recommended Future Preventative Measures**
To prevent recurrence, we modified the new instance types to emit metrics related to these OS thresholds and limits, enabling us to detect when these limits are approached or exceeded, regardless of instance type. This change allows us to scale proactively and properly route traffic based on instance type, ensuring we are more dynamic in heterogenous instance type deployment configuration.
Again, we apologize for the incidents outlined above and are committed to maintaining transparency when issues affect our customers. Should you have any questions regarding this analysis, please reach out to our support team at [[email protected]](mailto:[email protected]).
Increased global errors and latency for all services
开始时间 2025年6月29日 UTC 17:15 · 0m
Pending
resolved
On June 29, 2025 around 16:00 UTC we began seeing elevated latency and increased errors globally across all services. Our Engineers scaled-up resources and PubNub services returned to normal levels by 16:28 UTC.
postmortem
Beginning on Friday, June 27, 2025 at 08:15 UTC, there were occasional, intermittent increases in latency and errors in three of our services: Pub/Sub, History, and Presence. The root cause discussed in this analysis was identified and corrected on Monday, June 30.
### **Problem Description, Impact, and Resolution**
Recently, to ensure PubNub had access to more cloud server capacity across our many regions, we introduced new instance types to our system to provide a more heterogeneous set of instance types on which PubNub’s services run. Over time, PubNub has created many OS/kernel-level configurations to optimize the performance of each server. However, with the more heterogeneous instance types, an underlying setting that we were explicitly specifying, which controls limits on network connectivity, was being silently overridden by our upstream load balancers. When we introduced the new instance types, they would reach connectivity limits. Unfortunately, the errors we initially encountered pointed us in incorrect directions, causing the investigation to take longer than we normally strive for.
The issue was mitigated once we identified this issue and configured the affected services to run on other instance types and launched more capacity.
### **Mitigation Steps and Recommended Future Preventative Measures**
To prevent recurrence, we modified the new instance types to emit metrics related to these OS thresholds and limits, enabling us to detect when these limits are approached or exceeded, regardless of instance type. This change allows us to scale proactively and properly route traffic based on instance type, ensuring we are more dynamic in heterogenous instance type deployment configuration.
Again, we apologize for the incidents outlined above and are committed to maintaining transparency when issues affect our customers. Should you have any questions regarding this analysis, please reach out to our support team at [[email protected]](mailto:[email protected]).