Issue Discovered - Service disruption in Europe Region – Web User Interface
開始 2026年5月1日 8:45 UTC · 34m
Outage重大なインシデント
影響を受けたコンポーネント
Web Interface
investigating
xMatters have identified a potential issue with the xMatters Web User Interface for some clients located in the Europe region. Accessing some pages including Alert details produces an error. We are currently investigating the issue and will update as information becomes available.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team has identified the source of the issue and is working on a fix. We will update once a solution has been identified and implemented.
monitoring
The xMatters Incident Response team has deployed a fix for the issue. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
postmortem
**What happened?**
On May 1st, 2026, some customers reported an issue to xMatters Customer Support where attempting to send a message via the web user interface or viewing an alert on the Alerts report resulted in an error being displayed. The issue only affected the Reporting functions in the EMEA region; the system continued to accept signals, generate alerts, and send notifications across all regions.
**Why did it happen?**
The issue occurred when, during routine database maintenance, the database called a mismatched version of the library, resulting in an internal database error. The version mismatch within the cluster was traced to a prior database engine upgrade where a subset of replica nodes did not restart into the upgraded version. At no point was there any risk to data integrity.
**How did we respond?**
As soon as Customer Support confirmed the issue, they engaged the Engineering teams, who were able to identify the root cause and restore version consistency across all nodes. The teams validated stability and confirmed that all services were restored.
**What are we doing to prevent it from happening again?**
The Engineering Team has added explicit post-upgrade verification checks and monitoring to ensure node alignment is confirmed and maintained. This will provide safeguards to ensure node version alignment after upgrades and prevent this issue from reoccurring.
Issue Discovered - Service disruption in North American Region – API
開始 2026年4月1日 14:06 UTC · 30m
Issues軽微なインシデント
影響を受けたコンポーネント
API
investigating
xMatters monitoring tools have identified a potential issue with the API for some clients located in the North America region. We are currently investigating the issue and will update as information becomes available.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team has identified an issue with the API for some clients located in the North America region which is causing some intermittent API timeouts. This is also causing issues with running reports in the system. We are currently investigating the issue and will update as information becomes available.
monitoring
The xMatters Incident Response team has deployed a fix for the issue. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
postmortem
**What happened?**
On April 1st, 2026, some customers reported an issue to xMatters Customer Support where the Alerts or Notifications reports were timing out and failing to load.
**Why did it happen?**
This issue occurred because a backend service was experiencing significant unexpected load that caused report processing to be delayed. The ongoing resource constraint resulted in timeouts.
**How did we respond?**
As soon as customers reported an issue, Customer Support launched an investigation and escalated to the Engineering teams. The teams initiated rolling restart procedures for the applicable backend services to restore functionality. Once the rolling restart was completed, service was fully restored.
**What are we doing to prevent it from happening again?**
The Engineering teams are currently working to identify any possible bottlenecks that may have caused performance issues. In the interim, they have increased and adjusted resource allocation for several services to handle potential processing delays and prevent potential recurrences.
Issue Discovered - Service disruption in All Regions – SSO login to Web User Interface
開始 2026年3月10日 20:36 UTC · 17m
Issues軽微なインシデント
影響を受けたコンポーネント
Web InterfaceWeb InterfaceWeb Interface
investigating
We have reports that some users are not able to login to xMatters using SSO to login to the xMatters Web User Interface in All Regions. We are currently investigating the issue and will update as information becomes available.
Native Login still works to bypass SSO.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team believes they have identified the source of the issue and is working on a fix. We will update once a solution has been identified and implemented.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
postmortem
**What happened?**
On March 10th, 2026, some customers reported an issue to xMatters Customer Support where they were encountering a 404 error page when attempting to log in to their instances via SSO. Some users may also have encountered a “We’ve run into a problem while retrieving your data.” error message.
**Why did it happen?**
This issue occurred during a routine maintenance update to the xMatters platform. Although several components of the platform were updated, specific configurations of the component related to SSO-based authentication conflicted with the update and resulted in 404 errors. This issue was limited to those few customers that had specific criteria set for their SSO configuration.
**How did we respond?**
As soon as customers reported the issue, Customer Support verified the issue and escalated immediately to Engineering. The team traced the issue to the maintenance deployment and initiated rollback procedures to restore functionality. Once the rollback was completed, the 404 errors ceased and service was confirmed restored.
**What are we doing to prevent it from happening again?**
The Engineering teams have implemented additional testing for any configuration criteria related to the SSO-based authentication component. In addition, they have begun working on an improved maintenance plan to prevent further issues that could occur during similar deployments.
Issue Discovered - Service disruption in North American Region - Multiple Services
開始 2025年11月21日 13:54 UTC · 8m
Issues軽微なインシデント
影響を受けたコンポーネント
Mobile AppIntegration PlatformSMS NotificationsAPIConferencingEmail NotificationsVoice NotificationsWeb Interface
investigating
xMatters monitoring tools have identified a potential issue with xMatters On-Demand for some clients located in the North America region. This is affecting users located near the Asia Pacific Region attempting to access xMatters. We are currently investigating the issue and will update as information becomes available.
Please see incident details for specific services impacted.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
investigating
We are continuing to investigate this issue.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
postmortem
**What happened?**
On November 21, 2025, at 12:50 PM UTC, the xMatters internal monitoring tools detected irregular behavior in how internal traffic was being routed. Some customers in the APAC region communicating with services in North America \(specifically US-East\) may have encountered intermittent request failures or increased latency. Only traffic between these two regions was affected; all other systems and regions continued normal operations.
**Why did it happen?**
A temporary network disruption between Australia Southeast and US-East caused one internal routing node in Australia to lose accurate information about available backend systems in USEast. The node generated an incomplete routing configuration and temporarily stopped directing traffic to US-East. Under normal circumstances, routing updates refresh automatically when connectivity returns. In this case, the affected node did not recover cleanly and remained in a stale state until Engineering intervened.
**How did we respond?**
As soon as Engineering was alerted through internal monitoring, they engaged with the platform engineering team, service owners and Customer Support to launch an investigation. The teams reached out to impacted customers to validate issue symptoms and restarted routing components in both affected regions to force a configuration refresh. Once the restart completed, routing returned to normal levels while the teams continued to monitor and investigate the root cause. They were able to confirm that only one routing node and specific cross-region traffic was impacted.
**What are we doing to prevent it from happening again?**
While teams were mitigating this issue, they created new alerting rules to detect the routing patterns they observed during the incident and expanded internal monitoring to help identify when routing nodes fail to refresh their configuration or otherwise enter a ‘stale’ state. The teams also have planned and prepared infrastructure updates that will further reduce the risk of similar issues. These include improved configuration recovery behavior, enhanced stability for routing components, and additional logging and observability improvements for diagnosing routing anomalies. They will deploy these updates once the current code freeze window has elapsed.
Issue Discovered - Service disruption in All Regions – Conferencing
xMatters monitoring tools have identified a issue with xMatters Conferencing and Voice calls for clients in All Regions. The issue is due to downstream provider outages.
monitoring
Our downstream providers have deployed a fix and testing shows conferencing and voice systems are now working as expected. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
Our provider has confirmed the issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
postmortem
**What happened?**
On October 20th, 2025, at approximately 1:21 AM Pacific, customers began reporting an issue affecting live call routing, conferences, and voice notifications. During this issue, customers in all regions would have been affected.
**Why did it happen?**
This issue was caused by a global AWS outage that impacted one of our downstream providers, resulting in voice notifications, live call routing, and conferencing they were handling to fail.
**How did we respond?**
As soon as the first customer reported an issue, Customer Support engaged the Engineering teams and launched an investigation. Once they identified the root cause, the teams updated the primary provider for affected regions to an alternate provider that was not affected by the AWS outage. When the new provider was assigned, customers reported that all services were operating correctly.
**What are we doing to prevent it from happening again?**
While this issue was out of our control or that of our provider, we are working to identify potential ways to improve resilience in case of external factors such as this.
Issue Discovered - Degraded performance in North American Region – Integration Platform
開始 2025年10月10日 22:58 UTC · 3d 17h
Issues軽微なインシデント
影響を受けたコンポーネント
Integration Platform
monitoring
The xMatters monitoring tools have alerted Customer Support to a potential issue with the integration platform in the North America region. Our technical teams have identified intermittent request failures and occasional slowdowns, though the service appears to be responsive and recovering correctly. We are continuing to monitor the situation and mitigating where possible, though some users may encounter a 5xx response code when submitting a request.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
resolved
The issue has been addressed, and all services are running as expected. Thank you for your patience while we addressed this matter.
postmortem
**What happened?**
On October 10th, 2025, at approximately 4:14 PM Pacific, the xMatters internal monitoring tools alerted Customer Support to service degradation related to an issue that was already being internally monitored with the Integration Platform in the North American region. While this issue was being investigated and mitigated, customers may have experienced intermittent request failures and occasional slowdowns.
**Why did it happen?**
These particular issues were caused by a security update that caused conflicts with the underlying runtime environment, specifically the memory management routine. Slow performance of the routine was triggering frequent health checks for a request processing service and causing automatic restarts.
**How did we respond?**
As soon as the internal monitoring tools alerted Engineering to a potential issue, they began performing manual rolling restarts and deployed more forgiving liveness checks to avoid increasing error rates and to minimize any potential impact to customers. They also increased resources for the impacted service to improve responsiveness of the underlying routine and deployed a configuration fix to the Http Client Cache to try and stabilize the system. While mitigating the potential impact, they also increased monitoring levels as they continued to investigate the root cause and were able to deploy an update to the service and environment configuration on October 13 that resolved the issue. They continued monitoring and confirmed that the system was stable and all services were operational.
**What are we doing to prevent it from happening again?**
Engineering has updated the backend service and deployed additional updates and monitoring to the system configuration that will improve overall stability for the environment and prevent this issue from reoccurring.
Issue Discovered - Service disruption in All Regions – Mobile App
開始 2025年9月24日 3:59 UTC · 44m
Issues軽微なインシデント
影響を受けたコンポーネント
Mobile AppMobile AppMobile App
investigating
xMatters identified a issue with the Apple iOS Push Notifications for clients in All Regions. We are currently investigating the issue and will update as information becomes available.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
xMatters has identified the issue with the Apple Push notifications for clients located in the in All Regions. We are currently investigating a solution and will update as information becomes available.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
We are continuing to work on a fix for this issue.
investigating
xMatters has identified the issue with the Apple Push notifications for clients located in the in All Regions. We are currently deploying a solution and will update as information becomes available. Asia Pacific customers can now receive Apple push notifications.
resolved
The issue has been addressed, and Apple Push notifications have been restored in all regions. Thank you for your patience while we addressed this matter.
postmortem
**What happened?**
On September 23rd, 2025, at approximately 8:31 PM Pacific, the xMatters internal monitoring tools alerted the xMatters Support Team to an issue with Apple iOS notifications. Users were not receiving notifications on Apple iOS devices; other device types and alert and response processing were not impacted and continued to operate without interruption.
**Why did it happen?**
The issue occurred because of a backend service patch applied by the Engineering teams that affected necessary dependencies for Apple Push notification delivery.
**How did we respond?**
As soon as the monitoring tools alerted the Support team, they engaged the Engineering teams to launch an investigation. The teams quickly identified the source of the issue as a recent backend service patch and initiated a rollback of the patch. Once the rollback was complete, the teams confirmed that all services had been restored.
**What are we doing to prevent it from happening again?**
The Engineering team is working to enhance testing of future patches and improving visibility into the health of services. This will help to proactively resolve issues or, when necessary, initiate rollbacks more quickly without impacting customers.
Issue Discovered - Service disruption in North America – Multiple Services
xMatters monitoring tools have identified a potential issue with xMatters On-Demand for clients in North America that are hosted in the on us-central1 region. We are currently investigating the issue and will update as information becomes available.
Please see incident details for specific services impacted.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support team is waiting to help.
identified
The xMatters Incident Response team has identified an issue with our message routing system and is working on a fix. We will update once a solution has been identified and implemented.
monitoring
The xMatters Incident Response team has deployed a fix for the issue. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
postmortem
### What happened?
On August 5th, 2025, at approximately 7:53 PM Pacific, the xMatters internal monitoring tools alerted the xMatters Support team to an issue in the North America region with notification delivery. Some xMatters customers may have experienced delays when receiving notifications or noticed failed alerts in the system.
### Why did it happen?
The issue occurred because of a network problem that disrupted communication between parts of the queuing system. This caused some components to become temporarily out of sync, leading to timeouts, internal connectivity failures, and a small number of messages not being processed during the disruption. As the system recovered, some performance was degraded until normal operation could be restored.
### How did we respond?
As soon as the xMatters monitoring tools reported an issue, the xMatters Support Team initiated the internal Major Incident Management process and engaged the Engineering and incident response teams. The teams quickly mitigated the issue and restored performance by redirecting traffic from our message queuing system in the affected region \(us-central\) to our message queuing system in another local region \(us-east\). After the issue was resolved, the teams directed traffic back to the restored region.
### What are we doing to prevent it from happening again?
The Engineering team has determined that the best approach to prevent this issue from reoccurring is to replace the current message queuing system. Work on the replacement system is well underway and the teams will retire the current system as soon as they have finished the replacement.
xMatters - Support Center - support.xmatters.com
開始 2025年6月23日 20:14 UTC · 32m
Pending
investigating
We have identified a problem affecting the xMatters Support Center. Our team is investigating and working to resolve the issue.
Details: Users will be unable to access the xMatters Support Center. Tickets can be created via email or through Support line.
resolved
We have now resolved the issue with our support center.
Issue Discovered - Service disruption in North American Region – Integration Platform
開始 2025年6月12日 18:47 UTC · 3h 26m
Issues軽微なインシデント
影響を受けたコンポーネント
Integration Platform
investigating
xMatters monitoring tools have identified a potential issue with the xMatters Integration Platform for some clients located in the North America region. We are currently investigating the issue and will update as information becomes available.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team has identified the source of the issue with an upstream provider and is working with them on identifying a fix. We will update once a solution has been identified and implemented.
identified
We are still working with our upstream providers. Their engineers are continuing to mitigate the issue and we have confirmation that the issue is recovered in some locations. We do not have an ETA on full mitigation at this point. We continue to operate in a degraded state.
identified
We are continuing to work with our upstream providers. Their engineers have confirmed that the underlying dependency is recovered in most regions and all their engineering teams are actively engaged and working on service recovery.
We do not have an ETA for full service recovery.
monitoring
At this time, we are seeing all services recovering. We will continue to monitor until we have confirmation from our upstream provider that they have fully resolved the issue.
resolved
We are currently seeing error rates down to normal levels and we are considering this issue resolved at this time
postmortem
**What happened?**
On June 12th, 2025, at approximately 11:29 AM Pacific, the xMatters internal monitoring tools alerted to an issue in the North America region with the xMatters integration platform. While the issue was in progress, some xMatters customers may have experienced workflows not executing correctly or signals not being accepted as expected.
**Why did it happen?**
The issue occurred because multiple Google Cloud and Google Workspace products experienced external API issues for about three hours between the hours of 10:49 AM to 1:49 PM Pacific.
**How did we respond?**
As soon as the xMatters monitoring tools reported an issue with the xMatters integration platform, the xMatters Customer Support team initiated the internal Major Incident Management process and engaged the Engineering and incident response teams. The teams quickly determined that the issues were due to problems with an upstream provider and immediately opened a support case for them to investigate. They continued to work with the provider throughout the incident and confirmed with their engineers as they recovered the underlying dependencies.
**What are we doing to prevent it from happening again?**
The xMatters Support and Engineering teams requested an RCA and prevention methods from the upstream provider. In their response, they confirmed that they will be modularizing their Service Control architecture to prevent failures from affecting the entire service.
Issue Discovered - Service disruption in European Region - Multiple Services
開始 2025年5月13日 17:45 UTC · 3h 1m
Issues軽微なインシデント
影響を受けたコンポーネント
Mobile AppAPIVoice NotificationsIntegration PlatformWeb InterfaceEmail NotificationsConferencingSMS Notifications
investigating
xMatters monitoring tools have identified a potential issue with xMatters On-Demand for some clients located in the Europe region. We are currently investigating the issue and will update as information becomes available.
Please see incident details for specific services impacted.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team has identified the source of the issue and is working on a fix. We will update once a solution has been identified and implemented.
identified
We are continuing to work on a fix for this issue.
monitoring
The xMatters Incident Response team has deployed a fix for the issue. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
postmortem
**What happened?**
On May 13, 2025 at approximately 10:45 AM Pacific, xMatters internal monitoring tools identified an issue where customers in the EU region experienced intermittent web UI and API timeouts.
**Why did it happen?**
The issue occurred because a backend queueing service experienced network timeouts during an unpredictable rapid increase in usage. The increase in resource consumption due to the surge in network usage caused service timeouts and restarts, as well as higher latency which caused further delays in responses to backend requests.
**How did we respond?**
xMatters internal monitoring tools alerted the xMatters Incident Response Team to the issue, then the team launched the internal SEV-1 process. Due to early detection, Engineering teams were able to scale up the queueing services to prevent further service degradation and availability issues. The network timeouts were resolved after resources were scaled up to accommodate the increase in usage.
**What are we doing to prevent it from happening again?**
The Engineering teams have adjusted resources to better compensate for sudden usage increases and to prevent them from affecting backend services. The improvement in resource allocation and adaptability should prevent similar issues from occurring in the future.
Issue Discovered - Service disruption in All Regions – Integration Platform
xMatters monitoring tools have identified a potential issue with xMatters Integration Platform for some clients in All Regions. We are currently investigating the issue and will update as information becomes available.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team has identified the source of the issue and is still working on a fix. We will update once a solution has been identified and implemented.
identified
The xMatters Incident Response team has identified the source of the issue and is actively working on a fix, there is no current estimate on resolution time. We will update once a solution has been implemented.
identified
The xMatters Incident Response team has identified the source of the issue and is still actively working on a fix, there is no current estimate on resolution time. We will update once a solution has been implemented.
identified
The xMatters Incident Response team has identified the source of the issue and is working on a fix. We will update once a solution has been identified and implemented.
identified
The xMatters Incident Response team has identified the source of the issue and are currently testing a fix. We will provide another update shortly.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem
**What happened?**
On April 3, 2025, at approximately 9:32 AM Pacific, the xMatters internal monitoring systems identified an issue where the system was not processing events initiated via Flow Designer across multiple regions. Customers may have observed the system not processing events or creating alerts while the issue was in progress.
**Why did it happen?**
The issue occurred when a routine update to add new permissions to the xMatters' Google Cloud Platform \(GCP\) unexpectedly removed required permissions. When the teams performed the update, which should not have had any impact to customers, the policy used in the automation script was the authoritative resource at the GCP project level rather than authoritative at the individual resource level. Although the teams tested the change before deploying it and found no changes beyond what was included in the update, when the update was deployed to production GCP removed all permissions that were not in the policy in the background. Because the policy only included the new permissions, all other permissions were removed.
**How did we respond?**
As soon as the xMatters monitoring tools reported an issue with the system not processing events, the incident response teams initiated the internal Major Incident Management process and engaged the Engineering and Support teams. The teams were able to quickly identify the recent update as the root cause of the issue and reverted the change to restore the permissions that were managed by the automation process. This restored access and functionality for most of the xMatters services, but restoring permissions for Flow Designer proved to be more complicated.
The teams determined that missing permissions for the xMatters infrastructure were Google-generated permissions essential for specific xMatters services and engaged GCP Support to aid in the investigation. The teams generated a list of all permissions that existed prior to the update and designed a fix to re-apply them to the development environment. Once the teams had implemented the change and validated that the missing permissions and all services had been restored in the development environment, they moved to quickly apply the fix across all staging and production environments. Monitoring tools and customers confirmed that services were fully functional, and the teams continued to monitor the system as it processed all messages queued by the Flow Designer services. The system was fully restored at 2:46 PM Pacific.
**What are we doing to prevent it from happening again?**
The Engineering teams were able to identify all of the permissions created in the xMatters environment, including those that are created by Google, and are ensuring they are added to the management scripts. The teams are adding additional rigor to the application of these types of infrastructure changes to run idempotence tests after the changes are applied to ensure that there are no changes pending. Should a change be applied that fails this test, it will cause failures in development environments, which would catch and prevent a similar issue from occurring.
Issue Discovered - Service disruption in North American Region - Multiple Services
xMatters monitoring tools have identified a potential issue with xMatters On-Demand for some clients located in the North America region. We are currently investigating the issue and will update as information becomes available.
Please see incident details for specific services impacted.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem
**What happened?**
On July 5th, at approximately 8:35AM Pacific, the xMatters monitoring tools alerted Customer Support to an issue where alert notifications were not being sent out for some customers in the North America region. Some customers attempting to initiate alerts may have encountered long delays in processing or may have had requests time out.
**Why did it happen?**
This issue occurred due to a sudden spike in the number of resources required by our backend services. The resulting memory overload issue caused some request handlers to time out before they could properly process incoming alerts.
**How did we respond?**
As soon as the xMatters Customer Support team confirmed the issue from the monitoring tools, they initiate the internal major incident management process and engaged the xMatters Engineering teams. To immediately mitigate the issue and restore service quickly, the incident response teams performed a rolling restart for the affected services. As soon as the restart was completed, the system resumed processing alerts and all services were restored.
**What are we doing to prevent it from happening again?**
The Engineering teams have implemented a performance enhancement for backend service queries. In addition, the teams are evaluating and testing additional methods to help mitigate resource spikes and prevent them from impacting alert notifications in the future. Once development and testing are complete, we'll deploy these changes with our regularly scheduled maintenance.
**Timeline:**
July 5th, 2024
8:35AM PT - xMatters internal monitoring tools alert to potential issue.
9:22AM PT - Issue identified.
9:36AM PT - Rolling restart initiated.
10:16AM PT - Issue Resolved.
Issue Discovered - Service disruption in North America Region – Partial Outage
開始 2024年7月1日 6:23 UTC · 1h 58m
Outage重大なインシデント
影響を受けたコンポーネント
Web Interface
investigating
xMatters monitoring tools have identified a potential issue with xMatters for some clients in the North America region. We are currently investigating the issue and will update as information becomes available.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
monitoring
The xMatters Incident Response team has deployed a fix for the issue. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
Issue Discovered - Service disruption in North America Region – Partial Outage
開始 2024年6月18日 23:16 UTC · 22m
Issues軽微なインシデント
影響を受けたコンポーネント
Web Interface
investigating
xMatters monitoring tools have identified a potential issue with xMatters for some clients in the North America region. We are currently investigating the issue and will update as information becomes available.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team has identified the source of the issue and is working on a fix. We will update once a solution has been identified and implemented.
monitoring
The xMatters Incident Response team has deployed a fix for the issue. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
Intermittent scheduler services in APAC region
開始 2024年5月24日 1:00 UTC · 22m
Outage重大なインシデント
影響を受けたコンポーネント
Web InterfaceVoice NotificationsConferencingEmail NotificationsSMS NotificationsIntegration PlatformAPIMobile App
investigating
We are currently investigating an issue with one of our scheduler services impacting customers in Australia.
investigating
We are continuing to investigate this issue.
investigating
xMatters monitoring tools have identified a potential issue with xMatters On-Demand for some clients located in the Asia Pacific region. We are currently investigating the issue and will update as information becomes available.
Please see incident details for specific services impacted.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team has identified the source of the issue and is working on a fix. We will update once a solution has been identified and implemented.
monitoring
The xMatters Incident Response team has deployed a fix for the issue. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
Issue Discovered - Service disruption in Asia Pacific Region - Multiple Services
開始 2024年5月13日 17:04 UTC · 1h 30m
Issues軽微なインシデント
影響を受けたコンポーネント
Web InterfaceVoice NotificationsConferencingEmail NotificationsSMS NotificationsIntegration PlatformAPIMobile App
investigating
xMatters monitoring tools have identified a potential issue with xMatters On-Demand for some clients located in the Asia Pacific region. We are currently investigating the issue and will update as information becomes available.
Please see incident details for specific services impacted.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team has identified the source of the issue and is working on a fix. We will update once a solution has been identified and implemented.
identified
We are continuing to work on a fix for this issue.
monitoring
The xMatters Incident Response team has deployed a fix for the issue. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
Issue Discovered - Service disruption in All Regions – Mobile App (Android only)
開始 2024年3月21日 16:30 UTC · 6h 42m
Issues軽微なインシデント
影響を受けたコンポーネント
Mobile AppMobile AppMobile App
investigating
We are investigating an issue with the Android mobile app not working on some Android devices. Reports indicate that users are receiving an error message that their devices cannot connect to xMatters We are currently investigating the issue and will update as information becomes available.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Support at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
We've identified this issue as related to the Android certificate transparency library we use for enhanced security in xMatters. Certificate transparency is an important part of the mobile app's end-to-end security and we are currently testing mitigation strategies to ensure they will not cause any disruption of this feature.
We are continuing to investigate the issue and will provide more information as it becomes available.
identified
We are continuing to work toward mitigation and resolution of this issue, and believe we have identified the root cause. We are currently developing and testing a potential fix. In the meantime, push notifications are still being delivered and customers have reported that allowing pages within the app to fully load and waiting before attempting to navigate away or perform another action has significantly reduced occurrences of the error.
We will continue to update as more information becomes available, but note that it will likely be necessary to update the app to resolve the problem.
monitoring
The issue has been resolved; no app update is required.
The problem appears to have stemmed from a third-party conflict: the way the certificate transparency library handles updates from the Google log list conflicts with its own concurrency and job cancellation. We are continuing to monitor the situation and investigating ways to resolve the issue permanently.
resolved
The issue has been resolved, and no app update is required at this time.
We are continuing to investigate the root cause and develop permanent resolution to prevent this issue from reoccurring.
Issue Discovered - Service disruption in Asia Pacific Region - Multiple Services
開始 2023年11月8日 19:26 UTC · 19m
Issues軽微なインシデント
影響を受けたコンポーネント
Web InterfaceVoice NotificationsConferencingEmail NotificationsSMS NotificationsIntegration PlatformAPIMobile App
investigating
xMatters monitoring tools have identified a potential issue with xMatters On-Demand for some clients located in the Asia Pacific region. We are currently investigating the issue and will update as information becomes available.
Please see incident details for specific services impacted.
If you are also experiencing issues, or if you're not sure whether this issue impacts your service, please contact xMatters Client Assistance at https://support.xmatters.com/hc/en-us/requests/new - our support agents are waiting to help.
identified
The xMatters Incident Response team has identified the source of the issue and is working on a fix. We will update once a solution has been identified and implemented.
identified
We are continuing to work on a fix for this issue.
monitoring
The xMatters Incident Response team has deployed a fix for the issue. We are currently monitoring the situation to ensure the implementation is stable and that all services are restored.
resolved
The issue has been addressed, and all services have been restored. Thank you for your patience while we addressed this matter.
postmortem
**What happened?**
On November 9, 2023, at approximately 5:35 AM AEDT, some customers in the APAC region reported an issue to xMatters Customer Support where they were unable to add a new user. The Add User button was greyed out, and hovering over the button was showing the message "You've reached the maximum number of user licenses for your account" despite having additional licenses available. Some users may also have experienced an intermittent inability to log into the web user interface. Throughout this issue and the subsequent mitigation procedures, the system continued to accept events and generate alerts, and all notifications and responses were processed correctly.
**Why did it happen?**
During a regularly scheduled update to the backend services in the APAC region, a timing issue caused the service responsible for instance configuration and license tracking to be directed to a version that hadn't received the latest configuration data. This conflict caused the system to calculate allotted licenses incorrectly and caused intermittent login issues.
**How did we respond?**
The Engineering teams were monitoring the update and were not encountering any warnings or errors within the process that they considered outside acceptable levels for this specific operation. When customers reported the issue to xMatters Customer Support, however, the teams made the decision to roll back the deployment immediately to mitigate any potential problems. As soon as the rollback was completed, customers confirmed that all services had been restored. The Engineering team launched an internal review process and were able to identify some avenues of improvement and successfully redeployed the update without incident.
**What are we doing to prevent it from happening again?**
In addition to adding additional automated checks to ensure configuration data is always up to date across services prior to an update, the teams isolated the specific cause of the configuration data mismatch to a timeout issue and have updated the timing settings to ensure that it will not happen again.