La 19 august 2026, am identificat o problemă care afectează transmiterea notificărilor Microsoft Teams pentru clienții independenți ai Opsgenie. Cauza principală a fost identificată și echipele noastre lucrează cu partenerii noștri pentru a rezolva această problemă.
Această problemă nu afectează clienții Jira Service Management Operations.
Vom furniza următoarea actualizare în 24 de ore.
identified
Echipele noastre continuă să lucreze cu partenerii noștri pentru a rezolva problema de livrare a notificării Microsoft Teams pentru clienții de sine stătătoare Opsgenie.
Vom furniza următoarea actualizare în 72 de ore sau mai devreme dacă avem o actualizare materială.
identified
Echipele noastre continuă să lucreze cu partenerii noștri, pentru a rezolva problema de livrare a notificării Microsoft Teams pentru clienții de sine stătătoare Opsgenie.
Vom furniza următoarea actualizare în 24 de ore, sau mai devreme dacă există o actualizare semnificativă.
identified
Echipele noastre continuă să lucreze cu partenerii noștri pentru a rezolva problema de livrare a notificării Microsoft Teams pentru clienții de sine stătătoare Opsgenie.
Vom furniza următoarea actualizare în 24 de ore, sau mai devreme dacă există o actualizare semnificativă.
resolved
Pe 19 august am identificat o problemă care afectează transmiterea notificărilor Microsoft Teams pentru clienții independenți ai Opsgenie. Cauza principală a fost atribuită unei erori operaționale cauzate de partenerul nostru extern și au luat măsuri corective pentru restabilirea serviciilor.
În cazul în care vă confruntați încă cu probleme, vă rugăm să urmați acești pași pentru a restabili integrarea dumneavoastră: https://support.atlassian.com/opsgenie/docs/restore-the-opsgenie-microsoft-teams-v2-integration/. Dacă aveţi nevoie de asistenţă suplimentară, vă rugăm să contactaţi Atlassian Support prin intermediul suport.atlassian.com/contact.
Traducere automată din actualizarea oficială a incidentului.
Eşecuri vocale pentru nişte numere bazate pe India
Lucrăm cu furnizorul nostru de notificare pentru a rezolva problema. Vom furniza o altă actualizare în 24 de ore sau de îndată ce mai multe informații devin disponibile.
identified
Un număr mic de notificări vocale pentru Jira Service Management și Opsgenie clienții care utilizează numere de mobil bazate pe India ar putea eșua. Clienții pot vedea jurnalele de apeluri vocale eșuate și nu primesc apelul de paging așteptat.
Furnizorul nostru de notificare a celei de-a treia părți lucrează în mod activ la restabilirea livrării notificării vocale și vom furniza actualizări după rezolvarea problemei. Dacă aveţi nevoie de orice sprijin sau informaţii suplimentare, vă rugăm să contactaţi echipa noastră de sprijin.
resolved
Furnizorul nostru de notificare a identificat rădăcina problemei care cauzează eșecuri de apel vocal la unele numere de mobil din India.
Înţelegem inconvenientele pe care acest lucru le poate cauza şi aprecia răbdarea dumneavoastră în timp ce furnizorul nostru lucrează la restabilirea serviciului. Dacă aţi avut probleme cu notificările vocale, vă rugăm să contactaţi Asistenţa noastră pentru asistenţă.
Traducere automată din actualizarea oficială a incidentului.
Atlassian Cloud Services impacted
A început 20 octombrie 2025 la 07:56 UTC · 21h 28m
We have noticed that Atlassian Cloud services are impacted and our teams are actively investigating the same.
We shall keep you informed of the progress every hour.
investigating
Atlassian Cloud services are impacted and we are aware that our customers might not be able to create support tickets. Our teams are actively investigating the same.
We shall keep you informed of the progress every hour.
investigating
We are experiencing an outage due to some issue at the end of our public cloud provider. We are working closely with them to get this resolved or mitigated as quickly as possible. ETA of the same is not know at the moment.
We shall continue to share updates every hour, if not sooner.
identified
We understand that our public cloud provider has identified the cause of the issue. We are starting to see some recovery and is working towards mitigation. We appreciate your patience.
We shall continue to share updates every hour, if not sooner.
identified
We continue to work with our public cloud provider towards mitigating the issue at the earliest. We appreciate your patience. We shall continue to share updates every hour, if not sooner.
identified
Atlassian team is actively engaged and continues to work with our public cloud provider to mitigate this issue at the earliest. We are starting to see partial operations succeed. We appreciate your patience. We shall continue to share updates every hour, if not sooner.
identified
Our public cloud provider is working to mitigate this issue quickly. We are seeing some early positive indicators and are continuing to monitor. We appreciate your patience and will continue to provide updates every hour or sooner.
identified
We understand your pain and mitigating or fixing this issue is of utmost importance. Our public cloud provider is actively working to mitigate this issue on priority. We have been seeing partial operational success. We appreciate your patience and will continue to provide updates every hour or sooner.
identified
Update - Thank you for your continued patience. We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority. Our public cloud provider is still actively working to mitigate this issue with urgency and while we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation. We will be providing hourly updates on this issue.
identified
Update - We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority. Our public cloud provider is actively working to mitigate this issue with urgency. While we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation. We will continue to provide updates every hour or sooner as new information becomes available.
identified
We are currently aware of an ongoing incident impacting Atlassian Cloud services due to an outage with our public cloud provider, AWS. We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority and we are closely monitoring the health of AWS services. While we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation. We will continue to provide updates every hour or sooner as new information becomes available.
monitoring
There have been no changes since our last update. We will provide our next updated by 9:00PM UTC or sooner as new information becomes available.
We are currently aware of an ongoing incident impacting Atlassian Cloud services due to an outage with our public cloud provider, AWS. We understand the impact this issue is having on your operations and want to assure you that resolving this matter is our highest priority and we are closely monitoring the health of AWS services. While we do not have a definitive ETA at this time, we remain committed to full resolution and deeply appreciate your patience as we work through this situation.
monitoring
Monitoring - We've started seeing continued product experience improvement.
While we still have a backlog of event processing, we are seeing improvements in systems operational capabilities across all products. We estimate a significant improvement with the next few hours and will continue to monitor the health of AWS services and the effects on Atlassian customers. We appreciate your continued patience and remain committed to full resolution as we work through this situation. We will post our next update in two hours.
monitoring
Our teams are continuing to monitor the recovery of systems across Atlassian products. This update is to inform that the Atlassian Support portal is fully operational at this time for customers that wish to contact support.
monitoring
Our team is now seeing recovery across all impacted Atlassian products. We are continuing to monitor for individual products that may still be processing backlogged items now that services are restored.
The Atlassian Support portal is currently still displaying a message directing customers to our temporary support channel. Please note that our support portal is currently fully functional for those attempting to raise requests. We are continuing to look into this alert to remove this message.
We will provide further update on our recovery status in one hour.
monitoring
We continue to see recovery progressing across all impacted products as backlogged items continue to be processed.
The Atlassian Support portal is currently displaying a message directing customers to our temporary support channel. Please note that our support portal is currently fully functional for those attempting to raise requests. We are continuing to look into this alert to remove this message.
We will provide further update on our recovery status in two hours.
monitoring
The issue relating to the Atlassian Support portal displaying a message to customers to use our temporary support channel has now been resolved. The Atlassian Support portal is fully functional for any ongoing support issues.
With regards to other Atlassian products, we continue to see recovery continuing across all impacted products and our teams are continuing to monitor as the recovery continues.
We will provide further update on our recovery status within two hours.
resolved
Our team is now able to see full recovery across the vast majority of Atlassian products.
We are aware of some ongoing issues with specific components such as migrations and JSM virtual service agents, and our team is continuing to investigate with urgency.
We apologise for the inconvenience that this incident has caused and we will provide further information when the Post Incident Investigation has been completed.
postmortem
### Postmortem publish date: Nov 19th, 2025
### Summary
All dates and times below are in UTC unless stated otherwise.
Customers utilizing Atlassian products experienced elevated error rates and degraded performance between Oct 20, 2025 06:48 and Oct 21, 2025 04:05. The service disruptions were triggered due to an [AWS DynamoDB outage](https://aws.amazon.com/message/101925/#:~:text=1%3A50%20PM.-,DynamoDB,-Between%2011%3A48) and further affected by subsequent failures in [AWS EC2](https://aws.amazon.com/message/101925/#:~:text=service%20disruption%20event.-,Amazon%20EC2,-Between%2011%3A48) and [AWS Network Load Balancer](https://aws.amazon.com/message/101925/#:~:text=service%20disruption%20event.-,Amazon%20EC2,-Between%2011%3A48) within the us-east-1 region.
The incident started at Oct 20, 2025 06:48 and was detected within six minutes by our automated monitoring systems. Our teams worked to restore all core services by Oct 21, 2025 04:05. Final cleanup of backlogged processes and minor issues was completed on Oct 22, 2025.
We recognize the critical role our products play in your daily operations, and we offer our sincere apologies for any impact this incident had on your teams. We are taking immediate steps to enhance the reliability and performance of our services, so that you continue to receive the standard of service you have come to trust.
### IMPACT
Before examining product-level impacts, it's helpful to understand Atlassian's service topology and internal dependencies.
Products such as Jira and Confluence are deployed across multiple AWS regions. The data for each tenant is stored and processed exclusively within its designated host region. This design is intentional and represents the desired operational state, as it limits the impact of any regional outage strictly to tenants in-region, in this case us-east-1.
While in-scope application data is pinned to the region selected by the customer, there are times when systems need to call other internal services that may be based in a different region. If a problem occurs in the main region where these services operate, systems are designed to automatically fail over to a backup region, usually within three minutes.
However, if unexpected issues arise during this failover, it can take longer to restore services. In rare cases, this could affect customers in more than one region. It’s important to note that all in-scope application data for supported products is pinned according to a customer’s chosen region.
**Jira**
Between Oct 20, 2025 06:48 and Oct 20, 2025 20:00, customers with tenants hosted in the us-east-1 region experienced increased error rates when accessing core entities such as Issues, Boards, and Backlogs. This disruption was caused by AWS's inability to allocate AWS EC2 instances and elevated errors in AWS Network Load Balancer \(NLB\). During this window, users may also have observed intermittent timeouts, slow page loads, and failures when performing operations like creating or updating issues, loading board views, and executing workflow transitions.
Between Oct 20, 2025 08:36 and Oct 20, 2025 09:23, customers across all regions experienced elevated failure rates when attempting to load Jira pages. This disruption was caused by the regional frontend service entering an unhealthy state during this specific time interval.
Normally, the frontend service connects to the primary AWS DynamoDB instance located in the us-east-1 to retrieve the most recent configuration data necessary for proper operation. Additionally, the service is designed with a fallback mechanism that references static configuration data in the event that the primary database becomes inaccessible. Unfortunately, a latent bug existed in the local fallback path. When the frontend service nodes restarted, they were unable to load critical operational configuration data from primary or fallback sources, leading to the observed failures experienced by customers.
Between Oct 20, 2025 06:48 and Oct 21, 2025 06:30, customers experienced significant delays and missing Jira in-app notifications across all regions. The notification ingestion service, which is hosted exclusively in us-east-1, exhibited an increased failure rate when processing notification messages due to AWS EC2 and NLB issues. This issue resulted in notifications being delayed - and in some cases, not delivered at all - to users worldwide.
**Jira Service Management \(JSM\)**
JSM was impacted similarly to Jira above, with the same timeframes and for the same reasons.
Between Oct 20, 2025 08:36 and Oct 20, 2025 09:23, customers across all regions experienced significantly elevated failure rates when attempting to load JSM pages. This affected all JSM experiences including the Help Centre, Portal, Queues, Work Items, Operations, and Alerts.
**Confluence**
Between Oct 20, 2025 06:48 and Oct 21, 2025 02:45, customers using Confluence in the us-east-1 region experienced elevated failure rates when performing common operations such as editing pages or adding comments. The primary cause of this service degradation was the system's inability to auto-scale due to AWS EC2 issues to manage peak traffic load effectively.
Though the AWS outage ended at Oct 20, 21:09, a subset of customers continued to experience failures as some Confluence web server nodes across multiple clusters remained in an unhealthy state. This was ultimately mitigated by recycling the affected nodes.
To protect our systems while AWS recovered, we made a deliberate decision to enable node termination protection. This action successfully preserved our server capacity but, as a trade-off, it extended the time required for a full recovery once AWS services were restored.
**Automation**
Between Oct 20, 2025 06:55 and Oct 20, 2025 23:59, automation customers whose rules are processed in us-east-1 experienced delays of up to 23 hours in rule execution.
During this window, some events triggering rule executions were processed out of order because they arrived later during backlog processing. This caused potential inconsistencies in workflow executions, as rules were run in the order events were received, not when the action causing the event occurred. Additionally, some rule actions failed because they depend on first-party and third-party systems, which were also affected by the AWS outage. Customers can see most of these failures in their audit logs; however, a few updates were not logged due to the nature of the outage.
By Oct 21, 2025 5:30, the backlog of rule runs in us-east-1 was cleared. Although most of these delayed rules were successfully handled, there were some additional replays of events to ensure completeness. Our investigation confirmed that a few events may never have triggered their associated rules due to the outage.
Between Oct 20, 2025 06:55 and Oct 20, 2025 11:20, all non-us-east-1 regional automation services experienced delays of up to 4 hours in rule execution. This was caused by an upstream service that was unable to deliver events as expected. The delivery service encountered a failure due to a cross-region dependency call to a service hosted in the us-east-1 region. Because of this dependency issue, the delivery service was unable to successfully deliver events throughout this time frame, resulting in customer-defined rules not being executed in a timely manner.
**Bitbucket and Pipelines**
Between Oct 20, 2025 06:48 and Oct 20, 2025 09:33, Bitbucket experienced intermittent unavailability across core services. During this period, users faced increased error rates and latency when signing in, navigating repositories, and performing essential actions such as creating, updating, or approving pull requests. The primary cause was an AWS DynamoDB outage that impacted downstream services.
Between Oct 20, 2025 06:48 and Oct 20, 2025 22:46, numerous Bitbucket Pipeline steps failed to start, stalled mid-execution, or experienced significant queueing delays. Impact varied, with partial recoveries followed by degradation as downstream components re-synchronized. The primary cause was an AWS DynamoDB outage, compounded by instability in AWS EC2 instance availability and AWS Network Load Balancers.
Furthermore, Bitbucket Pipelines continued to experience a low but persistent rate of step timeouts and scheduling errors due to AWS bare-metal capacity shortages in select availability zones. Atlassian coordinated with AWS to provision additional bare-metal hosts and addressed a significant backlog of pending pods, successfully restoring services by 01:30 on Oct 21, 2025.
**Trello**
Between Oct 20, 2025 06:48 and Oct 20, 2025 15:25, users of Trello experienced widespread service degradation and intermittent failures due to upstream AWS issues affecting multiple components, including AWS DynamoDB and subsequent AWS EC2 capacity constraints. During this period, customers reported elevated error rates when loading boards, opening cards, adding comments or attachments.
**Login**
Between Oct 20, 2025 06:48 and Oct 20, 2025 09:30, a small subset of users experienced failures when attempting to initiate new login sessions using SAML tokens. This resulted in an inability for those users to access Atlassian products during that time period. However, users who already had valid active sessions were not affected by this issue and continued to have uninterrupted access.
The issue impacted all regions globally because regional identity services relied on a write replica located in the us-east-1 region to synchronize profile data. When the primary region became unavailable, the failover to a secondary database in another region failed, which delayed recovery. This failover defect has since been addressed.
**Statuspage**
Between Oct 20, 2025 06:48 and Oct 20, 2025 09:30, Statuspage customers who were not already logged in to the management portal were unable to log in to create or update incident statuses. This impact was restricted only to users who were not already logged in at the time. The root cause was the same as described in the Login section above, and it was resolved by the same remediation steps.
### REMEDIAL ACTION PLAN & NEXT STEPS
We have completed the following critical actions designed to help prevent cross-region impact from similar issues:
* Resolved the code defect in the fallback option to ensure that Jira Frontend Services in other regions remain unaffected during a region-wide outage.
* Fixed the issue that prevented timely failover of the identity service which impacted new login sessions.
* Resolved the code defect so that delivery services in unaffected regions remain operational during region-wide outages.
Additionally, we are prioritizing the following improvement actions:
* Implement mitigation strategies to strengthen resilience against region-wide outages in the notification ingestion service.
Although disruptions to our cloud services are sometimes unavoidable during outages of the underlying cloud provider, we continuously evaluate and improve test coverage to strengthen resilience of our cloud services against these issues.
We recognize the critical importance of our products to your daily operations and overall productivity, and we extend our sincere apologies for any disruptions this incident may have caused your teams. If you were impacted and require additional details for internal post-incident reviews, please reach out to your Atlassian support representative with affected timeframes and tenant identifiers so we can correlate logs and provide guidance.
Thanks,
Atlassian Customer Support
Delays in Opsgenie, JSM Ops and Compass Ops notifications
We are observing delays in Opsgenie, JSM ops and Compass ops notification flows. No alert has been lost and our team is actively working on it to mitigate the delays. We'll keep you posted with further updates.
identified
We continue to work on resolving the notification delays for Jira Service Management, Opsgenie, and Compass. We have identified the root cause and expect recovery shortly.
monitoring
We have identified the root cause of the notification delay for Opsgenie, JSM Ops and Compass Ops and have mitigated the problem. We are now monitoring closely.
resolved
Between 13:10 UTC to 14:00 UTC, we experienced notification delay for Jira Service Management, Opsgenie, and Compass in only EU-region. The issue has been resolved and the service is operating normally.
Delays observed at JSM, Opsgenie and Compass alert search functionality in US Region
We identified degraded performance in web experiences and Alert Rest API for some Jira Service Management and Opsgenie Cloud customers in the US Region. The team has taken actions to mitigate the issue and minimize the impact on search functionality.
identified
We continue to work on resolving the incident for Jira Service Management and Opsgenie. We have identified the root cause and taken actions to mitigate the issue and minimize the impact on search functionality.
identified
We are continuing to work on a fix for this issue.
identified
Our team is still actively working to resolve the issue and making progress toward a resolution. Thank you for your patience.
identified
Our team has been able to identify the root cause of these performance issues and have put a mitigation into place.
We are now continuing to see recovery of Jira Service Management operations and Opsgenie services.
Existing alerts are continuing to be processed at this time.
We will provide further update as soon as possible.
identified
We are continuing to see recovery of Jira Service Management operations and Opsgenie services.
There are a large number of missed alerts to be processed due to the prior service issue and our team is actively investigating methods to help expedite this processing.
Alerts should trigger in the meantime but the content of these alerts may not be visible at this time.
Our teams are continuing to investigate with urgency.
We will provide further updates within the hour.
identified
While new alerts should now correctly notify users, the content of those alerts for impacted customers will remain empty until missed alerts are fully processed.
Our team is continuing to investigate with urgency potential infrastructure options to expedite this process.
We will provide further updates within the hour.
identified
New alerts should still continue to notify users as they are generated, however the alert content will still be missing.
Our teams are still actively engaging with infrastructure teams to try and expedite historical processing of messages to fully restore services.
We will provide further updates within the hour.
identified
We have received reports that some customers are experiencing further issues relating to loading pages within Opsgenie and Jira Service Management which may be related to the infrastructure efforts underway to restore services to these products.
Please be aware we also have teams actively investigating these issues to ensure as fast a resolution as possible.
We want to reiterate that your alerts are still sending notifications correctly, and there is no data loss occurring. However, the underlying issue is affecting the propagation of alert details to notifications, and to the Web and Mobile UI.
Viewing schedules may also be impacted at this time while we continue to process the tasks required for recovery.
We will provide further updates as soon as possible.
identified
Our team is continuing to investigate with the highest level of urgency in order to restore Jira Service Management and Opsgenie services.
At this time we are prioritising infrastructure to try and restore new and active alerts to be populated with their alert content as soon as possible, while continuing to deliver existing messages as they are ready.
We will continue to provide updates as we progress, and will ensure we have an update posted within the hour.
identified
Our infrastructure teams are working diligently to try and ensure all services are restored to Jira Service Management and Opsgenie as soon as possible.
New alerts are continuing to notify users, however we are still prioritising infrastructure to restore the content inside these alerts for Web and Mobile UIs.
We will continue providing updates when available and will ensure we have further update within the hour.
identified
Our teams priority remains to restore full functionality to new alerts to users from Jira Service Management and Opsgenie.
Escalation teams are fully engaged on this incident to try and ensure we can get to a resolution as soon as possible.
At this time these alerts are still successfully notifying users, but without the intended content within the alerts.
We will continue providing specific updates when available and will ensure we have further update within the hour.
identified
All concerned teams are engaged and are working to restore full functionality to new alerts to users from Jira Service Management and Opsgenie.
Escalation teams are also engaged on this incident to try and ensure we can get to a resolution as soon as possible.
Currently, the alerts continue to successfully notify users, but without the intended content within the alerts.
We will continue providing specific updates when available and will ensure we have further update within the hour.
identified
Atlassian teams remain engaged with priority to restore full functionality to new alerts to users from Jira Service Management, Opsgenie and Compass.
Escalation teams are also engaged on this incident to ensure we can get to a resolution as soon as possible.
At this point, the alerts continue to successfully notify users, except that the responder name may be missing.
We will continue providing specific updates when available and will ensure we have further update within the hour.
identified
Atlassian teams continue to remain engaged with priority to restore full functionality to new alerts to users from Jira Service Management, Opsgenie and Compass.
Escalation teams also continue to remain engaged on this incident to ensure our focus is not diluted from a resolution as soon as possible.
Currently, the alerts continue to successfully notify users, except that the responder name may be missing.
We will continue providing specific updates when available and will ensure we have further update within the hour.
identified
Our teams have are identified some potential resolution steps for this issue and continue to test the same on priority to be able to restore full functionality to users from Jira Service Management, Opsgenie and Compass.
Currently, the alerts continue to successfully notify users, except that the responder name may be missing.
We will continue providing specific updates when available and will ensure we have further update within the hour.
identified
Atlassian teams continue to test the identified potential resolution steps for this issue. Our goal is to restore full functionality to users from Jira Service Management, Opsgenie and Compass.
Currently, the alerts continue to successfully notify users, except that the responder name may be missing.
We will continue providing specific updates when available and will ensure we have further update within the hour.
identified
We are progressing well with the testing of identified potential mitigation steps for this issue; and we are seeing positive early results. We will update with an ETA and further information in the next update.
Alerts continue to successfully notify users, in some cases the responder name may be missing. Our goal continues to restore full functionality to users from Jira Service Management, Opsgenie and Compass.
We will continue providing specific updates when available and will ensure we have further update within the hour.
identified
We have progressed further in preparing our system with a probable fix. Initial testing has been successful. We are in the process of completing further testing and finalising the fix. We expect the fix to be completed in next few hours.
Alerts continue to successfully notify users, in some cases the responder name may be missing. Our goal continues to restore full functionality to users from Jira Service Management, Opsgenie and Compass.
We will continue providing specific updates when available and will ensure we have further update within the hour.
monitoring
We have deployed the fix and all operations are back to normal. We will continue to monitor the operations.
We will continue providing specific updates when available.
resolved
Between September 10, 2025, 3:19 PM UTC and September 11, 2025, 2:44 PM UTC, there was degraded performance in web experiences and the REST APIs for some Jira Service Management, Opsgenie, and Compass Cloud customers in the US region. We have deployed a fix to mitigate the issue and have verified that the services have recovered. The issue has been resolved and the service is operating normally.
postmortem
### **SUMMARY**
Between September 10, 2025, at 13:56 UTC and September 11, 2025, at 14:40 UTC, Jira Service Management, Compass Cloud Operations, and Opsgenie customers in the U.S. region experienced degraded performance across web and mobile applications, as well as Alert REST APIs. Certain customers experienced difficulties with alert searches and page loading, particularly on alert-related pages. This incident was initiated by an EBS volume upgrade conducted within the cloud-based managed ElasticSearch clusters.
Our automated monitoring systems identified the incident within minutes, and it was resolved after 24 hours and 43 minutes by manually implementing vertical and horizontal scaling measures.
### **IMPACT**
Between September 10, 2025, 13:56 UTC and September 11, 2025, 14:40 UTC, some customers in the U.S. region experienced degraded performance in Jira Service Management, Opsgenie, and Compass Cloud. During this time, web and mobile applications, as well as the Alert REST APIs, were impacted. Some customers may have seen issues with alert searches and slow loading of alert pages. During the incident, 22% of customers in the US region experienced failures in web API functionality, 8.6% encountered failures with the REST Alert API, and 1.3% experienced delays in notification delivery. Importantly, there was no loss of data during this event.
### **ROOT CAUSE**
The degradation was caused by an EBS volume upgrade in the cloud-based managed ElasticSearch clusters, which necessitated a Blue/Green deployment strategy. One of the ElasticSearch nodes approached its shard size threshold, prompting the upgrade and subsequent deployment. This deployment resulted in elevated latency, increased 4xx HTTP response codes, and timeouts affecting both search and indexing operations. Recovery time exceeded expectations due to the prolonged blue/green deployment. After completion, the ElasticSearch cluster remained unhealthy and did not return to its normal state. Additionally, the switchover to the backup region failed because it was configured similarly in size and setup to the primary cluster.
### **REMEDIAL ACTIONS PLAN & NEXT STEPS**
We understand that outages can affect your productivity. We are prioritizing the following improvement actions to help prevent similar incidents in the future:
* Increase ElasticSearch cluster capacity and optimize settings to handle peak search loads efficiently.
* Strengthen our high availability \(HA\) and failover architecture to ensure rapid and reliable recovery in the event of primary region failures.
* Optimize cluster upgrade and shard rebalancing strategies to prevent similar issues.
We apologize to customers whose services were impacted during this incident. We are taking steps to improve the platform’s performance and availability.
Thanks,
Atlassian Customer Support
Degraded performance in Opsgenie and JSM operations
A început 19 iunie 2025 la 15:21 UTC · 1h 16m
Pending
investigating
We are investigating degraded performance issues in Jira Service Management, Opsgenie, and Compass Cloud customers. We will provide more details within the next hour.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
Degraded performance in Opsgenie and JSM operations
A început 17 iunie 2025 la 12:37 UTC · 10m
Pending
investigating
We are investigating cases of degraded performance for some Jira Service Management and Opsgenie Cloud customers. We will provide more details within the next hour.
monitoring
We have identified the root cause of the performance degraded service and have mitigated the problem. We are now monitoring closely.
resolved
This incident has been resolved, and Jira Service Management and Opsgenie are back in operational mode.
Delays observed at JSM and Opsgenie alert search functionality
A început 10 mai 2025 la 07:08 UTC · 3h 43m
IssuesIncident minor
Componente afectate
Web Application
identified
We identified degraded performance at alert search functionality for some Jira Service Management and Opsgenie Cloud customers due to the infrastructure issue from cloud provider. No impact has been observed at alert critical flows like notification. The team has taken actions to mitigate the issue and minimize the impact to search functionality
monitoring
The problem is mitigated, and we are now monitoring closely.
resolved
All search functionality is operational without any latency. Thank you for your patience.
Schedule API are getting timed out
A început 4 aprilie 2025 la 10:05 UTC · 47m
Pending
investigating
We are investigating cases of degraded performance for Alert Schedules experiencing timeouts and slowness for Opsgenie Cloud customers.
Requests have been taking more than 30s and some have been timing out.
We will provide more details within the next hour.
resolved
This incident has been resolved.
EU OpsGenie API calls having intermittent routing issues
A început 5 martie 2025 la 06:34 UTC · 50m
OutageIncident major
Componente afectate
Alert Flow
investigating
We are aware of an issue where some API calls configured to use api.opsgenie.com/eu are intermittently returning a 404 error with the message 'no Route matched with those values'.
API calls to api.eu.opsgenie.com and api.opsgenie.com (without /eu) are not affected.
Our team is investigating this issue.
resolved
Issues where some API calls configured to use api.opsgenie.com/eu were intermittently returning a 404 error with the message 'no Route matched with those values' should now be resolved.
API calls to api.eu.opsgenie.com and api.opsgenie.com (without /eu) were not affected at this time.
Services menu in Opsgenie is not responding
A început 14 ianuarie 2025 la 15:48 UTC · 59m
Pending
investigating
Our engineering team is actively investigating this incident and working to bring the Opsgenie service back up as quickly as possible.
Users affected by this incident may notice that Services functionality is slow or completely unavailable for the web page
We will update this page as we have additional information.
identified
Our team has identified the issue with the Services page in Opsgenie and is working to fix it.
resolved
Our engineers have been closely monitoring the platform and are declaring this incident resolved. Thank you for your patience.
Elevated 5XX errors in Schedule API at Opsgenie USA region
A început 16 decembrie 2024 la 15:30 UTC · 0m
IssuesIncident minor
resolved
Our team has identified the issue in Schedule API between 15:30 UTC and 17:00 UTC. We saw performance degradation and 5XX errors in response. Faulty deployment has been reverted quickly in the USA region and rapid recovery is seen. We are monitoring the system for a full recovery right now. The Schedule API is up and running again without any data loss.
Opsgenie Web UI is slow or unavailable in US region
A început 6 noiembrie 2024 la 19:47 UTC · 12m
IssuesIncident minor
Componente afectate
Web Application
investigating
We've noticed that Opsgenie Web UI is responding slowly or unavailable in US region.
Our engineering team is actively investigating this incident and working to bring Opsgenie back up to speed as quickly as possible.
We'll keep you posted with further updates on this page.
monitoring
Our engineering team has implemented fixes. We will continue to monitor all systems. Thank you for your patience.
resolved
Our engineers have been closely monitoring the platform and are declaring this incident resolved. Thank you for your patience.
We've been notified that a number of OEC clients have been failing to create Jira tickets. We're currently investigating the issue.
investigating
We are continuing to investigate this issue.
identified
We've identified a recent change that has broken some endpoint routing configurations and caused OEC endpoint requests to be directed to wrong service. We're reverting that change on production at the moment.
monitoring
We reverted the faulty routing configuration change and started getting traffic on OEC endpoints
resolved
We've verified that OEC endpoints are back online.
Users are experiencing reCaptcha errors while signing up
A început 12 septembrie 2024 la 02:08 UTC · 1h 21m
OutageIncident critic
Componente afectate
Signup, Login & Authorization
investigating
Users attempting to sign up are encountering reCaptcha errors that are preventing a successful signup.
monitoring
We have identified the root cause and the issue appears to be resolved.
resolved
This issue has been resolved.
Unable to edit Opsgenie rotations in US, EU and Sydney regions in Web Application
A început 28 august 2024 la 15:01 UTC · 16m
IssuesIncident minor
Componente afectate
Web ApplicationWeb Application
identified
Our team has identified the issue with Opsgenie Web Application / Edit Rotation feature in US, EU and Sydney regions and is working to fix it. Check back soon for another update! Our team is working hard to get the feature up and running again.
identified
We are continuing to work on a fix for this issue.
resolved
Our engineers have been closely monitoring the platform and are declaring this incident resolved. Thank you for your patience.
We are investigating an issue with that is impacting Atlassian, Atlassian Partners, Atlassian Support, Confluence, Jira Work Management, Jira Service Management, Jira, Opsgenie, Atlassian Developer, Atlassian (deprecated), Trello, Atlassian Bitbucket, Guard, Jira Align, Jira Product Discovery, Atlas, Atlassian Analytics, and Rovo Cloud customers. We will provide more details within the next hour.
monitoring
We have mitigated the problem and continue looking into the root cause.
The outage was between 8:08pm 03/07 UTC - 08:31pm 03/07 UTC
We are now monitoring closely.
resolved
Between 03-07-2024 20:08 UTC to 03-07-2024 20:31 UTC, we experienced downtime for Opsgenie. The issue has been resolved and the service is operating normally.
Intermittent error accessing content
A început 21 iunie 2024 la 00:17 UTC · 2h 30m
Pending
investigating
We are investigating an intermittent issue with accessing Atlassian Cloud services that is impacting some Atlassian Cloud customers. We will provide more details once we identify the root cause.
monitoring
We have identified the root cause of the intermittent errors and have mitigated the problem. We are now monitoring closely.
resolved
Between 2024-06-20 22:04 UTC to 2024-06-20 22:28 UTC, we experienced intermittent issue for users to access the services for some Atlassian Cloud customers. The issue has been resolved and the service is operating normally.
We are investigating an issue with error responses for some Cloud customers across multiple products. We have identified the root cause and expect recovery shortly.
resolved
Between 22:18 UTC to 22:56 UTC, we experienced errors for multiple Cloud products. The issue has been resolved and the service is operating normally.
postmortem
### Summary
On June 3rd, between 09:43pm and 10:58 pm UTC, Atlassian customers using multiple product\(s\) were unable to access their services. The event was triggered by a change to the infrastructure API Gateway, which is responsible for routing the traffic to the correct application backends.
The incident was detected by the automated monitoring system within five minutes and mitigated by correcting a faulty release feature flag, which put Atlassian systems into a known good state. The first communications were published on the Statuspage at 11:11pm UTC. The total time to resolution was about 75 minutes.
### **IMPACT**
The overall impact was between 09:43pm and 10:17pm UTC, with the system initially in a degraded state, followed by a total outage between 10:17pm and 10:58pm UTC.
_The Incident caused service disruption to customers in all regions and affected the following products:_
* Jira Software
* Jira Service Management
* Jira Work Management
* Jira Product Discovery
* Jira Align
* Confluence
* Trello
* Bitbucket
* Opsgenie
* Compass
### **ROOT CAUSE**
A policy used in the infrastructure API gateway was being updated in production via a feature flag. The combination of an erroneous value entered in a feature flag, and a bug in the code resulted in the API Gateway not processing any traffic.
This created a total outage, where all users started receiving 5XX errors for most Atlassian products.
Once the problem was identified and the feature flag updated to the correct values, all services started seeing recovery immediately.
### **REMEDIAL ACTIONS PLAN & NEXT STEPS**
We know that outages impact your productivity. While we have several testing and preventative processes in place, this specific issue wasn’t identified because the change did not go through our regular release process and instead was incorrectly applied through a feature flag.
We are prioritizing the following improvement actions to avoid repeating this type of incident:
* Prevent high-risk feature flags from being used in production
* Improve the policy changes testing
* Enforcing longer soak time for policy changes
* Any feature flags should go through progressive rollouts to minimize broad impact
* Review the infrastructure feature flags to ensure they all have appropriate defaults
* Improve our processes and internal tooling to provide faster communications to our customers
We apologize to customers whose services were affected by this incident and are taking immediate steps to address the above gaps.
Thanks,
Atlassian Customer Support
We are observing delays for our alert flow. No alert has been lost and our team is actively working on it to mitigate the delays.
We'll keep you posted with further updates.
identified
We are continuing to work on a fix for this issue.
monitoring
Our engineering team has implemented fixes. We will continue to monitor all systems. Thank you for your patience.