Flexera One – IT Asset Management – EU – Reconciliation Failures
开始时间 2026年8月13日 UTC 01:08 · 4h 49m
Outage重大事件
受影响的组件
IT Asset Management - EU Batch Processing System
investigating
Incident Description: We are currently investigating an issue impacting the reconciliation process within the IT Asset Management platform in the EU region. Affected customers may experience failures when attempting to run reconciliation jobs. As a result, reconciliation processes may not complete successfully, potentially delaying data updates and downstream processing.
Priority: P2
Restoration Activity: Our technical teams have been engaged and are actively working to identify the root cause and restore services. Further updates will be provided as we continue working toward resolution.
monitoring
Our teams identified a configuration issue introduced during the completion of a standard release earlier today, which prevented an instance from successfully communicating with the Reconcile service.
The affected instance's configuration process was successfully re-executed, restoring the required settings and re-establishing connectivity with the Reconcile service. Validation testing confirmed normal service operation, and reconciliation processing has resumed successfully. The environment remains under monitoring to ensure continued stability.
monitoring
We are continuing to monitor for any further issues.
resolved
Following extended monitoring, our teams have confirmed that the platform has remained stable and no recurrence of the issue has been observed since the remediation actions were completed. Reconciliation processing continues to operate normally, connectivity with the Reconcile service remains healthy, and all validation checks have been successful.
Based on the sustained stability of the environment and the absence of any further failures, this incident has been resolved
Flexera One - Cloud Cost Optimization (CCO) - NAM - Service Degradation
开始时间 2026年8月11日 UTC 17:26 · 3h 0m
Outage重大事件
受影响的组件
Cloud Cost Optimization - US
investigating
Incident Description: We are investigating an issue affecting Cloud Cost Optimization (CCO) services in the North America (NAM) region. Customers may experience pages failing to load, prolonged loading behavior, or difficulty accessing certain Cloud Cost Optimization functionality within Flexera One.
Priority: P2
Restoration Activity: Our technical teams are actively investigating the issue and working to restore normal service. Current efforts are focused on identifying the cause of the degradation and validating corrective actions. We will provide additional updates as more information becomes available.
identified
Our investigation has identified the issue affecting certain Cloud Cost Optimization (CCO) functionality in the North America region. Following a recent service migration, some application components were not operating with the expected resource configuration, resulting in certain requests failing to complete and impacted pages remaining in a continuous loading state.
Our technical teams have implemented configuration changes to restore the appropriate resource allocation for the affected services and are currently deploying those changes. We are monitoring recovery and validating service functionality as the update takes effect.
We will provide additional updates as progress continues.
monitoring
The configuration changes have been deployed, and the affected service components are now operating with the intended resource configuration. Our technical teams have observed recovery of the impacted Cloud Cost Optimization (CCO) functionality in the North America region and are performing final validation to confirm service stability.
We will continue to monitor the environment closely and provide a further update once validation activities are complete.
resolved
Following a period of monitoring, services have remained stable and no further issues have been observed.
This incident is now considered resolved.
postmortem
**Description:** Flexera One - Cloud Cost Optimization \(CCO\) - North America - Service Degradation Affecting Certain Functionality
**Timeframe:** August 11, 2026, 9:51 AM PDT – August 11, 2026, 10:45 AM PDT
**Incident Summary**
On August 11, 2026, customers may have experienced issues accessing certain Cloud Cost Optimization \(CCO\) functionality in the North America region. Impacted pages could remain in a continuous loading state and fail to render correctly.
Technical teams began investigating after identifying that certain application requests were not completing successfully. The impact was limited to specific functionality rather than the entire CCO service.
During the investigation, technical teams determined that dependent application services were unable to successfully retrieve required configuration information from a supporting service component, resulting in rendering issues for affected functionality.
Corrective configuration changes were implemented and updated service components were deployed. Following deployment, technical teams validated recovery and continued monitoring the environment to confirm stable operation. All affected functionality subsequently returned to normal operation.
**Root Cause**
During a recent service migration, a supporting application component was deployed with an incorrect resource configuration. As processing demand increased, the affected service components exhausted available resources and became unavailable.
As a result, dependent application services were unable to successfully retrieve required configuration information needed to process certain requests. This resulted in rendering issues affecting specific Cloud Cost Optimization \(CCO\) functionality.
**Remediation Actions**
The following actions were taken during the incident response:
1. Incident Investigation Initiated: Technical teams investigated reports of affected CCO functionality remaining in a continuous loading state and failing to render correctly.
2. Service Dependency Analysis Performed: Technical teams identified that dependent application services were unable to successfully retrieve required configuration information from a supporting service component.
3. Configuration Issue Identified: Investigation determined that a supporting application component had been deployed with an incorrect resource configuration following a recent service migration.
4. Corrective Configuration Changes Applied: Technical teams implemented configuration changes to restore the intended operating parameters for the affected service components.
5. Service Recovery Validation Performed: Updated service components were deployed, affected functionality was validated, and the environment was monitored to confirm stable operation before incident closure.
**Future Preventative Measures**
Following this incident, technical teams reviewed the migration and deployment process associated with the affected service component.
1. Configuration Validation Enhancements: An additional validation step has been added to verify that service configurations are appropriate for forecasted processing requirements following migration activities. This additional review is intended to help identify configuration discrepancies before they can affect production workloads.
Our teams have identified the cause of the issue and are currently working to resolve it.
identified
Our teams continue to work toward a resolution. During the ongoing investigation, we identified that in some instances emails from Flexera may be being blocked by SharePoint. Addressing this issue requires additional investigation and coordination with the relevant parties to ensure the appropriate corrective actions are taken. We will continue to provide updates as more information becomes available.
identified
Our teams continue working toward a resolution. Additional investigation and coordination with the appropriate teams are underway to restore normal email flow. We will provide further updates as more information becomes available.
monitoring
Our teams have implemented corrective measures and are observing positive results. Initial validation testing has been successful, indicating that the mitigation actions are effective. We will continue to closely monitor the service to ensure ongoing stability and will provide further updates as they become available.
resolved
After extended monitoring, our teams have confirmed that all email updates are being delivered successfully. This incident has been declared as resolved.
自动翻译自官方事件更新。
Flexera 1 - 信息技术 资产管理 - APAC - 数据加载失败/慢化
开始时间 2026年8月5日 UTC 08:28 · 0m
Outage重大事件
受影响的组件
IT Asset Management - APAC Inventory UploadIT Asset Management - APAC Batch Processing System
**Description:** Software Vulnerability Research \(SVR\) - Service Disruption
**Timeframe:** July 28, 2026, 2:00 PM PDT to July 29, 2026, 2:53 AM PDT
**Incident Summary**
On Tuesday, 28 July 2026, at 2:00 PM PDT, the Flexera Software Vulnerability Research\(SVR\) production environment experienced a service disruption that affected API availability and application functionality. The incident occurred when the primary database instance supporting the SVR platform became unavailable. Existing database connections were unexpectedly terminated, and application servers were unable to establish new connections. As a result, customers were unable to reliably access SVR services and APIs during the incident.
During the investigation, our teams determined that the disruption was caused by an unplanned failover initiated by the cloud service provider. According to the provider’s event history, an infrastructure issue was detected on the primary database host, prompting the provider to automatically initiate the failover process.
Our technical teams immediately engaged the cloud service provider to investigate the incident and restore service. Following recovery and validation activities, all production services were successfully restored and returned to normal operation by 2:53 AM PDT on 29 July 2026.
**Root Cause**
The incident was triggered by an unplanned failover of the production database, initiated by the cloud service provider in response to an infrastructure-level issue affecting the primary database host.
During the failover, the primary database instance became temporarily unavailable. This caused active database connections to terminate abruptly, preventing application services from establishing new connections. As a result, API requests failed and service availability across the SVR platform was impacted.
Flexera has formally requested a detailed Root Cause Analysis \(RCA\) from the cloud service provider to determine the precise infrastructure condition that triggered the failover and to identify opportunities to prevent recurrence.
**Contributing Factors**
During the investigation, the following factors were identified as potentially contributing to the overall impact of the incident:
* A significantly elevated volume of client connection attempts was observed originating from customer environments during the incident period.
* Connection volumes exceeded normal operating levels, increasing load on backend services while database services were recovering.
* Elevated connection activity may have amplified the impact of the database failover and increased recovery complexity.
* Existing application retry and connection behaviours generated additional connection demand during database recovery activities.
**Remediation Actions**
* Cloud provider engagement: Our teams immediately engaged the cloud service provider and worked directly with their on-call engineers throughout the recovery effort.
* Database recovery: Executed a controlled database failover/reboot in coordination with the cloud service provider.
* Restored connectivity: Re-established database availability and connectivity to the production environment.
* Application recovery validation: Confirmed that application services successfully reconnected to backend database services.
* Service validation: Verified recovery of APIs and critical SVR application functionality.
* Post-recovery monitoring: Implemented enhanced monitoring of database health, application performance, and connection activity following restoration.
**Future Preventative Measures**
* Continue monitoring customer connection volumes and backend database connection utilisation.
* Review and optimise database connection pooling, connection management, retry logic, and failover handling to improve resilience during future infrastructure events.
* Complete the cloud service provider support engagement and review the provider's detailed Root Cause Analysis once available.
* Conduct an internal post-incident review and implement additional preventive measures identified through the review process.
**Description:** Flexera One - APAC - Intermittent Service Disruption
## Timeframe
**Timeframe:** July 28, 2026, 3:16 PM PDT - July 28, 2026, 5:39 PM PDT
## Incident Summary
On Tuesday, July 28, 2026, at 3:49 PM PDT, customers using Flexera One services in the APAC region began experiencing intermittent service disruptions affecting multiple applications and platform capabilities. Customers may have encountered login issues, application errors, failed page loads, API errors, and intermittent access to certain Flexera One functionality.
During the incident, multiple services experienced intermittent communication failures with shared platform services. As a result, customers experienced inconsistent application behavior, with some requests succeeding while others failed.
Technical teams investigated the issue across affected applications and platform components to determine the scope and source of the failures. The investigation identified a connectivity issue affecting communication between application services and a shared platform dependency. Corrective configuration changes were implemented to restore service communication and stabilize affected functionality.
Following implementation of the corrective changes, technical teams validated functionality across impacted applications and confirmed that services had returned to normal operation. Continued monitoring showed stable service behavior, and the incident was resolved at 5:40 PM PDT on July 28, 2026.
## Root Cause
Investigation determined that the incident was caused by a connectivity issue introduced during a platform infrastructure migration.
Following the migration, certain application services continued using legacy connection ports when communicating with shared platform services. When affected services restarted, they attempted to communicate using endpoint and port combinations that were no longer aligned with the updated platform configuration.
This configuration mismatch resulted in intermittent communication failures between application services and shared platform components, causing customer-facing application errors and service disruptions across multiple Flexera One services in the APAC region.
## Remediation Actions
The following actions were taken during the incident response:
1. **Incident Investigation Initiated:** Technical teams investigated reports of intermittent failures affecting multiple Flexera One services in the APAC region and assessed the scope of customer impact.
2. **Dependency Analysis Performed:** Technical teams reviewed communication paths between impacted applications and shared platform services to identify the source of the connectivity failures.
3. **Configuration Mismatch Identified:** Investigation determined that certain services were attempting to communicate through legacy ports that were not aligned with the updated platform configuration following the migration.
4. **Platform Configuration Updated:** The required ports were added to the updated platform configuration, restoring connectivity for the affected services.
5. **Service Recovery Validation:** Technical teams validated functionality across impacted applications and confirmed that service communication and customer-facing functionality had returned to normal operation.
## Future Preventative Measures
The following improvements have been identified to further reduce the risk of similar incidents in the future:
1. **Platform Configuration Alignment:** Platform connectivity configurations were updated to ensure application services can communicate with shared platform components using the required connection ports. Maintaining alignment between service connectivity requirements and platform configurations helps reduce the risk of similar connectivity issues following future infrastructure changes.
**Description:** Snow Atlas - APAC - Login Failures and 404 Errors
**Timeframe:** July 26, 2026, 5:00 PM PDT to July 26, 2026, 6:15 PM PDT
**Incident Summary**
On Sunday, 26 July 2026, Flexera experienced a service disruption affecting Snow Atlas customers in the APAC Production environment, resulting in login failures and temporary inability to access Agreement pages. The issue was isolated to the APAc region, with no impact to other production regions.
Investigation by our technical teams determined that the disruption was caused by a backend service communication issue that prevented Agreement data from being retrieved successfully. The condition was consistent with a runtime synchronization issue following recent platform infrastructure maintenance.
Flexera teams validated platform health, restored the affected services, and confirmed successful recovery. Customer access was fully restored, and post-recovery monitoring confirmed stable operations with no recurrence of the underlying service communication errors.
A detailed post-incident review was completed, and corrective measures have been implemented to further strengthen platform resilience and recovery processes.
**Root Cause**
The incident was caused by an application communication failure within the APAC Production environment. A component responsible for processing requests between internal application services did not successfully handle Agreement-related requests, preventing those requests from being completed and resulting in customer-facing 404 errors and failures accessing Agreement pages.
**Contributing Factors**
* The application communication failure occurred following a recent infrastructure upgrade to the APAC Production environment.
* Although the application services and underlying infrastructure remained healthy, an internal service responsible for processing Agreement requests did not fully recover after the upgrade.
* Existing health checks validated application and infrastructure availability but did not verify the successful initialization of the internal communication path following the upgrade.
* The issue was therefore not detected until customers began experiencing login failures and HTTP 404 errors when accessing Agreement pages.
**Remediation Actions**
To restore service, technical teams:
* Confirmed the underlying infrastructure remained healthy throughout the incident.
* Identified the failed internal service communication affecting Agreement requests.
* Restarted the affected application services.
* Validated successful recovery across impacted tenants.
* Continued monitoring to confirm service stability.
**Future Preventative Measures**
* Enhanced Post-Upgrade Validation: Introduce additional validation checks following infrastructure upgrades to verify that critical application components and internal service communications are functioning as expected.
* Improved Monitoring and Alerting: Enhance monitoring and alerting to detect internal service communication failures and responder registration issues more quickly.
* Synthetic Health Checks: Implement synthetic health checks that validate critical customer workflows, including Agreement page functionality, following infrastructure changes.
**Description:** Flexera One – IT Asset Management – North America – Degraded Performance
**Timeframe:** July 20, 2026, 8:13 PM PDT – July 22, 2026, 8:41 AM PDT
**Incident Summary**
On July 20, 2026, at approximately 8:13 PM PDT, an issue began affecting Flexera One IT Asset Management customers in the North America production environment following completion of the production release.
During the affected period, some customers experienced intermittent application slowness, delayed processing, reconcile failures, timeout conditions, and degraded performance across inventory-related workflows. Technical teams investigated performance and reliability issues observed following the release and worked to determine the underlying cause and scope of impact.
The investigation determined that the observed behavior was associated with a defect previously identified in the release. Corrective changes intended to address the defect were included as part of the release activity; however, the corrective changes did not successfully apply across all North America production databases during the deployment process. As a result, some environments continued to experience significantly increased database processing activity during inventory-related operations, contributing to elevated resource consumption, processing delays, timeout conditions, and degraded application performance.
Technical teams worked to successfully deploy the corrective changes across the affected North America production databases. Following deployment, teams performed validation activities and monitored processing workflows to confirm recovery.
By July 22, 2026, at approximately 8:41 AM PDT, validation confirmed successful processing of previously impacted workloads and that the previously observed failure patterns were no longer occurring. The incident was considered resolved and technical teams continued monitoring to confirm service stability.
**Root Cause**
The incident was caused by corrective changes intended to address a previously identified defect not successfully applying across all North America production databases during the release deployment process.
As a result, the underlying defect remained active in affected environments and caused significantly more database processing activity than intended during inventory-related processing operations. This increased workload resulted in elevated resource consumption, processing delays, timeout conditions, and degraded platform performance.
These conditions contributed to intermittent application slowness, delayed processing, reconcile failures, and degraded performance affecting some IT Asset Management functionality within the North America production environment.
**Remediation Actions**
1. **Incident Investigation Initiated:** Technical teams investigated reports of application slowness, processing delays, reconcile failures, and degraded platform performance affecting the North America production environment.
2. **Impact Assessment Performed:** Teams reviewed customer-reported symptoms, system performance data, processing activity, and platform behavior to determine the scope and nature of the issue.
3. **Cause Identified:** Investigation determined that corrective changes intended to address a previously identified defect did not successfully apply across all North America production databases during the release deployment process.
4. **Processing Behavior Corrected:** Technical teams deployed corrections designed to eliminate the inefficient processing behavior that was contributing to elevated database workload, processing delays, timeout conditions, and degraded application performance.
5. **Recovery Validated:** Teams validated the effectiveness of the deployed changes through monitoring and successful execution of previously impacted processing activities.
6. **Post-Restoration Monitoring Performed:** Additional monitoring confirmed that the previously observed failure patterns and performance degradation were no longer occurring and that service stability had been restored.
**Future Preventative Measures**
Based on the investigation, the following follow-up activities have been identified:
* **Processing Logic Improvements:** Technical teams implemented and validated corrective changes to the inventory-processing logic responsible for the increased database workload observed following the release. These changes were designed to eliminate the inefficient processing behavior that contributed to processing delays, reconcile failures, timeout conditions, and degraded application performance.
* **Critical Hotfix Deployment Process Review:** Technical teams will review the deployment process for critical database hotfixes and evaluate alternative approaches to help ensure required corrective changes are successfully applied during future release activities.
**Description:** Flexera One - IT Asset Management - APAC - Missing Menu Items
**Timeframe:** July 12, 2026, 3:00 PM PDT - July 12, 2026, 7:16 PM PDT
**Incident Summary**
On Sunday, July 12, 2026, at 3:00 PM PDT, customers in the APAC production environment began reporting that several IT Asset Management \(ITAM\) menu items, including functionality such as Reports and All Applications, were no longer visible within the Flexera One user interface. Multiple customers were affected, impacting their ability to navigate and access portions of the ITAM application.
Technical teams immediately began investigating the issue and identified authentication-related errors affecting the ITAM user interface. During the investigation, teams reviewed infrastructure health, application behavior, and recent platform changes to determine the source of the problem.
The investigation determined that an application configuration change associated with a newly introduced authentication capability had been incorrectly applied to the production environment. The issue became apparent when application instances were refreshed, resulting in authentication-related failures that prevented certain ITAM menu items from loading correctly for affected users.
Technical teams reverted the affected configuration, refreshed the impacted application instances, and validated recovery across the environment. Following these actions, menu functionality was restored, affected customers confirmed recovery, and the incident was resolved on July 12, 2026, at 7:16 PM PDT after validation and monitoring confirmed normal operation had returned.
**Root Cause**
Investigation determined that an application configuration change associated with a new authentication capability was incorrectly applied to the production environment.
When application instances were subsequently refreshed, the incorrect configuration resulted in authentication-related errors within the ITAM user interface. These errors prevented certain navigation components from loading correctly, causing affected users to experience missing menu items until the configuration was corrected and the affected instances were refreshed.
**Remediation Actions**
The following actions were taken during the incident response:
1. Incident Investigation Initiated: Technical teams began investigating after receiving reports from multiple customers regarding missing ITAM menu items.
2. Application Configuration Reviewed: Recent application and configuration changes were reviewed to identify the source of the menu-loading failures.
3. Incorrect Configuration Identified: Technical teams determined that an authentication-related configuration change had been incorrectly applied to the production environment.
4. Configuration Corrected: The affected configuration was removed from the impacted production instances.
5. Application Instances Refreshed: Impacted application instances were refreshed to ensure the corrected configuration was consistently applied across the environment.
6. Recovery Validation Performed: Technical teams validated menu visibility and functionality across the affected environment and confirmed recovery with impacted customers.
**Future Preventative Measures**
Following the incident, corrective measures were implemented to prevent recurrence of the issue and improve detection of similar conditions in the future.
1. Health Check Monitoring Improvements: The existing health check was updated to detect the application behavior associated with this failure condition. The revised monitoring is designed to identify similar menu-loading and authentication-related application failures more effectively and accelerate detection should a similar issue occur in the future.
2. Deployment Process Reinforcement: The incident highlighted the importance of ensuring new authentication-related features are applied only to their intended environments. The deployment approach and expected application process for these changes have been reinforced with the technical team to reduce the risk of similar configuration issues in the future.