We are currently investigating reports of a collector affecting monitoring.
Impact:
Customers may experience issues with the SNMP monitoring protocol.
Other services are not affected.
Auvik will be performing an emergency rollback of the collector upgrade this weekend.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
identified
Our team has identified a suspected cause of the Collector SNMP monitoring issue and is taking steps to remediate.
Impact:
Customers may continue to experience collector SNMP monitoring issues
We will be having cluster maintenance windows as we roll back the collector upgrade on each cluster.
We are rolling back the EU2 cluster currently.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are implementing mitigation measures and will provide progress updates.
identified
We are continuing to rollback collector versions on each cluster.
Impact:
Customers may continue to experience collector SNMP monitoring issues
We will be having cluster maintenance windows as we roll back the collector upgrade on each cluster.
We are rolling back the remaining clusters currently.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are implementing mitigation measures and will provide progress updates.
resolved
The incident has been fully resolved. Regular service has been restored, and all systems are operating as expected.
Impact:
Users should no longer experience any issues related to this incident.
If you are still experiencing issues, please do not hesitate to reach out to the support team and update your ticket or report any problems you haven't reported yet.
Service has been fully restored. We apologize for the degradation in services. We thank you for your understanding. If you continue to experience issues, please don't hesitate to contact our support team.
We will post an RCA after an internal investigation.
postmortem
# Service Disruption - SNMPv3 Monitoring Loss Following Collector Upgrade
## Root Cause Analysis
### Duration of the incident
Discovered: Apr 13, 2026 20:00 - UTC
Resolved: Apr 13, 2026 22:32 - UTC
### Customer impact
Customers experienced a loss of monitoring data from only devices using specific SNMPv3 configurations. While devices remained online and reachable, monitoring data was not collected, resulting in reduced visibility across environments.
### Cause
A recent collector upgrade introduced changes to encryption handling that affected support for certain legacy SNMPv3 configurations. This resulted in failures when attempting to collect data from devices configured that way.
### Effect
Monitoring data collection failed for affected devices across multiple clusters. This led to a noticeable drop in available device metrics and visibility, despite no loss of connectivity to the devices themselves.
### Future consideation\(s\)
* Expand test coverage to include a broader range of SNMP configurations
* Improve monitoring to detect drops in data collection more proactively
* Strengthen validation processes for major upgrades and dependency changes
* Implement additional safeguards to identify compatibility issues prior to release
Service Degradation – Some Customers Experiencing Login Issues After Maintenance
Description:
During our recent maintenance window, a subset of user accounts were inadvertently disabled. We are actively working to identify affected users and restore access as quickly as possible.
Impact:
Some customers may experience login issues when accessing their sites.
Monitoring and alerting capabilities remain fully operational and are not impacted.
Next Steps:
Restoration efforts are ongoing. Updates will follow as more information becomes available.
We appreciate your patience as we work to restore full functionality.
identified
Our team has identified a suspected cause of the interruption of user authorization to sites and is taking steps to remediate the issue.
Impact:
A significant number of customers are receiving errors when they log into Auvik
The following services are not affected: Monitoring, alerting, or PSA integrations
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are implementing mitigation measures and will provide progress updates as they become available.
identified
Our team has identified a suspected cause of the interruption of user authorization to sites and is taking steps to remediate the issue.
Impact:
A significant number of customers are encountering errors when they log in to Auvik.
Access for these users is currently being restored.
The following services are not affected: Monitoring, alerting, or PSA integrations
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are implementing mitigation measures and will provide progress updates as they become available.
identified
Our team has identified a suspected cause of the interruption of user authorization to sites and is taking steps to remediate the issue.
Impact:
A significant number of customers are encountering errors when they log in to Auvik.
Access for users on the US1, US3, and US5 clusters has been restored.
We are working our way through the remaining clusters
The following services are not affected: Monitoring, alerting, or PSA integrations
Please report any related issues to Auvik Support so we can track and assist further.
identified
Our team has identified a suspected cause of the interruption of user authorization to sites and is taking steps to remediate the issue.
Impact:
A significant number of customers are encountering errors when they log in to Auvik.
Access for users on the US1, US2, US3, US5, AU1, CA1, EU1, EU2 clusters has been restored.
We are working our way through the remaining clusters US4 and US6
The following services are not affected: Monitoring, alerting, or PSA integrations
Please report any related issues to Auvik Support so we can track and assist further.
identified
Our team has identified a suspected cause of the interruption of user authorization to sites and is taking steps to remediate the issue.
Impact:
A significant number of customers are encountering errors when they log in to Auvik.
Access for users on the US1, US2, US3, US4, US5, AU1, CA1, EU1, EU2 clusters has been restored.
We are working our way through the remaining cluster, US6.
The following services are not affected: Monitoring, alerting, or PSA integrations
Please report any related issues to Auvik Support so we can track and assist further.
resolved
The incident has been fully resolved, and all services are operating normally.
Customers should no longer experience any login issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.
We will provide a Root Cause Analysis (RCA) once it is available.
postmortem
# Service Degraded - Some Customers Experiencing Login Issues After Maintenance
## Root Cause Analysis
### Duration of the incident
Discovered: Mar 7, 2026 18:42 - UTC
Resolved: Mar 9, 2026 21:10 - UTC
### Cause
Following a platform upgrade, an internal reconciliation process incorrectly identified a subset of user records as deleted. This occurred due to an error condition in a synchronization component that failed to process all expected user data during reconciliation.
### Effect
Affected users were able to log in to the platform but were unable to access their expected sites or tenants due to missing user authorizations. In some cases, API access using previously generated API keys was also affected.
Monitoring, alerting, and data collection services continued to function normally throughout the incident.
### Action taken
Engineering paused the process that spreads the changes across the platform to prevent further impact.
User access was then restored using backup data taken prior to the upgrade. This allowed the team to recover user access, permissions, and related access keys.
After restoration was completed and access was verified across all environments, the incident was declared resolved.
### Future consideration\(s\)
* Implement additional monitoring to detect unexpected changes in user or authorization records.
* Add safeguards within reconciliation processes to prevent large-scale unintended deletions.
* Improve post-upgrade validation checks to identify abnormal record changes earlier.
* Enhance disaster recovery tooling and backup retrieval processes to speed up restoration efforts.
Emergency maintenance for EU2
開始 2025年12月15日 23:28 UTC · 2h 29m
Issues軽微なインシデント
影響を受けたコンポーネント
eu2.my.auvik.com
investigating
Affected Services: Access
Cluster(s): EU2
We are currently performing emergency maintenance, which requires a restart of the EU2 cluster.
Impact:
During this time, users in the EU2 region may experience login issues or intermittent service disruptions.
Next Steps:
We will provide updates as we learn more.
We appreciate your patience as we work to resolve this issue.
monitoring
A restart of the EU2 cluster has resolved the issue. We will monitor the situation to ensure stability and confirm that the service remains fully functional.
Impact:
Services should be operating normally; however, we continue monitoring for irregularities.
If you are still experiencing issues, please do not hesitate to reach out to the support team and update your ticket or report any problems you haven't reported yet.
Next Steps:
We will provide a final update once the issue is resolved.
We appreciate your patience as we work through this issue.
resolved
The incident has been fully resolved. Regular service has been restored, and all systems operate as expected.
Impact:
Users should no longer experience any issues related to this service disruption.
If you are still experiencing issues, please do not hesitate to reach out to the support team and update your ticket or report any problems you haven't reported yet.
Service has been fully restored. We apologize for any disruption to our services. We thank you for your understanding. If you continue to experience issues, please don't hesitate to contact our support team.
We will post an RCA after an internal investigation.
postmortem
# Service Disruption - The EU2 cluster was unavailable
## Root Cause Analysis
### Duration of the incident
Discovered: Dec 15, 2025 20:00 – UTC
Resolved: Dec 16, 2025 02:00 – UTC
### Customer impact
During the incident window, customers hosted in the EU2 region experienced intermittent service degradation. This included slower system responsiveness, temporary inconsistencies in monitoring data, and brief periods where alerts may have been delayed or inaccurate.
Most customers regained access as services were progressively restored, and complete stability was confirmed before the incident was closed.
### Cause
The incident was caused by an elevated load in the EU2 service environment, resulting in an uneven workload distribution across backend resources. As the load increased, automated recovery mechanisms were unable to stabilize the environment fully, necessitating a controlled restart of the regional service to restore normal operations.
### Effect
The imbalance led to reduced service performance and temporary unavailability for some customers until recovery actions were completed. Engineering teams were required to intervene to safely restart the affected region and validate service health before returning operations to normal.
### Future consideation\(s\)
* Improve automated workload balancing to absorb regional load increases.
* Strengthen early indicators for backend saturation to enable earlier intervention.
* Refine operational procedures to further reduce recovery time in similar scenarios.
Meraki Switch stacks creating duplicate devices in the product
We are currently investigating reports of Meraki switch stacks creating duplicates.
Impact:
Customers may see duplicate inventory in the Meraki stacks.
The following services are not affected: Alerting and all other Device inventory.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
investigating
We are currently investigating reports of Meraki switch stacks creating duplicates.
Impact:
Customers may see duplicate inventory in the Meraki stacks.
No new devices are currently being created. Devices that have been duplicated cannot yet be deleted.
The following services are not affected: Alerting and all other Device inventory.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
identified
Our team has identified a suspected cause of the Meraki duplicate Stack switches and is taking steps to remediate the issue.
Impact:
Customers may continue to experience duplicated Meraki Stack switches
The following services are not affected: Alerting and all other Device inventory.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
monitoring
We have applied changes to address the issue. You'll now see the duplicate begin to be removed.
Impact:
Services should be operating normally; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
resolved
The incident has been fully resolved, and all services are operating normally.
Customers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.
We will provide a Root Cause Analysis (RCA) once it is available.
postmortem
# Service Degraded - Duplicated Devices for Meraki Switch Stacks
## Root Cause Analysis
### Duration of the incident
Discovered: Dec 15, 2025 08:34 – UTC
Resolved: Dec 15, 2025 13:20 – UTC
### Customer impact
Customers with Meraki switch stacks across all clusters experienced:
Duplicate Meraki switches are appearing in the device inventory, including entries without management IP addresses.
Inability to manually delete duplicates, as they were automatically recreated.
Confusing or inaccurate device and stack representations.
Temporary inflation of billable device counts for some tenants \(billing data captured and under review\).
No other device types were affected.
### Cause
A configuration update intended to improve how Meraki switch stacks are represented in the platform caused unintended behavior.
For tenants already using Meraki stacking, the system was unable to match newly processed stack data to existing switches consistently. This caused already-stacked switches to be incorrectly interpreted as new devices, generating duplicate inventory entries. Because the underlying synchronization logic reinforced these entries, any customer deletion of duplicates did not persist, and duplicates reappeared.
### Effect
The issue led to inaccurate device inventories and misleading switch-stack views for customers using Meraki stacking, causing confusion in device management and increasing support tickets. Internal teams experienced additional operational overhead as they investigated the duplication patterns and validated safe removal procedures. Complete remediation required coordinated cleanup across all clusters to restore accurate device records and ensure no further duplicates were generated.
### Future consideration\(s\)
* Strengthen validation of device and stack metadata before inventory updates.
* Improve testing coverage using a production-like stacked device configuration.
* Add monitoring and alerting for abnormal changes in device counts.
* Use safer rollout controls and feature flags for changes affecting device inventory logic.
Duplicate FortiNet firewalls created on the EU2 cluster.
開始 2025年12月11日 14:30 UTC · 9h 16m
Issues軽微なインシデント
影響を受けたコンポーネント
eu2.my.auvik.com
investigating
We are currently investigating reports of duplicate devices being created, which could cause alert flooding.
Impact:
Customers may experience a high false alert definitions
The following services are not affected by all other clusters and monitoring.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
identified
Our team has identified a suspected cause of the Fortinet firewall duplication and is taking steps to remediate the issue.
Impact:
Customers may continue to receive some excessive resolved alerts while duplicates are cleared from the system.
The following services are not affected: Any other devices or clusters.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are implementing mitigation measures and will provide progress updates.
identified
Our team has begun deleting duplicate Fortinet firewalls.
Impact:
Customers may continue to receive excessive resolved alerts while duplicates are cleared from the system.
The following services are not affected: Any other devices or clusters.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are implementing mitigation measures and will provide progress updates.
monitoring
We have applied changes to address the issue. Services are beginning to revert to normal, and we are monitoring closely for stability.
Impact:
Extraneous Fortinet firewalls will continue to be flushed from affected tenants; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
identified
Our team continues to investigate the suspected cause of the duplicate Fortinet firewalls and is taking steps to remediate the issue.
Impact:
Customers may continue to experience alerts from duplicate Fortinet firewalls.
The following services are not affected: All other clusters and vendors.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying new mitigation measures and will provide updates on progress.
identified
Our team continues to investigate the suspected cause of the duplicate Fortinet firewalls and is taking steps to remediate the issue.
Auvik will place the EU2 cluster in maintenance mode to continue working on the issue and prevent further false alerts.
Impact:
All alerting for the EU2 cluster has been placed in maintenance mode.
The following services are not affected: All other clusters.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying new mitigation measures and will provide updates on progress.
identified
Our team has begun isolating the devices requiring correction.
Auvik has placed the EU2 cluster in maintenance mode to continue working on the issue and prevent further false alerts.
Impact:
All alerting for the EU2 cluster has been placed in maintenance mode.
The following services are not affected: All other clusters.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying new mitigation measures and will provide updates on progress.
identified
We are continuing to work on a fix for this issue.
resolved
The incident has been fully resolved, and all services are operating normally.
Customers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.
We will provide a Root Cause Analysis (RCA) once it is available.
postmortem
# Service Degraded - FortiGate firewalls duplicated in the EU2 cluster, causing a flood of alerts.
## Root Cause Analysis
### Duration of the incident
Discovered: Dec 11, 2025 14:00 – UTC
Resolved: Dec 17, 2025 06:45 – UTC
### Customer impact
Clients with FortiGate firewalls in the EU2 region experienced:
* A significant increase in “device offline” alerts due to status flapping between duplicate and original device entries.
* Duplicate firewalls appear in inventory views, leading to confusion in device management.
* Inaccurate device statistics and monitoring gaps.
* In some cases, customers attempted to delete duplicates, which resulted in the loss of historical data on the original device.
No other device types or regions were affected.
### Cause
A system change intended to improve how high-availability firewall configurations were processed introduced unexpected behavior. Under certain conditions, the platform was unable to detect when multiple records referenced the same firewall device consistently. As a result, some firewalls were incorrectly detected as new devices, causing duplicate entries and inconsistent operational status reporting.
### Effect
The issue resulted in elevated alert volumes across impacted tenants and caused instability in device-status reporting for the affected firewalls. Customers experienced increased operational overhead as they reviewed unexpected alerts and duplicate device entries, and engineering teams were required to intervene to restore accurate device associations and remove the duplicated records.fFuture consideration\(s\)
* Strengthen validation of device-supplied information before it affects inventory or monitoring.
* Implement alerting for sudden spikes in offline/online transitions to detect similar issues earlier.
* Improve observability of device lifecycle behavior to identify anomalies proactively.
* Use feature-flagged or staged rollouts for changes that affect device processing logic.
* Incorporate production-like data into pre-deployment testing to better anticipate unexpected device behaviors.
Auvik Live Chat is Inaccessible
開始 2025年12月5日 12:50 UTC · 42m
Issues軽微なインシデント
影響を受けたコンポーネント
Auvik Website (www.auvik.com)
investigating
We are currently investigating reports of the Auvik chatbot not transferring the session to a live agent.
Impact:
Customers will not be able to reach a live agent until the service is restored.
The following services are not affected: The Knowledge Base, Support Portal, and Auvik site.
Please submit a web form ticket from the Support site in the meantime, if needed to contact support.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
identified
Our team has identified a suspected cause of the Auvik chatbot not transferring the session to a live agent.
Impact:
Customers will not be able to reach a live agent until the service is restored.
The following services are not affected: The Knowledge Base, Support Portal, and Auvik site.
Please submit a web form ticket from the Support site in the meantime, if needed to contact support.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
resolved
The incident has been fully resolved, and all services are operating normally.
Customers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.
We appreciate your understanding and patience.
Service Disruption - Login issues
開始 2025年12月3日 14:21 UTC · 3h 7m
Issues軽微なインシデント
identified
Affected Services: Login
We’re aware that an issue with Okta caused login failures for many users accessing Auvik. Okta has reported that the problem is resolved, but some users may still experience difficulties signing in.
You can follow their updates here: https://status.okta.com/#incident/a9CWR00000012Cg2AI
Recent status information is also available at: https://status.okta.com/#recent
Impact:
While we work on the resolution, users may continue to experience login issues
Next Steps:
Our team is actively working to resolve the issue and will provide updates as progress is made.
Your patience is greatly appreciated, and we regret any inconvenience you may have experienced.
monitoring
Latest update from Okta
December 3, 2025 at 7:50am PST
Our service is seeing recovery following successful mitigation efforts with our provider. We continue to closely monitor system stability. Our next update will be in 30 minutes, or sooner if new information is available.
We are watching our systems for recovery and performance.
resolved
Latest update from Okta
Incident Resolved
December 3, 2025 at 8:58am PST
An issue impacting the core authentication service for customers in Okta Cell US7 has been resolved. Additional root cause information will be available within five business days.
We will close the status page for Auvik's sites.
Service Disruption - Reduced functionality across all clusters
Affected Services: Hierarchy and permissions
Cluster(s): All
Our team has identified the root cause of the service disruption and is currently investigating a solution to restore normal service levels.
Impact:
While we work on the resolution, users may continue to experience reduced functionality in the product.
Next Steps:
Our team is actively working to resolve the issue and will provide updates as progress is made.
Your patience is greatly appreciated, and we regret any inconvenience you may have experienced.
identified
We are aware of an issue where customers can log in successfully but may be unable to view certain data, including maps, inventory, site lists, and a few other areas. Our team has identified the root cause and compiled a fix.
The estimated time to full recovery is approximately 30–60 minutes.
Thank you for your patience.
resolved
The incident has been fully resolved. Regular service has been restored, and all systems operate as expected.
Impact:
Users should no longer experience any issues related to this service disruption.
If you are still experiencing issues, please do not hesitate to reach out to the support team and update your ticket or report any problems you haven't reported yet.
Service has been fully restored. We apologize for any disruption to our services. We thank you for your understanding. If you continue to experience issues, please don't hesitate to contact our support team.
We will post an RCA after an internal investigation.
postmortem
# Service Degraded - Data Not Loading Across All Clusters
## Root Cause Analysis
### Duration of the incident
Discovered: Nov 21, 2025 17:43 – UTC
Resolved: Nov 21, 2025 19:18 – UTC
### Customer impact
Customers across all regions were able to log in, but many were unable to view key data within the platform. This included features such as:
* Network maps
* Inventory and site lists
* Dashboards
* Other views are dependent on permissions and hierarchical data.
This resulted in degraded usability and limited visibility into managed environments.
### Cause
During a routine maintenance task intended to clean up older data records, an incorrect piece of information was unintentionally added to the system. This faulty data prevented a core component—responsible for organizing how customer information is displayed in the platform—from functioning correctly.
Because this component could not run as expected, many areas of the product that rely on structured data \(such as maps, dashboards, and inventory views\) were unable to load. This led to widespread service degradation across all clusters.
### Effect
Because the system could not correctly load the information needed to display customer environments, several parts of the platform were unable to show data. As a result, many users could log in but were unable to see maps, inventory details, site lists, dashboards, or other information they usually rely on.
This created a degraded experience across all regions, limiting visibility and making it difficult for customers to perform routine monitoring and management tasks until the issue was resolved.
### Future consideation\(s\)
* Introduce stronger validation and safeguards to prevent malformed data from being accepted into critical systems.
* Improve the resilience of backend services so they fail gracefully rather than entering repeated restart cycles.
* Implement additional monitoring to provide earlier detection of issues affecting data-loading functions.
* Enhance internal tools used during maintenance and cleanup activities to reduce the risk of unintentional data corruption.
Auvik is currently being impacted by today’s AWS outage.
Some pages and APIs may be slow or have errors. No data loss is expected.
We will update here as we receive more information
monitoring
AWS services are recovering. Services appear to be operating normally, and we are monitoring closely for stability.
Impact:
Services should be operating normally; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
monitoring
AWS services are recovering. Services appear to be operating normally, and we are monitoring closely for stability. Amazon is recommending flushing your DNS cache if you are still experiencing issues.
Impact:
Services should be operating normally; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
resolved
The incident has been fully resolved, and all services are operating normally.
Customers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.
Clients on US4 cluster are experiencing 500 Errors
開始 2025年10月15日 15:22 UTC · 2h 12m
Issues軽微なインシデント
影響を受けたコンポーネント
us4.my.auvik.com
investigating
We are currently investigating reports of 500 errors affecting access to clients on the US4 cluster.
Impact:
Customers may experience access issues to their tenants.
Alerts and monitoring are not affected
The other clusters are not affected.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
investigating
We are continuing to investigate HTTP 500 errors affecting access to clients on the US4 cluster.
Impact:
Customers will experience access issues to their tenants.
Alerts and monitoring are not affected.
The other clusters are not affected.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
identified
Our team has identified a suspected cause of the HTTP 500 site access issue and is taking steps to remediate it.
Impact:
Customers should now be able to access their sites without encountering an HTTP 500 error.
Maps and site inventories are not rendering correctly in the UI.
The following services are not affected: Monitoring and Alerting.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
monitoring
We have applied changes to address the issues. Services appear to be operating normally, and we are monitoring closely for stability.
Impact:
Services should be operating normally; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
resolved
The incident has been fully resolved, and all services are operating normally.
Customers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.
We will provide a Root Cause Analysis (RCA) once it is available.
postmortem
# Service Degraded - Sites have lost settings after maintenance
## Root Cause Analysis
### Duration of the incident
Discovered: Oct 11, 2025 – 12:00 UTC
Resolved: Oct 15, 2025 – 22:00 UTC
### Cause
During a scheduled system update, a background maintenance process unintentionally removed reference files used to identify stored site configurations. When affected systems restarted after the update, they were unable to locate those configuration files and temporarily appeared as new, empty sites.
This occurred because the maintenance process was using outdated information when determining which data to clean up safely. The underlying data remained securely stored, but the missing reference files prevented normal access until they were restored.
### Effect
A subset of tenants across multiple clusters temporarily lost access to their site configurations and appeared as newly created environments.
Customers observed missing data and configurations, including previously defined network settings and device details.
### Action taken
_All times are in UTC_
**10/11/2025**
**12:00** – A scheduled system upgrade began across all clusters.
**15:25** – The support team received reports from customers that some sites appeared empty or missing data.
**16:00** – Engineering immediately began investigating and determined this was not related to normal data processing delays.
**10/12/2025**
Additional reports confirmed that several sites were missing configuration information. The engineering team confirmed that the original data was still securely stored, but was not being correctly loaded by the system.
**10/13/2025**
The issue was traced to missing metadata files that help identify stored configurations. The engineering team began restoring affected sites using the most recent valid configuration data.
Automated recovery tools were developed to safely restore additional sites and ensure consistent recovery across all clusters.
**10/14/2025**
Engineering verified that configuration data was fully restored and synchronized across supporting services.
The recovery process was extended to all remaining sites, with validation steps confirming successful restoration.
**10/15/2025**
Final recovery efforts for the remaining affected clusters were completed..
**22:00** – All affected sites were confirmed operational with their configurations restored and verified.
### Future consideration\(s\)
* Remove dependency of cleanup processes on outdated cluster data sources.
* Validate all automated cleanup jobs to ensure they do not operate on production clusters.
* Implement monitoring for missing or corrupted metadata files before deployments.
* Enhance post-deployment validation to verify the integrity of configuration data across all clusters.
We are currently investigating reports of a clients not being able to access their tenants on the CA1 cluster.
Impact:
Customers may experience an interruption of access to their sites.
The following services are not affected: Other clusters are not affected by this
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
investigating
We are continuing to investigate tenants on the CA1 cluster that are not accessible.
Impact:
Customers are experiencing an interruption of access to their sites.
The following services are not affected: Other clusters are not affected by this outage.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
identified
Our team has identified a suspected cause of the sites on the CA1 cluster being inaccessible and is taking steps to remediate the issue.
There have also been reports of some users experiencing issues accessing Auvik using their Okta setup. This is also being investigated.
Impact:
Customers may continue to experience their site being inaccessible on the CA1 cluster.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
identified
Our team has identified a suspected cause of the sites on the CA1 cluster being inaccessible and is taking steps to remediate the issue.
There have also been reports of some users experiencing issues accessing Auvik using their Okta setup. This is also being investigated.
Additionally, any new tenants or users created since 17:15 UTC (1:15 PM ET) are currently not functioning or accessible.
Impact:
Customers may continue to experience inaccessibility to their site on the CA1 cluster as the cluster comes back up.
Clients or clusters created since the incident began are not accessible.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
identified
Our team has identified a suspected cause of the sites on the CA1 cluster being inaccessible and is taking steps to remediate the issue.
Sites have been restored to yesterday. We are working to restore data that was present at the time of this incident.
There have also been reports of some users experiencing issues accessing Auvik using their Okta setup. We are working to restore access.
Additionally, any new tenants or users created since 17:15 UTC (1:15 PM ET) are currently not functioning or accessible. Tenant restoration is occurring.
Impact:
Customers should now be able to access their sites on the CA1 cluster, with the exceptions noted above. Data restoration is being worked on.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
monitoring
We have implemented changes to address the issue with services for clients on the CA1 clusters, and these clusters are now operating normally. We are monitoring closely for stability.
Impact:
Services should be operating normally for the tenants in the CA1 cluster.
Tenants created have also been restored.
There are still issues with some clients who use Okta to access Auvik. We are looking to remediate these.
If you continue to encounter problems, please report them to Auvik Support.
Next Steps:
We will be monitoring the situation during the evening and will report updates in the morning.
A final update will be posted once we confirm the resolution.
investigating
We are currently investigating reports of not being able to access new sites after they are created. This is under all clusters.
Impact:
All other services should be operating normally for the tenants in the CA1 cluster.
We are continuing to experience issues with a small percentage of Okta-enabled accounts that lack access. We are looking to remediate these.
If you continue to encounter problems, please report them to Auvik Support.
Next Steps:
We will be monitoring the situation during the evening and will report updates in the morning.
A final update will be posted once we confirm the resolution.
monitoring
We have applied changes to address the issue with new site creation. We continue to experience issues with some cases of child site creations and are working to rectify them.
Other services appear to be operating normally, and sites are becoming visible in the UI. We are monitoring closely for stability.
Impact:
All other services should be operating normally for the tenants in the CA1 cluster.
We are continuing to experience issues with a small percentage of Okta-enabled accounts that lack access. We are remediating these.
If you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
monitoring
We have applied changes to address the issue with new site creation. We continue to experience issues with sites not being visible promptly. We are continuing to work on this issue.
This is also contributing to a site redirection issue where sites are not accessible via xxx.my.auvik.com, but are accessible via xxx.us1.my.auvik.com, where us1 is the cluster on which the site is located.
Other services appear to be operating normally, and sites are becoming visible in the UI. We are monitoring closely for stability.
Impact:
We are continuing to experience issues with a small percentage of Okta-enabled accounts that lack access. We are remediating these accounts.
If you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
monitoring
We are currently experiencing some system slowness on clusters outside of CA1. This is related to residual cleanup from earlier issues on the CA1 cluster. Our team is actively addressing this to ensure ongoing platform stability.
Additionally, the cleanup contributed to a site redirection issue, where sites may not be accessible via xxx.my.auvik.com. These sites are accessible when using the fully qualified URL (e.g., xxx.us1.my.auvik.com), where us1 reflects the cluster hosting the site. In some cases, newly created sites may not display correctly.
A small percentage of clients using Okta authentication were also impacted. If you encounter a denied login message, you may need to reset your password by following the provided instructions.
Impact
Some clients may see slower performance.
Site access via xxx.my.auvik.com may fail; use the direct cluster URL (xxx.[cluster].my.auvik.com) as a workaround.
Limited Okta users may experience authentication issues requiring a password reset.
Next Steps
Our engineering team is continuing remediation efforts. A final update will be provided once the issues are fully resolved.
If you continue to encounter problems, don't hesitate to get in touch with Auvik Support
.
monitoring
We continue to experience system slowness on clusters outside of CA1, which is attributed to residual cleanup efforts related to earlier issues in the CA1 cluster. Although our team has made progress, some symptoms persist, and we continue to work on resolving them.
Impact:
Performance: Some clients may still see slower performance.
Site Access: Accessing sites via xxx.my.auvik.com may fail. Sites remain accessible using the fully qualified URL format (xxx.[cluster].my.auvik.com).
Login Issues: Some sites are still experiencing login issues.
Okta Authentication: A small subset of clients using Okta may continue to see denied login attempts. A password reset may be required.
Next Steps
We are focused on restoring full login functionality and addressing the residual impacts of the cleanup. A final update will be provided once all issues are fully resolved.
If you continue to experience problems, please contact Auvik Support.
monitoring
We are continuing to monitor for any further issues.
identified
Our team has identified a suspected cause of the slowness on the EU1 cluster and is taking steps to remediate the issue by performing an out-of-band maintenance on the EU1 cluster.
Impact:
Customers may experience issues accessing their sites on the EU1 cluster over the next two hours while this maintenance is in progress.
Sites on the other clusters are not affected by the action.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
monitoring
Our team has successfully completed the out-of-band maintenance on the EU1 cluster. We will continue to monitor EU1 closely throughout the evening to ensure stability.
The earlier slowness affecting clusters outside of CA1 now appears to be resolved. We will continue to validate system performance and client connections across all clusters.
Impact:
EU1: Sites remain stable following maintenance; we are monitoring for any recurrence.
Other Clusters: No ongoing slowness observed.
Login / Okta: A small subset of clients using Okta may still experience denied login attempts; a password reset may be required.
Next Steps:
Continue overnight monitoring of EU1 and client connection stability.
Perform additional validation tomorrow on site performance and login flows.
If you encounter related issues, please contact Auvik Support so we can track and assist further.
Next Steps:
A final update will be posted once we confirm the resolution.
monitoring
Services appear to be operating normally, and we are continuing to monitor closely for stability.
We are still receiving periodic reports of site slowness. This seems especially pertinent with the new discovery scans. We are currently investigating.
A small percentage of clients using Okta authentication were impacted. If you encounter a denied login message, you may need to reset your password by following the vendor's instructions. You will also need to re-enable your MFA.
Impact:
Services should be operating normally; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
investigating
We are currently investigating reports of 500 errors that occur periodically when accessing pages on sites.
We are still receiving periodic reports of site slowness. This seems especially pertinent with the new discovery scans. We are currently investigating.
A small percentage of clients using Okta authentication were impacted. If you encounter a denied login message, you may need to reset your password by following the vendor's instructions. You will also need to re-enable your MFA.
Impact:
Customers may experience intermittent issues with their sites
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
investigating
We are continuing to investigate this issue.
monitoring
We identified and addressed the root causes behind intermittent 500 errors, map loading issues for new tenants, and isolated EU1 access failures. Fixes have been implemented, including updates to API rate-limit handling, permissions validation, and service restarts. Impacted services are recovering as expected.
We are closely monitoring system stability and completing a sweep of collectors across clusters to confirm full health. No further impact is expected at this time.
Next Steps:
A final update will be posted once we confirm the resolution.
monitoring
We’ve resolved the issues that caused errors, map loading problems for new tenants, and access interruptions in EU1. Services are recovering and performing as usual.
We will be monitoring the platform closely throughout the evening, focusing on overall stability and conducting a comprehensive sweep of collectors across clusters to confirm their health. No further customer impact is expected at this time.
If you continue to encounter problems, please don't hesitate to contact Auvik Support.
Next Steps: A final update will be posted once the resolution is confirmed.
resolved
The incident has been fully resolved, and all services are operating normally.
Customers should no longer experience any related issues. If you continue to experience problems, please don't hesitate to contact Auvik Support.
We will provide a Root Cause Analysis (RCA) once it is available.
postmortem
# Service Disruption - Platform Availability and Login Access
## Root Cause Analysis
### Duration of the incident
Discovered: Sep 23, 2025 – 17:45- UTC
Resolved: Sep 26, 2025 – 13:58 - UTC
### Customer impact
Tenants on the CA1 cluster lost their settings, and Auvik was inaccessible for a time.
Intermittent platform errors and degraded performance for some customers.
A temporary issue prevented certain users from logging in via Okta.
Delays in data synchronization in the EU1 region necessitate restarting the cluster.
### Cause
A configuration update intended for specific tenants was applied globally due to a missing query constraint.
This caused the deletion of tenant and user records in one cluster, leading to cascading synchronization and workload impacts across other clusters.
As a result, user data was propagated as deletions to the identity provider, temporarily removing affected user accounts.
### Effect
The CA1 cluster required restoration from a backup, resulting in the temporary unavailability of some tenant data.
43 tenants and 22 user accounts required manual recreation and verification.
Authentication failures occurred for affected users due to the removal of identity records.
The EU1 cluster experienced performance degradation as it processed a large data synchronization backlog.
Increased load briefly impacted the responsiveness of other clusters’ regional services.
### Future consideration\(s\)
* Implement a dry-run and confirmation step in internal tooling to validate production commands before execution.
* Reinforce the change-management process for improved peer visibility and validation.
* Expand resource and capacity monitoring to identify anomalies earlier and respond proactively.
* Revisit release control lifecycle practices to remove obsolete configurations after rollout completion.
We are currently investigating reports of login issues affecting access to Auvik for its users.
Impact:
Customers may be unable to access their tenants.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
investigating
We are currently investigating reports of login issues affecting access to Auvik for its users.
Impact:
Customers may experience the inability to access their sites when using redirects to their site URL.
The following services are not affected: Monitoring and Alerting.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
identified
Our team has identified a suspected cause of the login access and is taking steps to remediate the issue.
Impact:
Customers may continue to experience login issues if using URL redirects to access their site(s).
The following services are not affected: Monitoring and Alerting.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
identified
Our team has identified a suspected cause of the login access and is taking steps to remediate the issue.
Impact:
Customers may continue to experience login issues if using URL redirects to access their site(s).
The following services are not affected: Monitoring and Alerting.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We continue applying mitigation measures and will provide updates on progress.
monitoring
We have applied changes to address the issue. Services appear to be operating normally, and we are monitoring closely for stability.
Impact:
Services should be operating normally; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
monitoring
The changes were applied to address the issue.. Services appear to be operating normally, and we are continuing to monitor closely for stability.
Impact:
Services should be operating normally; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
monitoring
Monitoring
The changes were applied to address the issue.. Services appear to be operating normally, and we are continuing to monitor closely for stability.
Impact:
We recently experienced a short disruption with URL redirects. This has been resolved, and services are working as expected.
Services should be operating normally; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
resolved
The incident has been fully resolved. Regular service has been restored, and all systems operate as expected.
Impact:
Users should no longer experience any issues related to this service disruption.
If you are still experiencing issues, please do not hesitate to reach out to the support team and update your ticket or report any problems you haven't reported yet.
Service has been fully restored. We apologize for any disruption to our services. We thank you for your understanding. If you continue to experience issues, please don't hesitate to contact our support team.
We will post an RCA after an internal investigation.
postmortem
# Service Degraded - Clients experienced login issues to their site because the URL redirect was not working.
## Root Cause Analysis
A recent update led to more traffic than expected, resulting in simultaneous overload of the same systems. This overloaded them, leading to delays and occasional failures when customers tried to log in. Some customers also experienced issues when trying to start new trials.
### Duration of the incident
Discovered: Sep 17, 2025 13:50 - UTC
Resolved: Sep 18, 2025 21:00 - UTC
### Cause
The update unintentionally created extra demand on shared systems. As a result, the login process and new trial creation sometimes failed or responded slowly.
### Effect
* Some customers could not log in after entering their password or completing MFA.
* Occasional slow responses and error messages \(404/500/502/504\).
* New trial sign-ups sometimes failed or were delayed.
* Internal tools that rely on the same login process also saw intermittent issues.
### Action taken
* Adjusted system settings to reduce pressure on overloaded services.
* Closely monitored traffic while making changes to keep the service stable.
* Applied a temporary workaround to allow new trials to be created reliably.
* Released a fix to stabilize the login redirect process.
* Made further tuning changes to spread out demand and reduce load.
* Continued monitoring until the login and sign-up flows were confirmed to be stable.
### Future consideration\(s\)
* Reduce reliance on a single system for login and trial flows by distributing the workload across multiple systems..
* Enhance monitoring to identify login errors and trial creation issues more promptly.
* Add safeguards to prevent overload, including traffic limits and fallback options.
* Test updates under heavier load conditions to catch these issues earlier.
We are currently investigating reports of sites receiving 500 errors affecting site access.
Impact:
Customers may experience an inability to connect to their site(s)or have collectors connect.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
identified
Our team has identified a suspected cause of the connection issues and is taking steps to remediate the issue.
Impact:
Customers who experienced the 500 errors are now remedied.
Connection issues after installing a Windows collector are still being investigated. Users may continue to experience the collector not connecting to the site.
The following services are not affected: monitoring and alerting.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
monitoring
We have applied changes to address the issue. Services appear to be operating normally, and we are monitoring closely for stability.
Impact:
Collector connection services should be operating normally; however, if you continue to encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
resolved
The connectivity issues with the site and collectors have been fully resolved, and services are operating as expected.
Impact:
Customers should no longer experience any related issues. If you continue to experience issues, please report them to Auvik Support.
postmortem
# Service Degraded - Login and Collector Installation Issues on US3 Cluster
## Root Cause Analysis
### Duration of the incident
Discovered: Sep 10, 2025 13:17 UTC
Resolved: Sep 12, 2025 15:08 UTC
### Cause
Two separate but overlapping issues contributed to this incident:
* Collector Installation Failures – Windows collectors were unable to install due to missing service principal credentials on the backend agent server, which prevented successful API calls for subscription data.
* Login and Redirect Failures – Following a restart of the US3 frontend, requests for user and tenant data from secondary clusters intermittently failed. This caused login attempts through Okta to hang and product redirects to fail.
### Effect
* Users attempting to log in via Okta were unable to complete authentication and access tenants.
* Some sites experienced 500 errors when attempting to access dashboards.
* Windows collector installations via GUI and CLI failed, preventing the deployment of new collectors.
### Action taken
_All times are in UTC_
**09/10/2025**
**13:17** – Users report inability to log in to Auvik Production through Okta.
**13:28** – Errors in US3 frontend logs identified relating to user/tenant data queries.
**13:37** – Engineering suspends frontend deployment in the US3 cluster.
**13:39** – Issues confirmed across secondary clusters; logs analyzed for root cause.
**14:54** – Identified that the frontend redirect service could not fetch required tenant data; feature flag disabled to restore functionality.
**15:01** – Engineering confirms that the workaround restores login functionality while monitoring tenant recovery. The incident is resolved on the status page.
**09/12/2025**
**15:08** – Feature flag re-enabled after system recovery; frontend services reconciled successfully. Incident fully resolved.
### Future consideration\(s\)
* Improve validation of service principal credentials to prevent collector installation failures.
* Enhance monitoring and alerting around login and redirect workflows to detect tenant query failures earlier.
* Review and refine feature flag rollout procedures to minimize dependency risks and ensure optimal deployment.
* Continue improving how customer data is distributed across clusters so that queries run more efficiently and reliably, even during high system load or maintenance events.
US4 clients are not accessible
開始 2025年8月27日 21:30 UTC · 1h 37m
Outage重大なインシデント
影響を受けたコンポーネント
us4.my.auvik.com
investigating
We are currently investigating reports of an outage affecting clients on US4.
Impact:
Customers may experience an inability to log in to their site
The following services are not affected: Other clusters, including alerting and monitoring.
Next Steps:
Our team is working to identify contributing factors. Updates will follow as more information becomes available.
identified
Our team has identified a suspected cause of the outage on the US4 cluster and is taking steps to remediate the issue.
Impact:
Customers will not have access to the sites on the US4 cluster for the next two hours while we bring them online. Some sites on the US4 may become operational before that time.
The following services are not affected: Sites on other clusters.. Please report any related issues to Auvik Support so we can track and assist further.
Next Steps:
We are applying mitigation measures and will provide updates on progress.
monitoring
We have applied changes to address the issue. Services appear to be starting up normally, and we are monitoring closely for stability.
Impact:
Services should be becoming available. The site will continue to come online over the next hour. If you encounter problems, please report them to Auvik Support.
Next Steps:
A final update will be posted once we confirm the resolution.
resolved
The access issues for clients on the US4 cluster have been fully resolved, and services are operating as expected.
Impact:
Customers should no longer experience any related issues. If you continue to experience issues, please report them to Auvik Support.
Service Disruption - Auvik clients are experiencing a disruption of services - Multiple Clusters
Affected Services: Clients are inaccessible
Cluster(s): All Cluster
We are currently experiencing a service disruption. Our team is actively investigating the root cause and working to resolve the issue as quickly as possible.
Impact:
It is still being determined.
Next Steps:
We will update this information as more details become available.
We appreciate your patience as we work to restore full functionality.
investigating
Affected Services: Clients are not accessible
Cluster(s): All Clusters
Description:
We are currently experiencing degraded services. Our team is actively investigating the root cause and working to resolve the issue as quickly as possible.
Impact:
Users may experience an inability to access their tenants.
We are currently performing a cluster restart on the US2 cluster. Down time is expected to be 1-1.5 hours.
We are investigating the other clusters.
Next Steps:
We will update this information as more details become available.
We appreciate your patience as we work to restore full functionality.
investigating
Affected Services: Clients are not accessible
Cluster(s): US2 and US5
Description:
We are currently experiencing degraded services. Our team is actively investigating the root cause and working to resolve the issue as quickly as possible.
Impact:
Users may experience an inability to access their tenants.
We are currently performing a cluster restart on the US2 cluster. Down time is expected to be 1-1.5 hours.
We are now also performing a cluster restart on the US5 cluster. Down time is expected to be 1-1.5 hours.
The other clusters appear not to be affected. Monitoring and alerting are working on them
Next Steps:
We will update this information as more details become available.
identified
Our team has identified a suspected cause of the service disruption on the US2 and US5 clusters and is taking steps to remediate the issue.
Impact: Customers may continue to experience connectivity issues as we bring the US2 and US5 cluster back into production.
Please report any possible related issues to Auvik support.
Next Steps: We are applying mitigation measures and will provide updates on progress.
identified
Our team has identified a suspected cause of the service disruption on the US2 and US5 clusters and is taking steps to remediate the issue.
Impact: Customers may continue to experience connectivity issues as we deliberately bring the US2 and US5 clusters back into production. The estimated time for recovery of these clusters has been extended.
Please report any possible related issues to Auvik support.
Next Steps: We are applying mitigation measures and will provide updates on progress.
monitoring
We have applied changes to address the issue on the US2 and US5 clusters. Site access should be restored. Services appear to be still recovering, but we are monitoring closely for stability.
Impact: Information for tenants in the US2 and US5 clusters is still experiencing a delay in the product.
If you continue to encounter problems, please report them to Auvik Support.
Next Steps: A final update will be posted once we confirm resolution.
monitoring
We have applied changes to address the issue on the US2 and US5 clusters. Site access is restored. Services appear to be still recovering, but we are monitoring closely for stability.
Impact: Information for tenants in the US2 and US5 clusters is still experiencing a delay in the product and will continue to recover. We will be monitoring the services into the evening.
If you continue to encounter problems, please report them to Auvik Support.
Next Steps: A final update will be posted once we confirm resolution.
monitoring
US2 and US5 clusters are running and stable.
The AU1 cluster had some slowness over the evening, which has been resolved.
Clients in the US4 cluster have had some reported lag in services that is currently under investigation.
Impact: US4 is experiencing possible slowness with load times and access to the lag.
If you continue to encounter problems, please report them to Auvik Support.
monitoring
We have implemented changes to address the outstanding issues, and services are currently operating as expected. As a precaution, all clusters are being closely monitored to ensure continued stability.
Impact:
Services should be functioning normally. If you continue to experience any problems, please contact Auvik Support.
Next Steps:
We will continue monitoring and will share any further updates if necessary.
monitoring
Several tenants in the EU2 cluster may have experienced an interruption in service, including Collector disconnects.
Impact:
Services should be returning to normal. If you continue to experience any problems, please contact Auvik Support.
Next Steps:
We will continue monitoring and will share any further updates if necessary.
identified
Our team has identified a suspected cause of the slowness and permission errors in EU2 and is taking steps to remediate the issue.
Impact:
Customers may continue to experience slowness and possible access to their sites.
Please report any related issues to Auvik Support so we can track and assist further.
Next Steps: We are applying mitigation measures and will provide updates on progress.
identified
The sites in the EU2 cluster are recovering.
Some sites on the US6 cluster are experiencing missing data in the UI (Maps, Devices, etc). This is being addressed.
Impact:
Customers may continue to experience slowness and may have limited access to their sites on the affected clusters.
Please report any related issues to Auvik Support so we can track and assist further.
monitoring
The sites in the EU2 cluster have recovered.
Some sites on the US6 cluster experienced missing data in the UI (e.g., Maps, Devices). This has also been addressed.
We will be performing rolling maintenance on sites on the AU1 cluster, during which the site may experience a momentary disconnection. Collectors may need to reconnect to Auvik.
Impact:
Customers may continue to experience slowness and may have limited access to their sites on the affected clusters.
Please report any related issues to Auvik Support so we can track and assist further.
monitoring
We have implemented changes to address the outstanding issues, and services are currently operating as expected. As a precaution, all clusters are being closely monitored to ensure continued stability. This will continue throughout the evening.
Impact:
Services should be functioning normally. If you continue to experience any issues, please contact Auvik Support.
Next Steps:
We will continue to monitor and share any further updates as necessary.
monitoring
We have implemented changes to address the outstanding issues, and services are currently operating as expected. As a precaution, all clusters are being closely monitored to ensure continued stability. This will continue throughout the day.
Impact:
Services should be functioning normally. If you continue to experience any issues, please contact Auvik Support.
Next Steps:
We will continue to monitor and share any further updates as necessary.
investigating
The incident has been fully resolved, and services are operating as expected.
Impact:
Customers should no longer experience any related issues. If you continue to experience problems, please report them to Auvik Support.
We will be posting an RCA as a follow-up.
resolved
The incident has been fully resolved, and services are operating as expected.
Impact:
Customers should no longer experience any related issues. If you continue to experience problems, please report them to Auvik Support.
We will be posting an RCA as a follow-up.
postmortem
# Service Disruption - Intermittent Availability & Performance Issues Across Multiple Clusters
## Root Cause Analysis
### Duration of the incident
Discovered: Aug 25, 2025 18:00 - UTC
Resolved: Aug 29, 2025 14:00 - UTC
### Cause
A configuration rollout unexpectedly generated a large number of configuration entries, which propagated across tenants. This resulted in excessive background processing and memory pressure in core services. The strain led to degraded performance, instability, and in some cases, brief service crashes across clusters.
### Effect
Customers experienced:
* Intermittent access and sign-in issues in several regions
* Slow page loads and missing/delayed alert notifications
* Errors or gaps in specific dashboard and visualization views
* Temporary unavailability for a small number of tenants
### Action taken
_All times are in UTC_
**08/25/2025**
**18:00** — Rollout halted after error rates increased.
**19:00** — Targeted service restarts restored partial availability.
**22:00** — Added backend capacity and began controlled rollouts.
**08/26–08/28/2025**
Continued staged rollouts with adjusted capacity.
Cleaned up configuration entries for affected tenants.
Tuned resource allocations for read/permissioning services
**08/29/2025**
**14:00** — All clusters stabilized; monitoring confirmed normal performance.
### Future consideration\(s\)
* Enhance autoscaling and resource thresholds for services under heavy background processing.
* Add scale-aware pre-deployment validation for configuration rollouts.
* Refine monitoring to surface customer-visible issues earlier.
* Expand operational runbooks for rollback and tenant recovery.
Site Dropdown and permission issues in the Auvik UI
開始 2025年8月25日 16:58 UTC · 0m
Pending
resolved
Auvik experienced an issue with the site dropdown missing sites and related permission issues in the Auvik UI. The event was an hour in length and has been remediated. All services are restored. If you have any problems related to this event, please open a ticket with Support.
postmortem
# Performance Degraded - Site dropdown not working for several clusters
## Root Cause Analysis
### Duration of the incident
Discovered: Aug 25, 2025 – 15:40 UTC
Resolved: Aug 25, 2025 – 16:45 UTC
### Cause
A recent system update introduced a change that depended on data not yet available in our production environment. As a result, the site selector \(drop-down\) was unable to load correctly, which prevented users from switching between sites until the update was rolled back.
### Effect
Customers were unable to switch sites in the platform, which limited access to some account information and historical data.
### Action taken
_All times are in UTC_
**08/25/2025**
**15:40** – Update deployed across production clusters.
**16:17** – Customer reports received that the site drop-down was not loading.
**16:26** – Engineering identified the change that introduced the issue.
16**:27** – Incident declared; response team assembled.
**16:40** – Deployment was rolled back.
**16:45** – Service restored; site drop-down working as expected. Incident resolved.
###
Future consideration\(s\)
* Add extra checks before updates to ensure all required data is available in production.
* Improve safeguards so that unrelated changes are not deployed together.
* Build fallback mechanisms to ensure that new features do not impact existing functionality.
* Strengthen monitoring to detect and alert on similar issues in the future quickly.