eu2: Data Center - Germanyus5: Cloud - US Eastus7: Cloud - USap1: Cloud - Japanus2: Data Center - US Westeu1: Data Center - Netherlandsus1: Data Center - US Easteu4: Cloud - Europe
eu2: Data Center - Germanyus5: Cloud - US Eastus7: Cloud - USap1: Cloud - Japanus2: Data Center - US Westeu1: Data Center - Netherlandsus1: Data Center - US Easteu4: Cloud - Europe
On July 29, 2026, between 10:57 UTC and 11:53 UTC, customers in our us1: Data Center - US East, us2: Data Center - US West, eu1: Data Center - Netherlands, eu2: Data Center - Germany, eu4: Cloud - Europe, us5: Cloud - US East, us7: Cloud - US, and ap1: Cloud - Japan regions experienced degraded performance. This resulted in delayed model and workspace access, slow page load times, and slower logins across the platform, as well as delayed execution of CloudWorks™ integrations. All other regions were unaffected during this time.
**Root Cause**
The incident was triggered by an automated security software update deployed across our platform servers. This simultaneous update caused an unexpected, short-lived spike in storage activity that temporarily exceeded the storage systems' processing capacity. The resulting latency disrupted communication between internal metadata services and core system servers, causing the service connection pools to become unresponsive and preventing them from automatically recovering.
**Recovery**
Our engineering team identified the issue and took immediate action. We performed rolling restarts of the affected metadata services to clear the unresponsive connection pools and stabilize the systems. By 11:30 UTC, affected systems had recovered. Following this, the CloudWorks™ scheduler was restarted to process and clear the backlog of integration tasks. The issue was fully resolved by 11:53 UTC.
**Corrective & Preventative Actions**
We are implementing the following actions to prevent recurrence:
* We are implementing enhanced storage performance tiers and traffic-prioritization controls to isolate key platform workloads from other background system activities.
* We are updating our internal service connection frameworks to automatically detect and gracefully recover from unexpected system connection interruptions.
* We are refining our security software deployment processes to stagger rollouts and limit simultaneous resource utilization.
* We are enhancing our synthetic monitoring dashboards to improve visibility of regional service performance deviations.
* We are conducting rigorous connection-recovery and system testing in our lower environments to validate application resilience under loaded states.
We apologize for any impact this issue may have had on your business operations. We are continuously strengthening our systems and procedures to ensure we avoid future disruptions to your business and users.
If you have further questions or concerns, please visit our [Support website](https://www.google.com/url?q=https%3A%2F%2Fsupport.anaplan.com%2F). We appreciate your patience during this incident and value the trust you place in Anaplan.
虽然自协调世界时7月25日星期六17:24起全面恢复了进入加拿大(CA1)地区的一般通道,但使用Bring your Own Key(BYOK)加密的工作空间仍然暂时下线。 我们的主要重点是在受控制的环境中复制这种干扰,目前我们与我们的供应商合作。 这是一个关键阶段,因为它将直接影响并确定BYOK剩余工作空间最安全、最有效的恢复路径,同时确保系统稳定性和数据完整性。 我们期望在协调世界时下午3:00之前提供下一次关于恢复BYOK工作空间的有针对性的最新情况。
我们诚恳地为这次行动受到的干扰和影响而道歉。 我们的支助小组随时随时可以协助你处理任何问题或迫切需要.
investigating
大多数客户仍然完全恢复进入CA1地区。
使用“Bring your Onn Key(BYOK)加密”的工作空间仍然暂时关闭,
我们正与第三方加密伙伴积极合作,以了解原因和潜在的解决办法.
identified
大多数客户继续全面恢复进入CA1地区。
使用 Bring your Onn Key( BYOK) 加密的办公空间仍然暂时关闭,
我们正与第三方加密伙伴积极合作,以了解原因和潜在的解决办法.
identified
进入加拿大(CA1)地区的普遍通道仍然完全恢复. 然而,使用 Bring your Onn Key( BYOK) 加密的工作空间将暂时关闭。
将在协调世界时2026年7月27日10:00进一步提供BYOK恢复状态的最新情况.
我们为这种干扰道歉,并感谢你的持续耐心.
加拿大(CA1)地区的总务仍然完全恢复并正常运行。
对于剩下的"Bring your Own Key"(BYOK)工作空间,我们的团队继续执行多个平行的工作流程. 正在有意采取这一并行办法,以验证BYOK剩余工作空间最有效的恢复路径,同时确保绝对系统稳定性。
我们赞赏你的耐心。 我们的下一次更新将在两个小时内提供,或者说,如果达到一个重要的里程碑,我们将更早地提供.
**Summary**
On July 24, 2026, at 20:22 UTC, our engineering team began investigating an issue affecting customers in our ca1: Cloud — Canada region. Customers with affected workspaces were unable to open their models. As a precaution, while we validated the scope of the issue, we restricted access to the region behind a maintenance page at 02:18 UTC on July 25. General access to CA1 was restored at 17:24 UTC on July 25. Bring Your Own Key \(BYOK\) workspaces remained offline for additional safeguards, and were re-enabled progressively as those safeguards were validated. The incident was fully resolved on July 31, 2026, at 13:47 UTC.
**Root cause**
The disruption was caused by a defect in a third-party component used within our BYOK service. The defect only surfaced under a very specific combination of events occurring in a particular order on the same host. When a BYOK workspace was unloaded, the component failed tofully clear one of its local resources, leaving behind a stale reference. When a BYOK workspace was subsequently loaded onto the same host, the component attempted to clean up that stale reference before proceeding. During this step, it incorrectly executed a removal that extended beyond the stale reference and deleted files that were still in active use.
The affected files were captured in our regular backups. Once our engineering team identified the source of the activity, isolated it, and applied protective controls to stop any further impact, restoration became a controlled process of returning each affected file to its most recent backup.
**Recovery**
Our engineering team identified the issue and isolated it at its source, then worked systematically to restore the affected files. This allowed us to bring non-BYOK workspaces back online in a controlled sequence, and general access to CA1 was restored at 17:24 UTC on July 25.
We deliberately kept BYOK workspaces offline while our team worked with the third-party vendor to reproduce the trigger in a controlled, non-production environment. This reproduction gave us the diagnostic evidence the vendor needed to build a fix and to confirm the exact cause. It also allowed us to develop and validate our own temporary safeguards — targeted changes to how BYOK workspaces are scheduled — that eliminated the specific combination of conditions required to trigger the defect. These safeguards act as compensating controls to bring BYOK workspaces back online while the third party completes the permanent fix to the underlying component. We re-enabled BYOK workspaces once those safeguards were validated, ensuring the trigger conditions couldn't recur. The incident was fully resolved on July 31, 2026, at 13:47 UTC.
**Corrective and preventative actions**
Our corrective actions follow two complementary tracks. The first removes the specific combination of conditions required to trigger the defect, using controls we have developed and deployed ourselves as compensating safeguards. The second is the permanent fix to the underlying component itself, which the third party is delivering. Together, these tracks address both the trigger and the defect, so that neither can produce another incident of this kind. We are implementing the following actions to prevent recurrence:
* We have deployed changes that prevent the specific combination of conditions required to trigger the defect. This is a temporary but effective control that removes the trigger today, ahead of the permanent fix.
* We are working with the third party to deploy their validated fix. This removes the defect at its source and closes the underlying cause of this incident.
* We are strengthening how we validate BYOK third-party components in non-production before they reach production, including reproducing a wider range of workspace lifecycle scenarios and event sequences. This gives us stronger assurance to surface these types of issues in non-production and are addressed before they can affect customers.
* We have deployed dedicated alerting on the specific event pattern that triggered this incident and are actively reviewing additional file-level alerting. These alerts provide an additional safety net and earlier warning, enabling faster preventative action before customers are affected.
**Closing**
We apologize for any impact this issue may have had on your business operations. We are continuously strengthening our systems and procedures to ensure we avoid future disruptions to your business and users.
If you have further questions or concerns, please visit our [Support](https://www.google.com/url?q=https%3A%2F%2Fsupport.anaplan.com%2F) website. We appreciate your patience during this incident and value the trust you place in Anaplan.
We are currently investigating an issue resulting in some customers not being able to load models.
We are working to resolve this issue as quickly as possible and will provide updates every 30 minutes or upon resolution.
investigating
Thank you for your patience as we continue to investigate this issue. Currently, we do not yet have a time to resolution. We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
investigating
We are actively investigating a service disruption currently affecting customers' ability to open Models. This issue is also affecting CloudWorks Integrations.
Our engineering teams are prioritizing this issue and evaluating immediate mitigation steps to restore full service as quickly and safely as possible.
We do not yet have an estimated time to resolution, but will provide progress updates every 30 minutes or sooner.
identified
We are progressing with active mitigation steps to resolve the disruption affecting Model opening and CloudWorks integrations.
Initial reports indicate positive outcomes from these activities. Our engineering teams are closely monitoring system stability while we complete the remaining mitigation steps to ensure a full and durable recovery.
We will continue to provide updates every 30 minutes as we work to bring all systems back to standard operations.
identified
We are currently proceeding with the final remediation steps while continuing our investigation into the root cause. Customers should now be able to open models, though some may still encounter temporary delays; however, CloudWorks integrations are not yet completing. We are monitoring these recovery steps closely to ensure full system stability and will continue to provide updates every 30 minutes.
identified
We are pleased to report that all issues affecting model loading have been fully resolved, and normal access has been restored. Additionally, CloudWorks integrations are running again, and our teams are currently processing the accumulated backlog of queued jobs.
We are monitoring the queue progression closely to ensure all delayed integrations complete successfully. We will provide our next update in 30 minutes or once the backlog is fully cleared.
identified
Our teams confirm that the CloudWorks integration backlog is actively processing and recovering. Access to opening models remains fully restored and stable. We are continuing to monitor the integration queues closely as jobs complete, and we will provide our next status update in 30 minutes.
monitoring
Service has now been restored; you should now be able to resume normal activities.
We will continue to monitor the platform to ensure no additional issues arise. If you have any questions, concerns, or continue to experience issues, please do not hesitate to contact Anaplan Support. We will provide a final update to you when we consider this situation fully resolved.
resolved
We have confirmed that the issue is now resolved.
We deeply apologize for any impact this issue may have caused. We appreciate your patience and partnership as we worked through this issue.
We will follow up within 7 business days with a detailed root cause analysis (RCA) that will be shared on our Status Page. If you have any question or concerns, please do not hesitate to contact us at Anaplan Support.
postmortem
On June 22, 2026, at 07:20 UTC, we became aware of an issue affecting our us7: Cloud – US region. Customers experienced difficulties loading models, with impact beginning at approximately 05:12 UTC. CloudWorks™ integrations in the region were also unable to run during this period. Customers with active workspace sessions weren't affected. However, any workspace that was unloaded during this time couldn't be loaded until service was restored. Full service was restored at 10:21 UTC.
**Root cause**
A component that manages active workspace resources in the us7 region encountered an unexpected fault and restarted. On restart, the component entered a state where it appeared to be operating normally but couldn't process new workspace requests. Our systems normally recover automatically from this kind of state, but in this case the fault wasn't detected by our recovery processes. We intervened manually and performed a corrective restart of the component, which restored normal operation.
**Recovery**
We conducted a thorough investigation of the platform and identified the source of the issue. We performed a targeted reset of the resource scheduling services to clear the pending connections and restore normal communication. Once communication was re-established at 08:48 UTC, the platform began successfully assigning resources and loading models.
To handle the accumulated backlog of scheduled integrations, we scaled up the processing capacity for CloudWorks™. By 10:21 UTC, the backlog had finished processing, and the issue was fully resolved.
**Corrective and preventative actions**
We're implementing the following actions to prevent recurrence:
* We're deploying enhanced automated monitoring specifically designed to detect the condition seen in this incident, where a component appears operational but isn't processing requests. This closes the detection gap that extended the impact of this issue.
* We're developing automated self-healing for the resource scheduling component, so that a corrective restart can be performed automatically without engineering intervention. This directly addresses the failure mode that caused this incident.
* We're streamlining our scaling procedures for CloudWorks™, so that integration processing capacity can be expanded more rapidly during recovery. This shortens the time to clear integration backlogs following incidents like this one.
We apologize for any impact this issue may have had on your business operations. We are continuously strengthening our systems and procedures to ensure we avoid future disruptions to your business and users.
If you have further questions or concerns, please visit our [Support](https://support.anaplan.com/) website. We appreciate your patience during this incident and value the trust you place in Anaplan.
Platform Alerts
开始时间 2026年6月17日 UTC 07:02 · 3h 20m
Outage重大事件
受影响的组件
eu4: Cloud - Europe
investigating
We are currently investigating an issue resulting in some customers not being able to load models.
We are working to resolve this issue as quickly as possible and will provide updates every 30 minutes or upon resolution.
investigating
Thank you for your patience as we continue to investigate this issue. Currently, we do not yet have a time to resolution. We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
investigating
Thank you for your patience as we continue to investigate this issue.
Our technical teams are fully engaged, and restoring normal service operations as quickly and safely as possible is our absolute priority.
Currently, we do not yet have a time to resolution. We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
investigating
Thank you for your patience as we continue to investigate this issue.
Our coordinated technical response has made progress in narrowing down the scope of our investigation. We have isolated the primary area of concern to our storage configuration layer and are evaluating mitigation steps to alleviate the issue.
Currently, we do not yet have a time to resolution. We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
identified
We have identified the likely cause of the issue, and we are focused right now on restoring service as quickly as possible.
We have identified mitigation steps to alleviate the issue and initial reports indicate positive outcomes of these activities.
Currently, we do not yet have a time to resolution. We will provide further updates in 30 minutes or upon resolution.
identified
We are pleased to report a significant milestone in our resolution efforts. Our cross-functional engineering teams have successfully validated a targeted configuration adjustment to restore access to the affected workspaces.
We have initiated a phased deployment of this mitigation across the region to safely and systematically restore full access to all affected workspaces.
Currently, we do not yet have a time to resolution. We will provide further updates in 30 minutes or upon resolution.
monitoring
Service has now been restored; you should now be able to resume normal activities.
We will continue to monitor the platform to ensure no additional issues arise. If you have any questions, concerns, or continue to experience issues, please do not hesitate to contact Anaplan Support. We will provide a final update to you when we consider this situation fully resolved.
resolved
We have confirmed that the issue is now resolved.
We deeply apologize for any impact this issue may have caused. We appreciate your patience and partnership as we worked through this issue.
We will follow up within 7 business days with a detailed root cause analysis (RCA) that will be shared on our Status Page. If you have any question or concerns, please do not hesitate to contact us at Anaplan Support.
postmortem
On June 17, 2026, at 06:50 UTC, we became aware of an issue affecting a subset of workspaces in our eu4: Cloud - Europe region. Customers with affected workspaces were unable to load models, seeing persistent loading screens or errors when opening their work. CloudWorks™ integrations in the region also experienced a period of degradation, from approximately 05:20 UTC to 06:26 UTC, because of the same underlying issue. The impact was limited to a specific subset of workspaces — all other regions, and many workspaces in eu4: Cloud - Europe, continued to operate normally. Full service for the affected workspaces was restored at 10:00 UTC.
Root cause
The affected workspaces ran on a previous storage configuration in the eu4 region. Over time, a storage directory used by these workspaces accumulated a large number of small temporary files. The directory reached its operational limit and could no longer accept new file operations. This prevented models from loading for customers whose workspaces were running on this configuration.
Recovery
Our engineering team cleared the accumulated temporary files, which provided immediate relief. We then moved all affected workspaces from the old configuration to the new configuration through an automated process. We checked that model loading was restored through testing. By 10:00 UTC, the issue was fully resolved. On the same day, we implemented a change across the eu4 region to make sure no workspace can be directed to the old configuration.
Corrective and preventative actions
We've taken the following actions to address this incident and prevent it from happening again:
1. We've deployed an update across all regions that prevents any workspace from being directed to the older storage configuration. This eliminates the specific configuration condition that caused this incident.
2. We're decommissioning the older storage configuration. This permanently removes the configuration that caused this incident.
3. We’re strengthening our automated validation checks to confirm that workspaces are running on the correct configuration. This catches any configuration mismatch before it can cause customer impact.
4. We're implementing automated cleanup rules for the temporary files involved in this incident. This prevents storage directories from reaching operational limits and supports faster automatic recovery.
Closing
We apologize for the impact this issue has had on your operations. We're committed to the improvements outlined above to prevent similar disruptions. If you have questions or concerns, please contact [Support](https://support.anaplan.com/).
Platform Alerts
开始时间 2026年6月11日 UTC 17:20 · 21h 49m
Outage严重事件
受影响的组件
eu2: Data Center - Germanyus5: Cloud - US Eastus7: Cloud - USap1: Cloud - Japanus2: Data Center - US Westeu1: Data Center - Netherlandsus1: Data Center - US Easteu4: Cloud - Europe
investigating
We are currently investigating an issue impacting customers’ ability to access the Anaplan Platform.
We are working to resolve this issue as quickly as possible and will provide updates every 30 minutes or upon resolution.
identified
We want to sincerely apologize to our customers for the continued disruption to the Anaplan Platform today. We understand how critical access to Anaplan is for your business, and we deeply regret the impact this is having on your teams.
Our engineers have identified the root cause of this issue and are actively working to restore service. We are currently executing remediation steps and will continue to provide updates here every 30 minutes until full service is restored.
identified
We want to start by offering our sincere apologies to every customer affected by today's disruption to the Anaplan Platform. We know your teams rely on Anaplan to do critical work.
Our engineering team has identified the root cause and is fully focused on resolution. Active remediation steps are underway, and we are committed to keeping you informed with updates every 30 minutes — or sooner if there is meaningful progress to share.
monitoring
We are pleased to share that the core Anaplan Platform has successfully recovered and all primary services are fully operational. Our engineering teams are actively monitoring platform stability while corrective actions to address the underlying root cause continue to advance.
We sincerely apologize for the disruption today and recognize that this has been a difficult week for many of our customers. Restoring your trust and ensuring platform stability remains our highest priority.
Next update: In 30 minutes or sooner
monitoring
Our engineering team has successfully implemented a targeted infrastructure patch and completed a controlled failover to the updated version. All primary services are fully restored and accessible to customers.
Our teams remain in active monitoring as we continue to validate the patch and advance corrective work across the remaining infrastructure. We do not anticipate further disruption at this time, and we will continue to monitor platform performance closely over the next several hours.
We sincerely apologize for the impact today's incidents have had on your teams. A full incident summary will be published once our corrective actions are complete.
monitoring
The issue impacting access to the Anaplan Platform across all affected regions has been mitigated. The platform is stable and all services have returned to normal operation.
Corrective measures to prevent recurrence are actively advancing and will continue to be applied to ensure this issue does not reoccur. We are continuing to monitor the platform to confirm stability. We do not anticipate any further customer impact from these activities.
We appreciate your patience and partnership. If you have any questions or concerns, please do not hesitate to contact us at Anaplan Support.
resolved
We would like to provide you with the latest update regarding the recent service disruption affecting our platform.
Working in close collaboration with our vendor, our engineering team were able to identify a software bug in our network infrastructure that was causing temporary connection instability.
To resolve the instability, we applied the fix recommended by the Vendor for the software bug, then successfully moved platform traffic to updated infrastructure at 18:17 UTC on 11 June. All connections stabilised with no disruption observed post the transition.
Since then, the platform has remained stable and our engineering teams have continued to monitor this closely.
This is our final update on this incident. We will provide a thorough Root Cause Analysis (RCA) within 7 business days.
Thank you for your patience and understanding throughout this incident.
postmortem
**June 8–12, 2026 Platform Disruptions**
On June 8, 2026, at 11:05 UTC, our monitoring detected a brief drop in incoming traffic across the Anaplan platform, with automated monitoring checks failing across all affected regions \(us1: Data Center - US East, us2: Data Center - US West, eu1: Data Center - Netherlands, eu2: Data Center - Germany, eu4: Cloud - Europe, us5: Cloud - US East, us7: Cloud - US, and ap1: Cloud - Japan\). Over the four days, a series of related disruptions affected the same regions. Customers experienced intermittent difficulties logging in, opening models, and running integrations. The platform was fully stabilized on June 11, 2026, at 18:12 UTC.
These disruptions are linked to our previously communicated infrastructure modernization program. On the weekend of June 6, we migrated our control plane, the part of the platform that directs how traffic is routed between services. This was the most complex of the planned migration weekends, and it was completed successfully, as had the two previous migration weekends. The issues described in this report emerged in the days following that migration.
This report covers the seven linked incidents that occurred between June 8 and June 11, 2026.
**Root cause**
As part of our infrastructure modernization program, on the weekend of June 6, we completed a successful migration of our control plane. In the days that followed, we observed two discrete network-hardware issues that interacted to drive these disruptions.
_Issue 1: Media Access Control \(MAC\) flapping \(June 8–9\)_
Every device on a network has a MAC address, a unique hardware identifier that network switches use to route traffic to the correct destination. In our environment, the MAC addresses involved are virtual — assigned to logical network gateways rather than to fixed physical hardware — which allows them to legitimately move between hosts as part of normal operation. MAC flapping occurs when a switch sees the same MAC address appearing on two different physical ports in rapid succession, forcing it repeatedly to update its routing tables. This briefly slows or interrupts traffic. For customers, this surfaced as a brief, self-recovering instability. Short bursts of load and intermittent errors appeared and cleared on their own within minutes. The trigger for this behavior was the specific way live production traffic interacted with the new post-migration network. To resolve the issue, we scaled capacity, engaged our vendor, and deployed configuration changes to contain the impact.
_Issue 2: network card driver defect \(June 10–11\)_
Once the first set of mitigations were in place, we identified a second issue: A defect in the network card driver, which is the software that controls how a network card sends and receives data. The defect affected fewer than 0.005% of the cards in our estate. Those cards failed randomly and unpredictably, dropping traffic while still appearing healthy to our monitoring systems. This is known as a "gray failure" condition, because the affected components don’t flag themselves as broken. The cards sat in the part of the network that routes traffic between the regions impacted. This resulted in failures that cascaded across the platform and drove disruptions on June 10 and June 11.
The defect hadn't been observed in previous migrations or in any other environment, and there were no indicators in pre-deployment testing. The gray failure pattern was also what initially masked the defect as a load issue, until further investigation pointed us to the driver itself.
We engaged our vendor, who confirmed the defect. Working closely with them, we rapidly prepared, tested, and deployed the patch to the affected hosts on the evening of June 11, UTC. After that, the platform stabilized.
**Recovery**
From the first occurrence on June 8 through to permanent resolution on June 11, our engineering team led a continuous, round-the-clock response, working closely with our vendor to identify, diagnose, and resolve the underlying defect. We deployed configuration mitigations, repeatedly rerouted traffic between hosts and across alternative network paths, and scaled out capacity to restore service to customers as quickly as possible while the underlying defect was being addressed. The seven linked incidents and their impact windows were:
* **June 8, 2026, 11:00–11:20 UTC** — A short network disruption from the initial MAC flapping event caused intermittent login failures and slow page loads. For most customers, this appeared as a brief 6-minute blip and recovered automatically. Some basic authentication users experienced a longer impact and needed to start a new browser session, with full recovery by 11:20 UTC.
* **June 9, 2026, 11:05–11:29 UTC** — A recurrence of the same MAC flapping condition caused a similar brief disruption. Again, most customers experienced only a short blip of around 6 minutes, with basic authentication users seeing a longer impact. We applied scaling changes to address the issue.
* **June 10, 2026, 11:02–12:40 UTC** — Customers across all affected regions experienced a loss of access to the platform for approximately 98 minutes. We worked with the vendor and applied configuration changes to address the issue.
* **June 10–11, 2026, 23:06–00:04 UTC** — A disruption of approximately 58 minutes, caused when the network card defect affected the backup host that traffic had been moved to. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 03:43–05:20 UTC** — A disruption of approximately 97 minutes, as the next backup host was affected by the same defect. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 11:10–13:02 UTC** — A further recurrence of the network traffic surge caused customers to experience slow or failed access to the platform. We restored service by applying an underlying configuration change.
* **June 11, 2026, 17:04–17:48 UTC** — A final disruption of approximately 44 minutes during which we moved traffic to a host running the updated network card driver.
The platform was fully stabilized at 18:12 UTC on June 11, 2026, once traffic was successfully moved to a host running the updated network card driver. We then applied the same update to a second host for resilience. Monitoring continued through to midday Friday before the incident was closed.
CloudWorks™ experienced a backlog of queued jobs during and immediately after each disruption. Our engineering team scaled out CloudWorks capacity to accelerate processing, and the backlog was fully cleared shortly after each recovery.
**Corrective and preventative actions**
The control plane migration that preceded these disruptions was a one-time, foundational piece of work. The actions below reflect the learnings we are carrying forward from the event to further strengthen the platform.
1. The updated network card driver has been rolled out across all hosts matching the affected hardware profile, in every region. This removes the underlying defect from our infrastructure, even though it hasn't been observed in any other region.
2. The conditions that allowed MAC flapping to surface as customer-visible instability have been addressed at multiple layers. Configuration changes have been deployed in the affected regions. The underlying network configuration has been standardized across the wider production estate, and we have tuned the network behavior that triggered the initial event.
3. During the incident, our engineering team deployed dedicated alerting in real time for the specific network conditions causing the disruptions, enabling faster detection and intervention throughout the response. We are continuing to strengthen proactive monitoring and early-warning detection across post-release windows, so that emerging anomalies are surfaced and investigated before they escalate into customer-visible disruptions.
4. The remaining infrastructure update that supports the network card fix is being completed across our other environments. This brings the full benefit of the fix to every part of our infrastructure.
**What's next**
One final migration weekend is scheduled for June 20, 2026. After this, the migration phase of the infrastructure modernization program is complete.
**Closing**
We apologize for the impact this issue has had on your operations. We're committed to the improvements outlined above to prevent similar disruptions. If you have questions or concerns, please contact [Support](https://support.anaplan.com/).
Platform Alerts
开始时间 2026年6月11日 UTC 04:32 · 9h 59m
Outage严重事件
受影响的组件
eu2: Data Center - Germanyus5: Cloud - US Eastus7: Cloud - USap1: Cloud - Japanus2: Data Center - US Westeu1: Data Center - Netherlandsus1: Data Center - US Easteu4: Cloud - Europe
investigating
We are currently investigating an issue impacting customers’ ability to access the Anaplan Platform.
We are working to resolve this issue as quickly as possible and will provide updates every 30 minutes or upon resolution.
investigating
Thank you for your patience as we continue to investigate this issue. Currently, we do not yet have a time to resolution. We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
identified
The symptoms currently being observed are similar to the platform issue experienced earlier today. Our engineering and operations teams have been immediately mobilized and are actively diagnosing the root cause. We treat any recurrence with the utmost urgency and priority.
We sincerely apologize for this disruption, we treat any recurrence with the utmost urgency and priority. We will provide another update in 30 minutes, or sooner as we obtain more actionable technical details.
identified
We are beginning to see signs of platform recovery, and primary services are starting to stabilize. As systems come back online, a backlog of queued CloudWorks jobs has accumulated and is currently processing.
Our engineering teams are closely monitoring the stability of the platform and tracking the queue clearance velocity. We will provide our next update in 30 minutes, or sooner as we confirm continued stability.
monitoring
The platform has successfully recovered and is exhibiting stable performance. We are now focused on clearing the accumulated backlog of CloudWorks jobs.
Our engineering teams are actively managing and verifying the processing queue to ensure all delayed jobs complete as quickly and safely as possible.
We will continue to supervise the queue clearance and will provide our next update in 30 minutes, or sooner as we approach full restoration.
If you have any questions, concerns, or continue to experience issues, please do not hesitate to contact Anaplan Support. We will provide a final update to you when we consider this situation fully resolved.
monitoring
The platform remains stable and fully operational. We are continuing to process the remaining backlog of CloudWorks jobs.
Our engineering teams are actively overseeing the processing queues to ensure all jobs complete successfully and system performance remains steady. We appreciate your continued patience as we work through this remaining queue.
We will provide our next update in 30 minutes, or sooner as queue clearance nears completion.
monitoring
Service has now been restored; you should now be able to resume normal activities.
We will continue to monitor the platform to ensure no additional issues arise. If you have any questions, concerns, or continue to experience issues, please do not hesitate to contact Anaplan Support. We will provide a final update to you when we consider this situation fully resolved.
monitoring
We want to keep you informed with the latest update on the service disruption affecting our platform.
What we've done
We have identified the source of the instability and are actively working to resolve it. Our engineering team, in collaboration with our vendor, has already implemented a configuration change that has reduced the frequency of errors. We are monitoring the platform closely and have seen improvement as a result.
We are aware that some customers are experiencing delays and interruptions with their CloudWorks integrations as a direct result of this incident. Our engineering team is actively identifying and resolving any affected integration jobs.
Current status
The platform is operational. Integration processing is being actively monitored and remediated. We will not consider this incident closed until we are fully satisfied that the platform is stable, all integration jobs are running normally, and the risk of recurrence has been addressed.
We will now provide updates every 2 hours, as our investigation progresses and we work with our Vendor.
We understand the impact this has had on your operations and we appreciate your patience.
monitoring
We want to keep you informed with the latest update on the service disruption affecting our platform.
After collaboration with our Vendor, we have identified two fixes to our underlying infrastructure.
The first fix has now been implemented successfully without further disruption and initial monitoring is proving positive.
We are continuing to liaise with our Vendor for the secondary fix, to ensure no further disruption is caused.
The platform remains stable and all Cloudworks integrations are working as expected and the backlog has been cleared.
The incident will remain open and we will provide updates every 2 hours or sooner, as we continue to investigate and have full resolution.
We understand the impact this has had on your operations and we appreciate your patience
resolved
We want to provide you with the latest update regarding the recent service disruption affecting our platform.
All agreed-upon fixes have been successfully implemented across our underlying infrastructure. Our ongoing monitoring shows that system performance remains stable and positive.
Our engineering teams will continue to monitor the platform closely to ensure continued stability. Additionally, we are conducting thorough Root Cause Analysis (RCA) and this will be provided within 7 business days.
We sincerely understand the impact this disruption has had on your daily operations, and we deeply appreciate your patience and continued partnership as we work to ensure a reliable experience.
postmortem
**June 8–12, 2026 Platform Disruptions**
On June 8, 2026, at 11:05 UTC, our monitoring detected a brief drop in incoming traffic across the Anaplan platform, with automated monitoring checks failing across all affected regions \(us1: Data Center - US East, us2: Data Center - US West, eu1: Data Center - Netherlands, eu2: Data Center - Germany, eu4: Cloud - Europe, us5: Cloud - US East, us7: Cloud - US, and ap1: Cloud - Japan\). Over the four days, a series of related disruptions affected the same regions. Customers experienced intermittent difficulties logging in, opening models, and running integrations. The platform was fully stabilized on June 11, 2026, at 18:12 UTC.
These disruptions are linked to our previously communicated infrastructure modernization program. On the weekend of June 6, we migrated our control plane, the part of the platform that directs how traffic is routed between services. This was the most complex of the planned migration weekends, and it was completed successfully, as had the two previous migration weekends. The issues described in this report emerged in the days following that migration.
This report covers the seven linked incidents that occurred between June 8 and June 11, 2026.
**Root cause**
As part of our infrastructure modernization program, on the weekend of June 6, we completed a successful migration of our control plane. In the days that followed, we observed two discrete network-hardware issues that interacted to drive these disruptions.
_Issue 1: Media Access Control \(MAC\) flapping \(June 8–9\)_
Every device on a network has a MAC address, a unique hardware identifier that network switches use to route traffic to the correct destination. In our environment, the MAC addresses involved are virtual — assigned to logical network gateways rather than to fixed physical hardware — which allows them to legitimately move between hosts as part of normal operation. MAC flapping occurs when a switch sees the same MAC address appearing on two different physical ports in rapid succession, forcing it repeatedly to update its routing tables. This briefly slows or interrupts traffic. For customers, this surfaced as a brief, self-recovering instability. Short bursts of load and intermittent errors appeared and cleared on their own within minutes. The trigger for this behavior was the specific way live production traffic interacted with the new post-migration network. To resolve the issue, we scaled capacity, engaged our vendor, and deployed configuration changes to contain the impact.
_Issue 2: network card driver defect \(June 10–11\)_
Once the first set of mitigations were in place, we identified a second issue: A defect in the network card driver, which is the software that controls how a network card sends and receives data. The defect affected fewer than 0.005% of the cards in our estate. Those cards failed randomly and unpredictably, dropping traffic while still appearing healthy to our monitoring systems. This is known as a "gray failure" condition, because the affected components don’t flag themselves as broken. The cards sat in the part of the network that routes traffic between the regions impacted. This resulted in failures that cascaded across the platform and drove disruptions on June 10 and June 11.
The defect hadn't been observed in previous migrations or in any other environment, and there were no indicators in pre-deployment testing. The gray failure pattern was also what initially masked the defect as a load issue, until further investigation pointed us to the driver itself.
We engaged our vendor, who confirmed the defect. Working closely with them, we rapidly prepared, tested, and deployed the patch to the affected hosts on the evening of June 11, UTC. After that, the platform stabilized.
**Recovery**
From the first occurrence on June 8 through to permanent resolution on June 11, our engineering team led a continuous, round-the-clock response, working closely with our vendor to identify, diagnose, and resolve the underlying defect. We deployed configuration mitigations, repeatedly rerouted traffic between hosts and across alternative network paths, and scaled out capacity to restore service to customers as quickly as possible while the underlying defect was being addressed. The seven linked incidents and their impact windows were:
* **June 8, 2026, 11:00–11:20 UTC** — A short network disruption from the initial MAC flapping event caused intermittent login failures and slow page loads. For most customers, this appeared as a brief 6-minute blip and recovered automatically. Some basic authentication users experienced a longer impact and needed to start a new browser session, with full recovery by 11:20 UTC.
* **June 9, 2026, 11:05–11:29 UTC** — A recurrence of the same MAC flapping condition caused a similar brief disruption. Again, most customers experienced only a short blip of around 6 minutes, with basic authentication users seeing a longer impact. We applied scaling changes to address the issue.
* **June 10, 2026, 11:02–12:40 UTC** — Customers across all affected regions experienced a loss of access to the platform for approximately 98 minutes. We worked with the vendor and applied configuration changes to address the issue.
* **June 10–11, 2026, 23:06–00:04 UTC** — A disruption of approximately 58 minutes, caused when the network card defect affected the backup host that traffic had been moved to. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 03:43–05:20 UTC** — A disruption of approximately 97 minutes, as the next backup host was affected by the same defect. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 11:10–13:02 UTC** — A further recurrence of the network traffic surge caused customers to experience slow or failed access to the platform. We restored service by applying an underlying configuration change.
* **June 11, 2026, 17:04–17:48 UTC** — A final disruption of approximately 44 minutes during which we moved traffic to a host running the updated network card driver.
The platform was fully stabilized at 18:12 UTC on June 11, 2026, once traffic was successfully moved to a host running the updated network card driver. We then applied the same update to a second host for resilience. Monitoring continued through to midday Friday before the incident was closed.
CloudWorks™ experienced a backlog of queued jobs during and immediately after each disruption. Our engineering team scaled out CloudWorks capacity to accelerate processing, and the backlog was fully cleared shortly after each recovery.
**Corrective and preventative actions**
The control plane migration that preceded these disruptions was a one-time, foundational piece of work. The actions below reflect the learnings we are carrying forward from the event to further strengthen the platform.
1. The updated network card driver has been rolled out across all hosts matching the affected hardware profile, in every region. This removes the underlying defect from our infrastructure, even though it hasn't been observed in any other region.
2. The conditions that allowed MAC flapping to surface as customer-visible instability have been addressed at multiple layers. Configuration changes have been deployed in the affected regions. The underlying network configuration has been standardized across the wider production estate, and we have tuned the network behavior that triggered the initial event.
3. During the incident, our engineering team deployed dedicated alerting in real time for the specific network conditions causing the disruptions, enabling faster detection and intervention throughout the response. We are continuing to strengthen proactive monitoring and early-warning detection across post-release windows, so that emerging anomalies are surfaced and investigated before they escalate into customer-visible disruptions.
4. The remaining infrastructure update that supports the network card fix is being completed across our other environments. This brings the full benefit of the fix to every part of our infrastructure.
**What's next**
One final migration weekend is scheduled for June 20, 2026. After this, the migration phase of the infrastructure modernization program is complete.
**Closing**
We apologize for the impact this issue has had on your operations. We're committed to the improvements outlined above to prevent similar disruptions. If you have questions or concerns, please contact [Support](https://support.anaplan.com/).
Platform Alerts
开始时间 2026年6月10日 UTC 23:14 · 2h 38m
Outage重大事件
受影响的组件
eu2: Data Center - Germanyus5: Cloud - US Eastus7: Cloud - USap1: Cloud - Japanus2: Data Center - US Westeu1: Data Center - Netherlandsus1: Data Center - US Easteu4: Cloud - Europe
investigating
We are currently investigating an issue impacting customers’ ability to access the Anaplan Platform.
We are working to resolve this issue as quickly as possible and will provide updates every 30 minutes or upon resolution.
investigating
Thank you for your patience as we continue to investigate this issue. Currently, we do not yet have a time to resolution. We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
identified
We are starting to observe the first initial signs of platform recovery. Our engineering teams are proceeding with extreme caution and are closely monitoring system stability and telemetry as services begin to stabilize.
We remain actively engaged in verifying that this early recovery is sustained. We will provide our next update in 30 minutes, or sooner if we detect any changes in performance.
identified
The core platform has successfully recovered, and all primary services are fully operational. We are currently processing a backlog of queued jobs within Cloudworks resulting from the incident.
While this backlog is being processed, some customers may experience delays in processing times.
We will continue to track performance and will provide our next update in 30 minutes, or sooner as the backlog clears.
monitoring
We are pleased to report that the CloudWorks backlog is actively processing, and we are now seeing jobs successfully and steadily completing. System throughput has returned to normal operational levels as the queue continues to clear.
Our engineering teams remain focused on monitoring the queue velocity until the backlog is fully exhausted and all services have returned to a completely nominal state.
We will provide our next update in 30 minutes, or sooner once queue clearance is complete.
monitoring
We are currently processing a backlog of queued jobs within Cloudworks for US1 and US2 region.
While this backlog is being processed, some customers may experience delays in processing times.
We will continue to track performance and will provide our next update in 30 minutes, or sooner as the backlog clears.
resolved
We have confirmed that the issue is now resolved.
We deeply apologize for any impact this issue may have caused. We appreciate your patience and partnership as we worked through this issue.
We will follow up within 7 business days with a detailed root cause analysis (RCA) that will be shared on our Status Page. If you have any question or concerns, please do not hesitate to contact us at Anaplan Support.
postmortem
**June 8–12, 2026 Platform Disruptions**
On June 8, 2026, at 11:05 UTC, our monitoring detected a brief drop in incoming traffic across the Anaplan platform, with automated monitoring checks failing across all affected regions \(us1: Data Center - US East, us2: Data Center - US West, eu1: Data Center - Netherlands, eu2: Data Center - Germany, eu4: Cloud - Europe, us5: Cloud - US East, us7: Cloud - US, and ap1: Cloud - Japan\). Over the four days, a series of related disruptions affected the same regions. Customers experienced intermittent difficulties logging in, opening models, and running integrations. The platform was fully stabilized on June 11, 2026, at 18:12 UTC.
These disruptions are linked to our previously communicated infrastructure modernization program. On the weekend of June 6, we migrated our control plane, the part of the platform that directs how traffic is routed between services. This was the most complex of the planned migration weekends, and it was completed successfully, as had the two previous migration weekends. The issues described in this report emerged in the days following that migration.
This report covers the seven linked incidents that occurred between June 8 and June 11, 2026.
**Root cause**
As part of our infrastructure modernization program, on the weekend of June 6, we completed a successful migration of our control plane. In the days that followed, we observed two discrete network-hardware issues that interacted to drive these disruptions.
_Issue 1: Media Access Control \(MAC\) flapping \(June 8–9\)_
Every device on a network has a MAC address, a unique hardware identifier that network switches use to route traffic to the correct destination. In our environment, the MAC addresses involved are virtual — assigned to logical network gateways rather than to fixed physical hardware — which allows them to legitimately move between hosts as part of normal operation. MAC flapping occurs when a switch sees the same MAC address appearing on two different physical ports in rapid succession, forcing it repeatedly to update its routing tables. This briefly slows or interrupts traffic. For customers, this surfaced as a brief, self-recovering instability. Short bursts of load and intermittent errors appeared and cleared on their own within minutes. The trigger for this behavior was the specific way live production traffic interacted with the new post-migration network. To resolve the issue, we scaled capacity, engaged our vendor, and deployed configuration changes to contain the impact.
_Issue 2: network card driver defect \(June 10–11\)_
Once the first set of mitigations were in place, we identified a second issue: A defect in the network card driver, which is the software that controls how a network card sends and receives data. The defect affected fewer than 0.005% of the cards in our estate. Those cards failed randomly and unpredictably, dropping traffic while still appearing healthy to our monitoring systems. This is known as a "gray failure" condition, because the affected components don’t flag themselves as broken. The cards sat in the part of the network that routes traffic between the regions impacted. This resulted in failures that cascaded across the platform and drove disruptions on June 10 and June 11.
The defect hadn't been observed in previous migrations or in any other environment, and there were no indicators in pre-deployment testing. The gray failure pattern was also what initially masked the defect as a load issue, until further investigation pointed us to the driver itself.
We engaged our vendor, who confirmed the defect. Working closely with them, we rapidly prepared, tested, and deployed the patch to the affected hosts on the evening of June 11, UTC. After that, the platform stabilized.
**Recovery**
From the first occurrence on June 8 through to permanent resolution on June 11, our engineering team led a continuous, round-the-clock response, working closely with our vendor to identify, diagnose, and resolve the underlying defect. We deployed configuration mitigations, repeatedly rerouted traffic between hosts and across alternative network paths, and scaled out capacity to restore service to customers as quickly as possible while the underlying defect was being addressed. The seven linked incidents and their impact windows were:
* **June 8, 2026, 11:00–11:20 UTC** — A short network disruption from the initial MAC flapping event caused intermittent login failures and slow page loads. For most customers, this appeared as a brief 6-minute blip and recovered automatically. Some basic authentication users experienced a longer impact and needed to start a new browser session, with full recovery by 11:20 UTC.
* **June 9, 2026, 11:05–11:29 UTC** — A recurrence of the same MAC flapping condition caused a similar brief disruption. Again, most customers experienced only a short blip of around 6 minutes, with basic authentication users seeing a longer impact. We applied scaling changes to address the issue.
* **June 10, 2026, 11:02–12:40 UTC** — Customers across all affected regions experienced a loss of access to the platform for approximately 98 minutes. We worked with the vendor and applied configuration changes to address the issue.
* **June 10–11, 2026, 23:06–00:04 UTC** — A disruption of approximately 58 minutes, caused when the network card defect affected the backup host that traffic had been moved to. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 03:43–05:20 UTC** — A disruption of approximately 97 minutes, as the next backup host was affected by the same defect. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 11:10–13:02 UTC** — A further recurrence of the network traffic surge caused customers to experience slow or failed access to the platform. We restored service by applying an underlying configuration change.
* **June 11, 2026, 17:04–17:48 UTC** — A final disruption of approximately 44 minutes during which we moved traffic to a host running the updated network card driver.
The platform was fully stabilized at 18:12 UTC on June 11, 2026, once traffic was successfully moved to a host running the updated network card driver. We then applied the same update to a second host for resilience. Monitoring continued through to midday Friday before the incident was closed.
CloudWorks™ experienced a backlog of queued jobs during and immediately after each disruption. Our engineering team scaled out CloudWorks capacity to accelerate processing, and the backlog was fully cleared shortly after each recovery.
**Corrective and preventative actions**
The control plane migration that preceded these disruptions was a one-time, foundational piece of work. The actions below reflect the learnings we are carrying forward from the event to further strengthen the platform.
1. The updated network card driver has been rolled out across all hosts matching the affected hardware profile, in every region. This removes the underlying defect from our infrastructure, even though it hasn't been observed in any other region.
2. The conditions that allowed MAC flapping to surface as customer-visible instability have been addressed at multiple layers. Configuration changes have been deployed in the affected regions. The underlying network configuration has been standardized across the wider production estate, and we have tuned the network behavior that triggered the initial event.
3. During the incident, our engineering team deployed dedicated alerting in real time for the specific network conditions causing the disruptions, enabling faster detection and intervention throughout the response. We are continuing to strengthen proactive monitoring and early-warning detection across post-release windows, so that emerging anomalies are surfaced and investigated before they escalate into customer-visible disruptions.
4. The remaining infrastructure update that supports the network card fix is being completed across our other environments. This brings the full benefit of the fix to every part of our infrastructure.
**What's next**
One final migration weekend is scheduled for June 20, 2026. After this, the migration phase of the infrastructure modernization program is complete.
**Closing**
We apologize for the impact this issue has had on your operations. We're committed to the improvements outlined above to prevent similar disruptions. If you have questions or concerns, please contact [Support](https://support.anaplan.com/).
Platform Alerts
开始时间 2026年6月10日 UTC 11:11 · 5h 24m
Outage重大事件
受影响的组件
eu2: Data Center - Germanyus5: Cloud - US Eastus7: Cloud - USap1: Cloud - Japanus2: Data Center - US Westeu1: Data Center - Netherlandsus1: Data Center - US Easteu4: Cloud - Europe
investigating
We are currently investigating an active incident that may affect Basic Authentication access.
If you encounter issues when attempting to access the platform please use the following workaround while we resolve the issue:
- Close your current session— Close the active browser tab or open a new Incognito/Private window.
- Re-authenticate— Navigate to the portal and log in again with your credentials to establish a fresh session.
We are actively monitoring the platform and will continue to do so until the issue is fully resolved. If you are still experiencing difficulties or have any questions, please reach out to Anaplan Support
investigating
Since our last update, we have detected a further degradation in general platform performance. Our engineering teams are actively investigating the root cause with the highest priority, and we will provide further updates as the situation develops
Currently, we do not yet have a time to resolution. We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
investigating
Following our previous notification, the ongoing performance issues have escalated, resulting in broader platform-wide degradation. Our senior engineering teams have successfully identified the root cause of the issue and are currently executing targeted remediation steps to restore normal operations as our absolute highest priority.
Currently, we do not yet have a time to resolution. We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
investigating
Our engineering teams continue to actively investigate the incident while simultaneously implementing targeted remediation steps to stabilize the platform. We are closely monitoring the system's response to these actions and will provide our updates every 30 minutes as we work to resolve this issue as quickly as possible.
identified
The general platform has successfully recovered and is operating normally; however, we are actively addressing a residual issue impacting Cloudworks that is currently preventing integration jobs from completing. Our technical teams remain focused on resolving this remaining component to restore full service across all integrations as quickly as possible.
Currently, we do not yet have a time to resolution. We will provide further updates in 30 minutes or upon resolution.
identified
We are pleased to report that integration jobs are now successfully completing, and our systems are actively processing the accumulated backlog across all regions. To accelerate this recovery, we have scaled out our service capacity, with a particular focus on the us7: Cloud - US region to expedite the clearance of its more significant backlog.
We will provide further updates in 30 minutes or upon resolution.
identified
We are pleased to report that integration processing has returned to normal operational levels across all global regions, with the sole exception of us7: Cloud - US. The us7: Cloud - US region continues to steadily work through its remaining queue utilizing our expanded capacity, and we are monitoring progress closely to ensure a complete return to baseline performance.
Currently, we do not yet have a time to resolution. We will provide further updates in 30 minutes or upon resolution.
identified
As we continue our active recovery efforts, we have identified localized capacity constraints in the us1: Data Center - US East region that may intermittently affect customers attempting to load models. Our engineering teams are addressing this resource bottleneck to restore full service.
The CloudWorks backlog in us7: Cloud - US has finished processing, and the service has returned to normal operations.
While an estimated time to resolution is not yet established, our teams are treating this with the highest urgency. We will provide our next update in 30 minutes, or sooner if a resolution is reached.
investigating
Our investigation indicates that the backend connectivity issues in US1, stemming from a localized network disruption, are specifically isolated to the loading of large models in excess of 200GB. Customers attempting to load smaller models should experience normal performance and remain unaffected as our network engineering teams actively work to restore stable pathways for larger model payloads.
Currently, we do not yet have a time to resolution. We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
identified
Our network engineering teams continue to actively implement remediation steps to address the localized network issue in the us1: Data Center - US East region. We remain focused on resolving the backend connectivity disruptions affecting the loading of larger models, while smaller models continue to perform normally. We are closely monitoring system behavior as we deploy these targeted fixes and will provide our next update in 30 minutes, or sooner if significant progress is made.
Currently, we do not yet have a time to resolution. We will provide further updates in 30 minutes or upon resolution.
monitoring
Service has now been restored; you should now be able to resume normal activities.
We will continue to monitor the platform to ensure no additional issues arise. If you have any questions, concerns, or continue to experience issues, please do not hesitate to contact Anaplan Support. We will provide a final update to you when we consider this situation fully resolved.
resolved
We have confirmed that the issue is now resolved.
We deeply apologize for any impact this issue may have caused. We appreciate your patience and partnership as we worked through this issue.
We will follow up within 7 business days with a detailed root cause analysis (RCA) that will be shared on our Status Page. If you have any question or concerns, please do not hesitate to contact us at Anaplan Support.
postmortem
**June 8–12, 2026 Platform Disruptions**
On June 8, 2026, at 11:05 UTC, our monitoring detected a brief drop in incoming traffic across the Anaplan platform, with automated monitoring checks failing across all affected regions \(us1: Data Center - US East, us2: Data Center - US West, eu1: Data Center - Netherlands, eu2: Data Center - Germany, eu4: Cloud - Europe, us5: Cloud - US East, us7: Cloud - US, and ap1: Cloud - Japan\). Over the four days, a series of related disruptions affected the same regions. Customers experienced intermittent difficulties logging in, opening models, and running integrations. The platform was fully stabilized on June 11, 2026, at 18:12 UTC.
These disruptions are linked to our previously communicated infrastructure modernization program. On the weekend of June 6, we migrated our control plane, the part of the platform that directs how traffic is routed between services. This was the most complex of the planned migration weekends, and it was completed successfully, as had the two previous migration weekends. The issues described in this report emerged in the days following that migration.
This report covers the seven linked incidents that occurred between June 8 and June 11, 2026.
**Root cause**
As part of our infrastructure modernization program, on the weekend of June 6, we completed a successful migration of our control plane. In the days that followed, we observed two discrete network-hardware issues that interacted to drive these disruptions.
_Issue 1: Media Access Control \(MAC\) flapping \(June 8–9\)_
Every device on a network has a MAC address, a unique hardware identifier that network switches use to route traffic to the correct destination. In our environment, the MAC addresses involved are virtual — assigned to logical network gateways rather than to fixed physical hardware — which allows them to legitimately move between hosts as part of normal operation. MAC flapping occurs when a switch sees the same MAC address appearing on two different physical ports in rapid succession, forcing it repeatedly to update its routing tables. This briefly slows or interrupts traffic. For customers, this surfaced as a brief, self-recovering instability. Short bursts of load and intermittent errors appeared and cleared on their own within minutes. The trigger for this behavior was the specific way live production traffic interacted with the new post-migration network. To resolve the issue, we scaled capacity, engaged our vendor, and deployed configuration changes to contain the impact.
_Issue 2: network card driver defect \(June 10–11\)_
Once the first set of mitigations were in place, we identified a second issue: A defect in the network card driver, which is the software that controls how a network card sends and receives data. The defect affected fewer than 0.005% of the cards in our estate. Those cards failed randomly and unpredictably, dropping traffic while still appearing healthy to our monitoring systems. This is known as a "gray failure" condition, because the affected components don’t flag themselves as broken. The cards sat in the part of the network that routes traffic between the regions impacted. This resulted in failures that cascaded across the platform and drove disruptions on June 10 and June 11.
The defect hadn't been observed in previous migrations or in any other environment, and there were no indicators in pre-deployment testing. The gray failure pattern was also what initially masked the defect as a load issue, until further investigation pointed us to the driver itself.
We engaged our vendor, who confirmed the defect. Working closely with them, we rapidly prepared, tested, and deployed the patch to the affected hosts on the evening of June 11, UTC. After that, the platform stabilized.
**Recovery**
From the first occurrence on June 8 through to permanent resolution on June 11, our engineering team led a continuous, round-the-clock response, working closely with our vendor to identify, diagnose, and resolve the underlying defect. We deployed configuration mitigations, repeatedly rerouted traffic between hosts and across alternative network paths, and scaled out capacity to restore service to customers as quickly as possible while the underlying defect was being addressed. The seven linked incidents and their impact windows were:
* **June 8, 2026, 11:00–11:20 UTC** — A short network disruption from the initial MAC flapping event caused intermittent login failures and slow page loads. For most customers, this appeared as a brief 6-minute blip and recovered automatically. Some basic authentication users experienced a longer impact and needed to start a new browser session, with full recovery by 11:20 UTC.
* **June 9, 2026, 11:05–11:29 UTC** — A recurrence of the same MAC flapping condition caused a similar brief disruption. Again, most customers experienced only a short blip of around 6 minutes, with basic authentication users seeing a longer impact. We applied scaling changes to address the issue.
* **June 10, 2026, 11:02–12:40 UTC** — Customers across all affected regions experienced a loss of access to the platform for approximately 98 minutes. We worked with the vendor and applied configuration changes to address the issue.
* **June 10–11, 2026, 23:06–00:04 UTC** — A disruption of approximately 58 minutes, caused when the network card defect affected the backup host that traffic had been moved to. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 03:43–05:20 UTC** — A disruption of approximately 97 minutes, as the next backup host was affected by the same defect. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 11:10–13:02 UTC** — A further recurrence of the network traffic surge caused customers to experience slow or failed access to the platform. We restored service by applying an underlying configuration change.
* **June 11, 2026, 17:04–17:48 UTC** — A final disruption of approximately 44 minutes during which we moved traffic to a host running the updated network card driver.
The platform was fully stabilized at 18:12 UTC on June 11, 2026, once traffic was successfully moved to a host running the updated network card driver. We then applied the same update to a second host for resilience. Monitoring continued through to midday Friday before the incident was closed.
CloudWorks™ experienced a backlog of queued jobs during and immediately after each disruption. Our engineering team scaled out CloudWorks capacity to accelerate processing, and the backlog was fully cleared shortly after each recovery.
**Corrective and preventative actions**
The control plane migration that preceded these disruptions was a one-time, foundational piece of work. The actions below reflect the learnings we are carrying forward from the event to further strengthen the platform.
1. The updated network card driver has been rolled out across all hosts matching the affected hardware profile, in every region. This removes the underlying defect from our infrastructure, even though it hasn't been observed in any other region.
2. The conditions that allowed MAC flapping to surface as customer-visible instability have been addressed at multiple layers. Configuration changes have been deployed in the affected regions. The underlying network configuration has been standardized across the wider production estate, and we have tuned the network behavior that triggered the initial event.
3. During the incident, our engineering team deployed dedicated alerting in real time for the specific network conditions causing the disruptions, enabling faster detection and intervention throughout the response. We are continuing to strengthen proactive monitoring and early-warning detection across post-release windows, so that emerging anomalies are surfaced and investigated before they escalate into customer-visible disruptions.
4. The remaining infrastructure update that supports the network card fix is being completed across our other environments. This brings the full benefit of the fix to every part of our infrastructure.
**What's next**
One final migration weekend is scheduled for June 20, 2026. After this, the migration phase of the infrastructure modernization program is complete.
**Closing**
We apologize for the impact this issue has had on your operations. We're committed to the improvements outlined above to prevent similar disruptions. If you have questions or concerns, please contact [Support](https://support.anaplan.com/).
Platform Alerts
开始时间 2026年6月10日 UTC 09:48 · 53m
Issues轻微事件
受影响的组件
us1: Data Center - US East
investigating
We are currently investigating an issue impacting customers’ ability to run Cloudworks integrations.
We are working to resolve this issue as quickly as possible and will provide updates every 30 minutes or upon resolution.
monitoring
Service has now been restored; you should now be able to resume normal activities.
We will continue to monitor the platform to ensure no additional issues arise. If you have any questions, concerns, or continue to experience issues, please do not hesitate to contact Anaplan Support. We will provide a final update to you when we consider this situation fully resolved.
resolved
We have confirmed that the issue is now resolved.
We deeply apologize for any impact this issue may have caused. We appreciate your patience and partnership as we worked through this issue.
We will follow up within 7 business days with a detailed root cause analysis (RCA) that will be shared on our Status Page. If you have any question or concerns, please do not hesitate to contact us at Anaplan Support.
postmortem
**Introduction**
On June 10, 2026, at 09:48 UTC, customers in us1: Data Center - US East experienced slight delays in the processing and successful completion of CloudWorks™ integrations.
**Root cause**
During routine database patching on June 10, all regions completed without issue except for the us1: Data Center - US East region, where CloudWorks encountered an unexpected loss of database connectivity. This caused deadlocks within the scheduler service. Our investigation determined that one remote connection had been dropped while a second remained active, resulting in connection conflicts and errors. The issue was remediated by performing a rolling restart of the scheduler service and clearing the database rules.
**Recovery**
Upon identifying the issue, our engineering team acted quickly to initiate a rolling restart of the scheduler service across the affected region. To fully resolve the connection conflict and errors, the team cleared the internal routing and caching rules associated with the database gateway. Following this action, the database connection errors ceased immediately, and manual validation confirmed that integration scheduling had been fully restored. The issue was completely resolved by 10:42 UTC.
**Corrective and preventative actions**
We are implementing the following actions to prevent recurrence and improve service reliability:
* We are developing enhanced system alerts to immediately detect database connectivity issues and abnormal error volumes.
* We are working to replicate the scenario in non-production environments to strengthen CloudWorks' ability to recover from database disconnections.
* We are currently reviewing the errors encountered, along with the relevant system settings, to determine whether any changes can be implemented to better handle this type of issue in the future.
* We are reviewing and optimizing the capacity and sizing of our integration services across all regions to ensure they are robust and resilient.
* We are establishing more detailed incident response runbooks and centralizing triage procedures to accelerate resolution of similar issues.
* We are strengthening our platform observability and system telemetry to improve our diagnostic capabilities and detect connection issues more rapidly.
We apologize for any impact this issue may have had on your business operations.` `We are continuously strengthening our systems and procedures to ensure we avoid future disruptions to your business and users.
If you have further questions or concerns, please visit our [Support website](https://www.google.com/url?q=https%3A%2F%2Fsupport.anaplan.com%2F). We appreciate your patience during this incident and value the trust you place in Anaplan.
Platform Alerts
开始时间 2026年6月9日 UTC 11:14 · 13m
Issues轻微事件
受影响的组件
eu2: Data Center - Germanyus5: Cloud - US Eastus7: Cloud - USap1: Cloud - Japanus2: Data Center - US Westeu1: Data Center - Netherlandsus1: Data Center - US Easteu4: Cloud - Europe
investigating
We are currently investigating an active incident that may affect Basic Authentication access.
If you encounter errors when attempting to log in please do the following:
- Close your current session— Close the active browser tab or open a new Incognito/Private window.
- Re-authenticate— Navigate to the portal and log in again with your credentials to establish a fresh session.
We are actively monitoring the platform and will continue to do so until the issue is fully resolved. If you are still experiencing difficulties or have any questions, please reach out to Anaplan Support
We are working to resolve this issue as quickly as possible and will provide updates every 30 minutes or upon resolution.
resolved
We have confirmed that the issue is now resolved.
We deeply apologize for any impact this issue may have caused. We appreciate your patience and partnership as we worked through this issue.
We will follow up within 7 business days with a detailed root cause analysis (RCA) that will be shared on our Status Page. If you have any question or concerns, please do not hesitate to contact us at Anaplan Support.
postmortem
**June 8–12, 2026 Platform Disruptions**
On June 8, 2026, at 11:05 UTC, our monitoring detected a brief drop in incoming traffic across the Anaplan platform, with automated monitoring checks failing across all affected regions \(us1: Data Center - US East, us2: Data Center - US West, eu1: Data Center - Netherlands, eu2: Data Center - Germany, eu4: Cloud - Europe, us5: Cloud - US East, us7: Cloud - US, and ap1: Cloud - Japan\). Over the four days, a series of related disruptions affected the same regions. Customers experienced intermittent difficulties logging in, opening models, and running integrations. The platform was fully stabilized on June 11, 2026, at 18:12 UTC.
These disruptions are linked to our previously communicated infrastructure modernization program. On the weekend of June 6, we migrated our control plane, the part of the platform that directs how traffic is routed between services. This was the most complex of the planned migration weekends, and it was completed successfully, as had the two previous migration weekends. The issues described in this report emerged in the days following that migration.
This report covers the seven linked incidents that occurred between June 8 and June 11, 2026.
**Root cause**
As part of our infrastructure modernization program, on the weekend of June 6, we completed a successful migration of our control plane. In the days that followed, we observed two discrete network-hardware issues that interacted to drive these disruptions.
_Issue 1: Media Access Control \(MAC\) flapping \(June 8–9\)_
Every device on a network has a MAC address, a unique hardware identifier that network switches use to route traffic to the correct destination. In our environment, the MAC addresses involved are virtual — assigned to logical network gateways rather than to fixed physical hardware — which allows them to legitimately move between hosts as part of normal operation. MAC flapping occurs when a switch sees the same MAC address appearing on two different physical ports in rapid succession, forcing it repeatedly to update its routing tables. This briefly slows or interrupts traffic. For customers, this surfaced as a brief, self-recovering instability. Short bursts of load and intermittent errors appeared and cleared on their own within minutes. The trigger for this behavior was the specific way live production traffic interacted with the new post-migration network. To resolve the issue, we scaled capacity, engaged our vendor, and deployed configuration changes to contain the impact.
_Issue 2: network card driver defect \(June 10–11\)_
Once the first set of mitigations were in place, we identified a second issue: A defect in the network card driver, which is the software that controls how a network card sends and receives data. The defect affected fewer than 0.005% of the cards in our estate. Those cards failed randomly and unpredictably, dropping traffic while still appearing healthy to our monitoring systems. This is known as a "gray failure" condition, because the affected components don’t flag themselves as broken. The cards sat in the part of the network that routes traffic between the regions impacted. This resulted in failures that cascaded across the platform and drove disruptions on June 10 and June 11.
The defect hadn't been observed in previous migrations or in any other environment, and there were no indicators in pre-deployment testing. The gray failure pattern was also what initially masked the defect as a load issue, until further investigation pointed us to the driver itself.
We engaged our vendor, who confirmed the defect. Working closely with them, we rapidly prepared, tested, and deployed the patch to the affected hosts on the evening of June 11, UTC. After that, the platform stabilized.
**Recovery**
From the first occurrence on June 8 through to permanent resolution on June 11, our engineering team led a continuous, round-the-clock response, working closely with our vendor to identify, diagnose, and resolve the underlying defect. We deployed configuration mitigations, repeatedly rerouted traffic between hosts and across alternative network paths, and scaled out capacity to restore service to customers as quickly as possible while the underlying defect was being addressed. The seven linked incidents and their impact windows were:
* **June 8, 2026, 11:00–11:20 UTC** — A short network disruption from the initial MAC flapping event caused intermittent login failures and slow page loads. For most customers, this appeared as a brief 6-minute blip and recovered automatically. Some basic authentication users experienced a longer impact and needed to start a new browser session, with full recovery by 11:20 UTC.
* **June 9, 2026, 11:05–11:29 UTC** — A recurrence of the same MAC flapping condition caused a similar brief disruption. Again, most customers experienced only a short blip of around 6 minutes, with basic authentication users seeing a longer impact. We applied scaling changes to address the issue.
* **June 10, 2026, 11:02–12:40 UTC** — Customers across all affected regions experienced a loss of access to the platform for approximately 98 minutes. We worked with the vendor and applied configuration changes to address the issue.
* **June 10–11, 2026, 23:06–00:04 UTC** — A disruption of approximately 58 minutes, caused when the network card defect affected the backup host that traffic had been moved to. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 03:43–05:20 UTC** — A disruption of approximately 97 minutes, as the next backup host was affected by the same defect. We restored connectivity by moving traffic onto an alternative network path.
* **June 11, 2026, 11:10–13:02 UTC** — A further recurrence of the network traffic surge caused customers to experience slow or failed access to the platform. We restored service by applying an underlying configuration change.
* **June 11, 2026, 17:04–17:48 UTC** — A final disruption of approximately 44 minutes during which we moved traffic to a host running the updated network card driver.
The platform was fully stabilized at 18:12 UTC on June 11, 2026, once traffic was successfully moved to a host running the updated network card driver. We then applied the same update to a second host for resilience. Monitoring continued through to midday Friday before the incident was closed.
CloudWorks™ experienced a backlog of queued jobs during and immediately after each disruption. Our engineering team scaled out CloudWorks capacity to accelerate processing, and the backlog was fully cleared shortly after each recovery.
**Corrective and preventative actions**
The control plane migration that preceded these disruptions was a one-time, foundational piece of work. The actions below reflect the learnings we are carrying forward from the event to further strengthen the platform.
1. The updated network card driver has been rolled out across all hosts matching the affected hardware profile, in every region. This removes the underlying defect from our infrastructure, even though it hasn't been observed in any other region.
2. The conditions that allowed MAC flapping to surface as customer-visible instability have been addressed at multiple layers. Configuration changes have been deployed in the affected regions. The underlying network configuration has been standardized across the wider production estate, and we have tuned the network behavior that triggered the initial event.
3. During the incident, our engineering team deployed dedicated alerting in real time for the specific network conditions causing the disruptions, enabling faster detection and intervention throughout the response. We are continuing to strengthen proactive monitoring and early-warning detection across post-release windows, so that emerging anomalies are surfaced and investigated before they escalate into customer-visible disruptions.
4. The remaining infrastructure update that supports the network card fix is being completed across our other environments. This brings the full benefit of the fix to every part of our infrastructure.
**What's next**
One final migration weekend is scheduled for June 20, 2026. After this, the migration phase of the infrastructure modernization program is complete.
**Closing**
We apologize for the impact this issue has had on your operations. We're committed to the improvements outlined above to prevent similar disruptions. If you have questions or concerns, please contact [Support](https://support.anaplan.com/).
Platform Alerts
开始时间 2026年6月8日 UTC 13:25 · 2h 26m
Outage重大事件
受影响的组件
eu2: Data Center - Germanyus5: Cloud - US Eastus7: Cloud - USap1: Cloud - Japanus2: Data Center - US Westeu1: Data Center - Netherlands
investigating
We are currently investigating an issue impacting customers’ ability to run Cloudworks integrations.
We are working to resolve this issue as quickly as possible and will provide updates every 30 minutes or upon resolution.
investigating
Thank you for your patience as we continue to investigate this issue. Currently, we do not yet have a time to full resolution.
Access to the UI has been restored and we are currently investigating the degradation to integrations.
We will continue to provide updates every 30 minutes as we work to resolve this issue as quickly as possible.
investigating
We are currently investigating an issue impacting impacting Cloudworks.
We are working to resolve this issue as quickly as possible and will provide updates every 30 minutes or upon resolution.
monitoring
Service has now been restored; you should now be able to resume normal activities.
We will continue to monitor the platform to ensure no additional issues arise. If you have any questions, concerns, or continue to experience issues, please do not hesitate to contact Anaplan Support. We will provide a final update to you when we consider this situation fully resolved.
resolved
We have confirmed that the issue is now resolved.
We deeply apologize for any impact this issue may have caused. We appreciate your patience and partnership as we worked through this issue.
We will follow up within 7 business days with a detailed root cause analysis (RCA) that will be shared on our Status Page. If you have any question or concerns, please do not hesitate to contact us at Anaplan Support.
postmortem
**Introduction**
On June 8, 2026, at 13:14 UTC, customers in our ap1: Cloud - Japan, eu1: Data Center - Netherlands, eu2: Data Center - Germany, us2: Data Center - US West, us5: Cloud - US East, us7: Cloud - US, and eu4: Cloud - Europe regions experienced delays and difficulties accessing the CloudWorks™ user interface. During this period, users experienced slow UI access and a temporary inability to create new integrations or run certain scheduled jobs. Existing workflows were unaffected, although some experienced minor delays in completion.
**Root cause**
The root cause of this incident was an incorrect configuration setting where certain backend components did not automatically pick up updated hostname configurations following a scheduled system migration. Consequently, these components continued to attempt connections using the outdated configuration, leading to connection failures. This prevented users from accessing the user interface and successfully initiating new integration runs. The issue was resolved by reloading the updated configuration on the affected services.
**Recovery**
Our engineering team identified the issue and took immediate action. We initiated a rolling restart of the job management service across the affected regions, which restored user interface access. To fully resolve the connection errors, the engineering team performed a comprehensive restart of all related service components in every affected region, forcing them to load the updated database configurations. Following these restarts, database connection errors immediately ceased, and manual validation confirmed that both user-interface access and integration scheduling were fully restored and functioning normally. By 15:52 UTC, the issue was fully resolved.
**Corrective and preventative actions**
We are implementing the following actions to prevent recurrence and improve service reliability:
* We are developing enhanced system alerts to immediately detect database connectivity issues and abnormal error volumes.
* We are reviewing and optimizing the capacity and sizing of our integration services across all regions to ensure they are robust and resilient.
* We are establishing more detailed incident response runbooks and centralizing triage procedures to accelerate resolution of similar configuration mismatches.
* We are strengthening our platform observability and system telemetry to improve our diagnostic capabilities and detect connection issues more rapidly.
* We are refining our strategy and APIs to proactively detect and manage delayed integration workflows.
We apologize for any impact this issue may have had on your business operations. We are continuously strengthening our systems and procedures to ensure we avoid future disruptions to your business and users.
If you have further questions or concerns, please visit our [Support website](https://www.google.com/url?q=https%3A%2F%2Fsupport.anaplan.com%2F). We appreciate your patience during this incident and value the trust you place in Anaplan.