我们已经在Octopus Server 2026.3.4658中发现了一个问题,或者说,当现有的Bicep部署可能开始失败时。
我们有一个公开问题,用一个工作环绕https://github.com/OctopusDeploy/Issues/issues/10129追踪这个问题。
我们正在研究一个解决方案,一旦解决,将更新这一事件和公开问题.
resolved
The fix has been live across all Octopus Cloud servers for several weeks now without further issues.
自动翻译自官方事件更新。
Ability to deploy in West US 2 might be affected
开始时间 2026年5月29日 UTC 06:10 · 18h 28m
Issues轻微事件
受影响的组件
Octopus Cloud
monitoring
Our upstream provider, Azure in West US2, is currently experiencing an issue which might be affecting your ability to deploy.
We will resolve this incident once we have received confirmation the issue has been resolved from our upstream provider.
resolved
Our upstream provider, Azure in West US2, has resolved their issue. While the issue was ongoing, some Octopus Cloud instances in West US2 may have experienced brief downtime and network communication issues. For more details, see Azure's incident report (Tracking ID: GHRP-84G) at https://azure.status.microsoft/en-us/status/history/.
Amazon ECS Service Steps fail with 'Container.image contains invalid cha...
开始时间 2026年5月15日 UTC 12:36 · 2d 22h
Outage重大事件
受影响的组件
Octopus Cloud
identified
We have identified the issue and are working on a fix. If you are experiencing this please see the Github issue below for workarounds - https://github.com/OctopusDeploy/Issues/issues/10026
resolved
We have pushed out a fix for this issue and it should be rolling out to Octopus Cloud customers over the next few days.
Please contact [email protected] if you are seeing this issue and we can arrange to get your instance upgraded.
Dynamic Worker issues in West US 2
开始时间 2026年5月5日 UTC 13:18 · 1d 17h
Outage重大事件
受影响的组件
Octopus Cloud
identified
We have observed issues provisioning Dynamic Workers in our West US 2 region. We are working on mitigations and monitoring.
A small number of customers have experienced lease failures.
identified
We are in the process of switching Octopus Cloud instances in the West US 2 region to use dynamic workers in our standby region (Central US). Customers who have experienced leasing failures should retry their deployments or runbook runs.
monitoring
We have completed switching Octopus Cloud instances in the West US 2 region to use dynamic workers in our standby region (Central US). This should mitigate this issue for all customers.
We will update this incident when dynamic workers in West US 2 are stable and we have reverted all Cloud instances back to using workers in West US 2.
monitoring
We are continuing to monitor for any further issues.
resolved
We have not observed any further issues in the West US 2 region so all Octopus Cloud instances in the West US 2 region are using dynamic workers in West US 2 once again.
AWS Deployment Failures for Cloud Customers
开始时间 2026年4月1日 UTC 08:06 · 6d 20h
Outage重大事件
受影响的组件
Octopus Cloud
investigating
We are aware of an issue affecting Octopus Cloud customers running deployments that interact with AWS, including EKS/Kubernetes targets and AWS health checks. Impacted customers may see authentication errors when running these deployments.
Our team is actively investigating and working to resolve the issue.
Workaround:
If you are affected, adding your AWS region as an environment variable in your script steps may resolve the issue:
export AWS_REGION= for Bash
$env:AWS_REGION = " " for Powershell
E.g
export AWS_REGION=ap-southeast-2
$env:AWS_REGION = "ap-southeast-2"
We apologise for the disruption and will share further updates as our investigation progresses. If you need help in the meantime, please contact us at [email protected]
investigating
We are aware of an issue affecting Octopus Cloud customers running deployments that interact with AWS, including EKS/Kubernetes targets and AWS health checks. Impacted customers may see authentication errors when running these deployments.
Our team is actively investigating and working to resolve the issue.
Workaround:
If you are affected, adding your AWS region as an environment variable in your script steps may resolve the issue:
export AWS_REGION=[REGION] for Bash
$env:AWS_REGION = "[REGION]" for Powershell
E.g
export AWS_REGION=ap-southeast-2
$env:AWS_REGION = "ap-southeast-2"
We apologise for the disruption and will share further updates as our investigation progresses. If you need help in the meantime, please contact us at [email protected]
resolved
This incident has been resolved.
Deployments using AWS-dependent resources may fail with “No RegionEndpoint or ServiceURL configured.”
开始时间 2026年3月30日 UTC 12:43 · 0m
Outage重大事件
受影响的组件
Octopus Cloud
resolved
We have identified the cause of this issue and have produced a fix in Octopus version 2026.2.3825. Please email [email protected] if you are seeing AWS region endpoint errors in your deployments similar to the one in the title and we can discuss upgrade options.
gRPC port (8443) shows Octopus Cloud instance is Undergoing Maintenance
开始时间 2026年3月24日 UTC 04:13 · 1d 1h
Outage重大事件
受影响的组件
Octopus Cloud
investigating
We have identified the cause and are working to complete it as quickly as possible. We will provide another update once the issue is resolved.
identified
The issue has been identified and a fix is being implemented.
monitoring
We have a resolution for the issue, and it should be rolled out to customer instances during their next maintenance windows. If this blocks your deployments in the meantime, please reach out to support.
resolved
This incident has been resolved. Please don't hesitate to reach out to our support team if you are still having any related problems.
Emails are currently not being delivered
开始时间 2026年2月26日 UTC 01:51 · 3h 56m
Outage重大事件
受影响的组件
Control Center (billing.octopus.com)Octopus CloudSign-in/Sign-up (octopus.com)
identified
Our upstream email provider is currently experiencing issues delivering emails from Octopus. This impacts all emails, including email verification during authentication, Octopus Cloud subscription invitations, and billing notifications.
We're currently investigating and will update this page when we know more.
identified
We are continuing to work on a fix for this issue.
identified
Sign-in/sign-up emails have been restored. Note: emails will now be sent from [email protected]
resolved
Email services have now been restored. Please contact [email protected] if you encounter any issues with email delivery.
Ubuntu Dynamic Workers are failing to lease
开始时间 2026年2月2日 UTC 21:31 · 1d 0h
Outage严重事件
受影响的组件
Octopus Cloud
investigating
Octopus Cloud instances are failing to lease Ubuntu Dynamic Workers due to an issue with our upstream provider. We are currently investigating and working on mitigating the issue.
identified
Azure have identified the following issue on their side that affects the Dynamic Workers: Virtual Machines and dependent services - Service management issues in multiple regions. They are actively working to mitigate impact and expect it to be resolved by approximately 00:00 UTC. See Azure status page for additional details: https://azure.status.microsoft/en-gb/status
monitoring
Azure have rolled out a fix for this issue and we are seeing Dynamic Workers return to normal operation across all regions.
monitoring
We are continuing to monitor for any further issues.
resolved
Azure has resolved the issue that was causing our Dynamic Workers lease failures and we haven't seen any additional failures since yesterday.
postmortem
# Summary
On 2 Feb 2026, between 20:13:34 to 22:56:04 UTC, Octopus Cloud customers in `West US 2` and `West Europe` may have experienced failed deployments or failed runbook runs due to `Ubuntu Dynamic Worker` steps failing on Leasing timeout.
This disruption was caused by Azure failing to provision Virtual Machines across multiple regions - see Azure Issue `FNJ8-VQZ` on [Azure Status History](https://azure.status.microsoft/en-us/status/history/).
# Background
Octopus Cloud [Dynamic Workers](https://octopus.com/docs/octopus-cloud/dynamic-worker) are isolated virtual machines that we provide as part of our Octopus Cloud Subscription offering as a way to execute deployment and runbook steps and scripts, without needing to run on the Octopus Server or deployment targets themselves. Customers can use both Windows and Ubuntu Dynamic Workers.
Octopus provides a [dynamic worker pool](https://octopus.com/docs/infrastructure/workers/dynamic-worker-pools) of these virtual machine types from which, as required by your deployment/runbook steps, your Octopus Cloud will exclusively lease a freshly provisioned dynamic worker VM for a limited time.
## Dynamic Workers Lifecycle
1. **Provisioning** - a new Azure Virtual Machine is provisioned, using the requested [Dynamic Worker image](https://octopus.com/docs/infrastructure/workers/dynamic-worker-pools#dynamic-worker-images) \(Windows/Ubuntu with a set of pre-installed tools\).
2. **Pool** - the newly created Dynamic Worker \(VM\) is placed in a worker pool until an instance requests a Dynamic Worker in one of its deployments/runbooks steps. Each region \(US, Europe and Australia\) has different pools for Windows and Ubuntu workers, with additional standby pools that can be turned on in cases of a temporary outage in an Azure region \(see below\). Octopus Cloud continuously monitors the pools’ levels and provision new workers automatically to keep them full.
3. **Leasing** - When a [deployment/runbook that uses a Dynamic Worker](https://octopus.com/docs/infrastructure/workers#where-steps-run) starts, if the instance doesn’t have a leased worker already \(in which case, it will continue using this worker, extending the worker’s lease for the new run\), the Octopus Server will request a new worker from the appropriate pool. This worker will be exclusively leased to this instance until it is no longer needed \(i.e. wasn’t used for an hour\) or until its maximum lifespan is reached \(3 days by default\).
4. **Deletion** - After a worker is no longer needed or it has reached its maximum lifespan, it is considered expired and will be deleted by the system automatically. The next time that the same instance will require a worker, it will lease a new one from the pool \(see above\).
## System Resilience & Safeguards
Octopus Cloud implements multiple layers of protection to ensure Dynamic Workers’ availability and minimize customer impact in cases of service disruptions:
* **Pre-provisioned Worker Pools:** We maintain multiple [dynamic worker pools](https://octopus.com/docs/infrastructure/workers/dynamic-worker-pools) \(for the different virtual machine OS and sizes\) with ready-to-use workers to provide immediate availability when deployments/runbooks are triggered, rather than waiting for on-demand provisioning. This also provides a safety buffer in cases where we can’t provision new Dynamic Workers due to temporary outages.
* **Standby Services in Multiple Regions:** We maintain standby Dynamic Worker services in alternate Azure regions that we activate to provide continuity when a primary region experiences issues
# Key Timing
# Timeline and Impact
All dates and times below are in UTC
**Feb 2 2026:**
**18:52:** 1st failed Dynamic Worker provisioning - At this point, our Dynamic Workers Service continued supplying workers successfully from the pools. However, since we couldn’t provision new workers, the pools started depleting.
**20:13:** 1st Ubuntu Dynamic Worker lease failed in `West US2` Azure region \(once the pool was depleted\) - start of customer impact
**20:47:** Octopus on-call was paged after 3 Dynamic Worker lease requests failed in `West US2`. The on-call then started the incident to investigate the issue
**20:47-21:20**:
* Octopus engineers started setting up a Dynamic Workers Service on a different Azure region in the US to mitigate the issue.
* During the investigation we saw that Dynamic Workers were also failing to provision in the `West Europe` and `East Australia` Azure regions. At this point we realized that this was a multi-region outage in Azure and decided to open a support ticket with them.
* Octopus Engineers turned off non-essential services \(e.g. Instance Upgrades\) to preserve the available Dynamic Workers for customer use.
**21:24:** an on-call engineer opened a Sev A support ticket with Azure
**21:31:** We published the initial partial outage alert for Octopus Cloud on [https://status.octopus.com/](https://status.octopus.com/)
**21:33:** 1st Dynamic Worker lease failed in `West Europe`
**21:42:** Azure acknowledged the multi-region issue and reported that they are investigating it.
**22:56:** We saw the last Dynamic Worker lease failure. After this time, we were able to provision the required Virtual Machines and return the Dynamic Workers successfully to all lease requests.
**23:46:** After verifying that all Dynamic Worker pools have been restored and we didn’t see any additional provisioning failures, we updated the incident status to “Mitigated”, and updated the Status page.
**Feb 3 2026:**
**6:05:** Azure confirmed that the issue was fully resolved on their side.
**21:38:** Incident was resolved and Status page updated
# Technical Details
Octopus Cloud uses Azure Virtual Machines in order to supply Dynamic Workers to customers. During this Azure outage, we couldn’t provision new Azure Virtual Machines for our Ubuntu Dynamic Workers. Our pre-provisioned Dynamic Workers’ pools continued supplying Dynamic Workers for additional
* 1:21 hours in `West US 2` and
* 2:41 hours in `West Europe`
before they were depleted and customer requests for new Dynamic Workers started failing. It’s worth noting that `East Australia` customers were not impacted because the pools in this region didn’t deplete during the incident. Additionally, customers that already had a Dynamic Worker leased at the time of the outage were not impacted \(unless their Dynamic Worker expired so a new one was requested during the incident\).
We made preparations to switch to our Dynamic Workers Standby Service on different Azure regions. However, once we realized that this was a multi-region outage, we decided to revert the switch since it wouldn’t have resolved the issue.
Once we saw that we couldn’t supply Dynamic Workers from our standby regions, we turned off non-essential services \(e.g. Instance Upgrades\) to preserve the available Dynamic Workers for customer use.
Once Azure mitigated the issue on their side, our system recovered automatically and resumed providing new Dynamic Workers successfully.
# Remediation
Octopus takes service availability seriously. Despite the difficulty with upstream cloud provider outages, especially ones that are widespread across multiple regions, we fully review and remediate any outages that occur. We do this so that we’re continuously improving and maintaining the best possible service we can.
Following our post mortem, we identified the following improvements to our system to help identify and mitigate \(where possible\) similar issues earlier:
* **Page an on-call earlier -** add an alert to page the on-call when multiple Dynamic Workers in the same region fail to provision. This will allow us to detect similar incidents quicker and give us more time to mitigate the issue, before any Dynamic Worker leases fail and customers are impacted.
* **Improve our Dynamic Workers Incident Playbook** to identify multi-regions incidents quicker in order to engage Azure support earlier to resolve the root cause of similar incidents.
# Conclusion
We apologize to our customers for any disruption and inconvenience as a result of this incident.
We have started work on the identified remediations to ensure that we can detect similar incidents more quickly and reduce the impact on our customers as much as possible.
OctopusID signin intermittent for cloud customers
开始时间 2025年11月25日 UTC 12:34 · 2h 23m
Outage重大事件
受影响的组件
Sign-in/Sign-up (octopus.com)
investigating
We are currently investigating issues with cloud customers accessing their instance via OctopusID. For some customers, this issue seems to be auto resolving after around 5 minutes.
Whilst we are investigating the root cause, we recommend attempting to sign in whenever convenient incase you are able to access the instance.
monitoring
We have performed some changes that based on customer feedback, has mitigated this issue.
We are still monitoring to ensure this issue does not reoccur, but customers should be able to sign in via OctopusID again.
monitoring
After monitoring the platform following our earlier mitigation, we can confirm that the issue has not resurfaced. OctopusID sign-in is functioning normally, and the incident has been resolved.
resolved
Resolved.
postmortem
# Intermittent sign-in issues for Octopus Cloud instances
## Summary
Between November 25, at 10:24 pm AEST, and November 26, at 12:58 am AEST, customers experienced sporadic issues signing into Octopus Cloud instances. These issues stemmed from a database connection leak, which caused some sign-in requests to time out. We apologize for the inconvenience this has caused our customers and are taking steps to prevent it from happening again.
### Timings
Time to detection: 9hrs 41mins
Time to incident declaration: 9hrs 41mins
Time to resolution: 12hrs 15mins
## What happened?
\(All dates and times below are in AEST\)
### November 25
**12:43 pm** We deployed a change to our authorization service. This change introduced a bug resulting in database connection leaks. The connection leak only became apparent in high-traffic scenarios, which our current test suite doesn't replicate. As a result, our test suite didn't detect the issue.
**10:24 pm** A DevOps Support Engineer declared an incident after receiving reports from customers that they were having difficulty signing into Octopus Cloud instances.
**10:24 pm - 11:41 pm** Investigations showed that the issue was due to database connection timeouts.
**11:42 pm** Temporary mitigations implemented, including restarting services to release database connections.
**11:42 pm - 11:59 pm** Effects of mitigations observed and deemed successful. Incident marked as mitigated.
### November 26
**12:00 am - 12:57 am** Incident responders continue to watch systems.
**12:58 am** Mitigation steps considered successful, and the incident marked as resolved.
**6:55 am** Root cause of database connection timeouts identified as a connection leak.
**6:55 am - 8:33 am** Permanent fix implemented and deployed.
## Remediation and next steps
We deployed a new service version to remove the database connection leak. We have also conducted an incident review to identify process improvements to prevent this from occurring in the future. We continue to work towards improving the reliability and security of our authentication and authorization services.
GitHub upstream outage causing issues with VCS-backed elements
开始时间 2025年11月18日 UTC 21:11 · 3h 36m
Outage重大事件
identified
GitHub is experiencing issue with git operations, resulting in in-product git operations to be degraded or fail. We're continuing to monitor this upstream outage.
monitoring
GitHub reports that git operations have been restored. We will continue to monitor this situation.
resolved
We have observed no further errors stemming from this issue.
Octopus.com and billing portal may be unavailable in certain regions
开始时间 2025年10月29日 UTC 16:31 · 8h 42m
Outage严重事件
受影响的组件
Octopus GitHub AppControl Center (billing.octopus.com)Octopus.comOctopus CloudSign-in/Sign-up (octopus.com)
investigating
We are currently investigating the issue.
identified
It appears this outage is due to a DNS issue with our upstream provider. We are currently awaiting for mitigation from the upstream provider to restore Octopus.com and related services.
identified
Customers may notice intermittent access to Octopus.com and Octopus services. Our upstream provider has made progress to mitigate the issue.
We will continue to monitor the situation. We have contingency plans we are prepared to deploy if service recovery progress from the upstream provider mitigation stalls.
identified
As we continue to monitor Octopus.com and Octopus services, metrics are showing services becoming more reliably available.
Our upstream provider is currently estimating full or near full mitigation approximately 2 hours from now (approximately 23:20 UTC on 29 October 2025).
We will provide additional updates as the situation necessitates.
identified
As we continue to monitor Octopus.com and Octopus services, metrics are showing services becoming more reliably available.
We are still waiting for our upstream provider's resolution and expect to give another update approximately 1 hour from now (approximately 00:55 UTC on 30 October 2025).
We will continue to provide additional updates as the situation necessitates.
resolved
All Octopus services affected by this outage are now operational.
Docker Hub is down
开始时间 2025年10月20日 UTC 07:34 · 17h 37m
Outage重大事件
受影响的组件
Octopus.comOctopus Cloud
identified
Docker Hub experienced issues that prevented access to container images, resulting in deployment failures for projects relying on those images.
monitoring
Docker Hub experienced issues that prevented access to container images, resulting in deployment failures for projects relying on those images.
monitoring
Most of the Docker Hub services seem to be operational but the incident is still on-going.
monitoring
Most of the Docker Hub services seem to be operational and we are continuing to monitor the situation
resolved
Docker Hub have resolved their incident and we have not observed any further issues
Kubernetes deployments using Kubernetes agent may fail
开始时间 2025年10月12日 UTC 23:43 · 41m
Outage重大事件
identified
We have identified an issue where Kubernetes deployments using the Kubernetes agent may fail due to a failure to pull the octopusdeploy/kubernetes-agent-tools-base docker image.
monitoring
We have fixed the issue and pushed the missing tags to DockerHub. Kubernetes agent deployments should be working again.
resolved
This has been resolved. All missing tags have been published to DockerHub and we are seeing successful deployments using these new tags.
Docker Hub authentication issues
开始时间 2025年9月25日 UTC 01:06 · 3h 12m
Outage重大事件
受影响的组件
Octopus.comOctopus Cloud
identified
Docker Hub is experiencing an active incident: https://www.dockerstatus.com/pages/incident/533c6539221ae15e3f000031/68d47a2f93c09e05486d93a9
Octopus customers may experience issues pulling images from Docker Hub, including:
1. Octopus Server self-host images
2. Deployment process container images including execution containers and reference packages hosted by Docker
Workarounds are available, please contact [email protected].
monitoring
Docker has implemented a fix and we are monitoring.
resolved
The Docker Hub incident is resolved and we have not observed any ongoing issues.
postmortem
# **Incident Summary**
Docker Hub experienced authentication issues that prevented access to container images, resulting in deployment failures for projects relying on those images.
## **What We're Doing**
Following this incident, we've identified several improvements to enhance service resilience and customer communication:
### **Visibility & Monitoring**
* Communicate third-party service outages more clearly on our status page
* Explore in-app messaging to alert customers during service disruptions
### **Resilience & Good Practices**
* Review image retrieval reliability and address identified risks
* Share good practices with customers to help reduce impacts when Docker Hub is unavailable
Intermittent Azure authentication issues
开始时间 2025年9月15日 UTC 03:11 · 2h 54m
Outage重大事件
受影响的组件
Octopus Cloud
investigating
We are investigating reports of intermittent Azure authentication issues during deployments and runbook runs. We are working with Microsoft to resolve this issue.
monitoring
The downstream issue with running `az login` commands in Azure steps appears to be working again. We will be continuing to monitor this.
resolved
The downstream issue with Azure CLI login commands has been resolved by Azure.
(Resolved) Some Deployments/Runbooks in Octopus Cloud may be failing to start
开始时间 2025年9月4日 UTC 02:31 · 1d 1h
Issues轻微事件
受影响的组件
Octopus Cloud
monitoring
Symptom: The Deployment or Runbook appears to have started but has not produced any logs after an extended period of time.
Mitigation: Manually cancel the Deployment or Runbook, then re-run the task. Please note that cancellation may take approximately 10 minutes.
After the next maintenance window all affected tasks will be automatically cancelled.
Please reach out to https://octopus.com/support if you would like tasks to be automatically cancelled earlier.
resolved
This issue is resolved, please reach out to https://octopus.com/support if you are still experiencing deployments that fail to start.
octopus.com/downloads degraded
开始时间 2025年5月22日 UTC 23:35 · 3h 59m
Issues轻微事件
受影响的组件
Octopus.com
investigating
octopus.com/downloads top navigation bar is failing to load correctly. We are investigating the cause. Links to downloads from this website are still correct. Please reload the page if you cannot see the content.
identified
We have identified the issue with the top navigation bar loading and are working on a fix.
monitoring
We have shipped a change to fix the navigation bar issue and are monitoring to ensure it works in all regions.
resolved
We have fixed the issue with the navigation bar on octopus.com/downloads