AWS Degraded – Multiple Services in US-East-1 Region
Started October 20, 2025 at 11:38 AM UTC · 11h 33m
OutageCritical incident
Affected components
RubyGems.org Gem Index APIec2-sa-east-1CloudFlare IAD - Ashburn, Virginia, USAAtlassian Bitbucket Git via HTTPSec2-ap-northeast-1ec2-ap-southeast-2ec2-us-east-1CloudFlare CloudFlare APIsec2-us-west-2CloudFlare SJC - San Jose, California, USAec2-us-west-1Help Centerec2-ap-southeast-1Engine Yard CloudAtlassian Bitbucket SSHRubyGems.org Gem Downloads
identified
AWS has confirmed increased error rates and latencies affecting multiple services in the US-EAST-1 region. The incident is currently classified as "Degraded" by AWS.
You may experience slowness, timeouts, or trouble accessing some parts of platform and services. We're closely monitoring the situation and will keep you updated as AWS provides more information.
resolved
AWS has confirmed that the outage in the US-EAST-1 region has now been resolved, and all services are operational.
403 Forbidden callback errors
Started August 28, 2025 at 3:15 AM UTC · 23h 33m
OutageMajor incident
Affected components
Engine Yard Cloud
investigating
We are currently investigating an issue causing 403 Forbidden callback errors, which are resulting in failed Chef runs and broken deployments. Our team is actively working to identify the root cause and restore full functionality.
monitoring
We have successfully implemented a fix for the issue that caused errors, and our systems are now operating normally. All impacted services are fully restored, and we have verified that the issue is no longer reproducible. We will continue to closely monitor system performance and provide another update within the next 24 hours
resolved
The incident has been resolved.
EngineYard Cloud login not working
Started August 4, 2025 at 9:23 AM UTC · 1h 49m
OutageMajor incident
Affected components
Engine Yard Cloud
investigating
We are currently investigating an issue related to the inability to login to https://login.engineyard.com/login.
resolved
Corrective actions were applied and the login to EngineYard Cloud should once again properly work.
eybackup process failing
Started May 26, 2025 at 5:51 AM UTC · 1d 23h
OutageMajor incident
Affected components
Engine Yard Cloud
investigating
We are currently investigating an issue with the eybackup process.
Our team is actively working to identify the root cause and implement a fix.
identified
We have identified the issue, and our team is working on a final fix.
monitoring
Our team has investigated and fixed the issue. If backups are still failing, please perform a Chef run and test again—this should resolve the issue.
resolved
This incident has been resolved.
Server provisioning failure - Failed to boot
Started May 15, 2025 at 9:29 PM UTC · 4d 3h
OutageMajor incident
Affected components
Engine Yard Cloud
investigating
We are currently investigating the issue where some customers are unable to make deployments with the “Failed to boot in” messages.
monitoring
A fix has been implemented and we are monitoring the situation.
resolved
The server provisioning failure issue has been resolved.
Access to backups using the UI not available
Started May 14, 2025 at 2:32 PM UTC · 1d 20h
OutageMajor incident
Affected components
Engine Yard Cloud
identified
We have been notified that certain UI functionalities, such as accessing backups, are currently degraded. We have identified the root of this issue and are working to have it addressed.
In the meantime, please reference this article to download your backups through a different method:
https://support.engineyard.com/article/45200-view-and-download-database-backups
monitoring
Our infrastructure team has implemented a fix. We continue to monitor the situation closely and will provide further updates as soon as possible. Thank you for your patience.
resolved
The backup-access issue using the UI has been successfully resolved
Service disruption preventing Engine Yard Cloud customers to execute certain actions affecting some AWS accounts managed by EngineYard.
Started January 11, 2024 at 10:20 PM UTC · 1d 0h
OutageMajor incident
Affected components
Engine Yard Cloud
identified
We're actively addressing a disruption affecting specific accounts, which is currently impacting deployment, change application, and SSH access. We've identified the cause and are working to implement a permanent fix.
identified
Dear Valued Customers,
We prioritize transparency and wish to keep you thoroughly informed during the ongoing service disruption. Please find the most recent update about our system's status below:
- Resolution Progress: Our technical experts have pinpointed the main cause of the issue. We are now working in close collaboration with Amazon Web Services (AWS) to expedite the restoration of our services.
- Anticipated Resolution Timeframe: We are hopeful that the service interruption will be brief. Our dedicated team is tirelessly working to reinstate our systems as swiftly as possible.
We acknowledge the significant effect this outage has on your operations and deeply regret any inconvenience it may cause. We are grateful for your patience and understanding during this period.
Be assured that we are diligently working to resume full services promptly and will continue to provide you with regular updates.
Thank you for your ongoing trust in the EngineYard service.
monitoring
Monitoring - We have addressed the underlying causes of the outage and are working on restoring functionality. We aim to progressively restore all service throughout the rest of the day, US hours.
We will continue to provide updates as the situation develops. As always, thank you for your continued patience.
resolved
Our systems are now operational and we expect no further issues. Regardless, to any customers who believe they are still affected by this incident please reach out to Support for immediate assistance.
Partial outage preventing some customers to access some EYK Containers managed by EngineYard.
Started January 9, 2024 at 9:31 PM UTC · 1d 0h
OutageMajor incident
identified
We have identified an issue affecting the access to EYK clusters. We have pinpointed the likely cause of the problem and are actively working towards a lasting solution. We will continue to update about our progress through this channel.
identified
Dear Valued Customers,
We are committed to keeping you fully informed during this service interruption. Here is the latest update regarding our system status:
- Current System Status: The EYK backend system is currently down. EYK clusters do present availability issues due to this. For EYC customers, on certain accounts, functionalities like deploying, applying changes, and accessing the instances via SSH continue to be unavailable. Work is ongoing to identify and fix the issues. We are aware of the inconvenience this causes and are working diligently to resolve it.
- Progress on Resolution: Our technical team has successfully identified the root cause of the issue. We are now collaborating closely with Amazon Web Services (AWS) to restore service as quickly as possible.
- Expected Timeline: We are optimistic that the delay in service will not be prolonged. Our team is fully engaged and working around the clock to bring our systems back online.
- Next Scheduled Update: We will provide our next update on January 10th at 10:00 AM EST. Please look out for our communication at that time for the latest information.
We understand the critical impact this disruption has on your operations and sincerely apologize for any inconvenience caused. Your patience and understanding during this time are greatly appreciated.
Rest assured, we are fully committed to restoring full service as soon as possible and will keep you updated every step of the way.
Thank you for your continued trust in the EngineYard product.
monitoring
We have addressed the underlying causes of the outage and are working on restoring functionality. We aim to progressively restore all service during the rest of the day, US hours.
We will continue to provide updates as the situation develops. As always, thank you for your continued patience.
resolved
The issue has been resolved.
Partial outage preventing customers to execute certain actions affecting some AWS accounts managed by EngineYard.
Started January 9, 2024 at 7:03 PM UTC · 1d 2h
OutageMajor incident
Affected components
Engine Yard Cloud
investigating
On certain accounts, functionalities like deploying, applying changes, and accessing the instances via SSH are currently unavailable. Work is ongoing to identify and fix the issues.
identified
We have identified the probable issue and we are working toward a permanent resolution. We will keep informing our customers via this channel.
identified
We have identified an issue affecting the access to EYK clusters. We have pinpointed the likely cause of the problem and are actively working towards a lasting solution. We will continue to update about our progress through this channel.
identified
Dear Valued Customers,
We are committed to keeping you fully informed during this service interruption. Here is the latest update regarding our system status:
- Current System Status: The EYK backend system is currently down. EYK clusters do present availability issues due to this. For EYC customers, on certain accounts, functionalities like deploying, applying changes, and accessing the instances via SSH continue to be unavailable. Work is ongoing to identify and fix the issues. We are aware of the inconvenience this causes and are working diligently to resolve it.
- Progress on Resolution: Our technical team has successfully identified the root cause of the issue. We are now collaborating closely with Amazon Web Services (AWS) to restore service as quickly as possible.
- Expected Timeline: We are optimistic that the delay in service will not be prolonged. Our team is fully engaged and working around the clock to bring our systems back online.
- Next Scheduled Update: We will provide our next update on January 10th at 10:00 AM EST. Please look out for our communication at that time for the latest information.
We understand the critical impact this disruption has on your operations and sincerely apologize for any inconvenience caused. Your patience and understanding during this time are greatly appreciated.
Rest assured, we are fully committed to restoring full service as soon as possible and will keep you updated every step of the way.
Thank you for your continued trust in the EngineYard product.
monitoring
We have addressed the underlying causes of the outage and are working on restoring functionality. We aim to progressively restore all service during the rest of the day, US hours.
We will continue to provide updates as the situation develops. As always, thank you for your continued patience.
resolved
The issue has been confirmed to be resolved.
Issues with deployment after updating to the latest version of the Engine Yard v7 stack
Started December 28, 2023 at 9:52 AM UTC · 6h 21m
OutageMajor incident
Affected components
Engine Yard Cloud
identified
An issue has been reported preventing deployment in the latest version of the Engine Yard v7 stack. We recommend that you refrain from upgrading when possible.
We are working on resolving this matter with the highest priority and will provide updates as they become available.
identified
The latest release causing the issue is withdrawn while we are looking for a fix to the observed behavior.
We are working on resolving this matter with the highest priority and will provide updates as they become available.
resolved
We are happy to confirm that the affected version has been removed. If you upgraded to it while it was available, please reach out to Customer Support so that our agents may roll back your environment to a working version.
Thank you for your patience and understanding.
Failure at running Chef
Started May 27, 2022 at 6:37 PM UTC · 2h 16m
IssuesMinor incident
Affected components
Engine Yard Cloud
identified
We Identified a problem on our backend when it comes to run Chef to configure instances. Staff on duty is currently working on this matter.
monitoring
A fix has been implemented and we are monitoring the status of Chef runs across the platform.
resolved
The deployed fix addressed the issue, allowing Chef runs to happen smoothly once again. As such we're marking the incident as Resolved. Feel free to open a ticket with Engine Yard Support for any questions, comments, or concerns that may arise.
AWS US-East-1 (N. Virginia) AZ issues impacting instances including EY platform instances
Started December 22, 2021 at 12:36 PM UTC · 12h 17m
OutageMajor incident
Affected components
ec2-us-east-1Engine Yard Cloud
investigating
We're seeing instance failures in one AZ in the N. Virginia AWS region. Customer environments and applications may be impacted by this at the current time. Engine Yard's own cloud.engineyard.com application has been impacted and at this time we are working on restoring service to this application. AWS have yet to official report any issues, but please monitor https://status.aws.amazon.com/ for updates.
identified
Functionality has now been restored to https://cloud.engineyard.com. AWS are seeing connectivity issues in one AZ in the US-East-1 (N. Virginia) region, which is causing instance reachability issues impacting both application functionality and instance control operations. Please open a ticket if you are having issues, and the issue can be monitored via https://status.aws.amazon.com/
monitoring
AWS have stabilised the impacted AZ and are restoring impacted instances. We are seeing the majority of impacted instances fully functional again. If you see any further issues please contact Support.
resolved
AWS has resolved the incident and we are not experiencing issues anymore. Please contact Support for any queries.
AWS US-West connectivity issues
Started December 15, 2021 at 4:11 PM UTC · 22m
IssuesMinor incident
Affected components
ec2-us-west-2ec2-us-west-1
identified
AWS are reporting internet connectivity issues in both US-West regions (N. California and N. Oregon), which may be impacting the reachablility of customer's applications and environments hosted in those regions. Amazon are currently investigating and working towards resolution, and the latest updates on the status of the issue can be viewed at https://status.aws.amazon.com/.
identified
We are continuing to work on a fix for this issue.
resolved
AWS have now resolved the internet connectivity issues in the two US-West regions and normal service has been resumed.
AWS operations failing
Started December 7, 2021 at 5:10 PM UTC · 16h 40m
IssuesMinor incident
Affected components
ec2-us-east-1
investigating
Amazon are reporting increased API error rates in the US-East-1 (North Virginia) region. Please hold off from making any infrastructure level changes to environments hosted in this region (instance restart, creation and termination) at this time. If you see any issues with your environments or applications please open a ticket.
resolved
Amazon have mitigated the issue causing the API errors and therefore the region is functioning as normal again and all Engine Yard infrastructure related operations should now run without issue.
Support portal sign-in issues
Started November 11, 2021 at 8:15 PM UTC · 1h 28m
OutageMajor incident
Affected components
Help Center
investigating
We are currently investigating issues with signing into our support portal https://support.cloud.engineyard.com/, as well as the failure to submit tickets via the "File a Ticket" functionality in the Engine Yard dashboard. On-going tickets can still be updated by replying to their corresponding email threads, but should you wish to raise a new ticket please contact us via the #engineyard IRC channel on irc.libera.chat (https://web.libera.chat/?nick=EYGuest%7C?#engineyard) or call us (+1-866-518-9273).
resolved
The issue causing the sign-in issues has now been identified and resolved, so all support portal functionality is now restored.
Let's Encrypt root certificate expiry
Started October 1, 2021 at 10:11 AM UTC · 3d 9h
Pending
monitoring
The expiry of the Let's Encrypt DST Root CA X3 certificate is causing failures for calls to servers running Let's Encrypt SSLs from client instances running the stable-v5 stack or lower. For more information and details of how to fix, please see this guide .
resolved
A fix has now been released on v5 and v6 stacks to resolve this issue, further information can be found at the following link
https://support.cloud.engineyard.com/hc/en-us/articles/4407649023515-Let-s-Encrypt-X3-root-certificate-expiration
Support portal not available
Started September 30, 2021 at 3:58 PM UTC · 40m
OutageMajor incident
Affected components
Help Center
identified
Due to an issue with SSL validations, the EY Support Portal may not be accessible. For urgent issues please call EY Support phone line.
Our staff on duty has identified the issue and is investigating a solution.
monitoring
A fix on the integration between EY Cloud and Zendesk has been put in place. The EY Support Portal should be accessible again.
Staff on duty continues to monitor the issue.
resolved
Integration between EYCloud and Zendesk is working properly. Monitoring, as well as customers, have not alerted of new occurrences of this case. This issue is deemed resolved.
Issue that prevents running chef currently blocking configuration changes and spinning up new instances on stacks up to V5
Started April 1, 2020 at 8:13 PM UTC · 1h 21m
OutageMajor incident
Affected components
Engine Yard Cloud
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem
**Incident Timeline:**
* 18:15 UTC: First customer reports issues with Chef _Apply_ runs on instances.
* 18:50 UTC: Investigation into the issue links the issue to failures with Portage \(Gentoo package manager\), specifically the failure of customer instances to access the platform’s Portage server.
* 20:15 UTC: Further testing has shown the issue to be wide-reaching and impacting all stacks \(stable v1 to v5\) aside from \(the Ubuntu based\) stable-v6. As such an Incident is created to inform customers.
* 21:00 UTC: The Portage server is identified and restarted, though remains inaccessible.
* 21:30 UTC: Portage server access is gained and direct investigation of connection failures from platform instances is performed.
* 21:50 UTC: Flushing of IPTables firewall rules on Portage server restores connectivity of customer instances to Portage.
* 22:00 UTC: Re-application of dynamically generated IPTables firewall rules is found to not impact connectivity.
* 22:15 UTC: Source of issue is tracked to a change in the format of the AWS IP Range list published shortly before the latest automatic update to the dynamically generated IPTables firewall rules on the Portage server. This change was found to be reverted in later published lists.
* 22:30 UTC: Incident is declared as Resolved.
**Incident Root Causes:**
* Engine Yard instances running stacks stable-v1 to stable-v5 run the Gentoo OS, which utilises Portage as its package manager. Engine Yard curates its own packages through a dedicated Portage server.
* Upon a Chef _Apply_ run instances synchronise the local Portage Tree from the Portage server.
* For security reasons the Portage servers is fire-walled from the wider internet, but needs to still allow access from AWS IP addresses.
* To keep this fire-walling up to date, the Portage server runs a regular task to download the [latest published AWS IP Ranges list](https://docs.aws.amazon.com/general/latest/gr/aws-ip-ranges.html) and utilise this to dynamically generate firewall rules, granting these ranges access.
* In the \(UTC\) evening of 1st April Amazon published a new IP list, which contained an additional field, not previously included in the list.
* This additional field led to a failure of the list to be parsed by the firewall rule generation script, leading to the Portage server no longer granting access to the AWS IP ranges, thus blocking access to Portage from customer instances.
**Incident Impact:**
Customers on stacks stable-v1 to stable-v5 saw Chef _Apply_ run failures on creation of new instances or configuration of existing instances when such runs required updating of the Portage Tree or installation of packages. This failure applied to all regions. The incident did impact the deployment or running of customer applications.
**Incident Corrective Actions:**
We have reached out to Amazon regarding the changes to the published IP Range list, with regards to why the list was published in such a state and if such changes will be published again in future.
Engine Yard will be undertaking a review of the dynamic fire-walling script, in order to prevent future parsing failures resulting in unwanted firewall lockdowns. We shall also be reviewing and improving internal documentation and knowledge sharing in order to improve investigation and resolution of any future Platform issues.
The Engine Yard Cloud Dashboard (https://cloud.engineyard.com/) is currently offline due to AWS issues in a particularly availability zone causing connectivity issues for a failing component. Actions are being undertaken to restore service at this time.
Customer's may also see issues of their own due to the connectivity issues which are affecting some instances in a single AZ of the US-EAST-1 region. If you see any issues please submit a support ticket via (https://support.cloud.engineyard.com) or should you be unable to login to your Engine Yard account please contact us in IRC (irc.freenode.net, channel #engineyard).
monitoring
The failing component of EY Cloud has now been replaced in order to restore connectivity and service, bringing the Cloud Dashboard back online. Any operations that were being attempted at the time of the failure should be checked and repeated as necessary.
AWS continue to face issues in the single US-East-1 AZ with instances being impaired and EBS volumes experiencing degraded performance, so some customers may still be impacted.
Should you see any issues or require further assistance please submit a support ticket.
resolved
Engine Yard hasn't seen any further occurrence of the problem. AWS has marked as resolved the issues affecting a single AZ on US-East-1.
As such, we're marking the issue as 'Resolved'.
postmortem
**Incident Timeline:**
* 13:20 UTC: Our Virtual NOC alerts us to the failure of the monitoring URLs for the Engine Yard Cloud Dashboard. At this point EY Cloud is returning errors and so is inaccessible for customers.
* 13:30 UTC: Investigation of the issue identifies the source as being a failed AWS instance, this instance is restarted at AWS in order to restore the instance. The restart hangs in the “stopping” state at AWS.
* 13:35 UTC: Amazon publish a Service Status message to inform customers that a connectivity issue is affecting some instances in a single Availability Zone in the US-East-1 Region.
* 13:35 - 14:45 UTC: Further attempts to restore the failed instance are unsuccessful due to the AWS issues, as are attempts to snapshot or reassign the instance’s EBS volume.
* 14:45 UTC: Another platform instance running in the Cloud Dashboard environment is repurposed in order to take the place of the failed instance.
* 15:00 UTC: Reconfiguration of the Cloud Dashboard application completes successfully and service is restored.
* 15:00 UTC and onwards: With the Cloud Dashboard restored Engine Yard Support engineers work with customers to replace affected instances with alternative instances in unaffected AWS AZs.
* 20:35 UTC: Amazon announce the issue is resolved and service is restored. Certain EC2 instances, RDS instances and EBS volumes remain in a failed state and cannot be restored by Amazon automatically, so EY Support continue to work with customers to restore these.
**Incident Root Causes:**
* **From AWS**: At 4:33 AM PDT one of ten data centers in one of the six Availability Zones in the US-EAST-1 Region saw a failure of utility power. Our backup generators came online immediately but began failing at around 6:00 AM PDT. This impacted 7.5% of EC2 instances and EBS volumes in the Availability Zone. Power was fully restored to the impacted data center at 7:45 AM PDT. By 10:45 AM PDT, all but 1% of instances had been recovered, and by 12:30 PM PDT only 0.5% of instances remained impaired. Since the beginning of the impact, we have been working to recover the remaining instances and volumes. A small number of remaining instances and volumes are hosted on hardware which was adversely affected by the loss of power. We continue to work to recover all affected instances and volumes and will be communicating to the remaining impacted customers via the Personal Health Dashboard. For immediate recovery, we recommend replacing any remaining affected instances or volumes if possible.
* **From Engine Yard:** The root cause of the issue was quickly identified to be the failure of an AWS instance, responsible for the running of a core component of the EY Cloud Platform . The inability for any other platform instances to communicate with this instance resulted in errors in the platform. It was initially desired to restore the existing instance, or failing that, utilise the snapshot to retain data, but the AWS issues prevented such actions, leaving the only option to be to repurpose another existing instance. Once this action was completed, service was restored.
**Incident Impact:**
For the period that the Engine Yard Cloud Dashboard was offline no customers were able to view or manage their environments through either the Dashboard or the API, so were unable to make environment or application changes. Running instances outside of the subset of failed instances in the single AZ of US-East-1 were unaffected, so the majority of customer applications were not impacted. For those with instances in the affected AZ, application impact was dependent on the role of the impacted instances, with database and application master instances resulting in application downtime, whilst slaves instances most likely not. EY Support staff worked with customers to restore failed instances where practically possible within the limitations of the AWS issues.
**Incident Corrective Actions:**
Engine Yard will be working to strengthen the platform environments in order to ensure the highest resilience across all components in order to minimise the disruption from any future infrastructure failures.
Inhability to boot and/or configure instances
Started August 15, 2019 at 9:31 PM UTC · 0m
Pending
resolved
On August 15, 2019 at 20:45:43 GMT, our monitoring system alerted that a core component of the platform was not responding. This component is responsible of providing information on stack versions and access to Chef recipes used to configure instances.
At 20:55:49 GMT our Support Team begun investigating the issue.
At 21:00:01 GMT it was identified that the application running this component couldn't connect to the database.
At 21:01:47 GMT it was detected that the db instance had restarted itself and was entering recovery mode.
At 21:02:48 GMT the component resumed services.
At 21:14:42 GMT, after both automatic and manual testing reported no failures, the issue was deemed as solved.