We are seeing elevated error rates for Real Device tests in EU-Central-1 data center. We are investigating.
identified
The root cause of the elevated error rates for Real Device tests in the EU-Central-1 data center has been identified. Remediation has been applied, and we are actively monitoring the situation.
resolved
After taking remedial action, Real Device test error rates have returned to normal in the EU-Central-1 data center. All services are fully operational.
2026-August-25 Service Incident
Started August 25, 2026 at 12:33 PM UTC · 46m
Pending
investigating
We are currently seeing elevated error rates for macOS and iOS tests in the US-West-1 data center. We are currently investigating.
resolved
After taking remedial action, all macOS and iOS tests are now performing as expected in the US-West-1 Data center. All services are fully operational.
postmortem
### **Dates:**
Tuesday August 25th 2026, 11:38 – 13:10 UTC
### **What happened:**
Customers running macOS and iOS tests in our US-West region experienced degraded service. Roughly 50% of the virtual Mac capacity in the region stopped accepting new tests, so tests either queued or failed to start. Remaining capacity came under additional pressure as work shifted onto it, which extended start times for both desktop and simulator tests.
### **Why it happened:**
An internal security certificate used by our Mac hosts to reach a supporting cloud service reached its expiry date. Once it lapsed, the hosts could no longer establish a trusted connection to that service and stopped provisioning new test machines. The certificate had been issued manually and had neither automated renewal nor expiry alerting.
### **How we fixed it:**
We issued and deployed a replacement certificate, which restored connectivity and returned Mac capacity to normal levels.
### **What we are doing to prevent it from happening again:**
We are moving these certificates onto automated renewal and adding alerting so they are replaced well ahead of expiry, along with reviewing the surrounding tooling to make certificate handling safer.
2026-August-07 Service Incident
Started August 7, 2026 at 12:08 PM UTC · 58m
Pending
investigating
We are currently seeing elevated error rates for virtual desktop tests in the US-West-1 datacenter. We are currently investigating.
resolved
After taking remedial action, Virtual Desktop tests are starting successfully in the US-West-1 Data center. All services are fully operational. We are closely monitoring the situation
postmortem
### **Dates:**
Friday, August 7th 2026, 11:00 UTC - 13:04 UTC.
### **What happened:**
Windows and Intel Mac jobs in `us-west1` failed to start due to virtual machine \(VM\) allocation starvation.
### **Why it happened:**
A service crash loop left VMs in an allocated but unclaimed state, while a cleanup bug prevented the system from releasing the orphaned capacity to boot new VMs.
### **How we fixed it:**
Restored service stability and cleared stale allocations to resume VM provisioning and clear queued jobs.
### **What we are doing to prevent it from happening again:**
Fixing allocator cleanup logic, strengthening deployment health checks, and improving capacity accounting for stale allocations.
2026-July-16 Resolved Service Incident - Error Reporting (Backtrace)
Started July 16, 2026 at 1:00 PM UTC · 0m
Pending
resolved
Between July 16th 13:03 UTC and July 16th 17:33 UTC, we experienced a service disruption where symbol archives utilizing multi-part upload protocols failed to process. Incident has been resolved and all services are operational.
postmortem
### **Dates:**
Thursday July 16th 2026, 13:03 – 17:33 UTC
### **What happened:**
Symbol archives uploaded in multiple parts failed to process. All regions were affected.
### **Why it happened:**
Our symbol processing service sent an upload-verification field that our cloud storage provider's API does not accept for multi-part uploads, so those uploads were rejected. A fix for this had already been developed, but it had not yet been included in a released build and the service was not configured to use it. As in the first incident, rejected uploads were retried and accumulated on local disk.
### **How we fixed it:**
We deployed a build containing the fix and corrected the service configuration on the affected workers. A large backlog of uploads then processed, which briefly re-filled the disks before draining completely.
### **What we are doing to prevent it from happening again:**
We are releasing the fix formally through our build pipeline and persisting the corrected configuration in our configuration management, so it can't be lost. The disk-utilization alerting added after the first incident also covers the disk-exhaustion pattern common to both.
2026-July-15 Resolved Service Incident - Error Reporting (Backtrace)
Started July 14, 2026 at 11:30 PM UTC · 0m
Pending
resolved
Between July 14th 23:45 UTC and July 15th 23:05 UTC, we experienced an issue where symbol archive uploads to Backtrace projects failed with HTTP 400 errors. Incident has been resolved and all services are operational.
postmortem
### **Dates:**
Tuesday July 14th 2026, 23:45 UTC – Wednesday July 15th 2026, 23:05 UTC
### **What happened:**
Symbol archive uploads to Error Reporting \(Backtrace\) projects failed with HTTP 400 errors. All regions were affected.
### **Why it happened:**
The credentials our symbol processing service used to write to cloud storage were no longer valid, so uploads could not be stored. Failed uploads were retried repeatedly and accumulated on local disk until the service ran out of space, at which point it also began rejecting new uploads.
### **How we fixed it:**
We reissued the storage credentials, increased disk capacity on the affected workers, and restarted the service. The queued uploads then processed successfully, and we confirmed recovery with affected customers.
### **What we are doing to prevent it from happening again:**
We've added disk-utilization monitoring and alerting to this service so we detect the condition ourselves before it affects uploads, rather than relying on customer reports.
2026-July-14 Service Incident
Started July 14, 2026 at 5:19 PM UTC · 3h 4m
OutageMajor incident
Affected components
EU-CentralEU-Central
investigating
We are experiencing an issue in the EU Central 1 data center, where iOS ARM Simulator app tests fail with infrastructure errors. We are investigating.
investigating
App downloads are taking longer than expected when these infrastructure errors are seen. We are continuing to investigate.
monitoring
We have deployed a fix for this issue. Tests should now be performing as expected in the EU Central 1 data center. We are monitoring.
monitoring
We are continuing to monitor for any further issues.
resolved
After taking remedial action and monitoring the situation, we are seeing iOS ARM application tests perform as expected in the EU Data Center. All services are fully operational.
postmortem
### **Dates:**
Tuesday July 14th 2026, 09:00 - 18:28 UTC
### **What happened:**
iOS ARM tests running in our EU data center failed.
### **Why it happened:**
A fault occurred with our primary network provider in our EU data center.
### **How we fixed it:**
We failed over to a secondary network provider from our EU data center.
### **What we are doing to prevent it from happening again:**
We've improved our monitoring and alerting to catch issues with third party network providers.
2026-July-1 Resolved Service Incident 1
Started July 1, 2026 at 5:02 PM UTC · 0m
IssuesMinor incident
resolved
On July 1st between 17:02 UTC and 21:31 UTC, macOS 14 tests were failing to start in the US-West and EU-Central data centers. This issue has been resolved. All services are fully operational.
postmortem
### **Dates:**
Wednesday July 1st 2026, 17:02 UTC - 21:31 UTC
### **What happened:**
macOS 14 tests in the US West and EU Central data centers were unable to start.
### **Why it happened:**
An internal datasource was unavailable due to a missing configuration entry.
### **How we fixed it:**
The missing entry was replaced.
### **What we are doing to prevent it from happening again:**
Safeguards have been put in place to prevent in-use datasource removal from configuration.
2026-July-01 Service Incident
Started July 1, 2026 at 12:12 PM UTC · 46m
Pending
investigating
We are experiencing an issue with accessing Appium Inspector in the US-West-1, EU-Central-1, and US-East-4 data centers. We are investigating.
resolved
We have identified the root cause and have deployed a fix for this issue. All services are fully operational.
postmortem
### **Dates:**
Wednesday July 1st 2026, 09:37 – 12:46 UTC
### **What happened:**
The Appium Inspector feature, used during real device live testing sessions, became unavailable after a scheduled UI deployment. Users who attempted to use the feature during a live test were presented with a 500 error, disrupting their active testing session. Automated test pipelines were not affected.
### **Why it happened:**
A routine upgrade of a core frontend library introduced an incompatibility with the Appium Inspector's rendering logic. The previous library version tolerated the pattern, but the updated version did not. The issue was not caught before deployment due to insufficient end-to-end test coverage for this specific feature.
### **How we fixed it:**
We rolled back the UI to the last known working version to restore the feature immediately, and then deployed a targeted fix for the incompatibility.
### **What we are doing to prevent it from happening again:**
We are adding end-to-end test coverage for the Appium Inspector feature to ensure it is validated automatically before future deployments.
2026-June-24 Resolved Service Incident
Started June 24, 2026 at 4:39 PM UTC · 0m
Pending
resolved
On June 24th between 08:28 UTC and 15:32 UTC, we experienced Real Device test session failures impacting Appium and Access API in the US East, US West and EU data centers. This issue has been resolved. All services are fully operational.
postmortem
### **Dates:**
Wednesday, June 24th 2026, 09:42 UTC - 15:45 UTC.
### **What happened:**
Real Device test sessions using Appium and Access API experienced increased error rates in the US East, US West, and EU data centers.
Sessions were timing out after approximately 90 seconds or becoming stuck in the "Connecting" state, preventing them from being closed.
### **Why it happened:**
A product defect was introduced, causing Real Device test sessions to fail to start and preventing active sessions from being closed.
### **How we fixed it:**
Rollback to a stable version.
### **What we are doing to prevent it from happening again:**
Improve monitoring & alerting and enhance post deployment validation.
2026-June-18 Service Incident
Started June 18, 2026 at 5:30 AM UTC · 1h 29m
OutageMajor incident
Affected components
EU-CentralEU-Central
investigating
Around 3:51 AM UTC we started experiencing lower iOS device availability in the EU-Central data center. Our team is actively investigating the root cause and working toward a resolution.
monitoring
We have been experiencing device unavailability in our EU-Central-1 data center and have identified the root cause. We have taken remedial action and are currently monitoring.
resolved
After taking remedial action, we are now seeing real devices available for testing in the EU-Central-1 data center. All services are fully operational.
postmortem
### **Dates:**
Thursday June 18 2026, 03:45 UTC - 06:45 UTC.
### **What happened:**
Approximately 7% of iOS devices in our EU data center were temporarily unavailable for customer test sessions after a loss of power to the rack hosting them.
### **Why it happened:**
The power circuit feeding the affected rack exceeded its capacity, and a protective breaker tripped to safeguard the line - cutting power to the devices on that rack until the circuit was restored.
### **How we fixed it:**
The affected circuit was reset and power restored to the rack, returning the devices to customer service.
### **What we are doing to prevent it from happening again:**
We are redistributing power load across affected racks, adding capacity monitoring with early-warning alerts ahead of circuit limits, and introducing a capacity review before new devices are deployed to a rack.
2026-June-1 Service Incident
Started June 1, 2026 at 2:02 PM UTC · 4h 22m
OutageCritical incident
Affected components
US-EastUS-EastUS-EastUS-EastUS-EastUS-EastUS-East
investigating
We are experiencing an issue where the US-EAST data center is currently unavailable. We are investigating.
investigating
We are continuing to investigate this issue.
investigating
We have identified the root cause and are working on implementing a fix.
monitoring
Access has been restored to the US-EAST data center. We are monitoring.
resolved
After taking remedial action, access to the US-EAST data center is fully restored. This incident is resolved.
postmortem
### **Dates:**
Monday, June 1st 2026, 13:06 UTC - 15:42 UTC.
### **What happened:**
Customers served by the US-EAST-4 region were unable to authenticate or start new test sessions because an incomplete TLS certificate chain was deployed to the core directory services.
### **Why it happened:**
A certificate extraction script defect silently truncated the certificate chain after the certificate authority transitioned to a longer hierarchy.
### **How we fixed it:**
Reverted the certificate rotation and re-applied the previous known ,good certificate to restore authentication services.
### **What we are doing to prevent it from happening again:**
Implementing pre-deployment certificate chain validation, adding active monitoring, and fixing the script's chain-length limitations.
2026-May-13 Resolved Service Incident
Started May 21, 2026 at 2:53 PM UTC · 0m
Pending
resolved
Between May 13th 18:47 UTC and May 15th 9:23 UTC, we experienced test details not appearing in the dashboard of the US-East Data Center. This issue has been resolved. All services are fully operational.
postmortem
### **Dates:**
Wednesday, May 13th 2026, 18:47 UTC - Friday, May 15th 2026, 09:23 UTC.
### **What happened:**
Customers were unable to access test run summaries for Real Device Cloud \(RDC\) jobs because events stopped publishing to the jobs Kafka topic in US-EAST.
### **Why it happened:**
An authentication key used by the message producer unexpectedly lost its permissions during an account cleanup.
### **How we fixed it:**
Manually restored the required permissions to re-establish the connection and resume service.
### **What we are doing to prevent it from happening again:**
Migrating to a permanent service account, implementing a Dead Letter Queue \(DLQ\) for the jobs Kafka topic, and replaying the missing events to restore customer data.
2026-April-23 Resolved Service Incident
Started April 24, 2026 at 4:33 PM UTC · 0m
OutageMajor incident
resolved
Between April 23rd 22:44 and April 24th 15:25 UTC, there was a technical issue that affected video recordings for tests running on macOS 15 and iOS within our EU and US-West Data Center. We identified the issue and deployed a fix. All systems are now fully operational.
postmortem
### **Dates:**
Thursday, April 23rd 2026, 22:43 UTC - Friday, April 24th 2026, 15:29 UTC
### **What happened:**
Video assets were missing for virtual iOS simulator tests on ARM and macOS ARM desktop tests in the US-West and EU data centers.
### **Why it happened:**
A product defect was introduced resulting in a screen capture failure.
### **How we fixed it:**
We performed a rollback to a stable version.
### **What we are doing to prevent it from happening again:**
We are improving monitoring & alerting to enhance our post deployment validation.
2026-April-16 Resolved Service Incident
Started April 16, 2026 at 10:10 AM UTC · 0m
Pending
resolved
Between 02:00 and 11:15 CEST, live and automated tests on iOS 17.0 simulators were failing to start in the EU and US-West Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
postmortem
### **Dates:**
Thursday, April 16th 2026, 00:00 UTC – 09:15 UTC
### **What happened:**
Live and automated tests on iOS 17.0 simulators failed to start in both the EU and US-West data centers. Customers running tests on iOS 17.0 Intel-based simulators were unable to execute their tests for approximately 9 hours.
### **Why it happened:**
A deployment introduced an incompatibility affecting iOS 17.0 on Intel-based infrastructure. The issue was not caught prior to release due to insufficient post-deployment test coverage for that specific simulator configuration.
### **How we fixed it:**
We performed a rollback to the previous deployment, which restored full iOS 17.0 simulator functionality.
### **What we are doing to prevent it from happening again:**
We are reviving and expanding automated post-deployment tests to cover a broader range of simulator configurations, including legacy Intel-based iOS versions, to catch incompatibilities before they reach production.
We are currently investigating reports of test failures affecting users running tests using SauceCtl in our US-West-1 and EU-Central-1 Data Center. We are investigating.
resolved
We have identified the root cause and have deployed a fix for this issue. All services are fully operational.
postmortem
### **Dates:**
Monday April 7th 2026, ~11:00 – 15:55 UTC
### **What happened:**
Some customers experienced 503 errors when running tests via saucectl. The test-composer service was intermittently unavailable, preventing framework-based test execution.
### **Why it happened:**
A stale Docker image was deployed to the test-composer service due to a packaging issue that arose during an internal container registry migration. This caused service pods to crash.
### **How we fixed it:**
We identified the stale image and redeployed the correct version, restoring the service.
### **What we are doing to prevent it from happening again:**
We are hardening our image deployment pipeline and adding validation checks to ensure container registry migrations do not result in stale or incorrect images being deployed to production.
2026-March-24 Resolved Service Incident
Started March 24, 2026 at 5:36 PM UTC · 0m
Pending
resolved
Between 09:32 and 15:13 UTC, we identified a technical issue affecting iOS tests when running with network capture enabled. We've resolved the underlying cause and tests are working as expected. All services are fully operational.
postmortem
### **Dates:**
Tuesday, March 24th 2026, 09:32 UTC – 15:13 UTC
### **What happened:**
Network calls failed on iOS devices during Real Device Cloud sessions where network capture was enabled. Approximately 12-13% of iOS sessions were affected. Android was not impacted.
### **Why it happened:**
A deployment introduced a DNS resolution change that was incompatible with the iOS platform, causing network capture to break.
### **How we fixed it:**
Rolled back the deployment to restore service.
### **What we are doing to prevent it from happening again:**
Adding synthetic tests to catch network capture regressions before production, and implementing monitoring alerts for faster detection after deployments.
2026-March-19 Service Incident
Started March 19, 2026 at 9:51 AM UTC · 1h 2m
OutageMajor incident
Affected components
US-WestUS-West
investigating
Around 4:45 AM UTC we started experiencing lower iOS device availability in the US-West data center. Our team is actively investigating the root cause and working toward a resolution.
resolved
This incident has been resolved and our services are fully operational.
postmortem
### **Dates:**
Wednesday, March 19 2026, 04:45 UTC - 10:47 UTC.
### **What happened:**
Approximately 15% of iOS devices in our US-West data center were temporarily unavailable for customer test sessions due to failed internet connectivity checks.
### **Why it happened:**
An automated wireless network optimization feature adjusted transmit power levels on access points serving the affected devices, degrading wireless connectivity and causing devices to fail their availability checks.
### **How we fixed it:**
The affected access points were identified and restarted, restoring normal wireless connectivity.
### **What we are doing to prevent it from happening again:**
Evaluation of the automated optimization tools and a monitoring improvement.
2026-March-13 Resolved Service Incident
Started March 13, 2026 at 2:30 PM UTC · 0m
Pending
resolved
Between 14:43 and 15:11 UTC on March 13, a small subset of Real Devices (iOS and Android) became unavailable across all our data centers. After taking remedial action, the issue was identified and resolved. All services are fully operational.
postmortem
### **Dates:**
Friday, March 13th 2026, 14:43 UTC - 15:11 UTC.
### **What happened:**
Real Devices \(iOS and Android\) availability gradually decreased across all data centers.
### **Why it happened:**
A product defect was introduced resulting in a small subset of Real Devices \(~10%\) failing to maintain required connectivity.
### **How we fixed it:**
Rollback to a stable version.
### **What we are doing to prevent it from happening again:**
Improve monitoring & alerting, enhance post deployment validation.
2026-March-10 Service Incident
Started March 10, 2026 at 6:46 PM UTC · 4h 49m
OutageMajor incident
Affected components
US-WestUS-WestEU-CentralEU-CentralUS-EastUS-East
investigating
We are experiencing device unavailability in the US West 1, EU Central 1, and US East 4 data centers and have found that the issue is caused by a 3rd party service disruption. We are investigating.
resolved
This incident has been resolved.
postmortem
### **Dates:**
Tuesday March 10th 2026, 17:52 - 23:34 UTC
### **What happened:**
The majority of iOS devices across all regions became unavailable.
### **Why it happened:**
Apple's [ppq.apple.com](http://ppq.apple.com) app verification endpoint was down, causing internal device monitoring checks to fail, bringing devices offline.
### **How we fixed it:**
We temporarily disabled these device monitoring checks.
### **What we are doing to prevent it from happening again:**
Improved external monitoring to catch outages of apple’s [ppq.apple.com](http://ppq.apple.com) endpoint, loosened device monitoring to not take down live iOS devices if [ppq.apple.com](http://ppq.apple.com) is down.
2026-March-6 Resolved Service Incident
Started March 6, 2026 at 9:38 PM UTC · 0m
Pending
resolved
Between 21:38 UTC and 23:11 UTC, our virtual iOS and MacOS live and automated device tests were failing to start in the EU Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
postmortem
### **Dates:**
Friday, March 6th 2026, 21:38 UTC - 23:11 UTC
### **What happened:**
During the incident timeline, customers running virtual iOS simulator tests on ARM or macOS ARM desktop tests in the EU Data Center were unable to start new sessions for either live or automated.
### **Why it happened:**
There was a sequencing issue on the release of the ARM side disk images in the EU.
### **How we fixed it:**
The image reference for the ARM side disk was rolled back to the previous reference to restore service.
### **What we are doing to prevent it from happening again:**
The tests that run to validate the image syncing have been completed in each region.
Sauce Labs outage history and incident timeline | Uptimus