2026-9月-02 服务事件
- investigating
我们看到欧盟-中央-1数据中心Real Devicement测试错误率上升。 我们正在调查.
- identified
已查明欧盟-中央-1数据中心Real Devicement测试错误率高的根本原因。 已经采取了补救措施,我们正在积极监测局势.
- resolved
在采取补救行动后,Real Deviceit测试出错率在EU-Central-1数据中心已恢复正常. 所有服务都全面投入运作.
自动翻译自官方事件更新。
60 Sauce Labs incidents · 2024年11月 — official updates, affected components, duration and resolution details.
我们看到欧盟-中央-1数据中心Real Devicement测试错误率上升。 我们正在调查.
已查明欧盟-中央-1数据中心Real Devicement测试错误率高的根本原因。 已经采取了补救措施,我们正在积极监测局势.
在采取补救行动后,Real Deviceit测试出错率在EU-Central-1数据中心已恢复正常. 所有服务都全面投入运作.
自动翻译自官方事件更新。
目前我们看到美国-West-1数据中心的macOS和iOS测试错误率上升。 我们正在调查.
在采取补救行动后,所有macOS和iOS测试现在都在美国-西一数据中心按预期进行。 所有服务都全面投入运作.
** 日期:** 2026年8月25日星期二,11:38 – 13:10 UTC 发生了什么事? 在美国-西部地区进行macOS和iOS测试的客户的服务下降。 该地区大约50%的虚拟Mac容量不再接受新的测试,因此测试要么排队,要么启动失败. 随着工作转向,剩余能力承受了额外压力,从而延长了桌面和模拟测试的起动时间。 原因:** 我们的Mac主机用来达到支持云服务的内部安全证书已经到期. 一旦失效,主机就再也无法建立与这项服务可信赖的连接,并停止提供新的测试机. 该证书是人工签发的,既没有自动更新,也没有过期警报。 我们如何修复它: ** 我们发放并部署了替换证书,恢复了连接,使Mac容量恢复到正常水平. 我们正采取什么措施来防止它再次发生:** 我们正在将这些证书移到自动更新上,并增加警示,以便在过期前及早更换,同时审查周围的工具,使证书处理更加安全.
自动翻译自官方事件更新。
We are currently seeing elevated error rates for virtual desktop tests in the US-West-1 datacenter. We are currently investigating.
After taking remedial action, Virtual Desktop tests are starting successfully in the US-West-1 Data center. All services are fully operational. We are closely monitoring the situation
### **Dates:** Friday, August 7th 2026, 11:00 UTC - 13:04 UTC. ### **What happened:** Windows and Intel Mac jobs in `us-west1` failed to start due to virtual machine \(VM\) allocation starvation. ### **Why it happened:** A service crash loop left VMs in an allocated but unclaimed state, while a cleanup bug prevented the system from releasing the orphaned capacity to boot new VMs. ### **How we fixed it:** Restored service stability and cleared stale allocations to resume VM provisioning and clear queued jobs. ### **What we are doing to prevent it from happening again:** Fixing allocator cleanup logic, strengthening deployment health checks, and improving capacity accounting for stale allocations.
在协调世界时7月16日13:03至协调世界时7月16日17:33之间,我们经历了服务中断,使用多段上传协议的符号档案未能处理. 事件已经得到解决,所有服务都已投入运作.
### **Dates:** Thursday July 16th 2026, 13:03 – 17:33 UTC ### **What happened:** Symbol archives uploaded in multiple parts failed to process. All regions were affected. ### **Why it happened:** Our symbol processing service sent an upload-verification field that our cloud storage provider's API does not accept for multi-part uploads, so those uploads were rejected. A fix for this had already been developed, but it had not yet been included in a released build and the service was not configured to use it. As in the first incident, rejected uploads were retried and accumulated on local disk. ### **How we fixed it:** We deployed a build containing the fix and corrected the service configuration on the affected workers. A large backlog of uploads then processed, which briefly re-filled the disks before draining completely. ### **What we are doing to prevent it from happening again:** We are releasing the fix formally through our build pipeline and persisting the corrected configuration in our configuration management, so it can't be lost. The disk-utilization alerting added after the first incident also covers the disk-exhaustion pattern common to both.
自动翻译自官方事件更新。
在协调世界时7月14日23:45至协调世界时7月15日23:05之间,我们经历了一个符号存档上传到Backtrace项目的HTTP 400出错失败的问题. 事件已经得到解决,所有服务都已投入运作.
### **Dates:** Tuesday July 14th 2026, 23:45 UTC – Wednesday July 15th 2026, 23:05 UTC ### **What happened:** Symbol archive uploads to Error Reporting \(Backtrace\) projects failed with HTTP 400 errors. All regions were affected. ### **Why it happened:** The credentials our symbol processing service used to write to cloud storage were no longer valid, so uploads could not be stored. Failed uploads were retried repeatedly and accumulated on local disk until the service ran out of space, at which point it also began rejecting new uploads. ### **How we fixed it:** We reissued the storage credentials, increased disk capacity on the affected workers, and restarted the service. The queued uploads then processed successfully, and we confirmed recovery with affected customers. ### **What we are doing to prevent it from happening again:** We've added disk-utilization monitoring and alerting to this service so we detect the condition ourselves before it affects uploads, rather than relying on customer reports.
自动翻译自官方事件更新。
IOS ARM模拟器应用测试因基础设施出错而失败。 我们正在调查.
当看到这些基础设施错误时,App下载需要的时间比预期的要长. 我们正在继续调查.
我们已经为这一问题制定了解决办法。 现在应该按照欧盟中央1数据中心的预期进行测试。 我们正在监测.
我们正在继续监测任何其他问题.
在采取补救行动并监测情况后,我们看到iOS ARM应用测试如欧盟数据中心所预期的那样进行。 所有服务都全面投入运作.
** 日期:** 2026年7月14日,星期二 UTC凌晨9:00 - 18:28 发生了什么事? 欧盟数据中心运行的iOS ARM测试失败. 原因:** 我们欧盟数据中心的主要网络供应商发生了故障。 我们如何修复它: ** 我们输给了欧盟数据中心的二级网络供应商 我们正采取什么措施来防止它再次发生:** 我们改进了对第三方网络供应商的监控和警报,以发现问题.
自动翻译自官方事件更新。
在协调世界时7月1日17:02至协调世界时21:31,macOS 14测试未能在美国-西部和欧盟-中央数据中心开始. 这个问题已经得到解决。 所有服务都全面投入运作.
** 日期:** 2026年7月1日星期三 UTC下午17:02 - UTC下午21:31 发生了什么事? 美国西部和欧盟中央数据中心的macOS 14测试未能启动。 原因:** 由于缺少配置条目,无法获取内部数据源。 我们如何修复它: ** 所缺条目被替换. 我们正采取什么措施来防止它再次发生:** 已经制定了保障措施,以防止从配置中移除使用中的数据源.
自动翻译自官方事件更新。
我们正在经历一个问题,即访问美国-西部-1、欧盟-中部-1和美国-东部-4数据中心的Appium检查员。 我们正在调查.
我们已查明了问题的根源,并已对这个问题采取了对策。 所有服务都全面投入运作.
** 日期:** 2026年7月1日星期三 09:37 – 12:46 UTC 发生了什么事? Appium 检查器功能,在实际设备现场测试过程中使用,在预定的UI部署后就变得无法使用. 在直播测试中试图使用该功能的用户被呈现出500个出错,干扰了他们的主动测试会话. 自动化试验管没有受到影响。 原因:** 一个核心前端库的例行升级引入了与 Appium 督察的渲染逻辑不相容. 上一个库版容忍了模式,但更新后版本没有. 由于这一具体特征的端到端测试覆盖面不足,这个问题在部署前没有被抓住. 我们如何修复它: ** 我们把UI回滚到最后一个已知的工作版本,以立即恢复功能,然后为不相容性部署定向固定. 我们正采取什么措施来防止它再次发生:** 我们正在增加Appium检查员功能的端到端测试范围,以确保在未来部署之前自动验证.
自动翻译自官方事件更新。
6月24日,在协调世界时08:28至协调世界时15:32之间,我们经历了美国东部,美国西部和欧盟数据中心的Appium和Access API的Real December测试会话故障. 这个问题已经得到解决。 所有服务都全面投入运作.
** 日期:** 2026年6月24日星期三 UTC09:42 - UTC15:45 (英语). 发生了什么事? 使用 Appium 和 Access API 的真正设备测试会在美国东部、美国西部和欧盟的数据中心中出现了更高的出错率。 会话在大约90秒后即被时间淘汰,或者被卡在"接通"状态下,防止被关闭. 原因:** 引入了产品缺陷,导致Real Devicember测试会话无法启动并阻止活动会话关闭. 我们如何修复它: ** 回滚到稳定的版本。 我们正采取什么措施来防止它再次发生:** 完善监测警示,强化部署后论证.
自动翻译自官方事件更新。
在协调世界时3:51左右,我们开始在欧盟-中央数据中心看到iOS设备可用率较低。 我们的小组正在积极调查根源并努力寻求解决办法.
我们一直在我们的欧盟-中央-1数据中心遇到设备无法使用的问题,并查明了根源。 我们已经采取了补救行动,目前正在监测.
在采取补救行动后,我们现在看到在欧盟-中央-1数据中心可以测试的真正设备。 所有服务都全面投入运作.
** 日期:** 2026年6月18日星期四,协调世界时03:45 - 协调世界时06:45. 发生了什么事? 我们欧盟数据中心大约7%的iOS设备在主机架断电后暂时无法用于客户测试会话. 原因:** 向受影响的架子提供电路超出其容量,一个防护断路器被绊倒来保护线路——切断该架子上装置的电能,直到电路被恢复. 我们如何修复它: ** 受影响的电路被重新设置并恢复了机架的电能,使设备回归客户服务. 我们正采取什么措施来防止它再次发生:** 我们正在通过受影响的架子重新分配电力负荷,增加能力监测,提前发出预警警报,并在新装置被部署到架子之前进行能力审查.
自动翻译自官方事件更新。
We are experiencing an issue where the US-EAST data center is currently unavailable. We are investigating.
We are continuing to investigate this issue.
We have identified the root cause and are working on implementing a fix.
Access has been restored to the US-EAST data center. We are monitoring.
After taking remedial action, access to the US-EAST data center is fully restored. This incident is resolved.
### **Dates:** Monday, June 1st 2026, 13:06 UTC - 15:42 UTC. ### **What happened:** Customers served by the US-EAST-4 region were unable to authenticate or start new test sessions because an incomplete TLS certificate chain was deployed to the core directory services. ### **Why it happened:** A certificate extraction script defect silently truncated the certificate chain after the certificate authority transitioned to a longer hierarchy. ### **How we fixed it:** Reverted the certificate rotation and re-applied the previous known ,good certificate to restore authentication services. ### **What we are doing to prevent it from happening again:** Implementing pre-deployment certificate chain validation, adding active monitoring, and fixing the script's chain-length limitations.
Between May 13th 18:47 UTC and May 15th 9:23 UTC, we experienced test details not appearing in the dashboard of the US-East Data Center. This issue has been resolved. All services are fully operational.
### **Dates:** Wednesday, May 13th 2026, 18:47 UTC - Friday, May 15th 2026, 09:23 UTC. ### **What happened:** Customers were unable to access test run summaries for Real Device Cloud \(RDC\) jobs because events stopped publishing to the jobs Kafka topic in US-EAST. ### **Why it happened:** An authentication key used by the message producer unexpectedly lost its permissions during an account cleanup. ### **How we fixed it:** Manually restored the required permissions to re-establish the connection and resume service. ### **What we are doing to prevent it from happening again:** Migrating to a permanent service account, implementing a Dead Letter Queue \(DLQ\) for the jobs Kafka topic, and replaying the missing events to restore customer data.
Between April 23rd 22:44 and April 24th 15:25 UTC, there was a technical issue that affected video recordings for tests running on macOS 15 and iOS within our EU and US-West Data Center. We identified the issue and deployed a fix. All systems are now fully operational.
### **Dates:** Thursday, April 23rd 2026, 22:43 UTC - Friday, April 24th 2026, 15:29 UTC ### **What happened:** Video assets were missing for virtual iOS simulator tests on ARM and macOS ARM desktop tests in the US-West and EU data centers. ### **Why it happened:** A product defect was introduced resulting in a screen capture failure. ### **How we fixed it:** We performed a rollback to a stable version. ### **What we are doing to prevent it from happening again:** We are improving monitoring & alerting to enhance our post deployment validation.
Between 02:00 and 11:15 CEST, live and automated tests on iOS 17.0 simulators were failing to start in the EU and US-West Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
### **Dates:** Thursday, April 16th 2026, 00:00 UTC – 09:15 UTC ### **What happened:** Live and automated tests on iOS 17.0 simulators failed to start in both the EU and US-West data centers. Customers running tests on iOS 17.0 Intel-based simulators were unable to execute their tests for approximately 9 hours. ### **Why it happened:** A deployment introduced an incompatibility affecting iOS 17.0 on Intel-based infrastructure. The issue was not caught prior to release due to insufficient post-deployment test coverage for that specific simulator configuration. ### **How we fixed it:** We performed a rollback to the previous deployment, which restored full iOS 17.0 simulator functionality. ### **What we are doing to prevent it from happening again:** We are reviving and expanding automated post-deployment tests to cover a broader range of simulator configurations, including legacy Intel-based iOS versions, to catch incompatibilities before they reach production.
We are currently investigating reports of test failures affecting users running tests using SauceCtl in our US-West-1 and EU-Central-1 Data Center. We are investigating.
We have identified the root cause and have deployed a fix for this issue. All services are fully operational.
### **Dates:** Monday April 7th 2026, ~11:00 – 15:55 UTC ### **What happened:** Some customers experienced 503 errors when running tests via saucectl. The test-composer service was intermittently unavailable, preventing framework-based test execution. ### **Why it happened:** A stale Docker image was deployed to the test-composer service due to a packaging issue that arose during an internal container registry migration. This caused service pods to crash. ### **How we fixed it:** We identified the stale image and redeployed the correct version, restoring the service. ### **What we are doing to prevent it from happening again:** We are hardening our image deployment pipeline and adding validation checks to ensure container registry migrations do not result in stale or incorrect images being deployed to production.
Between 09:32 and 15:13 UTC, we identified a technical issue affecting iOS tests when running with network capture enabled. We've resolved the underlying cause and tests are working as expected. All services are fully operational.
### **Dates:** Tuesday, March 24th 2026, 09:32 UTC – 15:13 UTC ### **What happened:** Network calls failed on iOS devices during Real Device Cloud sessions where network capture was enabled. Approximately 12-13% of iOS sessions were affected. Android was not impacted. ### **Why it happened:** A deployment introduced a DNS resolution change that was incompatible with the iOS platform, causing network capture to break. ### **How we fixed it:** Rolled back the deployment to restore service. ### **What we are doing to prevent it from happening again:** Adding synthetic tests to catch network capture regressions before production, and implementing monitoring alerts for faster detection after deployments.
Around 4:45 AM UTC we started experiencing lower iOS device availability in the US-West data center. Our team is actively investigating the root cause and working toward a resolution.
This incident has been resolved and our services are fully operational.
### **Dates:** Wednesday, March 19 2026, 04:45 UTC - 10:47 UTC. ### **What happened:** Approximately 15% of iOS devices in our US-West data center were temporarily unavailable for customer test sessions due to failed internet connectivity checks. ### **Why it happened:** An automated wireless network optimization feature adjusted transmit power levels on access points serving the affected devices, degrading wireless connectivity and causing devices to fail their availability checks. ### **How we fixed it:** The affected access points were identified and restarted, restoring normal wireless connectivity. ### **What we are doing to prevent it from happening again:** Evaluation of the automated optimization tools and a monitoring improvement.
Between 14:43 and 15:11 UTC on March 13, a small subset of Real Devices (iOS and Android) became unavailable across all our data centers. After taking remedial action, the issue was identified and resolved. All services are fully operational.
### **Dates:** Friday, March 13th 2026, 14:43 UTC - 15:11 UTC. ### **What happened:** Real Devices \(iOS and Android\) availability gradually decreased across all data centers. ### **Why it happened:** A product defect was introduced resulting in a small subset of Real Devices \(~10%\) failing to maintain required connectivity. ### **How we fixed it:** Rollback to a stable version. ### **What we are doing to prevent it from happening again:** Improve monitoring & alerting, enhance post deployment validation.
We are experiencing device unavailability in the US West 1, EU Central 1, and US East 4 data centers and have found that the issue is caused by a 3rd party service disruption. We are investigating.
This incident has been resolved.
### **Dates:** Tuesday March 10th 2026, 17:52 - 23:34 UTC ### **What happened:** The majority of iOS devices across all regions became unavailable. ### **Why it happened:** Apple's [ppq.apple.com](http://ppq.apple.com) app verification endpoint was down, causing internal device monitoring checks to fail, bringing devices offline. ### **How we fixed it:** We temporarily disabled these device monitoring checks. ### **What we are doing to prevent it from happening again:** Improved external monitoring to catch outages of apple’s [ppq.apple.com](http://ppq.apple.com) endpoint, loosened device monitoring to not take down live iOS devices if [ppq.apple.com](http://ppq.apple.com) is down.
Between 21:38 UTC and 23:11 UTC, our virtual iOS and MacOS live and automated device tests were failing to start in the EU Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
### **Dates:** Friday, March 6th 2026, 21:38 UTC - 23:11 UTC ### **What happened:** During the incident timeline, customers running virtual iOS simulator tests on ARM or macOS ARM desktop tests in the EU Data Center were unable to start new sessions for either live or automated. ### **Why it happened:** There was a sequencing issue on the release of the ARM side disk images in the EU. ### **How we fixed it:** The image reference for the ARM side disk was rolled back to the previous reference to restore service. ### **What we are doing to prevent it from happening again:** The tests that run to validate the image syncing have been completed in each region.