Мы наблюдаем повышенный уровень ошибок при тестировании Real Device в дата-центре EU-Central-1. Мы проводим расследование.
identified
Выявлена первопричина повышенной частоты ошибок при тестировании Real Device в дата-центре EU-Central-1. Проведена реабилитация, и мы активно следим за ситуацией.
resolved
После принятия мер по исправлению ошибок в дата-центре EU-Central-1 показатели ошибок в тестировании Real Device вернулись к норме. Все услуги полностью работоспособны.
Автоматический перевод официального обновления инцидента.
2026-25 Август Служебный инцидент
Начало 25 августа 2026 г. в 12:33 UTC · 46m
Pending
investigating
В настоящее время в дата-центре US-West-1 наблюдается повышенный уровень ошибок при тестировании macOS и iOS. В настоящее время мы проводим расследование.
resolved
После принятия мер по исправлению ситуации все тесты macOS и iOS теперь выполняются в дата-центре US-West-1. Все услуги полностью работоспособны.
postmortem
###** Даты:**
Вторник, 25 августа 2026, 11:38 - 13:10 UTC
###**Что случилось?**
Клиенты, проводившие тесты macOS и iOS в нашем регионе, испытали ухудшение обслуживания. Около 50% виртуальных мощностей Mac в регионе перестали принимать новые тесты, поэтому тесты либо стояли в очереди, либо не запускались. Оставшаяся емкость оказалась под дополнительным давлением, поскольку работа перешла на нее, что продлило время запуска как для настольных, так и для симуляторов.
### ** Почему это произошло
Сертификат внутренней безопасности, используемый нашими хостами Mac для доступа к поддерживаемому облачному сервису, истек. После того, как он истек, хосты больше не могли установить надежное соединение с этой службой и прекратили предоставление новых тестовых машин. Сертификат выдан вручную и не имеет автоматического продления или оповещения о истечении срока действия.
### ** Как мы это исправили
Мы выпустили и развернули сертификат замены, который восстановил связь и вернул емкость Mac на нормальный уровень.
### ** Что мы делаем, чтобы предотвратить это снова: **
Мы переводим эти сертификаты на автоматическое обновление и добавляем оповещения, чтобы они были заменены задолго до истечения срока действия, а также проверяем окружающие инструменты, чтобы сделать обработку сертификатов более безопасной.
Автоматический перевод официального обновления инцидента.
2026-07 августа Служебный инцидент
Начало 7 августа 2026 г. в 12:08 UTC · 58m
Pending
investigating
В настоящее время в дата-центре US-West-1 наблюдается повышенный уровень ошибок при тестировании виртуальных настольных компьютеров. В настоящее время мы проводим расследование.
resolved
After taking remedial action, Virtual Desktop tests are starting successfully in the US-West-1 Data center. All services are fully operational. We are closely monitoring the situation
postmortem
### **Dates:**
Friday, August 7th 2026, 11:00 UTC - 13:04 UTC.
### **What happened:**
Windows and Intel Mac jobs in `us-west1` failed to start due to virtual machine \(VM\) allocation starvation.
### **Why it happened:**
A service crash loop left VMs in an allocated but unclaimed state, while a cleanup bug prevented the system from releasing the orphaned capacity to boot new VMs.
### **How we fixed it:**
Restored service stability and cleared stale allocations to resume VM provisioning and clear queued jobs.
### **What we are doing to prevent it from happening again:**
Fixing allocator cleanup logic, strengthening deployment health checks, and improving capacity accounting for stale allocations.
Автоматический перевод официального обновления инцидента.
2026-16 июля - Решенный инцидент с обслуживанием - Отчет об ошибках (Backtrace)
Начало 16 июля 2026 г. в 13:00 UTC · 0m
Pending
resolved
Между 16 июля 13:03 UTC и 16 июля 17:33 UTC произошел сбой в обслуживании, когда архивы символов с использованием протоколов загрузки из нескольких частей не обрабатывались. Инцидент урегулирован, все службы работают.
postmortem
### **Dates:**
Thursday July 16th 2026, 13:03 – 17:33 UTC
### **What happened:**
Symbol archives uploaded in multiple parts failed to process. All regions were affected.
### **Why it happened:**
Our symbol processing service sent an upload-verification field that our cloud storage provider's API does not accept for multi-part uploads, so those uploads were rejected. A fix for this had already been developed, but it had not yet been included in a released build and the service was not configured to use it. As in the first incident, rejected uploads were retried and accumulated on local disk.
### **How we fixed it:**
We deployed a build containing the fix and corrected the service configuration on the affected workers. A large backlog of uploads then processed, which briefly re-filled the disks before draining completely.
### **What we are doing to prevent it from happening again:**
We are releasing the fix formally through our build pipeline and persisting the corrected configuration in our configuration management, so it can't be lost. The disk-utilization alerting added after the first incident also covers the disk-exhaustion pattern common to both.
Автоматический перевод официального обновления инцидента.
2026-15 июля Решенный инцидент с обслуживанием - Отчет об ошибках
Начало 14 июля 2026 г. в 23:30 UTC · 0m
Pending
resolved
Между 14 июля 23:45 UTC и 15 июля 23:05 UTC мы столкнулись с проблемой, когда архив символов загружался в проекты Backtrace с ошибками HTTP 400. Инцидент урегулирован, все службы работают.
postmortem
### **Dates:**
Tuesday July 14th 2026, 23:45 UTC – Wednesday July 15th 2026, 23:05 UTC
### **What happened:**
Symbol archive uploads to Error Reporting \(Backtrace\) projects failed with HTTP 400 errors. All regions were affected.
### **Why it happened:**
The credentials our symbol processing service used to write to cloud storage were no longer valid, so uploads could not be stored. Failed uploads were retried repeatedly and accumulated on local disk until the service ran out of space, at which point it also began rejecting new uploads.
### **How we fixed it:**
We reissued the storage credentials, increased disk capacity on the affected workers, and restarted the service. The queued uploads then processed successfully, and we confirmed recovery with affected customers.
### **What we are doing to prevent it from happening again:**
We've added disk-utilization monitoring and alerting to this service so we detect the condition ourselves before it affects uploads, rather than relying on customer reports.
Автоматический перевод официального обновления инцидента.
2026-14 июля Служебный инцидент
Начало 14 июля 2026 г. в 17:19 UTC · 3h 4m
OutageСерьёзный инцидент
Затронутые компоненты
EU-CentralEU-Central
investigating
Мы столкнулись с проблемой в центре обработки данных EU Central 1, где тесты приложений iOS ARM Simulator не справляются с ошибками инфраструктуры. Мы проводим расследование.
investigating
Загрузка приложений занимает больше времени, чем ожидалось, когда эти ошибки инфраструктуры видны. Мы продолжаем расследование.
monitoring
Мы внедрили исправление этого вопроса. Тесты теперь должны выполняться так, как ожидается, в центре обработки данных EU Central 1. Мы проводим мониторинг.
monitoring
Мы продолжаем отслеживать любые дополнительные проблемы.
resolved
После принятия мер по исправлению ситуации мы видим, что тесты приложений iOS ARM выполняются так, как ожидалось в Центре обработки данных ЕС. Все услуги полностью работоспособны.
postmortem
###** Даты:**
14 июля 2026, 09:00 - 18:28 UTC
###**Что случилось?**
Тесты iOS ARM в нашем дата-центре в ЕС провалились.
### ** Почему это произошло
Неисправность произошла с нашим основным сетевым провайдером в центре обработки данных ЕС.
### ** Как мы это исправили
Мы не смогли связаться со вторичным сетевым провайдером из нашего центра обработки данных в ЕС.
### ** Что мы делаем, чтобы предотвратить это снова: **
Мы улучшили наш мониторинг и оповещение о проблемах со сторонними сетевыми провайдерами.
Автоматический перевод официального обновления инцидента.
2026 - 1 июля Решен Служебный инцидент 1
Начало 1 июля 2026 г. в 17:02 UTC · 0m
IssuesНезначительный инцидент
resolved
1 июля между 17:02 UTC и 21:31 UTC тесты macOS 14 не смогли начаться в центрах обработки данных США-Запад и ЕС-Центральный. Этот вопрос решен. Все услуги полностью работоспособны.
postmortem
###** Даты:**
Среда 1 июля 2026, 17:02 UTC - 21:31 UTC
###**Что случилось?**
Испытания macOS 14 в центрах обработки данных США и ЕС не смогли начаться.
### ** Почему это произошло
Внутренний источник данных был недоступен из-за отсутствующей записи конфигурации.
### ** Как мы это исправили
Недостающий вход был заменен.
### ** Что мы делаем, чтобы предотвратить это снова: **
Были введены меры предосторожности для предотвращения удаления используемых источников данных из конфигурации.
Автоматический перевод официального обновления инцидента.
2026-01 июля Служебный инцидент
Начало 1 июля 2026 г. в 12:12 UTC · 46m
Pending
investigating
Мы испытываем проблемы с доступом к Appium Inspector в дата-центрах США-Запад-1, ЕС-Центральный-1 и США-Восток-4. Мы проводим расследование.
resolved
Мы определили первопричину и внедрили исправление этой проблемы. Все услуги полностью работоспособны.
postmortem
###** Даты:**
Среда 1 июля 2026, 09:37 - 12:46 UTC
###**Что случилось?**
Функция Appium Inspector, используемая во время реальных сессий тестирования устройств, стала недоступной после запланированного развертывания пользовательского интерфейса. Пользователям, которые пытались использовать эту функцию во время живого теста, была представлена ошибка 500, нарушающая их активную сессию тестирования. Автоматизированные испытательные трубопроводы не пострадали.
### ** Почему это произошло
Обычная модернизация основной интерфейсной библиотеки ввела несовместимость с логикой рендеринга Appium Inspector. Предыдущая библиотечная версия переносила шаблон, но обновлённая версия — нет. Проблема не была выявлена до развертывания из-за недостаточного охвата сквозных испытаний для этой конкретной функции.
### ** Как мы это исправили
Мы откатили пользовательский интерфейс до последней известной рабочей версии, чтобы немедленно восстановить функцию, а затем развернули целевое исправление несовместимости.
### ** Что мы делаем, чтобы предотвратить это снова: **
Мы добавляем сквозное тестовое покрытие для функции Appium Inspector, чтобы обеспечить ее автоматическую проверку перед будущими развертываниями.
Автоматический перевод официального обновления инцидента.
2026-24 Июнь Раскрытый Служебный Инцидент
Начало 24 июня 2026 г. в 16:39 UTC · 0m
Pending
resolved
24 июня между 08:28 UTC и 15:32 UTC мы испытали сбои тестового сеанса Real Device, влияющие на API Appium и Access на востоке США, западе США и в центрах обработки данных ЕС. Этот вопрос решен. Все услуги полностью работоспособны.
postmortem
###** Даты:**
Среда, 24 июня 2026, 09:42 UTC - 15:45 UTC.
###**Что случилось?**
Сеансы тестирования реальных устройств с использованием API Appium и Access показали повышенный уровень ошибок в центрах обработки данных на востоке США, западе США и в ЕС.
Сеансы были отсрочены примерно через 90 секунд или застряли в состоянии «Соединения», предотвращая их закрытие.
### ** Почему это произошло
Был введен дефект продукта, что привело к тому, что сеансы тестирования Real Device не начинались и не допускались активные сеансы.
### ** Как мы это исправили
Откат к стабильной версии.
### ** Что мы делаем, чтобы предотвратить это снова: **
Улучшить мониторинг и оповещение и повысить валидацию после развертывания.
Автоматический перевод официального обновления инцидента.
Служебный инцидент 2026-18 июня
Начало 18 июня 2026 г. в 05:30 UTC · 1h 29m
OutageСерьёзный инцидент
Затронутые компоненты
EU-CentralEU-Central
investigating
Примерно в 3:51 утра UTC мы начали испытывать более низкую доступность устройств iOS в центре обработки данных EU-Central. Наша команда активно исследует первопричину и работает над решением.
monitoring
Мы столкнулись с недоступностью устройств в нашем центре обработки данных EU-Central-1 и определили первопричину. Мы приняли меры по исправлению положения и в настоящее время осуществляем мониторинг.
resolved
После принятия мер по исправлению ситуации мы видим реальные устройства, доступные для тестирования в дата-центре EU-Central-1. Все услуги полностью работоспособны.
postmortem
###** Даты:**
Четверг 18 июня 2026, 03:45 UTC - 06:45 UTC.
###**Что случилось?**
Примерно 7% устройств iOS в нашем дата-центре в ЕС были временно недоступны для сессий тестирования клиентов после потери питания на стойке их размещения.
### ** Почему это произошло
Силовая цепь, питающая пораженную стойку, превышала ее емкость, и защитный выключатель споткнулся, чтобы защитить линию - мощность резки к устройствам на этой стойке, пока цепь не была восстановлена.
### ** Как мы это исправили
Пораженная схема была сброшена и мощность восстановлена в стойке, вернув устройства в службу поддержки клиентов.
### ** Что мы делаем, чтобы предотвратить это снова: **
Мы перераспределяем нагрузку на мощности по пораженным стойкам, добавляем мониторинг емкости с предупреждениями о раннем предупреждении перед ограничениями схемы и вводим обзор емкости до того, как новые устройства будут развернуты в стойке.
Автоматический перевод официального обновления инцидента.
2026-June-1 Service Incident
Начало 1 июня 2026 г. в 14:02 UTC · 4h 22m
OutageКритический инцидент
Затронутые компоненты
US-EastUS-EastUS-EastUS-EastUS-EastUS-EastUS-East
investigating
We are experiencing an issue where the US-EAST data center is currently unavailable. We are investigating.
investigating
We are continuing to investigate this issue.
investigating
We have identified the root cause and are working on implementing a fix.
monitoring
Access has been restored to the US-EAST data center. We are monitoring.
resolved
After taking remedial action, access to the US-EAST data center is fully restored. This incident is resolved.
postmortem
### **Dates:**
Monday, June 1st 2026, 13:06 UTC - 15:42 UTC.
### **What happened:**
Customers served by the US-EAST-4 region were unable to authenticate or start new test sessions because an incomplete TLS certificate chain was deployed to the core directory services.
### **Why it happened:**
A certificate extraction script defect silently truncated the certificate chain after the certificate authority transitioned to a longer hierarchy.
### **How we fixed it:**
Reverted the certificate rotation and re-applied the previous known ,good certificate to restore authentication services.
### **What we are doing to prevent it from happening again:**
Implementing pre-deployment certificate chain validation, adding active monitoring, and fixing the script's chain-length limitations.
2026-May-13 Resolved Service Incident
Начало 21 мая 2026 г. в 14:53 UTC · 0m
Pending
resolved
Between May 13th 18:47 UTC and May 15th 9:23 UTC, we experienced test details not appearing in the dashboard of the US-East Data Center. This issue has been resolved. All services are fully operational.
postmortem
### **Dates:**
Wednesday, May 13th 2026, 18:47 UTC - Friday, May 15th 2026, 09:23 UTC.
### **What happened:**
Customers were unable to access test run summaries for Real Device Cloud \(RDC\) jobs because events stopped publishing to the jobs Kafka topic in US-EAST.
### **Why it happened:**
An authentication key used by the message producer unexpectedly lost its permissions during an account cleanup.
### **How we fixed it:**
Manually restored the required permissions to re-establish the connection and resume service.
### **What we are doing to prevent it from happening again:**
Migrating to a permanent service account, implementing a Dead Letter Queue \(DLQ\) for the jobs Kafka topic, and replaying the missing events to restore customer data.
2026-April-23 Resolved Service Incident
Начало 24 апреля 2026 г. в 16:33 UTC · 0m
OutageСерьёзный инцидент
resolved
Between April 23rd 22:44 and April 24th 15:25 UTC, there was a technical issue that affected video recordings for tests running on macOS 15 and iOS within our EU and US-West Data Center. We identified the issue and deployed a fix. All systems are now fully operational.
postmortem
### **Dates:**
Thursday, April 23rd 2026, 22:43 UTC - Friday, April 24th 2026, 15:29 UTC
### **What happened:**
Video assets were missing for virtual iOS simulator tests on ARM and macOS ARM desktop tests in the US-West and EU data centers.
### **Why it happened:**
A product defect was introduced resulting in a screen capture failure.
### **How we fixed it:**
We performed a rollback to a stable version.
### **What we are doing to prevent it from happening again:**
We are improving monitoring & alerting to enhance our post deployment validation.
2026-April-16 Resolved Service Incident
Начало 16 апреля 2026 г. в 10:10 UTC · 0m
Pending
resolved
Between 02:00 and 11:15 CEST, live and automated tests on iOS 17.0 simulators were failing to start in the EU and US-West Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
postmortem
### **Dates:**
Thursday, April 16th 2026, 00:00 UTC – 09:15 UTC
### **What happened:**
Live and automated tests on iOS 17.0 simulators failed to start in both the EU and US-West data centers. Customers running tests on iOS 17.0 Intel-based simulators were unable to execute their tests for approximately 9 hours.
### **Why it happened:**
A deployment introduced an incompatibility affecting iOS 17.0 on Intel-based infrastructure. The issue was not caught prior to release due to insufficient post-deployment test coverage for that specific simulator configuration.
### **How we fixed it:**
We performed a rollback to the previous deployment, which restored full iOS 17.0 simulator functionality.
### **What we are doing to prevent it from happening again:**
We are reviving and expanding automated post-deployment tests to cover a broader range of simulator configurations, including legacy Intel-based iOS versions, to catch incompatibilities before they reach production.
We are currently investigating reports of test failures affecting users running tests using SauceCtl in our US-West-1 and EU-Central-1 Data Center. We are investigating.
resolved
We have identified the root cause and have deployed a fix for this issue. All services are fully operational.
postmortem
### **Dates:**
Monday April 7th 2026, ~11:00 – 15:55 UTC
### **What happened:**
Some customers experienced 503 errors when running tests via saucectl. The test-composer service was intermittently unavailable, preventing framework-based test execution.
### **Why it happened:**
A stale Docker image was deployed to the test-composer service due to a packaging issue that arose during an internal container registry migration. This caused service pods to crash.
### **How we fixed it:**
We identified the stale image and redeployed the correct version, restoring the service.
### **What we are doing to prevent it from happening again:**
We are hardening our image deployment pipeline and adding validation checks to ensure container registry migrations do not result in stale or incorrect images being deployed to production.
2026-March-24 Resolved Service Incident
Начало 24 марта 2026 г. в 17:36 UTC · 0m
Pending
resolved
Between 09:32 and 15:13 UTC, we identified a technical issue affecting iOS tests when running with network capture enabled. We've resolved the underlying cause and tests are working as expected. All services are fully operational.
postmortem
### **Dates:**
Tuesday, March 24th 2026, 09:32 UTC – 15:13 UTC
### **What happened:**
Network calls failed on iOS devices during Real Device Cloud sessions where network capture was enabled. Approximately 12-13% of iOS sessions were affected. Android was not impacted.
### **Why it happened:**
A deployment introduced a DNS resolution change that was incompatible with the iOS platform, causing network capture to break.
### **How we fixed it:**
Rolled back the deployment to restore service.
### **What we are doing to prevent it from happening again:**
Adding synthetic tests to catch network capture regressions before production, and implementing monitoring alerts for faster detection after deployments.
2026-March-19 Service Incident
Начало 19 марта 2026 г. в 09:51 UTC · 1h 2m
OutageСерьёзный инцидент
Затронутые компоненты
US-WestUS-West
investigating
Around 4:45 AM UTC we started experiencing lower iOS device availability in the US-West data center. Our team is actively investigating the root cause and working toward a resolution.
resolved
This incident has been resolved and our services are fully operational.
postmortem
### **Dates:**
Wednesday, March 19 2026, 04:45 UTC - 10:47 UTC.
### **What happened:**
Approximately 15% of iOS devices in our US-West data center were temporarily unavailable for customer test sessions due to failed internet connectivity checks.
### **Why it happened:**
An automated wireless network optimization feature adjusted transmit power levels on access points serving the affected devices, degrading wireless connectivity and causing devices to fail their availability checks.
### **How we fixed it:**
The affected access points were identified and restarted, restoring normal wireless connectivity.
### **What we are doing to prevent it from happening again:**
Evaluation of the automated optimization tools and a monitoring improvement.
2026-March-13 Resolved Service Incident
Начало 13 марта 2026 г. в 14:30 UTC · 0m
Pending
resolved
Between 14:43 and 15:11 UTC on March 13, a small subset of Real Devices (iOS and Android) became unavailable across all our data centers. After taking remedial action, the issue was identified and resolved. All services are fully operational.
postmortem
### **Dates:**
Friday, March 13th 2026, 14:43 UTC - 15:11 UTC.
### **What happened:**
Real Devices \(iOS and Android\) availability gradually decreased across all data centers.
### **Why it happened:**
A product defect was introduced resulting in a small subset of Real Devices \(~10%\) failing to maintain required connectivity.
### **How we fixed it:**
Rollback to a stable version.
### **What we are doing to prevent it from happening again:**
Improve monitoring & alerting, enhance post deployment validation.
2026-March-10 Service Incident
Начало 10 марта 2026 г. в 18:46 UTC · 4h 49m
OutageСерьёзный инцидент
Затронутые компоненты
US-WestUS-WestEU-CentralEU-CentralUS-EastUS-East
investigating
We are experiencing device unavailability in the US West 1, EU Central 1, and US East 4 data centers and have found that the issue is caused by a 3rd party service disruption. We are investigating.
resolved
This incident has been resolved.
postmortem
### **Dates:**
Tuesday March 10th 2026, 17:52 - 23:34 UTC
### **What happened:**
The majority of iOS devices across all regions became unavailable.
### **Why it happened:**
Apple's [ppq.apple.com](http://ppq.apple.com) app verification endpoint was down, causing internal device monitoring checks to fail, bringing devices offline.
### **How we fixed it:**
We temporarily disabled these device monitoring checks.
### **What we are doing to prevent it from happening again:**
Improved external monitoring to catch outages of apple’s [ppq.apple.com](http://ppq.apple.com) endpoint, loosened device monitoring to not take down live iOS devices if [ppq.apple.com](http://ppq.apple.com) is down.
2026-March-6 Resolved Service Incident
Начало 6 марта 2026 г. в 21:38 UTC · 0m
Pending
resolved
Between 21:38 UTC and 23:11 UTC, our virtual iOS and MacOS live and automated device tests were failing to start in the EU Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
postmortem
### **Dates:**
Friday, March 6th 2026, 21:38 UTC - 23:11 UTC
### **What happened:**
During the incident timeline, customers running virtual iOS simulator tests on ARM or macOS ARM desktop tests in the EU Data Center were unable to start new sessions for either live or automated.
### **Why it happened:**
There was a sequencing issue on the release of the ARM side disk images in the EU.
### **How we fixed it:**
The image reference for the ARM side disk was rolled back to the previous reference to restore service.
### **What we are doing to prevent it from happening again:**
The tests that run to validate the image syncing have been completed in each region.