W centrum danych EU- Central-1 widzimy podwyższony poziom błędów w testach Real Device. Prowadzimy śledztwo.
identified
Zidentyfikowano podstawową przyczynę podwyższonych wskaźników błędów dla testów Real Device w centrum danych EU- Central-1. Zastosowano środki zaradcze i aktywnie monitorujemy sytuację.
resolved
Po podjęciu działań zaradczych, wskaźnik błędów w teście Real Device powrócił do normy w centrum danych EU- Central-1. Wszystkie usługi są w pełni sprawne.
Obecnie w centrum danych US- West-1 obserwujemy podwyższony poziom błędów w testach na macOS i iOS. Prowadzimy śledztwo.
resolved
Po podjęciu działań zaradczych wszystkie testy w systemie Mac OS i iOS są obecnie wykonywane zgodnie z oczekiwaniami w centrum danych US- West-1. Wszystkie usługi są w pełni sprawne.
postmortem
/ Daty:
Wtorek 25 sierpnia 2026, 11: 38 - 13: 10 UTC
Co się stało?
Klienci przeprowadzający testy na macOS i iOS w naszym regionie US- West doświadczyli zdegradowanej usługi. Około 50% wirtualnej pojemności Mac w regionie przestało przyjmować nowe testy, więc testy albo w kolejce lub nie rozpoczęły się. Pozostała pojemność była pod dodatkowym naciskiem podczas przesuwania prac na nią, co wydłużało czas startu zarówno dla testów na pulpicie jak i symulatorze.
* * * Why it happen: * *
Wewnętrzny certyfikat bezpieczeństwa używany przez naszych Mac hostów w celu osiągnięcia usługi wspierającej chmurę osiągnął jej datę ważności. Po jej wygaśnięciu, gospodarze nie mogli już ustanowić zaufanego połączenia z tą usługą i zaprzestali dostarczania nowych maszyn testowych. Certyfikat został wydany ręcznie i nie miał automatycznego odnowienia ani ostrzeżenia o wygaśnięciu.
Jak to naprawiliśmy?
Wydaliśmy i uruchomiliśmy certyfikat zastępczy, który przywrócił łączność i przywrócił pojemność Mac do normalnego poziomu.
* * * What we are doing to do not it from happing again: * *
Przenosimy te certyfikaty na automatyczną odnowę i dodamy ostrzeżenie, tak aby zostały one zastąpione znacznie przed wygaśnięciem, wraz z przeglądem otaczającego oprzyrządowania w celu zapewnienia bezpieczniejszej obsługi certyfikatu.
We are currently seeing elevated error rates for virtual desktop tests in the US-West-1 datacenter. We are currently investigating.
resolved
After taking remedial action, Virtual Desktop tests are starting successfully in the US-West-1 Data center. All services are fully operational. We are closely monitoring the situation
postmortem
### **Dates:**
Friday, August 7th 2026, 11:00 UTC - 13:04 UTC.
### **What happened:**
Windows and Intel Mac jobs in `us-west1` failed to start due to virtual machine \(VM\) allocation starvation.
### **Why it happened:**
A service crash loop left VMs in an allocated but unclaimed state, while a cleanup bug prevented the system from releasing the orphaned capacity to boot new VMs.
### **How we fixed it:**
Restored service stability and cleared stale allocations to resume VM provisioning and clear queued jobs.
### **What we are doing to prevent it from happening again:**
Fixing allocator cleanup logic, strengthening deployment health checks, and improving capacity accounting for stale allocations.
2026-July-16 Resoluted Service Incident - Zgłaszanie błędów (Backtrace)
Początek 16 lipca 2026 13:00 UTC · 0m
Pending
resolved
Między 16 lipca 13: 03 UTC a 16 lipca 17: 33 UTC, doświadczyliśmy zakłóceń w obsłudze, w których archiwa symboli wykorzystujące wieloczęściowe protokoły przesyłania nie przetworzyły się. Incydent został rozwiązany i wszystkie usługi są operacyjne.
postmortem
### **Dates:**
Thursday July 16th 2026, 13:03 – 17:33 UTC
### **What happened:**
Symbol archives uploaded in multiple parts failed to process. All regions were affected.
### **Why it happened:**
Our symbol processing service sent an upload-verification field that our cloud storage provider's API does not accept for multi-part uploads, so those uploads were rejected. A fix for this had already been developed, but it had not yet been included in a released build and the service was not configured to use it. As in the first incident, rejected uploads were retried and accumulated on local disk.
### **How we fixed it:**
We deployed a build containing the fix and corrected the service configuration on the affected workers. A large backlog of uploads then processed, which briefly re-filled the disks before draining completely.
### **What we are doing to prevent it from happening again:**
We are releasing the fix formally through our build pipeline and persisting the corrected configuration in our configuration management, so it can't be lost. The disk-utilization alerting added after the first incident also covers the disk-exhaustion pattern common to both.
2026-July-15 Resoluted Service Incident - Zgłaszanie błędów (Backtrace)
Początek 14 lipca 2026 23:30 UTC · 0m
Pending
resolved
Między 14 lipca 23: 45 UTC a 15 lipca 23: 05 UTC, doświadczyliśmy problemu, w którym archiwa symboli przesyłane do Backtrace projektów nie powiodły się z błędami HTTP 400. Incydent został rozwiązany i wszystkie usługi są operacyjne.
postmortem
### **Dates:**
Tuesday July 14th 2026, 23:45 UTC – Wednesday July 15th 2026, 23:05 UTC
### **What happened:**
Symbol archive uploads to Error Reporting \(Backtrace\) projects failed with HTTP 400 errors. All regions were affected.
### **Why it happened:**
The credentials our symbol processing service used to write to cloud storage were no longer valid, so uploads could not be stored. Failed uploads were retried repeatedly and accumulated on local disk until the service ran out of space, at which point it also began rejecting new uploads.
### **How we fixed it:**
We reissued the storage credentials, increased disk capacity on the affected workers, and restarted the service. The queued uploads then processed successfully, and we confirmed recovery with affected customers.
### **What we are doing to prevent it from happening again:**
We've added disk-utilization monitoring and alerting to this service so we detect the condition ourselves before it affects uploads, rather than relying on customer reports.
Mamy do czynienia z problemem w centrum danych EU Central 1, gdzie testy aplikacji iOS ARM Simulator nie powiodły się z błędami infrastrukturalnymi. Prowadzimy śledztwo.
investigating
Pobieranie aplikacji trwa dłużej niż się spodziewano, gdy widoczne są błędy w infrastrukturze. Kontynuujemy śledztwo.
monitoring
Zastosowaliśmy rozwiązanie tego problemu. Obecnie badania powinny być przeprowadzane zgodnie z oczekiwaniami w centrum danych UE Central 1. Monitorujemy.
monitoring
Nadal monitorujemy wszelkie dalsze problemy.
resolved
Po podjęciu działań naprawczych i monitorowaniu sytuacji, widzimy testy aplikacji iOS ARM wykonywane zgodnie z oczekiwaniami w Centrum Danych UE. Wszystkie usługi są w pełni sprawne.
postmortem
/ Daty:
Wtorek 14 lipca 2026, 09: 00 - 18: 28 UTC
Co się stało?
Testy iOS ARM prowadzone w naszym centrum danych UE nie powiodły się.
* * * Why it happen: * *
Wystąpił błąd z naszym głównym dostawcą sieci w naszym centrum danych UE.
Jak to naprawiliśmy?
Zawiodliśmy drugorzędnego dostawcę sieci z naszego unijnego centrum danych.
* * * What we are doing to do not it from happing again: * *
Ulepszyliśmy nasze monitorowanie i ostrzegamy o problemach z dostawcami sieci osób trzecich.
1 lipca między 17: 02 UTC a 21: 31 UTC nie rozpoczęły się testy MAKOS 14 w centrach danych USA-Zachód i UE-Central. Kwestia ta została rozwiązana. Wszystkie usługi są w pełni sprawne.
postmortem
/ Daty:
Środa 1 lipca 2026, 17: 02 UTC - 21: 31 UTC
Co się stało?
MacOS 14 testów w amerykańskich centrach danych Zachodu i UE nie były w stanie rozpocząć.
* * * Why it happen: * *
Wewnętrzne dane były niedostępne ze względu na brak konfiguracji.
Jak to naprawiliśmy?
Brak wpisu został zastąpiony.
* * * What we are doing to do not it from happing again: * *
Wprowadzono zabezpieczenia zapobiegające usuwaniu danych z konfiguracji.
Mamy problem z dostępem do Inspektora Appium w centrach danych US- West-1, EU- Central-1 i US- East- 4. Prowadzimy śledztwo.
resolved
Zidentyfikowaliśmy przyczynę i wprowadziliśmy rozwiązanie tej kwestii. Wszystkie usługi są w pełni sprawne.
postmortem
/ Daty:
Środa 1 lipca 2026, 09: 37 - 12: 46 UTC
Co się stało?
Funkcja Inspektora Appium, używana podczas prawdziwych sesji testowych na żywo, stała się niedostępna po planowanym wdrożeniu interfejsu użytkownika. Użytkownicy, którzy próbowali skorzystać z tej funkcji podczas testu na żywo, popełnili błąd 500, zakłócając ich aktywną sesję testową. Automatyczne rurociągi testowe nie zostały naruszone.
* * * Why it happen: * *
Rutynowa modernizacja podstawowej biblioteki frontend wprowadziła niezgodność z logiką renderowania inspektora Appium. Poprzednia wersja biblioteki tolerowała wzór, ale zaktualizowana wersja nie. Kwestia ta nie została wykryta przed wdrożeniem z powodu niewystarczającego pokrycia testu końcowego dla tej specyfiki.
Jak to naprawiliśmy?
Cofnęliśmy UI do ostatniej znanej wersji roboczej, aby natychmiast przywrócić funkcję, a następnie wdrożyliśmy ukierunkowane rozwiązanie dla niezgodności.
* * * What we are doing to do not it from happing again: * *
Dodajemy pokrycie testowe end-to@-@ end dla funkcji Inspektora Appium, aby upewnić się, że jest on zatwierdzany automatycznie przed wdrożeniem w przyszłości.
24 czerwca, między 08: 28 UTC a 15: 32 UTC, doświadczyliśmy błędów sesji testowych Real Device wpływających na Appium i Access API w amerykańskich centrach danych Wschodu, USA Zachodu i UE. Kwestia ta została rozwiązana. Wszystkie usługi są w pełni sprawne.
postmortem
/ Daty:
Środa, 24 czerwca 2026, 09: 42 UTC - 15: 45 UTC.
Co się stało?
Sesje testowe Real Device przy użyciu API Appium i Access doświadczyły zwiększonego poziomu błędów w amerykańskich centrach danych na wschodzie, na zachodzie USA i w UE.
Sesje trwały w czasie po około 90 sekundach lub utknęły w stanie "Connecting", uniemożliwiając ich zamknięcie.
* * * Why it happen: * *
Wprowadzono defekt produktu powodujący, że sesje testowe Real Device nie rozpoczęły się i uniemożliwiały zamknięcie aktywnych sesji.
Jak to naprawiliśmy?
Powrót do stabilnej wersji.
* * * What we are doing to do not it from happing again: * *
Poprawa monitorowania i ostrzegania oraz poprawa walidacji po wdrożeniu.
Około 3: 51 AM UTC zaczęliśmy doświadczać niższej dostępności urządzenia iOS w centrum danych EU- Central. Nasz zespół aktywnie bada przyczynę i pracuje nad rozwiązaniem.
monitoring
Doświadczyliśmy niedostępności urządzenia w naszym centrum danych EU- Central-1 i zidentyfikowaliśmy przyczynę. Podjęliśmy działania naprawcze i obecnie monitorujemy.
resolved
Po podjęciu działań naprawczych, widzimy teraz prawdziwe urządzenia dostępne do testowania w centrum danych EU- Central-1. Wszystkie usługi są w pełni sprawne.
postmortem
/ Daty:
Czwartek 18 czerwca 2026, 03: 45 UTC - 06: 45 UTC.
Co się stało?
Około 7% urządzeń z systemem iOS w naszym unijnym centrum danych było tymczasowo niedostępnych na sesjach testowych dla klientów po utracie mocy przez hosting ich stojaka.
* * * Why it happen: * *
Obwód zasilania zasilający dotknięty stojak przekroczył jego pojemność, a wyłącznik ochronny przełączył się w celu zabezpieczenia linii - mocy cięcia do urządzeń na stojaku do czasu przywrócenia obwodu.
Jak to naprawiliśmy?
Zagrożony obwód został zresetowany i zasilanie przywrócone do stojaka, zwracając urządzenia do obsługi klienta.
* * * What we are doing to do not it from happing again: * *
Redystrybuujemy obciążenie energetyczne w obrębie dotkniętych stojaków, dodając monitorowanie wydajności z ostrzeżeniami ostrzegającymi przed ograniczeniami obwodów i wprowadzając przegląd wydajności przed umieszczeniem nowych urządzeń na stojaku.
We are experiencing an issue where the US-EAST data center is currently unavailable. We are investigating.
investigating
We are continuing to investigate this issue.
investigating
We have identified the root cause and are working on implementing a fix.
monitoring
Access has been restored to the US-EAST data center. We are monitoring.
resolved
After taking remedial action, access to the US-EAST data center is fully restored. This incident is resolved.
postmortem
### **Dates:**
Monday, June 1st 2026, 13:06 UTC - 15:42 UTC.
### **What happened:**
Customers served by the US-EAST-4 region were unable to authenticate or start new test sessions because an incomplete TLS certificate chain was deployed to the core directory services.
### **Why it happened:**
A certificate extraction script defect silently truncated the certificate chain after the certificate authority transitioned to a longer hierarchy.
### **How we fixed it:**
Reverted the certificate rotation and re-applied the previous known ,good certificate to restore authentication services.
### **What we are doing to prevent it from happening again:**
Implementing pre-deployment certificate chain validation, adding active monitoring, and fixing the script's chain-length limitations.
2026-May-13 Resolved Service Incident
Początek 21 maja 2026 14:53 UTC · 0m
Pending
resolved
Between May 13th 18:47 UTC and May 15th 9:23 UTC, we experienced test details not appearing in the dashboard of the US-East Data Center. This issue has been resolved. All services are fully operational.
postmortem
### **Dates:**
Wednesday, May 13th 2026, 18:47 UTC - Friday, May 15th 2026, 09:23 UTC.
### **What happened:**
Customers were unable to access test run summaries for Real Device Cloud \(RDC\) jobs because events stopped publishing to the jobs Kafka topic in US-EAST.
### **Why it happened:**
An authentication key used by the message producer unexpectedly lost its permissions during an account cleanup.
### **How we fixed it:**
Manually restored the required permissions to re-establish the connection and resume service.
### **What we are doing to prevent it from happening again:**
Migrating to a permanent service account, implementing a Dead Letter Queue \(DLQ\) for the jobs Kafka topic, and replaying the missing events to restore customer data.
2026-April-23 Resolved Service Incident
Początek 24 kwietnia 2026 16:33 UTC · 0m
OutagePoważny incydent
resolved
Between April 23rd 22:44 and April 24th 15:25 UTC, there was a technical issue that affected video recordings for tests running on macOS 15 and iOS within our EU and US-West Data Center. We identified the issue and deployed a fix. All systems are now fully operational.
postmortem
### **Dates:**
Thursday, April 23rd 2026, 22:43 UTC - Friday, April 24th 2026, 15:29 UTC
### **What happened:**
Video assets were missing for virtual iOS simulator tests on ARM and macOS ARM desktop tests in the US-West and EU data centers.
### **Why it happened:**
A product defect was introduced resulting in a screen capture failure.
### **How we fixed it:**
We performed a rollback to a stable version.
### **What we are doing to prevent it from happening again:**
We are improving monitoring & alerting to enhance our post deployment validation.
2026-April-16 Resolved Service Incident
Początek 16 kwietnia 2026 10:10 UTC · 0m
Pending
resolved
Between 02:00 and 11:15 CEST, live and automated tests on iOS 17.0 simulators were failing to start in the EU and US-West Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
postmortem
### **Dates:**
Thursday, April 16th 2026, 00:00 UTC – 09:15 UTC
### **What happened:**
Live and automated tests on iOS 17.0 simulators failed to start in both the EU and US-West data centers. Customers running tests on iOS 17.0 Intel-based simulators were unable to execute their tests for approximately 9 hours.
### **Why it happened:**
A deployment introduced an incompatibility affecting iOS 17.0 on Intel-based infrastructure. The issue was not caught prior to release due to insufficient post-deployment test coverage for that specific simulator configuration.
### **How we fixed it:**
We performed a rollback to the previous deployment, which restored full iOS 17.0 simulator functionality.
### **What we are doing to prevent it from happening again:**
We are reviving and expanding automated post-deployment tests to cover a broader range of simulator configurations, including legacy Intel-based iOS versions, to catch incompatibilities before they reach production.
We are currently investigating reports of test failures affecting users running tests using SauceCtl in our US-West-1 and EU-Central-1 Data Center. We are investigating.
resolved
We have identified the root cause and have deployed a fix for this issue. All services are fully operational.
postmortem
### **Dates:**
Monday April 7th 2026, ~11:00 – 15:55 UTC
### **What happened:**
Some customers experienced 503 errors when running tests via saucectl. The test-composer service was intermittently unavailable, preventing framework-based test execution.
### **Why it happened:**
A stale Docker image was deployed to the test-composer service due to a packaging issue that arose during an internal container registry migration. This caused service pods to crash.
### **How we fixed it:**
We identified the stale image and redeployed the correct version, restoring the service.
### **What we are doing to prevent it from happening again:**
We are hardening our image deployment pipeline and adding validation checks to ensure container registry migrations do not result in stale or incorrect images being deployed to production.
2026-March-24 Resolved Service Incident
Początek 24 marca 2026 17:36 UTC · 0m
Pending
resolved
Between 09:32 and 15:13 UTC, we identified a technical issue affecting iOS tests when running with network capture enabled. We've resolved the underlying cause and tests are working as expected. All services are fully operational.
postmortem
### **Dates:**
Tuesday, March 24th 2026, 09:32 UTC – 15:13 UTC
### **What happened:**
Network calls failed on iOS devices during Real Device Cloud sessions where network capture was enabled. Approximately 12-13% of iOS sessions were affected. Android was not impacted.
### **Why it happened:**
A deployment introduced a DNS resolution change that was incompatible with the iOS platform, causing network capture to break.
### **How we fixed it:**
Rolled back the deployment to restore service.
### **What we are doing to prevent it from happening again:**
Adding synthetic tests to catch network capture regressions before production, and implementing monitoring alerts for faster detection after deployments.
2026-March-19 Service Incident
Początek 19 marca 2026 09:51 UTC · 1h 2m
OutagePoważny incydent
Dotknięte komponenty
US-WestUS-West
investigating
Around 4:45 AM UTC we started experiencing lower iOS device availability in the US-West data center. Our team is actively investigating the root cause and working toward a resolution.
resolved
This incident has been resolved and our services are fully operational.
postmortem
### **Dates:**
Wednesday, March 19 2026, 04:45 UTC - 10:47 UTC.
### **What happened:**
Approximately 15% of iOS devices in our US-West data center were temporarily unavailable for customer test sessions due to failed internet connectivity checks.
### **Why it happened:**
An automated wireless network optimization feature adjusted transmit power levels on access points serving the affected devices, degrading wireless connectivity and causing devices to fail their availability checks.
### **How we fixed it:**
The affected access points were identified and restarted, restoring normal wireless connectivity.
### **What we are doing to prevent it from happening again:**
Evaluation of the automated optimization tools and a monitoring improvement.
2026-March-13 Resolved Service Incident
Początek 13 marca 2026 14:30 UTC · 0m
Pending
resolved
Between 14:43 and 15:11 UTC on March 13, a small subset of Real Devices (iOS and Android) became unavailable across all our data centers. After taking remedial action, the issue was identified and resolved. All services are fully operational.
postmortem
### **Dates:**
Friday, March 13th 2026, 14:43 UTC - 15:11 UTC.
### **What happened:**
Real Devices \(iOS and Android\) availability gradually decreased across all data centers.
### **Why it happened:**
A product defect was introduced resulting in a small subset of Real Devices \(~10%\) failing to maintain required connectivity.
### **How we fixed it:**
Rollback to a stable version.
### **What we are doing to prevent it from happening again:**
Improve monitoring & alerting, enhance post deployment validation.
2026-March-10 Service Incident
Początek 10 marca 2026 18:46 UTC · 4h 49m
OutagePoważny incydent
Dotknięte komponenty
US-WestUS-WestEU-CentralEU-CentralUS-EastUS-East
investigating
We are experiencing device unavailability in the US West 1, EU Central 1, and US East 4 data centers and have found that the issue is caused by a 3rd party service disruption. We are investigating.
resolved
This incident has been resolved.
postmortem
### **Dates:**
Tuesday March 10th 2026, 17:52 - 23:34 UTC
### **What happened:**
The majority of iOS devices across all regions became unavailable.
### **Why it happened:**
Apple's [ppq.apple.com](http://ppq.apple.com) app verification endpoint was down, causing internal device monitoring checks to fail, bringing devices offline.
### **How we fixed it:**
We temporarily disabled these device monitoring checks.
### **What we are doing to prevent it from happening again:**
Improved external monitoring to catch outages of apple’s [ppq.apple.com](http://ppq.apple.com) endpoint, loosened device monitoring to not take down live iOS devices if [ppq.apple.com](http://ppq.apple.com) is down.
2026-March-6 Resolved Service Incident
Początek 6 marca 2026 21:38 UTC · 0m
Pending
resolved
Between 21:38 UTC and 23:11 UTC, our virtual iOS and MacOS live and automated device tests were failing to start in the EU Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
postmortem
### **Dates:**
Friday, March 6th 2026, 21:38 UTC - 23:11 UTC
### **What happened:**
During the incident timeline, customers running virtual iOS simulator tests on ARM or macOS ARM desktop tests in the EU Data Center were unable to start new sessions for either live or automated.
### **Why it happened:**
There was a sequencing issue on the release of the ARM side disk images in the EU.
### **How we fixed it:**
The image reference for the ARM side disk was rolled back to the previous reference to restore service.
### **What we are doing to prevent it from happening again:**
The tests that run to validate the image syncing have been completed in each region.