Vedem rate ridicate de eroare pentru testele Real Device în centrul de date UE-Central-1. Investigăm.
identified
Cauza principală a ratelor ridicate de eroare pentru testele Real Device în centrul de date UE-Central-1 a fost identificată. Remediarea a fost aplicată și monitorizăm în mod activ situația.
resolved
După ce au luat măsuri de remediere, ratele de eroare de încercare Real Device au revenit la normal în centrul de date UE-Central-1. Toate serviciile sunt pe deplin operaționale.
Traducere automată din actualizarea oficială a incidentului.
2026-august-25 Incident de serviciu
A început 25 august 2026 la 12:33 UTC · 46m
Pending
investigating
În prezent vedem rate ridicate de eroare pentru testele MacOS și iOS în centrul de date US-West-1. În prezent investigăm.
resolved
După ce au luat măsuri de remediere, toate testele MacOS şi iOS sunt acum efectuate conform aşteptărilor în centrul de date US-West-1. Toate serviciile sunt pe deplin operaționale.
postmortem
Traducerea şi adaptarea:
Marţi, 25 august 2026, 11:38, 13:10
Ce s-a întâmplat?
Clienţii care efectuează teste MacOS şi iOS în regiunea noastră SUA-Vest au avut servicii degradate. Aproximativ 50% din capacitatea virtuală Mac în regiune a încetat să mai accepte noi teste, astfel încât testele fie coadă sau nu a reușit să înceapă. Capacitatea rămasă a fost supusă unei presiuni suplimentare pe măsură ce munca a trecut pe ea, ceea ce a prelungit timpul de pornire atât pentru testele de birou cât și pentru simulatoare.
De ce s-a întâmplat:
Un certificat intern de securitate folosit de gazdele noastre Mac pentru a ajunge la un serviciu de sprijin cloud a ajuns la data de expirare. Odată ce a expirat, gazdele nu mai puteau stabili o legătură de încredere cu acest serviciu și nu mai puteau furniza noi mașini de testare. Certificatul a fost eliberat manual și nu a avut nici reînnoire automată și nici alertă de expirare.
Cum am reparat-o:
Am emis și implementat un certificat de înlocuire, care a restaurat conectivitatea și a returnat capacitatea Mac la niveluri normale.
Ce facem pentru a preveni să se întâmple din nou:
Mutăm aceste certificate în reînnoirea automată și adăugăm alertarea astfel încât acestea să fie înlocuite cu mult înainte de expirare, împreună cu revizuirea instrumentelor din jur pentru a asigura siguranța manipulării certificatelor.
Traducere automată din actualizarea oficială a incidentului.
2026-august-07 Incident de serviciu
A început 7 august 2026 la 12:08 UTC · 58m
Pending
investigating
Vedem în prezent rate ridicate de eroare pentru testele virtuale desktop în US-West-1 datacenter. În prezent investigăm.
resolved
After taking remedial action, Virtual Desktop tests are starting successfully in the US-West-1 Data center. All services are fully operational. We are closely monitoring the situation
postmortem
### **Dates:**
Friday, August 7th 2026, 11:00 UTC - 13:04 UTC.
### **What happened:**
Windows and Intel Mac jobs in `us-west1` failed to start due to virtual machine \(VM\) allocation starvation.
### **Why it happened:**
A service crash loop left VMs in an allocated but unclaimed state, while a cleanup bug prevented the system from releasing the orphaned capacity to boot new VMs.
### **How we fixed it:**
Restored service stability and cleared stale allocations to resume VM provisioning and clear queued jobs.
### **What we are doing to prevent it from happening again:**
Fixing allocator cleanup logic, strengthening deployment health checks, and improving capacity accounting for stale allocations.
Traducere automată din actualizarea oficială a incidentului.
2026-Iulie-16 Incident de serviciu rezolvat - Raportarea erorilor (Backtrace)
A început 16 iulie 2026 la 13:00 UTC · 0m
Pending
resolved
Între 16 iulie 13:03 UTC și 16 iulie 17:33 UTC, am experimentat o perturbare a serviciilor în care arhivele cu simboluri care utilizau protocoale de încărcare cu mai multe părți nu au reușit să proceseze. Incidentul a fost rezolvat și toate serviciile sunt operaționale.
postmortem
### **Dates:**
Thursday July 16th 2026, 13:03 – 17:33 UTC
### **What happened:**
Symbol archives uploaded in multiple parts failed to process. All regions were affected.
### **Why it happened:**
Our symbol processing service sent an upload-verification field that our cloud storage provider's API does not accept for multi-part uploads, so those uploads were rejected. A fix for this had already been developed, but it had not yet been included in a released build and the service was not configured to use it. As in the first incident, rejected uploads were retried and accumulated on local disk.
### **How we fixed it:**
We deployed a build containing the fix and corrected the service configuration on the affected workers. A large backlog of uploads then processed, which briefly re-filled the disks before draining completely.
### **What we are doing to prevent it from happening again:**
We are releasing the fix formally through our build pipeline and persisting the corrected configuration in our configuration management, so it can't be lost. The disk-utilization alerting added after the first incident also covers the disk-exhaustion pattern common to both.
Traducere automată din actualizarea oficială a incidentului.
2026-Iulie-15 Incident al serviciului rezolvat - Raportarea erorilor (Backtrace)
A început 14 iulie 2026 la 23:30 UTC · 0m
Pending
resolved
Între 14 iulie 23:45 UTC şi 15 iulie 23:05 UTC, am avut o problemă în care arhiva de simboluri se încarcă la proiectele Backtrace a eşuat cu erori HTTP 400. Incidentul a fost rezolvat și toate serviciile sunt operaționale.
postmortem
### **Dates:**
Tuesday July 14th 2026, 23:45 UTC – Wednesday July 15th 2026, 23:05 UTC
### **What happened:**
Symbol archive uploads to Error Reporting \(Backtrace\) projects failed with HTTP 400 errors. All regions were affected.
### **Why it happened:**
The credentials our symbol processing service used to write to cloud storage were no longer valid, so uploads could not be stored. Failed uploads were retried repeatedly and accumulated on local disk until the service ran out of space, at which point it also began rejecting new uploads.
### **How we fixed it:**
We reissued the storage credentials, increased disk capacity on the affected workers, and restarted the service. The queued uploads then processed successfully, and we confirmed recovery with affected customers.
### **What we are doing to prevent it from happening again:**
We've added disk-utilization monitoring and alerting to this service so we detect the condition ourselves before it affects uploads, rather than relying on customer reports.
Traducere automată din actualizarea oficială a incidentului.
2026-Iulie-14 Incident de serviciu
A început 14 iulie 2026 la 17:19 UTC · 3h 4m
OutageIncident major
Componente afectate
EU-CentralEU-Central
investigating
Ne confruntăm cu o problemă în centrul de date central al UE 1, unde testele de aplicații iOS ARM Simulator nu reușesc să facă față erorilor de infrastructură. Investigăm.
investigating
Descărcările de aplicații durează mai mult decât se preconizează atunci când se observă aceste erori de infrastructură. Continuăm să investigăm.
monitoring
Am implementat o soluţie pentru această problemă. Testele ar trebui să aibă loc acum, după cum se prevede în centrul de date central al UE 1. Suntem de monitorizare.
monitoring
Continuăm să monitorizăm orice alte probleme.
resolved
După ce am luat măsuri de remediere şi am monitorizat situaţia, vedem că testele de aplicaţie iOS ARM se desfăşoară conform aşteptărilor din cadrul Centrului de Date al UE. Toate serviciile sunt pe deplin operaționale.
postmortem
Traducerea şi adaptarea:
Marţi, 14 iulie 2026, 09:00 - 18:28 UTC
Ce s-a întâmplat?
Testele IOS ARM efectuate în centrul nostru de date UE au eșuat.
De ce s-a întâmplat:
S-a produs o eroare în ceea ce privește furnizorul principal de rețele din centrul nostru de date din UE.
Cum am reparat-o:
Am eşuat în faţa unui furnizor secundar de reţea de la centrul nostru de date din UE.
Ce facem pentru a preveni să se întâmple din nou:
Ne-am îmbunătăţit monitorizarea şi alertarea cu privire la problemele de captură cu furnizorii de reţea terţi.
Traducere automată din actualizarea oficială a incidentului.
2026-Iulie-1 Incident de serviciu rezolvat 1
A început 1 iulie 2026 la 17:02 UTC · 0m
IssuesIncident minor
resolved
La 1 iulie între orele 17:02 UTC şi 21:31 UTC, testele MacOS 14 nu au început în centrele de date US-Vest şi UE-Central. Această problemă a fost rezolvată. Toate serviciile sunt pe deplin operaționale.
postmortem
### **Dates:**
Wednesday July 1st 2026, 17:02 UTC - 21:31 UTC
### **What happened:**
macOS 14 tests in the US West and EU Central data centers were unable to start.
### **Why it happened:**
An internal datasource was unavailable due to a missing configuration entry.
### **How we fixed it:**
The missing entry was replaced.
### **What we are doing to prevent it from happening again:**
Safeguards have been put in place to prevent in-use datasource removal from configuration.
Traducere automată din actualizarea oficială a incidentului.
2026-Iulie-01 Incident de serviciu
A început 1 iulie 2026 la 12:12 UTC · 46m
Pending
investigating
Avem o problemă cu accesarea Inspectorului Appium în centrele de date US-Vest-1, UE-Central-1, și SUA-Est-4. Investigăm.
resolved
Am identificat cauza principală şi am implementat o soluţie pentru această problemă. Toate serviciile sunt pe deplin operaționale.
postmortem
Traducerea şi adaptarea:
Miercuri, 1 iulie 2026, 09:37, 12:46
Ce s-a întâmplat?
Funcţia Inspectorului Appium, folosită în timpul sesiunilor de testare în direct a dispozitivului real, a devenit indisponibilă după o desfăşurare programată a UI. Utilizatorii care au încercat să folosească funcția în timpul unui test live au fost prezentați cu o eroare de 500, perturbăndu-și sesiunea de testare activă. Conductele automate de testare nu au fost afectate.
De ce s-a întâmplat:
O actualizare de rutină a unei biblioteci de bază a introdus o incompatibilitate cu logica de redare a inspectorului Appium. Versiunea anterioară a bibliotecii tolera modelul, dar versiunea actualizată nu. Problema nu a fost abordată înainte de desfășurare din cauza acoperirii insuficiente a testelor la sfârșit pentru această caracteristică specifică.
Cum am reparat-o:
Ne-am rostogolit înapoi UI la ultima versiune de lucru cunoscut pentru a restabili funcția imediat, și apoi a implementat un fix vizat pentru incompatibilitate.
Ce facem pentru a preveni să se întâmple din nou:
Adăugăm o acoperire de testare de la un capăt la altul pentru caracteristica Inspectorului Appium pentru a se asigura că aceasta este validată automat înainte de implementarea viitoare.
Traducere automată din actualizarea oficială a incidentului.
2026-iunie-24 Incident de serviciu rezolvat
A început 24 iunie 2026 la 16:39 UTC · 0m
Pending
resolved
Pe 24 iunie, între 08:28 UTC şi 15:32 UTC, am experimentat eşecuri la sesiunea de testare Real Device care au afectat Appium şi Access API în centrele de date din estul SUA, vestul SUA şi UE. Această problemă a fost rezolvată. Toate serviciile sunt pe deplin operaționale.
postmortem
Traducerea şi adaptarea:
Miercuri, 24 iunie 2026, 09:42 UTC - 15:45 UTC.
Ce s-a întâmplat?
Sesiuni de testare Real Device folosind Appium and Access API au înregistrat rate de eroare crescute în centrele de date din Estul SUA, Vestul SUA și UE.
Sesiunile s-au sincronizat după aproximativ 90 de secunde sau au rămas blocate în starea "Conectare," împiedicându-le să fie închise.
De ce s-a întâmplat:
A fost introdus un defect al produsului, ceea ce a făcut ca sesiunile de testare Real Device să nu înceapă și să împiedice închiderea sesiunilor active.
Cum am reparat-o:
Întoarce-te la o versiune stabilă.
Ce facem pentru a preveni să se întâmple din nou:
Să îmbunătățească monitorizarea și alertarea și să îmbunătățească validarea post-punere în aplicare.
Traducere automată din actualizarea oficială a incidentului.
2026-iunie-18 Incident de serviciu
A început 18 iunie 2026 la 05:30 UTC · 1h 29m
OutageIncident major
Componente afectate
EU-CentralEU-Central
investigating
În jur de 3:51 AM UTC am început să experimentăm o disponibilitate mai scăzută a dispozitivelor iOS în centrul de date central UE. Echipa noastră investighează în mod activ cauza fundamentală și lucrează spre o rezoluție.
monitoring
Ne-am confruntat cu indisponibilitatea dispozitivelor în centrul nostru de date UE-Central-1 și am identificat cauza principală. Am luat măsuri de remediere și în prezent monitorizăm.
resolved
După ce luăm măsuri de remediere, vedem acum dispozitive reale disponibile pentru testare în centrul de date UE-Central-1. Toate serviciile sunt pe deplin operaționale.
postmortem
Traducerea şi adaptarea:
Joi, 18 iunie 2026, 03:45 UTC - 06:45 UTC.
Ce s-a întâmplat?
Aproximativ 7% din dispozitivele iOS din centrul nostru de date din UE nu au fost disponibile temporar pentru sesiunile de testare a clienților, după o pierdere de putere la raftul care le găzduiește.
De ce s-a întâmplat:
Circuitul de alimentare cu energie al raftului afectat i - a depăşit capacitatea, iar un întrerupător de protecţie s - a împiedicat pentru a proteja puterea de tăiere a liniei de dispozitivele de pe acel raft până când circuitul a fost restabilit.
Cum am reparat-o:
Circuitul afectat a fost resetat, iar alimentarea a fost repusă în rack, returnând dispozitivele serviciului clienţilor.
Ce facem pentru a preveni să se întâmple din nou:
Redistribuim sarcina electrică pe rafturile afectate, adăugând monitorizarea capacității cu alerte de avertizare timpurie înaintea limitelor de circuit și introducând o revizuire a capacității înainte ca noile dispozitive să fie instalate pe un suport.
Traducere automată din actualizarea oficială a incidentului.
2026-June-1 Service Incident
A început 1 iunie 2026 la 14:02 UTC · 4h 22m
OutageIncident critic
Componente afectate
US-EastUS-EastUS-EastUS-EastUS-EastUS-EastUS-East
investigating
We are experiencing an issue where the US-EAST data center is currently unavailable. We are investigating.
investigating
We are continuing to investigate this issue.
investigating
We have identified the root cause and are working on implementing a fix.
monitoring
Access has been restored to the US-EAST data center. We are monitoring.
resolved
After taking remedial action, access to the US-EAST data center is fully restored. This incident is resolved.
postmortem
### **Dates:**
Monday, June 1st 2026, 13:06 UTC - 15:42 UTC.
### **What happened:**
Customers served by the US-EAST-4 region were unable to authenticate or start new test sessions because an incomplete TLS certificate chain was deployed to the core directory services.
### **Why it happened:**
A certificate extraction script defect silently truncated the certificate chain after the certificate authority transitioned to a longer hierarchy.
### **How we fixed it:**
Reverted the certificate rotation and re-applied the previous known ,good certificate to restore authentication services.
### **What we are doing to prevent it from happening again:**
Implementing pre-deployment certificate chain validation, adding active monitoring, and fixing the script's chain-length limitations.
2026-May-13 Resolved Service Incident
A început 21 mai 2026 la 14:53 UTC · 0m
Pending
resolved
Between May 13th 18:47 UTC and May 15th 9:23 UTC, we experienced test details not appearing in the dashboard of the US-East Data Center. This issue has been resolved. All services are fully operational.
postmortem
### **Dates:**
Wednesday, May 13th 2026, 18:47 UTC - Friday, May 15th 2026, 09:23 UTC.
### **What happened:**
Customers were unable to access test run summaries for Real Device Cloud \(RDC\) jobs because events stopped publishing to the jobs Kafka topic in US-EAST.
### **Why it happened:**
An authentication key used by the message producer unexpectedly lost its permissions during an account cleanup.
### **How we fixed it:**
Manually restored the required permissions to re-establish the connection and resume service.
### **What we are doing to prevent it from happening again:**
Migrating to a permanent service account, implementing a Dead Letter Queue \(DLQ\) for the jobs Kafka topic, and replaying the missing events to restore customer data.
2026-April-23 Resolved Service Incident
A început 24 aprilie 2026 la 16:33 UTC · 0m
OutageIncident major
resolved
Between April 23rd 22:44 and April 24th 15:25 UTC, there was a technical issue that affected video recordings for tests running on macOS 15 and iOS within our EU and US-West Data Center. We identified the issue and deployed a fix. All systems are now fully operational.
postmortem
### **Dates:**
Thursday, April 23rd 2026, 22:43 UTC - Friday, April 24th 2026, 15:29 UTC
### **What happened:**
Video assets were missing for virtual iOS simulator tests on ARM and macOS ARM desktop tests in the US-West and EU data centers.
### **Why it happened:**
A product defect was introduced resulting in a screen capture failure.
### **How we fixed it:**
We performed a rollback to a stable version.
### **What we are doing to prevent it from happening again:**
We are improving monitoring & alerting to enhance our post deployment validation.
2026-April-16 Resolved Service Incident
A început 16 aprilie 2026 la 10:10 UTC · 0m
Pending
resolved
Between 02:00 and 11:15 CEST, live and automated tests on iOS 17.0 simulators were failing to start in the EU and US-West Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
postmortem
### **Dates:**
Thursday, April 16th 2026, 00:00 UTC – 09:15 UTC
### **What happened:**
Live and automated tests on iOS 17.0 simulators failed to start in both the EU and US-West data centers. Customers running tests on iOS 17.0 Intel-based simulators were unable to execute their tests for approximately 9 hours.
### **Why it happened:**
A deployment introduced an incompatibility affecting iOS 17.0 on Intel-based infrastructure. The issue was not caught prior to release due to insufficient post-deployment test coverage for that specific simulator configuration.
### **How we fixed it:**
We performed a rollback to the previous deployment, which restored full iOS 17.0 simulator functionality.
### **What we are doing to prevent it from happening again:**
We are reviving and expanding automated post-deployment tests to cover a broader range of simulator configurations, including legacy Intel-based iOS versions, to catch incompatibilities before they reach production.
We are currently investigating reports of test failures affecting users running tests using SauceCtl in our US-West-1 and EU-Central-1 Data Center. We are investigating.
resolved
We have identified the root cause and have deployed a fix for this issue. All services are fully operational.
postmortem
### **Dates:**
Monday April 7th 2026, ~11:00 – 15:55 UTC
### **What happened:**
Some customers experienced 503 errors when running tests via saucectl. The test-composer service was intermittently unavailable, preventing framework-based test execution.
### **Why it happened:**
A stale Docker image was deployed to the test-composer service due to a packaging issue that arose during an internal container registry migration. This caused service pods to crash.
### **How we fixed it:**
We identified the stale image and redeployed the correct version, restoring the service.
### **What we are doing to prevent it from happening again:**
We are hardening our image deployment pipeline and adding validation checks to ensure container registry migrations do not result in stale or incorrect images being deployed to production.
2026-March-24 Resolved Service Incident
A început 24 martie 2026 la 17:36 UTC · 0m
Pending
resolved
Between 09:32 and 15:13 UTC, we identified a technical issue affecting iOS tests when running with network capture enabled. We've resolved the underlying cause and tests are working as expected. All services are fully operational.
postmortem
### **Dates:**
Tuesday, March 24th 2026, 09:32 UTC – 15:13 UTC
### **What happened:**
Network calls failed on iOS devices during Real Device Cloud sessions where network capture was enabled. Approximately 12-13% of iOS sessions were affected. Android was not impacted.
### **Why it happened:**
A deployment introduced a DNS resolution change that was incompatible with the iOS platform, causing network capture to break.
### **How we fixed it:**
Rolled back the deployment to restore service.
### **What we are doing to prevent it from happening again:**
Adding synthetic tests to catch network capture regressions before production, and implementing monitoring alerts for faster detection after deployments.
2026-March-19 Service Incident
A început 19 martie 2026 la 09:51 UTC · 1h 2m
OutageIncident major
Componente afectate
US-WestUS-West
investigating
Around 4:45 AM UTC we started experiencing lower iOS device availability in the US-West data center. Our team is actively investigating the root cause and working toward a resolution.
resolved
This incident has been resolved and our services are fully operational.
postmortem
### **Dates:**
Wednesday, March 19 2026, 04:45 UTC - 10:47 UTC.
### **What happened:**
Approximately 15% of iOS devices in our US-West data center were temporarily unavailable for customer test sessions due to failed internet connectivity checks.
### **Why it happened:**
An automated wireless network optimization feature adjusted transmit power levels on access points serving the affected devices, degrading wireless connectivity and causing devices to fail their availability checks.
### **How we fixed it:**
The affected access points were identified and restarted, restoring normal wireless connectivity.
### **What we are doing to prevent it from happening again:**
Evaluation of the automated optimization tools and a monitoring improvement.
2026-March-13 Resolved Service Incident
A început 13 martie 2026 la 14:30 UTC · 0m
Pending
resolved
Between 14:43 and 15:11 UTC on March 13, a small subset of Real Devices (iOS and Android) became unavailable across all our data centers. After taking remedial action, the issue was identified and resolved. All services are fully operational.
postmortem
### **Dates:**
Friday, March 13th 2026, 14:43 UTC - 15:11 UTC.
### **What happened:**
Real Devices \(iOS and Android\) availability gradually decreased across all data centers.
### **Why it happened:**
A product defect was introduced resulting in a small subset of Real Devices \(~10%\) failing to maintain required connectivity.
### **How we fixed it:**
Rollback to a stable version.
### **What we are doing to prevent it from happening again:**
Improve monitoring & alerting, enhance post deployment validation.
2026-March-10 Service Incident
A început 10 martie 2026 la 18:46 UTC · 4h 49m
OutageIncident major
Componente afectate
US-WestUS-WestEU-CentralEU-CentralUS-EastUS-East
investigating
We are experiencing device unavailability in the US West 1, EU Central 1, and US East 4 data centers and have found that the issue is caused by a 3rd party service disruption. We are investigating.
resolved
This incident has been resolved.
postmortem
### **Dates:**
Tuesday March 10th 2026, 17:52 - 23:34 UTC
### **What happened:**
The majority of iOS devices across all regions became unavailable.
### **Why it happened:**
Apple's [ppq.apple.com](http://ppq.apple.com) app verification endpoint was down, causing internal device monitoring checks to fail, bringing devices offline.
### **How we fixed it:**
We temporarily disabled these device monitoring checks.
### **What we are doing to prevent it from happening again:**
Improved external monitoring to catch outages of apple’s [ppq.apple.com](http://ppq.apple.com) endpoint, loosened device monitoring to not take down live iOS devices if [ppq.apple.com](http://ppq.apple.com) is down.
2026-March-6 Resolved Service Incident
A început 6 martie 2026 la 21:38 UTC · 0m
Pending
resolved
Between 21:38 UTC and 23:11 UTC, our virtual iOS and MacOS live and automated device tests were failing to start in the EU Data Center. We executed a deployment rollback, which restored services. All systems are now fully operational.
postmortem
### **Dates:**
Friday, March 6th 2026, 21:38 UTC - 23:11 UTC
### **What happened:**
During the incident timeline, customers running virtual iOS simulator tests on ARM or macOS ARM desktop tests in the EU Data Center were unable to start new sessions for either live or automated.
### **Why it happened:**
There was a sequencing issue on the release of the ARM side disk images in the EU.
### **How we fixed it:**
The image reference for the ARM side disk was rolled back to the previous reference to restore service.
### **What we are doing to prevent it from happening again:**
The tests that run to validate the image syncing have been completed in each region.