14:30 ve 15:30 UTC arasında, CA bölgesinde müşteriler Public API'ye erişim yaşadı. Kontrol katmanında bir hata dağıtımı, başarısız API talepleri ile sonuçlanan doğru şekilde reddedilmelidir.
Sorun tamamen çözüldü. Sonunda hiçbir eylem gerekli değildir. Bozukluk için özür dileriz.
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
Web SDK Outage (tüm bölgeler)
Başlangıç 1 Eylül 2026 13:21 UTC · 0m
OutageBüyük çaplı olay
resolved
Yaklaşık 1:21 PM UTC'den 1:46 PM UTC'ye kadar, Web SDK'dan gelen Doküman taleplerini kabul etmek için bir kesinti oldu ve Web modüllerini içeren akışlar için Güvenilir IDV SDK'lar.
Onfido Android & iOS SDKs etkilenmedi.
Bu rahatsızlıktan özür dileriz. A detailed postmortem soruşturmamızı sonlandırdığımız bir kez takip edecektir.
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
Şu anda ABD kümemizde yüz Benzerlik, Doküman ve Bilinen Faces raporlarının işlenmesi için gecikmeleri araştırıyoruz.
investigating
Arama grubumuzda bazı düğümlerin olağanüstü performanslarını indekslenen yüzler için görüyoruz. Kök nedenini ve olası remediasyonları araştırmaya devam ediyoruz.
identified
ABD arama grubumuzda bozulan performansın kök nedenini tespit ettik ve bunu çözmek için aktif olarak çalışıyoruz. Tamamen restore edilmiş bir kez takip edeceğiz.
monitoring
Bir düzeltme uyguladık ve şimdi geliştirilmiş olan işleme süresini izlemeye devam ettik.
monitoring
İşleme zamanı gelişiyor, tüm sistemler artık operasyonel.
resolved
İşleme süreleri normale geri döndü. Bu konu şimdi çözüldü. Yakında takip etmek için Post-mortem.
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
Kimlik Geliştirilmiş sorunlar
Başlangıç 24 Temmuz 2026 04:18 UTC · 2h 50m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Identity EnhancedIdentity Enhanced
investigating
Kimlik geliştirmesi, İngiltere dışında tüm dünyada başvuranlar için büyük bir düşüş yaşıyor
identified
Sorunu çözmek için üçüncü taraf sağlayıcımızla çalışıyoruz ve mevcut oldukları gibi güncelleştirmeleri paylaşacağız.
identified
Bu sorun için bir düzeltme üzerinde çalışmaya devam ediyoruz.
monitoring
Sağlayıcımızın hizmet kurtarmasını görüyoruz ve başarı oranlarının iyileştirilmesini talep ediyoruz. Durumu yakından izlemeye devam ediyoruz.
resolved
Sağlayıcı sorunu çözdü ve hizmet artık normal olarak çalışıyor
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
QES tasks cannot complete because our provider is encountering issues.
Başlangıç 23 Haziran 2026 02:38 UTC · 1h 11m
OutageKritik olay
Etkilenen bileşenler
QES
investigating
We cannot complete QES tasks because our electronic signature provider is encountering issues.
identified
Our provider is experiencing issues. All QES tasks cannot complete for now.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
monitoring
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
resolved
This incident has been resolved.
Severe service disruption across the platform in EU
Başlangıç 22 Haziran 2026 17:44 UTC · 12m
OutageKritik olay
Etkilenen bileşenler
QESFacial SimilarityDevice IntelligenceWebhooksAPIKnown facesIdentity EnhancedDashboardAutofillDocument VerificationWatchlistApplicant Form
monitoring
We're currently monitoring the EU cluster after an Amazon RDS issue.
resolved
We’re seeing recovery across our internal metrics, and processing has now returned to full capacity. At this time, the issue appears to be resolved.
Our current leading hypothesis is resource contention on a shared Amazon RDS instance that several services depend on, potentially related to a VACUUM operation running alongside a long-running job deleting a large volume of accumulated historical data. We have not yet confirmed the root cause and will continue investigating as follow-up, but service has been restored for now.
We'll be following up with a public post-mortem.
postmortem
**Incident date:** 22 June 2026
**Region:** EU \(eu-west-1\)
**Affected EU services**: API, Dashboard, Applicant Form, Document Verification,
Facial Similarity, Watchlist, Identity Enhanced, Webhooks, Known Faces, Autofill,
QES and Device Intelligence.
**Customer impact:** ~16:50–17:00 UTC \(acute degradation\); ~17:00–17:20 UTC \(backlog recovery\)
## Summary
On 22 June 2026, from approximately 16:50 UTC, a shared database cluster serving our EU region came under severe load and could not reliably serve queries for about 10 minutes. EU services returned elevated errors, and processing throughput briefly fell to ~20–35% of normal levels, with many subcomponents of our system \(e.g., Facial Similarity report processing\) being entirely disrupted, some others less heavily impacted \(e.g., Document report processing\). Service recovered by 17:01 UTC, the database fully stabilizing after an automatic failover \(~17:05–17:07 UTC\). A resultant report backlog was cleared by ~17:20 UTC.
Requests in flight during the acute degradation window may have failed unless retried; queued background work was processed automatically once the database recovered.
## Root cause
The incident was triggered by a routine database storage-reclamation task following standard scheduled data-deletion processing. This task normally completes without issue; why it failed on this occasion remains under investigation, although we observed that it was processing a larger-than-usual backlog. We have a support case open with our cloud provider to confirm a definitive root cause.
The reclamation task began to compete with normal application queries, which slowed as the database struggled to keep up. Applications opened more and more connections, leading to connection saturation and causing queries across the affected services to fail.
The database stabilized when an automatic failover to a healthy standby was triggered; the contention fully resolving with the failover to a new instance.
## Timeline \(UTC\)
* **16:50** — Our monitoring detected errors and elevated latency across EU services caused by resource contention on a shared database cluster.
* **16:53** — We start to see improvements, but system still not acting at normal levels of performance.
* **~17:00** — Customer-facing errors subsided and processing resumed as the contention eased.
* **17:01** — On-call engineers opened an incident and continued investigations.
* **~17:05–17:07** — The database performed an automatic failover to a healthy instance, which reset the overloaded writer and fully stabilised the cluster. The failover was triggered because of resource contention \(out of memory\) caused by the heavy vacuuming in the preceding minutes of the incident. Once the impacting vacuum operations had finished freeing up resources, we had started to see signs of improvement \(16:53—17:01\), but added latency in the feedback loop and aggregation window at AWS still decided to trigger the failover, even though we were already in a recovering state.
* **17:15–17:20** — Requests that had queued during the incident were worked through and the backlog returned to normal.
* **17:21–17:56** — We monitored the recovery and confirmed processing remained at full capacity.
## Remedies
* Reviewing connection limits and pooling so a single service cannot saturate a shared database, and evaluating dedicated database clusters per product to remove cross-service impact.
* Changing large historical-data deletions to run in smaller, throttled batches, and tuning database maintenance to avoid large catch-up operations.
* Adding earlier, proactive alerting on database memory, connections and load so we can intervene before customer impact.
* Continue working with our cloud provider on a definitive root cause.
Electronic signature increased TaT
Başlangıç 17 Haziran 2026 04:59 UTC · 4h 24m
IssuesKüçük çaplı olay
Etkilenen bileşenler
QES
identified
One of our providers is still experiencing issues, resulting in higher TaT for electronic signature tasks.
identified
Our provider is actively working on fixing the issue.
identified
Our provider is still working on fixing the issue.
We're monitoring the impact and we're making sure to keep the TaT as low as possible, considering the situation.
identified
Our provider's error rate dropped significantly. The TaT of QES tasks is now close to the usual value.
identified
We're working closely with our provider to find the reason of the remaining errors.
The error rate contacting our provider is stable and the impact on the QES task TaT is under control.
monitoring
Our provider was able to fix the issue and we don't get any error when contacting them.
The TaT of QES tasks is back to normal.
Monitoring the situation to make sure the errors don't come back.
resolved
The QES tasks TaT is back to normal.
Electronic signature high TaT
Başlangıç 16 Haziran 2026 22:19 UTC · 58m
IssuesKüçük çaplı olay
Etkilenen bileşenler
QES
identified
Our electronic signature tasks are taking more time to complete because one of our providers is doing some maintenance.
identified
Our provider is still not able to answer to all requests, resulting in longer-than-usual electronic signature task processing time.
resolved
The service is operational again.
Sporadic PDF generation issue
Başlangıç 15 Haziran 2026 09:30 UTC · 2d 1h
Pending
resolved
Incident was solved.
postmortem
## Summary
Between May and June 2026, around 0.037% of the evidence files generated on the platform were incomplete and included in the corresponding evidence folders.
There was no impact on workflow results, and no incorrect information was ever displayed — the affected files were simply empty rather than wrong.
A small number of requests to generate timeline files were also affected \(around 0.073%\).
## Root Causes
The issue was introduced during an initiative to improve how timeline and evidence PDFs are generated — making them significantly smaller, more efficient, and more resilient to produce. As part of these changes, in rare cases the new flow attempted to generate the PDF before its content had fully loaded, resulting in a blank or nearly empty file.
Because the problem was intermittent and occurred only in rare timing conditions, it was not consistently reproduced during rollout testing.
## Timeline
* 15 June 2026: Engineering investigation found an incomplete PDF and started the root cause investigation
* 17 June 2026: Root cause investigation completed and a safeguard to prevent incomplete PDFs was deployed
## Remedies
* We corrected the PDF generation flow so PDFs are only printed after content has fully loaded, and added a safeguard that stops generation when an incomplete PDF is detected.
* We also improved our monitoring and test coverage for PDF generation, so similar issues are caught earlier in the future.
Partial outage for Watchlist reports
Başlangıç 6 Haziran 2026 14:08 UTC · 1h 41m
OutageBüyük çaplı olay
Etkilenen bileşenler
WatchlistWatchlist
investigating
Some Watchlist search search profiles are not working correctly the impacted reports are not being processed. We are investigation the root cause.
identified
We identified the issue and are implementing a fix.
monitoring
The issue has been fixed, and the reports are being processed correctly. We are still monitoring the service.
resolved
This incident has been resolved.
Degraded performance for Identity Reports in UK jurisdiction
Başlangıç 11 Mayıs 2026 23:11 UTC · 14h 3m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Identity Enhanced
identified
We are facing issues with one of our providers and we see a slight decrease in clear rate for Identity Reports in UK jurisdiction
monitoring
A fix has been implemented and we are monitoring the results. All clear rates should be back to normal
resolved
All reports are back to normal.
QES tasks cannot complete because our provider is encountering issues.
Başlangıç 9 Mayıs 2026 11:27 UTC · 37m
OutageBüyük çaplı olay
Etkilenen bileşenler
QES
identified
We cannot complete QES tasks because our electronic signature provider is encountering issues.
identified
The provider is working on a fix.
monitoring
Our provider fixed the issue. QES tasks are completing now.
Tasks started during the incident were resumed.
resolved
The service is operational, no error where detected after the fix.
QES tasks outage
Başlangıç 5 Mayıs 2026 06:52 UTC · 1h 47m
OutageKritik olay
Etkilenen bileşenler
QES
investigating
We noticed an issue with our QES tasks, we're investigating.
identified
Our provider is experiencing issues. All QES tasks cannot complete for now.
identified
Our provider is fixing the issue. QES tasks still cannot complete.
monitoring
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
resolved
QES is operational again.
QES tasks degraded performance
Başlangıç 21 Nisan 2026 14:28 UTC · 2h 45m
OutageKritik olay
Etkilenen bileşenler
QES
identified
We identified an issue with one of our provider and some QES tasks may encounter some problems.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
identified
Our provider continues experiencing issues and is working on a fix.
identified
Our provider is slowly recovering, we expect some QES tasks to complete.
resolved
QES is operational again.
QES tasks partial outage
Başlangıç 3 Şubat 2026 10:42 UTC · 2h 34m
OutageBüyük çaplı olay
Etkilenen bileşenler
QES
identified
One of our provider is experiencing issues. QES capture tasks will fail and QES verification tasks will have an increase in TaT.
identified
The provider is working on a fix, QES is still suffering a partial outage.
identified
The provider is still working on a fix.
identified
The provider is still working on a fix.
QES verification tasks may start to time out, depending on their configuration, as we passed 90 minutes of downtime.
monitoring
The provider fixed the issue. We can confirm our QES tasks are now processed correctly.
We'll monitor the situation to make sure the error is indeed fixed.
resolved
We can confirm the error stopped and services are stable and operational.
QES tasks degraded performance
Başlangıç 2 Şubat 2026 15:36 UTC · 1h 48m
OutageBüyük çaplı olay
Etkilenen bileşenler
QES
identified
We identified an issue with one of our provider and some QES tasks may encounter some problems.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
monitoring
Our provider is slowly recovering, we expect some QES tasks to complete.
monitoring
We're now seeing the error rate decrease. Most of the QES tasks should complete.
resolved
QES is operational again.
Service Degradation - Manual Tasks
Başlangıç 28 Ocak 2026 11:16 UTC · 4h 7m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Document Verification
investigating
We are currently investigating this issue.
monitoring
Issue found and fixed.
Increased turn around time for manual reports.
Estimated time to live manual processing is 4h.
resolved
Incident is fully resolved, manual processing is now working normally.
Manual reports will keep having an increased turn around time for a few more hours while it works through the task backlog.
postmortem
### Summary
Manual task assignment for all EU customers stopped working between 10h50 UTC and 11:50 UTC. This led to an increase in manual processing Turnaround Time \(TaT\) affecting approximately 20% of our document verification volumes with all customers recovering to TaT SLA by 18h00 UTC.
During this period:
* All checks that required **manual review** showed an increase in TaT.
* **Fully automated reports were not affected** and continued to run as normal.
The issue was caused by a **configuration error in our internal task management system**, which prevented it from correctly assigning tasks to our analysts.
We fixed the configuration and **restored normal processing** by 28 Jan 2026 11h50 AM UTC, and cleared all manual task backlogs by 18h00 UTC.
We have updated our validation and deployment checks to prevent similar issues in the future.
### Root Causes
_Manual processing queue assignment was affected by an invalid manual configuration input. This single queue configuration parameter resulted in an error that affected assignments in all queues._
### Timeline
* all times in UTC:
_10:50: Configuration manually updated and errors started, no more tasks assigned._
_11:02: On-call is notified of a spike in manual system assignment errors through our monitoring_
_11:07: The error responsible for the spike is identified \(invalid UUID\)_
_11:10: Incident declared_
_11:45: Origin of invalid UUID is found_
_11:50: Bad configuration parameter is deleted_
_11:56: Configuration reintroduced correctly_
_18:00: Recovered from manual task backlog_
### Remedies
* Adding appropriate input configuration value validation
* Improve task assignment resilience to these types of errors
* Review configuration guidance and post-release monitoring
Document report processing disrupted in EU
Başlangıç 26 Ocak 2026 12:00 UTC · 0m
OutageBüyük çaplı olay
resolved
Between 12:15 UTC and 12:40 UTC there was disruption to document report processing in EU. Autofill requests also saw a disruption between 12:16 UTC and 12:27 UTC. A postmortem will follow.
postmortem
## Summary
On 26 January 2026 between 12.16 UTC and 12.35 UTC, our document reports processing was severely degraded with customers experiencing extended processing time to about 85% of their traffic. The incident also caused extended processing time on Biometric Authentications and Biometric Verifications for 20% of the traffic between 12.25 UTC and 12.35 UTC.
## Root Causes
A core service for fraud prevention on documents came under elevated load, causing kubernetes pods to go down in quick sequence. Our upstream service retry policy proved too aggressive to let the service recover and required manual scaling-up.
The retry policy also caused elevated load on a shared database which in turn also affected the biometrics service.
## Timeline
* **12:16 UTC** – Our monitoring detected a sharp increase in errors when processing document reports.
* **12:19** **UTC** – The on‑call team was alerted and began investigating the affected document‑processing service.
* **12:25 UTC** – We identified that the incident was also affecting a small portion of biometric checks, leading to some failures and short delays.
* **12:29** **UTC** – On-call engineers manually increased capacity for the impacted document fraud‑prevention service.
* **12:35** **UTC** – Error rates for both document reports and biometric checks returned to normal and new requests were being processed successfully.
* **12:45–13:06** **UTC** – We processed document reports that were impacted during the incident window.
* **13:15** **UTC** – We confirmed that all affected document and biometric requests had completed successfully and marked the incident as resolved.
## Remedies
* We have continued to fine-tune our auto-scaling parameters of the affected service to scale up with lower CPU targets.
* We modified the retry policy of fraud services to avoid overloading an already struggling service so that it can auto-recover.
* We made our load-tests on this service more representative of production traffic \(similar image sizes/document type distribution…\) and test for accelerated traffic spikes.
* We lowered the total amount of shared database connections the document processing can take to avoid noisy neighbour impact on biometrics processing.
Document Reports not being processed in US
Başlangıç 21 Ocak 2026 11:05 UTC · 1h 29m
OutageKritik olay
Etkilenen bileşenler
AutofillDocument Verification
investigating
We have a high error rate in document report processing and autofill. Our team is investigating
identified
We have identified the issue as a configuration issue between services, and are working on a fix.
monitoring
We have deployed a fix. New document reports and autofill requests are being processed normally. We will rerun the earlier impacted reports.
resolved
This issue is now resolved:
Between 10:30 and 12:00 UTC, document reports could not be processed and autofill requests were failing in the US region. Starting from 12:00, live traffic was unaffected. From 12:00 to 12:30 UTC, pending reports were processed.
We take a lot of pride in running a robust, reliable service and we're working hard to make sure this does not happen again. A detailed postmortem will follow once we've concluded our investigation.
postmortem
### Summary
Around 10:29 UTC, around 10 minutes after restoring service in the EU and CA from the 10:05-10:17 UTC outage, we go alerted for high error rate in the US region and we noticed the new ML model version was back online in that region. This impacted Turnaround Time \(TaT\) for all Document Report and Studio Autofill until 12:11 UTC.
Autofill Classic had an average 100% error rate during the same time frame.
All impacted Document Reports and Studio Autofill tasks were completed successfully with an average TaT of ~40 minutes.
### Root Causes
A previously rolled back version of an ML model went back live automatically in the US region. The automated canary reversal did not succeed in that region, leaving the deployment in an inconsistent state. As a result, the model routing labels were not updated as expected during the next forced deployment, and traffic continued to be served by pods running the incorrect model version until we manually rolled back the model in the cluster. This was the result of a bug in our version of Helm.
### Timeline
_10:29 UTC: The new model release went back live in the US region only, after the previous automated canary rollback_
_11:44 UTC: We manually re-deployed a previous version via CI/CD pipeline. This was not successful and errors continued._
1_2:00 UTC: We manually rolled back the version of the model directly in the cluster._
1_2:11 UTC: All services went back to normal_
### Remedies
* We will fix our deployment tooling to reliably apply model-routing label changes to these resources by upgrading Helm.
Document report processing disrupted in EU
Başlangıç 21 Ocak 2026 10:30 UTC · 0m
Pending
resolved
Between 10:30 UTC and 10:47 UTC there was disruption to document report processing in EU, as a side effect of reports delayed from the earlier incident being re-processed. Further details will follow in a postmortem.
postmortem
### Summary
For the EU region, one critical service struggled to reprocess the traffic affected by a previous faulty release which led to higher Turnaround Time \(TaT\) for all Document reports created between 10:32 and 10:50 UTC.
All impacted Document Reports were completed successfully with an average TaT of ~6 minutes.
### Root Causes
The re-processing batch caused a spike in traffic and the auto-scaling didn’t work as expected for one critical service. The service entered a crash loop state and had to be manually up-scaled to recover. An unbounded number of in-flight requests were accepted leading to memory exhaustion and I/O event loop non-responsiveness, while waiting for a downstream ML inference service to scale up.
### Timeline
_10:32 UTC: The critical service went up to a 100% error rate_
_10:33 UTC: Engineers who initiated the report backlog reprocessing become aware of the issue through our monitoring and start investigating_
_10:48 UTC: We manually scaled up the critical service_
1_0:50 UTC: The service went back to normal and errors stopped_
### Remedies
* Investigate how to reduce memory footprint on this service, allowing bigger request queues while it waits for downstream ML model serving to scale up
* Change auto-scaling parameters to be more aggressive \(i.e. scale with lower CPU targets\)
* Add concurrent requests monitoring
* Reduce ML model serving image sizes for faster scaling of inference services
* Improve back-pressure mechanisms to be able to sustain minimum traffic levels independent of spikes while auto-scaling kicks-in
* Change our weekly load testing scripts to specifically test for accelerated traffic spikes