Between 14:30 and 15:30 UTC, customers in the CA region experienced degraded access to the Public API. A faulty deployment in the authentication layer caused valid tokens to be incorrectly rejected, resulting in failed API requests.
The issue has been fully resolved. No action is required on your end. We apologise for the disruption.
Web SDK Outage (all regions)
Started September 1, 2026 at 1:21 PM UTC · 0m
OutageMajor incident
resolved
From approximately 1:21 PM UTC to 1:46 PM UTC, there was a disruption to accepting Document requests from the Web SDK as well as Entrust IDV SDKs for flows containing web modules.
Onfido Android & iOS SDKs were not impacted.
We apologize for this disruption. A detailed postmortem will follow once we've concluded our investigation.
Increase in Known Faces consider rate
Started August 26, 2026 at 3:30 PM UTC · 0m
OutageMajor incident
resolved
Known Faces experiencing higher consider rates across all regions.
postmortem
## Summary
On August 26, 2026, an increase in the consider rate of Known Faces was observed during the period between 15:00 and 18:00 UTC. The overall observed consider rate of Known Faces increased from ~6% to ~23% across all regions. Perceived increase may have been higher or lower, depending on the client’s specific geography or vertical.
The was driven by a change to how matching scores are calculated for Known Faces matches, which inflated scores for non-matching faces and thus causing many reports to return false matches and a **consider** result.
## Root Causes
The change introduced a new algorithm for scoring matching faces, which is intended to improve Known Faces recall performance. This was extensively evaluated and validated by internal benchmarks before being released. Inadvertently, the transition from benchmarks to production partially failed to account for particular production environment characteristics. At the root, a single constant wasn’t properly converted in the transition, resulting in wrong behavior of the new scoring algorithm.
## Timeline
15:20 UTC: Change to Known Faces is released
17:35 UTC: Increased consider rate on Known Faces is noticed.
17:50 UTC: Change is rolled back. Consider rate drops soon after.
## Remedies
We are reviewing internal benchmarks and testing procedures which can increase the safety of this kind of change. Specifically, ensuring internal testing reflects the real production landscape when it comes to datasets and that transitions from internal benchmarks and experiments to production are thoroughly tested.
In addition, change procedures will be carried out with tighter monitoring so it allows us to notice and act faster.
Spike in Visual Authenticity failures on Facial Similarity Motion randomness checks
Started August 24, 2026 at 5:00 PM UTC · 0m
OutageMajor incident
resolved
Facial Similarity Motion checks using randomness challenges are failing at a much higher rate than normal in the EU region.
postmortem
## Summary
On 24 August 2026, between 15:00 UTC and 06:53 UTC the following morning, Facial Similarity Motion checks using randomness challenges failed at a much higher rate than normal in the EU region. A stricter validation was applied to all customers in the region when it should have been limited to a single customer, causing verifications that would otherwise have passed to be rejected. Affected reports were returned with a "Visual Authenticity" failure.
Because Motion checks are fully automated, the end users behind these reports were rejected with no manual review, and some attempted the check again and were rejected a second time.
The clear rate for these checks fell from around 97% to around 54% for approximately 1.67% of Facial Similarity Motion reports were affected. Disabling the change restored normal clear rates immediately. Other Motion checks and other regions were unaffected.
## Root Causes
A stricter validation of the motion randomness challenge was being trialled with a single customer, controlled by a per-customer configuration setting. A bug in our code meant the setting was evaluated without the customer identifier, so the stricter validation was applied to all customers in the EU region running Motion with randomness checks rather than to the intended one. Verifications that would have passed under the standard validation were rejected instead, and reported under the "Visual Authenticity" breakdown. No customer configuration or submitted data was at fault.
## Timeline
All times UTC.
* 24 August, 15:00: we enabled a stricter validation for motion randomness challenges in the EU region. Motion checks using randomness began failing at a much higher rate.
* 24 August, 16:36: a customer reported a drop in workflow success rates and we began investigating.
* 25 August, 04:18: we confirmed the elevated failure rate and escalated it.
* 25 August, 06:53: we reverted the change and clear rates returned to normal within minutes.
* 26 August: we deployed a permanent fix.
## Remedies
* Strengthen the controls around customer-specific configuration changes to ensure they cannot be applied more broadly than intended.
* Improve monitoring and alerting for Motion randomness clear rate checks so that unexpected changes in rejection rates are detected automatically and escalated more quickly.
We are currently investigating delays in our US cluster for the processing of Facial Similarity, Document and Known Faces reports.
investigating
We are seeing degraded performance of some nodes in our search cluster for indexed faces. We continue to investigate the root cause and possible remediations.
identified
We've identified the root cause of degraded performance in our US search cluster and are actively working to resolve it. We'll follow up once fully restored.
monitoring
We implemented a fix and continue to monitor the quality of processing time, which is now improved.
monitoring
Processing time keeps improving, all systems are now operational.
resolved
Processing times are back to normal. This issue is now resolved. Post-mortem to follow soon.
Identity Enhanced issues
Started July 24, 2026 at 4:18 AM UTC · 2h 50m
IssuesMinor incident
Affected components
Identity EnhancedIdentity Enhanced
investigating
Identity Enhanced is experiencing a major decrease in clear rates for applicants all other the world except UK
identified
We are working with our third-party provider to resolve the issue and will share updates as they become available.
identified
We are continuing to work on a fix for this issue.
monitoring
We are seeing our provider's service recovery and request success rates improving. We continue to monitor the situation closely.
resolved
The provider has resolved the issue, and the service is now operating normally
QES tasks cannot complete because our provider is encountering issues.
Started June 23, 2026 at 2:38 AM UTC · 1h 11m
OutageCritical incident
Affected components
QES
investigating
We cannot complete QES tasks because our electronic signature provider is encountering issues.
identified
Our provider is experiencing issues. All QES tasks cannot complete for now.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
monitoring
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
resolved
This incident has been resolved.
Severe service disruption across the platform in EU
Started June 22, 2026 at 5:44 PM UTC · 12m
OutageCritical incident
Affected components
QESFacial SimilarityDevice IntelligenceWebhooksAPIKnown facesIdentity EnhancedDashboardAutofillDocument VerificationWatchlistApplicant Form
monitoring
We're currently monitoring the EU cluster after an Amazon RDS issue.
resolved
We’re seeing recovery across our internal metrics, and processing has now returned to full capacity. At this time, the issue appears to be resolved.
Our current leading hypothesis is resource contention on a shared Amazon RDS instance that several services depend on, potentially related to a VACUUM operation running alongside a long-running job deleting a large volume of accumulated historical data. We have not yet confirmed the root cause and will continue investigating as follow-up, but service has been restored for now.
We'll be following up with a public post-mortem.
postmortem
**Incident date:** 22 June 2026
**Region:** EU \(eu-west-1\)
**Affected EU services**: API, Dashboard, Applicant Form, Document Verification,
Facial Similarity, Watchlist, Identity Enhanced, Webhooks, Known Faces, Autofill,
QES and Device Intelligence.
**Customer impact:** ~16:50–17:00 UTC \(acute degradation\); ~17:00–17:20 UTC \(backlog recovery\)
## Summary
On 22 June 2026, from approximately 16:50 UTC, a shared database cluster serving our EU region came under severe load and could not reliably serve queries for about 10 minutes. EU services returned elevated errors, and processing throughput briefly fell to ~20–35% of normal levels, with many subcomponents of our system \(e.g., Facial Similarity report processing\) being entirely disrupted, some others less heavily impacted \(e.g., Document report processing\). Service recovered by 17:01 UTC, the database fully stabilizing after an automatic failover \(~17:05–17:07 UTC\). A resultant report backlog was cleared by ~17:20 UTC.
Requests in flight during the acute degradation window may have failed unless retried; queued background work was processed automatically once the database recovered.
## Root cause
The incident was triggered by a routine database storage-reclamation task following standard scheduled data-deletion processing. This task normally completes without issue; why it failed on this occasion remains under investigation, although we observed that it was processing a larger-than-usual backlog. We have a support case open with our cloud provider to confirm a definitive root cause.
The reclamation task began to compete with normal application queries, which slowed as the database struggled to keep up. Applications opened more and more connections, leading to connection saturation and causing queries across the affected services to fail.
The database stabilized when an automatic failover to a healthy standby was triggered; the contention fully resolving with the failover to a new instance.
## Timeline \(UTC\)
* **16:50** — Our monitoring detected errors and elevated latency across EU services caused by resource contention on a shared database cluster.
* **16:53** — We start to see improvements, but system still not acting at normal levels of performance.
* **~17:00** — Customer-facing errors subsided and processing resumed as the contention eased.
* **17:01** — On-call engineers opened an incident and continued investigations.
* **~17:05–17:07** — The database performed an automatic failover to a healthy instance, which reset the overloaded writer and fully stabilised the cluster. The failover was triggered because of resource contention \(out of memory\) caused by the heavy vacuuming in the preceding minutes of the incident. Once the impacting vacuum operations had finished freeing up resources, we had started to see signs of improvement \(16:53—17:01\), but added latency in the feedback loop and aggregation window at AWS still decided to trigger the failover, even though we were already in a recovering state.
* **17:15–17:20** — Requests that had queued during the incident were worked through and the backlog returned to normal.
* **17:21–17:56** — We monitored the recovery and confirmed processing remained at full capacity.
## Remedies
* Reviewing connection limits and pooling so a single service cannot saturate a shared database, and evaluating dedicated database clusters per product to remove cross-service impact.
* Changing large historical-data deletions to run in smaller, throttled batches, and tuning database maintenance to avoid large catch-up operations.
* Adding earlier, proactive alerting on database memory, connections and load so we can intervene before customer impact.
* Continue working with our cloud provider on a definitive root cause.
Electronic signature increased TaT
Started June 17, 2026 at 4:59 AM UTC · 4h 24m
IssuesMinor incident
Affected components
QES
identified
One of our providers is still experiencing issues, resulting in higher TaT for electronic signature tasks.
identified
Our provider is actively working on fixing the issue.
identified
Our provider is still working on fixing the issue.
We're monitoring the impact and we're making sure to keep the TaT as low as possible, considering the situation.
identified
Our provider's error rate dropped significantly. The TaT of QES tasks is now close to the usual value.
identified
We're working closely with our provider to find the reason of the remaining errors.
The error rate contacting our provider is stable and the impact on the QES task TaT is under control.
monitoring
Our provider was able to fix the issue and we don't get any error when contacting them.
The TaT of QES tasks is back to normal.
Monitoring the situation to make sure the errors don't come back.
resolved
The QES tasks TaT is back to normal.
Electronic signature high TaT
Started June 16, 2026 at 10:19 PM UTC · 58m
IssuesMinor incident
Affected components
QES
identified
Our electronic signature tasks are taking more time to complete because one of our providers is doing some maintenance.
identified
Our provider is still not able to answer to all requests, resulting in longer-than-usual electronic signature task processing time.
resolved
The service is operational again.
Sporadic PDF generation issue
Started June 15, 2026 at 9:30 AM UTC · 2d 1h
Pending
resolved
Incident was solved.
postmortem
## Summary
Between May and June 2026, around 0.037% of the evidence files generated on the platform were incomplete and included in the corresponding evidence folders.
There was no impact on workflow results, and no incorrect information was ever displayed — the affected files were simply empty rather than wrong.
A small number of requests to generate timeline files were also affected \(around 0.073%\).
## Root Causes
The issue was introduced during an initiative to improve how timeline and evidence PDFs are generated — making them significantly smaller, more efficient, and more resilient to produce. As part of these changes, in rare cases the new flow attempted to generate the PDF before its content had fully loaded, resulting in a blank or nearly empty file.
Because the problem was intermittent and occurred only in rare timing conditions, it was not consistently reproduced during rollout testing.
## Timeline
* 15 June 2026: Engineering investigation found an incomplete PDF and started the root cause investigation
* 17 June 2026: Root cause investigation completed and a safeguard to prevent incomplete PDFs was deployed
## Remedies
* We corrected the PDF generation flow so PDFs are only printed after content has fully loaded, and added a safeguard that stops generation when an incomplete PDF is detected.
* We also improved our monitoring and test coverage for PDF generation, so similar issues are caught earlier in the future.
Partial outage for Watchlist reports
Started June 6, 2026 at 2:08 PM UTC · 1h 41m
OutageMajor incident
Affected components
WatchlistWatchlist
investigating
Some Watchlist search search profiles are not working correctly the impacted reports are not being processed. We are investigation the root cause.
identified
We identified the issue and are implementing a fix.
monitoring
The issue has been fixed, and the reports are being processed correctly. We are still monitoring the service.
resolved
This incident has been resolved.
Degraded performance for Identity Reports in UK jurisdiction
Started May 11, 2026 at 11:11 PM UTC · 14h 3m
IssuesMinor incident
Affected components
Identity Enhanced
identified
We are facing issues with one of our providers and we see a slight decrease in clear rate for Identity Reports in UK jurisdiction
monitoring
A fix has been implemented and we are monitoring the results. All clear rates should be back to normal
resolved
All reports are back to normal.
QES tasks cannot complete because our provider is encountering issues.
Started May 9, 2026 at 11:27 AM UTC · 37m
OutageMajor incident
Affected components
QES
identified
We cannot complete QES tasks because our electronic signature provider is encountering issues.
identified
The provider is working on a fix.
monitoring
Our provider fixed the issue. QES tasks are completing now.
Tasks started during the incident were resumed.
resolved
The service is operational, no error where detected after the fix.
QES tasks outage
Started May 5, 2026 at 6:52 AM UTC · 1h 47m
OutageCritical incident
Affected components
QES
investigating
We noticed an issue with our QES tasks, we're investigating.
identified
Our provider is experiencing issues. All QES tasks cannot complete for now.
identified
Our provider is fixing the issue. QES tasks still cannot complete.
monitoring
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
resolved
QES is operational again.
QES tasks degraded performance
Started April 21, 2026 at 2:28 PM UTC · 2h 45m
OutageCritical incident
Affected components
QES
identified
We identified an issue with one of our provider and some QES tasks may encounter some problems.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
identified
Our provider continues experiencing issues and is working on a fix.
identified
Our provider is slowly recovering, we expect some QES tasks to complete.
resolved
QES is operational again.
QES tasks partial outage
Started February 3, 2026 at 10:42 AM UTC · 2h 34m
OutageMajor incident
Affected components
QES
identified
One of our provider is experiencing issues. QES capture tasks will fail and QES verification tasks will have an increase in TaT.
identified
The provider is working on a fix, QES is still suffering a partial outage.
identified
The provider is still working on a fix.
identified
The provider is still working on a fix.
QES verification tasks may start to time out, depending on their configuration, as we passed 90 minutes of downtime.
monitoring
The provider fixed the issue. We can confirm our QES tasks are now processed correctly.
We'll monitor the situation to make sure the error is indeed fixed.
resolved
We can confirm the error stopped and services are stable and operational.
QES tasks degraded performance
Started February 2, 2026 at 3:36 PM UTC · 1h 48m
OutageMajor incident
Affected components
QES
identified
We identified an issue with one of our provider and some QES tasks may encounter some problems.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
monitoring
Our provider is slowly recovering, we expect some QES tasks to complete.
monitoring
We're now seeing the error rate decrease. Most of the QES tasks should complete.
resolved
QES is operational again.
Service Degradation - Manual Tasks
Started January 28, 2026 at 11:16 AM UTC · 4h 7m
IssuesMinor incident
Affected components
Document Verification
investigating
We are currently investigating this issue.
monitoring
Issue found and fixed.
Increased turn around time for manual reports.
Estimated time to live manual processing is 4h.
resolved
Incident is fully resolved, manual processing is now working normally.
Manual reports will keep having an increased turn around time for a few more hours while it works through the task backlog.
postmortem
### Summary
Manual task assignment for all EU customers stopped working between 10h50 UTC and 11:50 UTC. This led to an increase in manual processing Turnaround Time \(TaT\) affecting approximately 20% of our document verification volumes with all customers recovering to TaT SLA by 18h00 UTC.
During this period:
* All checks that required **manual review** showed an increase in TaT.
* **Fully automated reports were not affected** and continued to run as normal.
The issue was caused by a **configuration error in our internal task management system**, which prevented it from correctly assigning tasks to our analysts.
We fixed the configuration and **restored normal processing** by 28 Jan 2026 11h50 AM UTC, and cleared all manual task backlogs by 18h00 UTC.
We have updated our validation and deployment checks to prevent similar issues in the future.
### Root Causes
_Manual processing queue assignment was affected by an invalid manual configuration input. This single queue configuration parameter resulted in an error that affected assignments in all queues._
### Timeline
* all times in UTC:
_10:50: Configuration manually updated and errors started, no more tasks assigned._
_11:02: On-call is notified of a spike in manual system assignment errors through our monitoring_
_11:07: The error responsible for the spike is identified \(invalid UUID\)_
_11:10: Incident declared_
_11:45: Origin of invalid UUID is found_
_11:50: Bad configuration parameter is deleted_
_11:56: Configuration reintroduced correctly_
_18:00: Recovered from manual task backlog_
### Remedies
* Adding appropriate input configuration value validation
* Improve task assignment resilience to these types of errors
* Review configuration guidance and post-release monitoring
Document report processing disrupted in EU
Started January 26, 2026 at 12:00 PM UTC · 0m
OutageMajor incident
resolved
Between 12:15 UTC and 12:40 UTC there was disruption to document report processing in EU. Autofill requests also saw a disruption between 12:16 UTC and 12:27 UTC. A postmortem will follow.
postmortem
## Summary
On 26 January 2026 between 12.16 UTC and 12.35 UTC, our document reports processing was severely degraded with customers experiencing extended processing time to about 85% of their traffic. The incident also caused extended processing time on Biometric Authentications and Biometric Verifications for 20% of the traffic between 12.25 UTC and 12.35 UTC.
## Root Causes
A core service for fraud prevention on documents came under elevated load, causing kubernetes pods to go down in quick sequence. Our upstream service retry policy proved too aggressive to let the service recover and required manual scaling-up.
The retry policy also caused elevated load on a shared database which in turn also affected the biometrics service.
## Timeline
* **12:16 UTC** – Our monitoring detected a sharp increase in errors when processing document reports.
* **12:19** **UTC** – The on‑call team was alerted and began investigating the affected document‑processing service.
* **12:25 UTC** – We identified that the incident was also affecting a small portion of biometric checks, leading to some failures and short delays.
* **12:29** **UTC** – On-call engineers manually increased capacity for the impacted document fraud‑prevention service.
* **12:35** **UTC** – Error rates for both document reports and biometric checks returned to normal and new requests were being processed successfully.
* **12:45–13:06** **UTC** – We processed document reports that were impacted during the incident window.
* **13:15** **UTC** – We confirmed that all affected document and biometric requests had completed successfully and marked the incident as resolved.
## Remedies
* We have continued to fine-tune our auto-scaling parameters of the affected service to scale up with lower CPU targets.
* We modified the retry policy of fraud services to avoid overloading an already struggling service so that it can auto-recover.
* We made our load-tests on this service more representative of production traffic \(similar image sizes/document type distribution…\) and test for accelerated traffic spikes.
* We lowered the total amount of shared database connections the document processing can take to avoid noisy neighbour impact on biometrics processing.
Onfido outage history and incident timeline | Uptimus