Comenzó 7 de septiembre de 2026 a las 2:30 UTC · 12h 0m
IssuesIncidente menor
resolved
Entre 14:30 y 15:30 UTC, los clientes de la región de CA experimentaron un acceso degradado a la API pública. Un despliegue defectuoso en la capa de autenticación causó que se rechazaran las fichas válidas, lo que dio lugar a solicitudes de API fallidas.
La cuestión se ha resuelto plenamente. No se requiere acción en su extremo. Nos disculpamos por la interrupción.
Traducido automáticamente desde la actualización oficial del incidente.
Web SDK Outage (todas las regiones)
Comenzó 1 de septiembre de 2026 a las 13:21 UTC · 0m
OutageIncidente mayor
resolved
De aproximadamente 1:21 PM UTC a 1:46 PM UTC, hubo una interrupción para aceptar solicitudes de documentos del SDK Web, así como SDKs Entrust para flujos que contienen módulos web.
Onfido Android & iOS SDKs no fueron impactados.
Nos disculpamos por esta perturbación. Un postmortem detallado seguirá una vez que hayamos concluido nuestra investigación.
Traducido automáticamente desde la actualización oficial del incidente.
Aumento de caras conocidas consideran la tasa
Comenzó 26 de agosto de 2026 a las 15:30 UTC · 0m
OutageIncidente mayor
resolved
Las caras conocidas que experimentan tasas de mayor consideración en todas las regiones.
postmortem
## Summary
El 26 de agosto de 2026, se observó un aumento en la tasa de consideración de caras conocidas durante el período comprendido entre las 15:00 y las 18:00 UTC. La tasa global observada de caras conocidas aumentó de ~6% a ~23% en todas las regiones. El aumento percibido puede haber sido mayor o menor, dependiendo de la geografía o vertical específica del cliente.
El fue impulsado por un cambio a cómo se calculan las puntuaciones correspondientes a los partidos de Known Faces, que inflaron las puntuaciones para caras no palpadoras y causando así muchos informes para devolver partidos falsos y un resultado **consider**.
## Root Causes
El cambio introdujo un nuevo algoritmo para marcar caras iguales, que está destinado a mejorar el rendimiento de memoria de caras conocidas. Esto fue ampliamente evaluado y validado por parámetros internos antes de ser liberado. Inadvertidamente, la transición de puntos de referencia a la producción no tuvo en cuenta parcialmente las características específicas del entorno de producción. En la raíz, una única constante no se convirtió correctamente en la transición, lo que resulta en un comportamiento equivocado del nuevo algoritmo de puntuación.
## Timeline
15:20 UTC: El cambio a caras conocidas es liberado
17:35 UTC: Aumento de la tasa de consideración en caras conocidas se nota.
17:50 UTC: El cambio se hace retroceder. Considere caídas de tarifas poco después.
## Remedios
Estamos revisando los parámetros internos y los procedimientos de prueba que pueden aumentar la seguridad de este tipo de cambio. Específicamente, garantizar las pruebas internas refleja el panorama de producción real cuando se trata de conjuntos de datos y que las transiciones de parámetros y experimentos internos a la producción se prueban a fondo.
Además, los procedimientos de cambio se llevarán a cabo con un control más estricto para que nos permita notar y actuar más rápido.
Traducido automáticamente desde la actualización oficial del incidente.
Spike en fallas de autenticidad visual en cheques de simulación facial
Comenzó 24 de agosto de 2026 a las 17:00 UTC · 0m
OutageIncidente mayor
resolved
Semejanza Facial Las comprobaciones de movimiento que utilizan retos de azar están fallando a un ritmo mucho más alto de lo normal en la región de la UE.
postmortem
## Summary
El 24 de agosto de 2026, entre las 15:00 UTC y las 06:53 UTC a la mañana siguiente, los cheques de Moción de Similitud Facial con problemas de azar fallaron a un ritmo mucho más alto que lo normal en la región de la UE. Se aplicó una validación más estricta a todos los clientes de la región cuando debía haberse limitado a un solo cliente, lo que causaría que se rechazaran las verificaciones que de otro modo habrían pasado. Los informes afectados fueron devueltos con un fallo de " Autenticidad Visual".
Debido a que los controles de Motion son totalmente automatizados, los usuarios finales detrás de estos informes fueron rechazados sin revisión manual, y algunos intentaron volver a comprobar y fueron rechazados por segunda vez.
La tasa clara de estos cheques disminuyó de alrededor del 97% a alrededor del 54% para aproximadamente el 1.67% de los informes de Moción de Similitud Facial fueron afectados. Desactivar el cambio restaurado tasas normales claras inmediatamente. Otros controles de movimiento y otras regiones no fueron afectados.
## Root Causes
Se estaba juzgando una validación más estricta del desafío de aleatoriedad de movimiento con un solo cliente, controlado por un ajuste de configuración per-clómero. Un error en nuestro código significaba que el ajuste se evaluó sin el identificador del cliente, por lo que la validación más estricta se aplicó a todos los clientes de la región de la UE ejecutando Motion con cheques de aleatoriedad en lugar de al indicado. Las comprobaciones que habrían pasado bajo la validación estándar fueron rechazadas y reportadas bajo el desglose de " Autenticidad Visual". Ninguna configuración del cliente o datos enviados fue culpado.
## Timeline
Todas las veces UTC.
* 24 de agosto, 15:00: habilitamos una validación más estricta para los desafíos de azar de movimiento en la región de la UE. Los cheques de movimiento usando aleatoriedad comenzaron a fallar a un ritmo mucho mayor.
* 24 August, 16:36: a customer reported a drop in workflow success rates and we started investigating.
* 25 agosto, 04:18: confirmamos la elevada tasa de fracaso y la escalamos.
* 25 August, 06:53: revertimos el cambio y las tasas claras retornaron a la normalidad en cuestión de minutos.
* 26 de agosto: implementamos una solución permanente.
## Remedios
* Fortalecer los controles en torno a los cambios de configuración específicos del cliente para asegurar que no puedan aplicarse más ampliamente de lo previsto.
* Mejorar el monitoreo y alerta para cheques de velocidad clara de Motion para que los cambios inesperados en las tasas de rechazo se detecten automáticamente y se escalan más rápidamente.
Traducido automáticamente desde la actualización oficial del incidente.
Retrasos en los informes de procesamiento en el grupo estadounidense
Comenzó 27 de julio de 2026 a las 18:12 UTC · 1h 9m
Actualmente estamos investigando retrasos en nuestro clúster estadounidense para el procesamiento de informes de la similitud facial, documentos y caras conocidas.
investigating
Estamos viendo el rendimiento degradado de algunos nodos en nuestro grupo de búsqueda de caras indexadas. Seguimos investigando la causa raíz y posibles remediaciones.
identified
Hemos identificado la causa raíz del rendimiento degradado en nuestro grupo de búsqueda estadounidense y estamos trabajando activamente para resolverlo. Seguiremos de nuevo.
monitoring
Implementamos una solución y seguimos monitoreando la calidad del tiempo de procesamiento, que ahora se mejora.
monitoring
El tiempo de procesamiento sigue mejorando, todos los sistemas están en funcionamiento.
resolved
Los tiempos de procesamiento vuelven a la normalidad. Esta cuestión se resuelve ahora. Post-mortem a seguir pronto.
Traducido automáticamente desde la actualización oficial del incidente.
Cuestiones relacionadas con la identidad
Comenzó 24 de julio de 2026 a las 4:18 UTC · 2h 50m
IssuesIncidente menor
Componentes afectados
Identity EnhancedIdentity Enhanced
investigating
Identity Enhanced está experimentando una disminución importante de las tasas claras para los solicitantes todo el mundo excepto Reino Unido
identified
Trabajamos con nuestro proveedor de terceros para resolver el problema y compartiremos las actualizaciones cuando estén disponibles.
identified
Seguimos trabajando en una solución para esta cuestión.
monitoring
Estamos viendo la recuperación de servicios de nuestro proveedor y solicitamos tasas de éxito mejorando. Seguimos vigilando de cerca la situación.
resolved
El proveedor ha resuelto el problema, y el servicio está funcionando normalmente
Traducido automáticamente desde la actualización oficial del incidente.
QES tasks cannot complete because our provider is encountering issues.
Comenzó 23 de junio de 2026 a las 2:38 UTC · 1h 11m
OutageIncidente crítico
Componentes afectados
QES
investigating
We cannot complete QES tasks because our electronic signature provider is encountering issues.
identified
Our provider is experiencing issues. All QES tasks cannot complete for now.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
monitoring
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
resolved
This incident has been resolved.
Severe service disruption across the platform in EU
Comenzó 22 de junio de 2026 a las 17:44 UTC · 12m
OutageIncidente crítico
Componentes afectados
QESFacial SimilarityDevice IntelligenceWebhooksAPIKnown facesIdentity EnhancedDashboardAutofillDocument VerificationWatchlistApplicant Form
monitoring
We're currently monitoring the EU cluster after an Amazon RDS issue.
resolved
We’re seeing recovery across our internal metrics, and processing has now returned to full capacity. At this time, the issue appears to be resolved.
Our current leading hypothesis is resource contention on a shared Amazon RDS instance that several services depend on, potentially related to a VACUUM operation running alongside a long-running job deleting a large volume of accumulated historical data. We have not yet confirmed the root cause and will continue investigating as follow-up, but service has been restored for now.
We'll be following up with a public post-mortem.
postmortem
**Incident date:** 22 June 2026
**Region:** EU \(eu-west-1\)
**Affected EU services**: API, Dashboard, Applicant Form, Document Verification,
Facial Similarity, Watchlist, Identity Enhanced, Webhooks, Known Faces, Autofill,
QES and Device Intelligence.
**Customer impact:** ~16:50–17:00 UTC \(acute degradation\); ~17:00–17:20 UTC \(backlog recovery\)
## Summary
On 22 June 2026, from approximately 16:50 UTC, a shared database cluster serving our EU region came under severe load and could not reliably serve queries for about 10 minutes. EU services returned elevated errors, and processing throughput briefly fell to ~20–35% of normal levels, with many subcomponents of our system \(e.g., Facial Similarity report processing\) being entirely disrupted, some others less heavily impacted \(e.g., Document report processing\). Service recovered by 17:01 UTC, the database fully stabilizing after an automatic failover \(~17:05–17:07 UTC\). A resultant report backlog was cleared by ~17:20 UTC.
Requests in flight during the acute degradation window may have failed unless retried; queued background work was processed automatically once the database recovered.
## Root cause
The incident was triggered by a routine database storage-reclamation task following standard scheduled data-deletion processing. This task normally completes without issue; why it failed on this occasion remains under investigation, although we observed that it was processing a larger-than-usual backlog. We have a support case open with our cloud provider to confirm a definitive root cause.
The reclamation task began to compete with normal application queries, which slowed as the database struggled to keep up. Applications opened more and more connections, leading to connection saturation and causing queries across the affected services to fail.
The database stabilized when an automatic failover to a healthy standby was triggered; the contention fully resolving with the failover to a new instance.
## Timeline \(UTC\)
* **16:50** — Our monitoring detected errors and elevated latency across EU services caused by resource contention on a shared database cluster.
* **16:53** — We start to see improvements, but system still not acting at normal levels of performance.
* **~17:00** — Customer-facing errors subsided and processing resumed as the contention eased.
* **17:01** — On-call engineers opened an incident and continued investigations.
* **~17:05–17:07** — The database performed an automatic failover to a healthy instance, which reset the overloaded writer and fully stabilised the cluster. The failover was triggered because of resource contention \(out of memory\) caused by the heavy vacuuming in the preceding minutes of the incident. Once the impacting vacuum operations had finished freeing up resources, we had started to see signs of improvement \(16:53—17:01\), but added latency in the feedback loop and aggregation window at AWS still decided to trigger the failover, even though we were already in a recovering state.
* **17:15–17:20** — Requests that had queued during the incident were worked through and the backlog returned to normal.
* **17:21–17:56** — We monitored the recovery and confirmed processing remained at full capacity.
## Remedies
* Reviewing connection limits and pooling so a single service cannot saturate a shared database, and evaluating dedicated database clusters per product to remove cross-service impact.
* Changing large historical-data deletions to run in smaller, throttled batches, and tuning database maintenance to avoid large catch-up operations.
* Adding earlier, proactive alerting on database memory, connections and load so we can intervene before customer impact.
* Continue working with our cloud provider on a definitive root cause.
Electronic signature increased TaT
Comenzó 17 de junio de 2026 a las 4:59 UTC · 4h 24m
IssuesIncidente menor
Componentes afectados
QES
identified
One of our providers is still experiencing issues, resulting in higher TaT for electronic signature tasks.
identified
Our provider is actively working on fixing the issue.
identified
Our provider is still working on fixing the issue.
We're monitoring the impact and we're making sure to keep the TaT as low as possible, considering the situation.
identified
Our provider's error rate dropped significantly. The TaT of QES tasks is now close to the usual value.
identified
We're working closely with our provider to find the reason of the remaining errors.
The error rate contacting our provider is stable and the impact on the QES task TaT is under control.
monitoring
Our provider was able to fix the issue and we don't get any error when contacting them.
The TaT of QES tasks is back to normal.
Monitoring the situation to make sure the errors don't come back.
resolved
The QES tasks TaT is back to normal.
Electronic signature high TaT
Comenzó 16 de junio de 2026 a las 22:19 UTC · 58m
IssuesIncidente menor
Componentes afectados
QES
identified
Our electronic signature tasks are taking more time to complete because one of our providers is doing some maintenance.
identified
Our provider is still not able to answer to all requests, resulting in longer-than-usual electronic signature task processing time.
resolved
The service is operational again.
Sporadic PDF generation issue
Comenzó 15 de junio de 2026 a las 9:30 UTC · 2d 1h
Pending
resolved
Incident was solved.
postmortem
## Summary
Between May and June 2026, around 0.037% of the evidence files generated on the platform were incomplete and included in the corresponding evidence folders.
There was no impact on workflow results, and no incorrect information was ever displayed — the affected files were simply empty rather than wrong.
A small number of requests to generate timeline files were also affected \(around 0.073%\).
## Root Causes
The issue was introduced during an initiative to improve how timeline and evidence PDFs are generated — making them significantly smaller, more efficient, and more resilient to produce. As part of these changes, in rare cases the new flow attempted to generate the PDF before its content had fully loaded, resulting in a blank or nearly empty file.
Because the problem was intermittent and occurred only in rare timing conditions, it was not consistently reproduced during rollout testing.
## Timeline
* 15 June 2026: Engineering investigation found an incomplete PDF and started the root cause investigation
* 17 June 2026: Root cause investigation completed and a safeguard to prevent incomplete PDFs was deployed
## Remedies
* We corrected the PDF generation flow so PDFs are only printed after content has fully loaded, and added a safeguard that stops generation when an incomplete PDF is detected.
* We also improved our monitoring and test coverage for PDF generation, so similar issues are caught earlier in the future.
Partial outage for Watchlist reports
Comenzó 6 de junio de 2026 a las 14:08 UTC · 1h 41m
OutageIncidente mayor
Componentes afectados
WatchlistWatchlist
investigating
Some Watchlist search search profiles are not working correctly the impacted reports are not being processed. We are investigation the root cause.
identified
We identified the issue and are implementing a fix.
monitoring
The issue has been fixed, and the reports are being processed correctly. We are still monitoring the service.
resolved
This incident has been resolved.
Degraded performance for Identity Reports in UK jurisdiction
Comenzó 11 de mayo de 2026 a las 23:11 UTC · 14h 3m
IssuesIncidente menor
Componentes afectados
Identity Enhanced
identified
We are facing issues with one of our providers and we see a slight decrease in clear rate for Identity Reports in UK jurisdiction
monitoring
A fix has been implemented and we are monitoring the results. All clear rates should be back to normal
resolved
All reports are back to normal.
QES tasks cannot complete because our provider is encountering issues.
Comenzó 9 de mayo de 2026 a las 11:27 UTC · 37m
OutageIncidente mayor
Componentes afectados
QES
identified
We cannot complete QES tasks because our electronic signature provider is encountering issues.
identified
The provider is working on a fix.
monitoring
Our provider fixed the issue. QES tasks are completing now.
Tasks started during the incident were resumed.
resolved
The service is operational, no error where detected after the fix.
QES tasks outage
Comenzó 5 de mayo de 2026 a las 6:52 UTC · 1h 47m
OutageIncidente crítico
Componentes afectados
QES
investigating
We noticed an issue with our QES tasks, we're investigating.
identified
Our provider is experiencing issues. All QES tasks cannot complete for now.
identified
Our provider is fixing the issue. QES tasks still cannot complete.
monitoring
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
resolved
QES is operational again.
QES tasks degraded performance
Comenzó 21 de abril de 2026 a las 14:28 UTC · 2h 45m
OutageIncidente crítico
Componentes afectados
QES
identified
We identified an issue with one of our provider and some QES tasks may encounter some problems.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
identified
Our provider continues experiencing issues and is working on a fix.
identified
Our provider is slowly recovering, we expect some QES tasks to complete.
resolved
QES is operational again.
QES tasks partial outage
Comenzó 3 de febrero de 2026 a las 10:42 UTC · 2h 34m
OutageIncidente mayor
Componentes afectados
QES
identified
One of our provider is experiencing issues. QES capture tasks will fail and QES verification tasks will have an increase in TaT.
identified
The provider is working on a fix, QES is still suffering a partial outage.
identified
The provider is still working on a fix.
identified
The provider is still working on a fix.
QES verification tasks may start to time out, depending on their configuration, as we passed 90 minutes of downtime.
monitoring
The provider fixed the issue. We can confirm our QES tasks are now processed correctly.
We'll monitor the situation to make sure the error is indeed fixed.
resolved
We can confirm the error stopped and services are stable and operational.
QES tasks degraded performance
Comenzó 2 de febrero de 2026 a las 15:36 UTC · 1h 48m
OutageIncidente mayor
Componentes afectados
QES
identified
We identified an issue with one of our provider and some QES tasks may encounter some problems.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
monitoring
Our provider is slowly recovering, we expect some QES tasks to complete.
monitoring
We're now seeing the error rate decrease. Most of the QES tasks should complete.
resolved
QES is operational again.
Service Degradation - Manual Tasks
Comenzó 28 de enero de 2026 a las 11:16 UTC · 4h 7m
IssuesIncidente menor
Componentes afectados
Document Verification
investigating
We are currently investigating this issue.
monitoring
Issue found and fixed.
Increased turn around time for manual reports.
Estimated time to live manual processing is 4h.
resolved
Incident is fully resolved, manual processing is now working normally.
Manual reports will keep having an increased turn around time for a few more hours while it works through the task backlog.
postmortem
### Summary
Manual task assignment for all EU customers stopped working between 10h50 UTC and 11:50 UTC. This led to an increase in manual processing Turnaround Time \(TaT\) affecting approximately 20% of our document verification volumes with all customers recovering to TaT SLA by 18h00 UTC.
During this period:
* All checks that required **manual review** showed an increase in TaT.
* **Fully automated reports were not affected** and continued to run as normal.
The issue was caused by a **configuration error in our internal task management system**, which prevented it from correctly assigning tasks to our analysts.
We fixed the configuration and **restored normal processing** by 28 Jan 2026 11h50 AM UTC, and cleared all manual task backlogs by 18h00 UTC.
We have updated our validation and deployment checks to prevent similar issues in the future.
### Root Causes
_Manual processing queue assignment was affected by an invalid manual configuration input. This single queue configuration parameter resulted in an error that affected assignments in all queues._
### Timeline
* all times in UTC:
_10:50: Configuration manually updated and errors started, no more tasks assigned._
_11:02: On-call is notified of a spike in manual system assignment errors through our monitoring_
_11:07: The error responsible for the spike is identified \(invalid UUID\)_
_11:10: Incident declared_
_11:45: Origin of invalid UUID is found_
_11:50: Bad configuration parameter is deleted_
_11:56: Configuration reintroduced correctly_
_18:00: Recovered from manual task backlog_
### Remedies
* Adding appropriate input configuration value validation
* Improve task assignment resilience to these types of errors
* Review configuration guidance and post-release monitoring
Document report processing disrupted in EU
Comenzó 26 de enero de 2026 a las 12:00 UTC · 0m
OutageIncidente mayor
resolved
Between 12:15 UTC and 12:40 UTC there was disruption to document report processing in EU. Autofill requests also saw a disruption between 12:16 UTC and 12:27 UTC. A postmortem will follow.
postmortem
## Summary
On 26 January 2026 between 12.16 UTC and 12.35 UTC, our document reports processing was severely degraded with customers experiencing extended processing time to about 85% of their traffic. The incident also caused extended processing time on Biometric Authentications and Biometric Verifications for 20% of the traffic between 12.25 UTC and 12.35 UTC.
## Root Causes
A core service for fraud prevention on documents came under elevated load, causing kubernetes pods to go down in quick sequence. Our upstream service retry policy proved too aggressive to let the service recover and required manual scaling-up.
The retry policy also caused elevated load on a shared database which in turn also affected the biometrics service.
## Timeline
* **12:16 UTC** – Our monitoring detected a sharp increase in errors when processing document reports.
* **12:19** **UTC** – The on‑call team was alerted and began investigating the affected document‑processing service.
* **12:25 UTC** – We identified that the incident was also affecting a small portion of biometric checks, leading to some failures and short delays.
* **12:29** **UTC** – On-call engineers manually increased capacity for the impacted document fraud‑prevention service.
* **12:35** **UTC** – Error rates for both document reports and biometric checks returned to normal and new requests were being processed successfully.
* **12:45–13:06** **UTC** – We processed document reports that were impacted during the incident window.
* **13:15** **UTC** – We confirmed that all affected document and biometric requests had completed successfully and marked the incident as resolved.
## Remedies
* We have continued to fine-tune our auto-scaling parameters of the affected service to scale up with lower CPU targets.
* We modified the retry policy of fraud services to avoid overloading an already struggling service so that it can auto-recover.
* We made our load-tests on this service more representative of production traffic \(similar image sizes/document type distribution…\) and test for accelerated traffic spikes.
* We lowered the total amount of shared database connections the document processing can take to avoid noisy neighbour impact on biometrics processing.