公開 API の認可エラー
- resolved
14:30~15:30 UTC の間、CA 地域の顧客はパブリック API へのアクセスを劣化させました。 認証層の不正な展開により、有効なトークンが誤って拒否され、失敗した API リクエストが発生した。 問題は完全に解決しました。 終了時にアクションは必要ありません。 わたしたちは、破壊のために報じています.
公式のインシデント更新を自動翻訳しています。
55 Onfido incidents · 2025年5月 — official updates, affected components, duration and resolution details.
14:30~15:30 UTC の間、CA 地域の顧客はパブリック API へのアクセスを劣化させました。 認証層の不正な展開により、有効なトークンが誤って拒否され、失敗した API リクエストが発生した。 問題は完全に解決しました。 終了時にアクションは必要ありません。 わたしたちは、破壊のために報じています.
公式のインシデント更新を自動翻訳しています。
1:21 PM UTC から 1:46 PM UTC まで、Web SDK からのドキュメントリクエストの受け付けや、Web モジュールを含むフローの Entrust IDV SDK の受け入れが中断されました。 Onfido AndroidとiOS SDKは影響を受けません。 大変申し訳ございません。 調査を終えると、詳細な姿勢が追随します.
公式のインシデント更新を自動翻訳しています。
現在、顔の類似性、文書および既知の顔のレポートの処理のための米国のクラスターの遅延を調査しています.
インデックスされた面の検索クラスターで、一部のノードの劣化性能が確認されています。 根本原因を調査し、是正を可能とします.
米国の検索クラスターで劣化した性能の根本原因を特定し、解決に積極的に取り組んでいます。 復旧完了後、フォローアップします.
修正を実施し、処理時間の品質を監視し続けます.
処理時間は改善し続けます。すべてのシステムが稼働しています.
処理時間は正常に戻ります。 この問題は解決しました。 すぐに続くためにポスト・モリテム.
公式のインシデント更新を自動翻訳しています。
アイデンティティの強化は、英国を除いて、他のすべての世界の申請者に明確な料金の大きな減少を経験しています
当社は、第三者のプロバイダと協力して、問題を解決し、アップデートを共有します.
この問題の修正には引き続き取り組んでいます。
提供者のサービスの回復および要求の成功率の改善を見ています。 今後も状況を監視し続けます.
プロバイダは問題を解決し、サービスは現在正常に動作しています
公式のインシデント更新を自動翻訳しています。
We cannot complete QES tasks because our electronic signature provider is encountering issues.
Our provider is experiencing issues. All QES tasks cannot complete for now.
Our provider is still experiencing issues. All QES tasks cannot complete for now.
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
This incident has been resolved.
We're currently monitoring the EU cluster after an Amazon RDS issue.
We’re seeing recovery across our internal metrics, and processing has now returned to full capacity. At this time, the issue appears to be resolved. Our current leading hypothesis is resource contention on a shared Amazon RDS instance that several services depend on, potentially related to a VACUUM operation running alongside a long-running job deleting a large volume of accumulated historical data. We have not yet confirmed the root cause and will continue investigating as follow-up, but service has been restored for now. We'll be following up with a public post-mortem.
**Incident date:** 22 June 2026 **Region:** EU \(eu-west-1\) **Affected EU services**: API, Dashboard, Applicant Form, Document Verification, Facial Similarity, Watchlist, Identity Enhanced, Webhooks, Known Faces, Autofill, QES and Device Intelligence. **Customer impact:** ~16:50–17:00 UTC \(acute degradation\); ~17:00–17:20 UTC \(backlog recovery\) ## Summary On 22 June 2026, from approximately 16:50 UTC, a shared database cluster serving our EU region came under severe load and could not reliably serve queries for about 10 minutes. EU services returned elevated errors, and processing throughput briefly fell to ~20–35% of normal levels, with many subcomponents of our system \(e.g., Facial Similarity report processing\) being entirely disrupted, some others less heavily impacted \(e.g., Document report processing\). Service recovered by 17:01 UTC, the database fully stabilizing after an automatic failover \(~17:05–17:07 UTC\). A resultant report backlog was cleared by ~17:20 UTC. Requests in flight during the acute degradation window may have failed unless retried; queued background work was processed automatically once the database recovered. ## Root cause The incident was triggered by a routine database storage-reclamation task following standard scheduled data-deletion processing. This task normally completes without issue; why it failed on this occasion remains under investigation, although we observed that it was processing a larger-than-usual backlog. We have a support case open with our cloud provider to confirm a definitive root cause. The reclamation task began to compete with normal application queries, which slowed as the database struggled to keep up. Applications opened more and more connections, leading to connection saturation and causing queries across the affected services to fail. The database stabilized when an automatic failover to a healthy standby was triggered; the contention fully resolving with the failover to a new instance. ## Timeline \(UTC\) * **16:50** — Our monitoring detected errors and elevated latency across EU services caused by resource contention on a shared database cluster. * **16:53** — We start to see improvements, but system still not acting at normal levels of performance. * **~17:00** — Customer-facing errors subsided and processing resumed as the contention eased. * **17:01** — On-call engineers opened an incident and continued investigations. * **~17:05–17:07** — The database performed an automatic failover to a healthy instance, which reset the overloaded writer and fully stabilised the cluster. The failover was triggered because of resource contention \(out of memory\) caused by the heavy vacuuming in the preceding minutes of the incident. Once the impacting vacuum operations had finished freeing up resources, we had started to see signs of improvement \(16:53—17:01\), but added latency in the feedback loop and aggregation window at AWS still decided to trigger the failover, even though we were already in a recovering state. * **17:15–17:20** — Requests that had queued during the incident were worked through and the backlog returned to normal. * **17:21–17:56** — We monitored the recovery and confirmed processing remained at full capacity. ## Remedies * Reviewing connection limits and pooling so a single service cannot saturate a shared database, and evaluating dedicated database clusters per product to remove cross-service impact. * Changing large historical-data deletions to run in smaller, throttled batches, and tuning database maintenance to avoid large catch-up operations. * Adding earlier, proactive alerting on database memory, connections and load so we can intervene before customer impact. * Continue working with our cloud provider on a definitive root cause.
One of our providers is still experiencing issues, resulting in higher TaT for electronic signature tasks.
Our provider is actively working on fixing the issue.
Our provider is still working on fixing the issue. We're monitoring the impact and we're making sure to keep the TaT as low as possible, considering the situation.
Our provider's error rate dropped significantly. The TaT of QES tasks is now close to the usual value.
We're working closely with our provider to find the reason of the remaining errors. The error rate contacting our provider is stable and the impact on the QES task TaT is under control.
Our provider was able to fix the issue and we don't get any error when contacting them. The TaT of QES tasks is back to normal. Monitoring the situation to make sure the errors don't come back.
The QES tasks TaT is back to normal.
Our electronic signature tasks are taking more time to complete because one of our providers is doing some maintenance.
Our provider is still not able to answer to all requests, resulting in longer-than-usual electronic signature task processing time.
The service is operational again.
Incident was solved.
## Summary Between May and June 2026, around 0.037% of the evidence files generated on the platform were incomplete and included in the corresponding evidence folders. There was no impact on workflow results, and no incorrect information was ever displayed — the affected files were simply empty rather than wrong. A small number of requests to generate timeline files were also affected \(around 0.073%\). ## Root Causes The issue was introduced during an initiative to improve how timeline and evidence PDFs are generated — making them significantly smaller, more efficient, and more resilient to produce. As part of these changes, in rare cases the new flow attempted to generate the PDF before its content had fully loaded, resulting in a blank or nearly empty file. Because the problem was intermittent and occurred only in rare timing conditions, it was not consistently reproduced during rollout testing. ## Timeline * 15 June 2026: Engineering investigation found an incomplete PDF and started the root cause investigation * 17 June 2026: Root cause investigation completed and a safeguard to prevent incomplete PDFs was deployed ## Remedies * We corrected the PDF generation flow so PDFs are only printed after content has fully loaded, and added a safeguard that stops generation when an incomplete PDF is detected. * We also improved our monitoring and test coverage for PDF generation, so similar issues are caught earlier in the future.
Some Watchlist search search profiles are not working correctly the impacted reports are not being processed. We are investigation the root cause.
We identified the issue and are implementing a fix.
The issue has been fixed, and the reports are being processed correctly. We are still monitoring the service.
This incident has been resolved.
We are facing issues with one of our providers and we see a slight decrease in clear rate for Identity Reports in UK jurisdiction
A fix has been implemented and we are monitoring the results. All clear rates should be back to normal
All reports are back to normal.
We cannot complete QES tasks because our electronic signature provider is encountering issues.
The provider is working on a fix.
Our provider fixed the issue. QES tasks are completing now. Tasks started during the incident were resumed.
The service is operational, no error where detected after the fix.
We noticed an issue with our QES tasks, we're investigating.
Our provider is experiencing issues. All QES tasks cannot complete for now.
Our provider is fixing the issue. QES tasks still cannot complete.
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
QES is operational again.
We identified an issue with one of our provider and some QES tasks may encounter some problems.
Our provider is still experiencing issues. All QES tasks cannot complete for now.
Our provider continues experiencing issues and is working on a fix.
Our provider is slowly recovering, we expect some QES tasks to complete.
QES is operational again.
One of our provider is experiencing issues. QES capture tasks will fail and QES verification tasks will have an increase in TaT.
The provider is working on a fix, QES is still suffering a partial outage.
The provider is still working on a fix.
The provider is still working on a fix. QES verification tasks may start to time out, depending on their configuration, as we passed 90 minutes of downtime.
The provider fixed the issue. We can confirm our QES tasks are now processed correctly. We'll monitor the situation to make sure the error is indeed fixed.
We can confirm the error stopped and services are stable and operational.
We identified an issue with one of our provider and some QES tasks may encounter some problems.
Our provider is still experiencing issues. All QES tasks cannot complete for now.
Our provider is slowly recovering, we expect some QES tasks to complete.
We're now seeing the error rate decrease. Most of the QES tasks should complete.
QES is operational again.
We are currently investigating this issue.
Issue found and fixed. Increased turn around time for manual reports. Estimated time to live manual processing is 4h.
Incident is fully resolved, manual processing is now working normally. Manual reports will keep having an increased turn around time for a few more hours while it works through the task backlog.
### Summary Manual task assignment for all EU customers stopped working between 10h50 UTC and 11:50 UTC. This led to an increase in manual processing Turnaround Time \(TaT\) affecting approximately 20% of our document verification volumes with all customers recovering to TaT SLA by 18h00 UTC. During this period: * All checks that required **manual review** showed an increase in TaT. * **Fully automated reports were not affected** and continued to run as normal. The issue was caused by a **configuration error in our internal task management system**, which prevented it from correctly assigning tasks to our analysts. We fixed the configuration and **restored normal processing** by 28 Jan 2026 11h50 AM UTC, and cleared all manual task backlogs by 18h00 UTC. We have updated our validation and deployment checks to prevent similar issues in the future. ### Root Causes _Manual processing queue assignment was affected by an invalid manual configuration input. This single queue configuration parameter resulted in an error that affected assignments in all queues._ ### Timeline * all times in UTC: _10:50: Configuration manually updated and errors started, no more tasks assigned._ _11:02: On-call is notified of a spike in manual system assignment errors through our monitoring_ _11:07: The error responsible for the spike is identified \(invalid UUID\)_ _11:10: Incident declared_ _11:45: Origin of invalid UUID is found_ _11:50: Bad configuration parameter is deleted_ _11:56: Configuration reintroduced correctly_ _18:00: Recovered from manual task backlog_ ### Remedies * Adding appropriate input configuration value validation * Improve task assignment resilience to these types of errors * Review configuration guidance and post-release monitoring
Between 12:15 UTC and 12:40 UTC there was disruption to document report processing in EU. Autofill requests also saw a disruption between 12:16 UTC and 12:27 UTC. A postmortem will follow.
## Summary On 26 January 2026 between 12.16 UTC and 12.35 UTC, our document reports processing was severely degraded with customers experiencing extended processing time to about 85% of their traffic. The incident also caused extended processing time on Biometric Authentications and Biometric Verifications for 20% of the traffic between 12.25 UTC and 12.35 UTC. ## Root Causes A core service for fraud prevention on documents came under elevated load, causing kubernetes pods to go down in quick sequence. Our upstream service retry policy proved too aggressive to let the service recover and required manual scaling-up. The retry policy also caused elevated load on a shared database which in turn also affected the biometrics service. ## Timeline * **12:16 UTC** – Our monitoring detected a sharp increase in errors when processing document reports. * **12:19** **UTC** – The on‑call team was alerted and began investigating the affected document‑processing service. * **12:25 UTC** – We identified that the incident was also affecting a small portion of biometric checks, leading to some failures and short delays. * **12:29** **UTC** – On-call engineers manually increased capacity for the impacted document fraud‑prevention service. * **12:35** **UTC** – Error rates for both document reports and biometric checks returned to normal and new requests were being processed successfully. * **12:45–13:06** **UTC** – We processed document reports that were impacted during the incident window. * **13:15** **UTC** – We confirmed that all affected document and biometric requests had completed successfully and marked the incident as resolved. ## Remedies * We have continued to fine-tune our auto-scaling parameters of the affected service to scale up with lower CPU targets. * We modified the retry policy of fraud services to avoid overloading an already struggling service so that it can auto-recover. * We made our load-tests on this service more representative of production traffic \(similar image sizes/document type distribution…\) and test for accelerated traffic spikes. * We lowered the total amount of shared database connections the document processing can take to avoid noisy neighbour impact on biometrics processing.
We have a high error rate in document report processing and autofill. Our team is investigating
We have identified the issue as a configuration issue between services, and are working on a fix.
We have deployed a fix. New document reports and autofill requests are being processed normally. We will rerun the earlier impacted reports.
This issue is now resolved: Between 10:30 and 12:00 UTC, document reports could not be processed and autofill requests were failing in the US region. Starting from 12:00, live traffic was unaffected. From 12:00 to 12:30 UTC, pending reports were processed. We take a lot of pride in running a robust, reliable service and we're working hard to make sure this does not happen again. A detailed postmortem will follow once we've concluded our investigation.
### Summary Around 10:29 UTC, around 10 minutes after restoring service in the EU and CA from the 10:05-10:17 UTC outage, we go alerted for high error rate in the US region and we noticed the new ML model version was back online in that region. This impacted Turnaround Time \(TaT\) for all Document Report and Studio Autofill until 12:11 UTC. Autofill Classic had an average 100% error rate during the same time frame. All impacted Document Reports and Studio Autofill tasks were completed successfully with an average TaT of ~40 minutes. ### Root Causes A previously rolled back version of an ML model went back live automatically in the US region. The automated canary reversal did not succeed in that region, leaving the deployment in an inconsistent state. As a result, the model routing labels were not updated as expected during the next forced deployment, and traffic continued to be served by pods running the incorrect model version until we manually rolled back the model in the cluster. This was the result of a bug in our version of Helm. ### Timeline _10:29 UTC: The new model release went back live in the US region only, after the previous automated canary rollback_ _11:44 UTC: We manually re-deployed a previous version via CI/CD pipeline. This was not successful and errors continued._ 1_2:00 UTC: We manually rolled back the version of the model directly in the cluster._ 1_2:11 UTC: All services went back to normal_ ### Remedies * We will fix our deployment tooling to reliably apply model-routing label changes to these resources by upgrading Helm.
Between 10:30 UTC and 10:47 UTC there was disruption to document report processing in EU, as a side effect of reports delayed from the earlier incident being re-processed. Further details will follow in a postmortem.
### Summary For the EU region, one critical service struggled to reprocess the traffic affected by a previous faulty release which led to higher Turnaround Time \(TaT\) for all Document reports created between 10:32 and 10:50 UTC. All impacted Document Reports were completed successfully with an average TaT of ~6 minutes. ### Root Causes The re-processing batch caused a spike in traffic and the auto-scaling didn’t work as expected for one critical service. The service entered a crash loop state and had to be manually up-scaled to recover. An unbounded number of in-flight requests were accepted leading to memory exhaustion and I/O event loop non-responsiveness, while waiting for a downstream ML inference service to scale up. ### Timeline _10:32 UTC: The critical service went up to a 100% error rate_ _10:33 UTC: Engineers who initiated the report backlog reprocessing become aware of the issue through our monitoring and start investigating_ _10:48 UTC: We manually scaled up the critical service_ 1_0:50 UTC: The service went back to normal and errors stopped_ ### Remedies * Investigate how to reduce memory footprint on this service, allowing bigger request queues while it waits for downstream ML model serving to scale up * Change auto-scaling parameters to be more aggressive \(i.e. scale with lower CPU targets\) * Add concurrent requests monitoring * Reduce ML model serving image sizes for faster scaling of inference services * Improve back-pressure mechanisms to be able to sustain minimum traffic levels independent of spikes while auto-scaling kicks-in * Change our weekly load testing scripts to specifically test for accelerated traffic spikes