QES tasks cannot complete because our provider is encountering issues.
์์ 2026๋ 6์ 23์ผ AM 2:38 UTC ยท 1h 11m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
QES
investigating
We cannot complete QES tasks because our electronic signature provider is encountering issues.
identified
Our provider is experiencing issues. All QES tasks cannot complete for now.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
monitoring
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
resolved
This incident has been resolved.
Severe service disruption across the platform in EU
์์ 2026๋ 6์ 22์ผ PM 5:44 UTC ยท 12m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
QESFacial SimilarityDevice IntelligenceWebhooksAPIKnown facesIdentity EnhancedDashboardAutofillDocument VerificationWatchlistApplicant Form
monitoring
We're currently monitoring the EU cluster after an Amazon RDS issue.
resolved
Weโre seeing recovery across our internal metrics, and processing has now returned to full capacity. At this time, the issue appears to be resolved.
Our current leading hypothesis is resource contention on a shared Amazon RDS instance that several services depend on, potentially related to a VACUUM operation running alongside a long-running job deleting a large volume of accumulated historical data. We have not yet confirmed the root cause and will continue investigating as follow-up, but service has been restored for now.
We'll be following up with a public post-mortem.
postmortem
**Incident date:** 22 June 2026
**Region:** EU \(eu-west-1\)
**Affected EU services**: API, Dashboard, Applicant Form, Document Verification,
Facial Similarity, Watchlist, Identity Enhanced, Webhooks, Known Faces, Autofill,
QES and Device Intelligence.
**Customer impact:** ~16:50โ17:00 UTC \(acute degradation\); ~17:00โ17:20 UTC \(backlog recovery\)
## Summary
On 22 June 2026, from approximately 16:50 UTC, a shared database cluster serving our EU region came under severe load and could not reliably serve queries for about 10 minutes. EU services returned elevated errors, and processing throughput briefly fell to ~20โ35% of normal levels, with many subcomponents of our system \(e.g., Facial Similarity report processing\) being entirely disrupted, some others less heavily impacted \(e.g., Document report processing\). Service recovered by 17:01 UTC, the database fully stabilizing after an automatic failover \(~17:05โ17:07 UTC\). A resultant report backlog was cleared by ~17:20 UTC.
Requests in flight during the acute degradation window may have failed unless retried; queued background work was processed automatically once the database recovered.
## Root cause
The incident was triggered by a routine database storage-reclamation task following standard scheduled data-deletion processing. This task normally completes without issue; why it failed on this occasion remains under investigation, although we observed that it was processing a larger-than-usual backlog. We have a support case open with our cloud provider to confirm a definitive root cause.
The reclamation task began to compete with normal application queries, which slowed as the database struggled to keep up. Applications opened more and more connections, leading to connection saturation and causing queries across the affected services to fail.
The database stabilized when an automatic failover to a healthy standby was triggered; the contention fully resolving with the failover to a new instance.
## Timeline \(UTC\)
* **16:50** โ Our monitoring detected errors and elevated latency across EU services caused by resource contention on a shared database cluster.
* **16:53** โ We start to see improvements, but system still not acting at normal levels of performance.
* **~17:00** โ Customer-facing errors subsided and processing resumed as the contention eased.
* **17:01** โ On-call engineers opened an incident and continued investigations.
* **~17:05โ17:07** โ The database performed an automatic failover to a healthy instance, which reset the overloaded writer and fully stabilised the cluster. The failover was triggered because of resource contention \(out of memory\) caused by the heavy vacuuming in the preceding minutes of the incident. Once the impacting vacuum operations had finished freeing up resources, we had started to see signs of improvement \(16:53โ17:01\), but added latency in the feedback loop and aggregation window at AWS still decided to trigger the failover, even though we were already in a recovering state.
* **17:15โ17:20** โ Requests that had queued during the incident were worked through and the backlog returned to normal.
* **17:21โ17:56** โ We monitored the recovery and confirmed processing remained at full capacity.
## Remedies
* Reviewing connection limits and pooling so a single service cannot saturate a shared database, and evaluating dedicated database clusters per product to remove cross-service impact.
* Changing large historical-data deletions to run in smaller, throttled batches, and tuning database maintenance to avoid large catch-up operations.
* Adding earlier, proactive alerting on database memory, connections and load so we can intervene before customer impact.
* Continue working with our cloud provider on a definitive root cause.
Electronic signature increased TaT
์์ 2026๋ 6์ 17์ผ AM 4:59 UTC ยท 4h 24m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
QES
identified
One of our providers is still experiencing issues, resulting in higher TaT for electronic signature tasks.
identified
Our provider is actively working on fixing the issue.
identified
Our provider is still working on fixing the issue.
We're monitoring the impact and we're making sure to keep the TaT as low as possible, considering the situation.
identified
Our provider's error rate dropped significantly. The TaT of QES tasks is now close to the usual value.
identified
We're working closely with our provider to find the reason of the remaining errors.
The error rate contacting our provider is stable and the impact on the QES task TaT is under control.
monitoring
Our provider was able to fix the issue and we don't get any error when contacting them.
The TaT of QES tasks is back to normal.
Monitoring the situation to make sure the errors don't come back.
resolved
The QES tasks TaT is back to normal.
Electronic signature high TaT
์์ 2026๋ 6์ 16์ผ PM 10:19 UTC ยท 58m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
QES
identified
Our electronic signature tasks are taking more time to complete because one of our providers is doing some maintenance.
identified
Our provider is still not able to answer to all requests, resulting in longer-than-usual electronic signature task processing time.
resolved
The service is operational again.
Sporadic PDF generation issue
์์ 2026๋ 6์ 15์ผ AM 9:30 UTC ยท 2d 1h
Pending
resolved
Incident was solved.
postmortem
## Summary
Between May and June 2026, around 0.037% of the evidence files generated on the platform were incomplete and included in the corresponding evidence folders.
There was no impact on workflow results, and no incorrect information was ever displayed โ the affected files were simply empty rather than wrong.
A small number of requests to generate timeline files were also affected \(around 0.073%\).
## Root Causes
The issue was introduced during an initiative to improve how timeline and evidence PDFs are generated โ making them significantly smaller, more efficient, and more resilient to produce. As part of these changes, in rare cases the new flow attempted to generate the PDF before its content had fully loaded, resulting in a blank or nearly empty file.
Because the problem was intermittent and occurred only in rare timing conditions, it was not consistently reproduced during rollout testing.
## Timeline
* 15 June 2026: Engineering investigation found an incomplete PDF and started the root cause investigation
* 17 June 2026: Root cause investigation completed and a safeguard to prevent incomplete PDFs was deployed
## Remedies
* We corrected the PDF generation flow so PDFs are only printed after content has fully loaded, and added a safeguard that stops generation when an incomplete PDF is detected.
* We also improved our monitoring and test coverage for PDF generation, so similar issues are caught earlier in the future.
Partial outage for Watchlist reports
์์ 2026๋ 6์ 6์ผ PM 2:08 UTC ยท 1h 41m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
WatchlistWatchlist
investigating
Some Watchlist search search profiles are not working correctly the impacted reports are not being processed. We are investigation the root cause.
identified
We identified the issue and are implementing a fix.
monitoring
The issue has been fixed, and the reports are being processed correctly. We are still monitoring the service.
resolved
This incident has been resolved.
Degraded performance for Identity Reports in UK jurisdiction
์์ 2026๋ 5์ 11์ผ PM 11:11 UTC ยท 14h 3m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Identity Enhanced
identified
We are facing issues with one of our providers and we see a slight decrease in clear rate for Identity Reports in UK jurisdiction
monitoring
A fix has been implemented and we are monitoring the results. All clear rates should be back to normal
resolved
All reports are back to normal.
QES tasks cannot complete because our provider is encountering issues.
์์ 2026๋ 5์ 9์ผ AM 11:27 UTC ยท 37m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
QES
identified
We cannot complete QES tasks because our electronic signature provider is encountering issues.
identified
The provider is working on a fix.
monitoring
Our provider fixed the issue. QES tasks are completing now.
Tasks started during the incident were resumed.
resolved
The service is operational, no error where detected after the fix.
QES tasks outage
์์ 2026๋ 5์ 5์ผ AM 6:52 UTC ยท 1h 47m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
QES
investigating
We noticed an issue with our QES tasks, we're investigating.
identified
Our provider is experiencing issues. All QES tasks cannot complete for now.
identified
Our provider is fixing the issue. QES tasks still cannot complete.
monitoring
The issue has been fixed. We're monitoring and resuming blocked QES tasks, if they can be.
resolved
QES is operational again.
QES tasks degraded performance
์์ 2026๋ 4์ 21์ผ PM 2:28 UTC ยท 2h 45m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
QES
identified
We identified an issue with one of our provider and some QES tasks may encounter some problems.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
identified
Our provider continues experiencing issues and is working on a fix.
identified
Our provider is slowly recovering, we expect some QES tasks to complete.
resolved
QES is operational again.
QES tasks partial outage
์์ 2026๋ 2์ 3์ผ AM 10:42 UTC ยท 2h 34m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
QES
identified
One of our provider is experiencing issues. QES capture tasks will fail and QES verification tasks will have an increase in TaT.
identified
The provider is working on a fix, QES is still suffering a partial outage.
identified
The provider is still working on a fix.
identified
The provider is still working on a fix.
QES verification tasks may start to time out, depending on their configuration, as we passed 90 minutes of downtime.
monitoring
The provider fixed the issue. We can confirm our QES tasks are now processed correctly.
We'll monitor the situation to make sure the error is indeed fixed.
resolved
We can confirm the error stopped and services are stable and operational.
QES tasks degraded performance
์์ 2026๋ 2์ 2์ผ PM 3:36 UTC ยท 1h 48m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
QES
identified
We identified an issue with one of our provider and some QES tasks may encounter some problems.
identified
Our provider is still experiencing issues. All QES tasks cannot complete for now.
monitoring
Our provider is slowly recovering, we expect some QES tasks to complete.
monitoring
We're now seeing the error rate decrease. Most of the QES tasks should complete.
resolved
QES is operational again.
Service Degradation - Manual Tasks
์์ 2026๋ 1์ 28์ผ AM 11:16 UTC ยท 4h 7m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Document Verification
investigating
We are currently investigating this issue.
monitoring
Issue found and fixed.
Increased turn around time for manual reports.
Estimated time to live manual processing is 4h.
resolved
Incident is fully resolved, manual processing is now working normally.
Manual reports will keep having an increased turn around time for a few more hours while it works through the task backlog.
postmortem
### Summary
Manual task assignment for all EU customers stopped working between 10h50 UTC and 11:50 UTC. This led to an increase in manual processing Turnaround Time \(TaT\) affecting approximately 20% of our document verification volumes with all customers recovering to TaT SLA by 18h00 UTC.
During this period:
* All checks that required **manual review** showed an increase in TaT.
* **Fully automated reports were not affected** and continued to run as normal.
The issue was caused by a **configuration error in our internal task management system**, which prevented it from correctly assigning tasks to our analysts.
We fixed the configuration and **restored normal processing** by 28 Jan 2026 11h50 AM UTC, and cleared all manual task backlogs by 18h00 UTC.
We have updated our validation and deployment checks to prevent similar issues in the future.
### Root Causes
_Manual processing queue assignment was affected by an invalid manual configuration input. This single queue configuration parameter resulted in an error that affected assignments in all queues._
### Timeline
* all times in UTC:
_10:50: Configuration manually updated and errors started, no more tasks assigned._
_11:02: On-call is notified of a spike in manual system assignment errors through our monitoring_
_11:07: The error responsible for the spike is identified \(invalid UUID\)_
_11:10: Incident declared_
_11:45: Origin of invalid UUID is found_
_11:50: Bad configuration parameter is deleted_
_11:56: Configuration reintroduced correctly_
_18:00: Recovered from manual task backlog_
### Remedies
* Adding appropriate input configuration value validation
* Improve task assignment resilience to these types of errors
* Review configuration guidance and post-release monitoring
Document report processing disrupted in EU
์์ 2026๋ 1์ 26์ผ PM 12:00 UTC ยท 0m
Outage์ค๋ํ ์ธ์๋ํธ
resolved
Between 12:15 UTC and 12:40 UTC there was disruption to document report processing in EU. Autofill requests also saw a disruption between 12:16 UTC and 12:27 UTC. A postmortem will follow.
postmortem
## Summary
On 26 January 2026 between 12.16 UTC and 12.35 UTC, our document reports processing was severely degraded with customers experiencing extended processing time to about 85% of their traffic. The incident also caused extended processing time on Biometric Authentications and Biometric Verifications for 20% of the traffic between 12.25 UTC and 12.35 UTC.
## Root Causes
A core service for fraud prevention on documents came under elevated load, causing kubernetes pods to go down in quick sequence. Our upstream service retry policy proved too aggressive to let the service recover and required manual scaling-up.
The retry policy also caused elevated load on a shared database which in turn also affected the biometrics service.
## Timeline
* **12:16 UTC** โ Our monitoring detected a sharp increase in errors when processing document reports.
* **12:19** **UTC** โ The onโcall team was alerted and began investigating the affected documentโprocessing service.
* **12:25 UTC** โ We identified that the incident was also affecting a small portion of biometric checks, leading to some failures and short delays.
* **12:29** **UTC** โ On-call engineers manually increased capacity for the impacted document fraudโprevention service.
* **12:35** **UTC** โ Error rates for both document reports and biometric checks returned to normal and new requests were being processed successfully.
* **12:45โ13:06** **UTC** โ We processed document reports that were impacted during the incident window.
* **13:15** **UTC** โ We confirmed that all affected document and biometric requests had completed successfully and marked the incident as resolved.
## Remedies
* We have continued to fine-tune our auto-scaling parameters of the affected service to scale up with lower CPU targets.
* We modified the retry policy of fraud services to avoid overloading an already struggling service so that it can auto-recover.
* We made our load-tests on this service more representative of production traffic \(similar image sizes/document type distributionโฆ\) and test for accelerated traffic spikes.
* We lowered the total amount of shared database connections the document processing can take to avoid noisy neighbour impact on biometrics processing.