Degraded-performance in MLS Observer thruough DU - US
Started September 3, 2026 at 3:34 PM UTC · 43m
Pending
Affected components
Document Understanding
investigating
We are currently investigaing the issue
identified
The issue was identified, and the appropriate mitigation measures were implemented to address it.
monitoring
The issue has been successfully resolved, and the service has been fully restored. The service is operating normally at this time.
resolved
The issue has been successfully resolved, and the service has been fully restored. The service is operating normally at this time.
Insights Dashboard: Delayed Maestro Run Data
Started September 2, 2026 at 8:00 PM UTC · 0m
Pending
resolved
The Insights dashboard in Looker experienced an issue that prevented the latest Maestro runs, from appearing in the dashboard
Impacted regions : US, SEA, IND, CA, AUE,JP ,UK
Incident timeline
Started: September 2, 2026 at 20:00:00 UTC
Resolved: September 4, 2026 at 16:30:00 UTC
The issue has been currently resolved, and the Insights dashboard is displaying the latest Maestro run data.
US - Document Understanding and IXP - Degraded Performance
Started September 1, 2026 at 1:49 PM UTC · 1h 16m
Pending
Affected components
Document UnderstandingIXP
investigating
We are investigating reports of degraded performance impacting digitization and extraction for Document Understanding and IXP in US.
Impact: Users may experience slowness and failed Document Understanding runtime operations.
Our teams are working to identify the cause and will share more details as the investigation progresses.
monitoring
We have identified the issue and applied a mitigation. Services are returned back to healthy state and we are monitoring the services.
resolved
This issue has been fully resolved and services are stable now. We'll post more details of the incident to status page soon.
degraded-performance in MLS Observer thruough DU - US
Started August 31, 2026 at 4:54 PM UTC · 2h 44m
Pending
Affected components
Document UnderstandingDocument Understanding
investigating
We are currently investigation the issue.
identified
We have identified the issue and implemented the necessary fix
monitoring
The issue has been mitigated, and the service is currently operational. We will continue to closely monitor the service health and take further action if needed.
resolved
The issue has been mitigated, and the service is currently operational. We will continue to closely monitor the service health and take further action if needed.
Orchestrator - Customers are experiencing issues in accessing queue item
The issue has been identified, and the team is actively working on deploying a fix. We expect the deployment to begin in the coming hours.
monitoring
We have identified the root cause of the degraded performance and we are in the process of deploying a fix in the coming hours. It affects linked queries where the user does not have access to the original folder. API surface is unaffected. Export via CSV can be used as a workaround.Additional updates will be provided as we move toward resolution.
identified
We have identified the root cause of the degraded performance and we are in the process of deploying a fix in the coming hours. It affects linked queries where the user does not have access to the original folder. API surface is unaffected. Export via CSV can be used as a workaround. Additional updates will be provided as we move toward resolution.
identified
The fix is currently being applied. We will provide another update shortly once the deployment has been completed across all affected regions.
resolved
The fix has been successfully deployed across all regions, and we have validated that the issue is resolved. The service is operating as expected.
postmortem
## Customer impact
Between August 28, 2026 at 3:23 pm UTC and August 28, 2026 at 10:57 pm UTC, a subset of customers experienced errors when opening queue item details panels in Orchestrator. The affected path returned a 404 page for LINKED queues when the user did NOT have access to the original folder in which the items were created.
The impact affected only Orchestrator UI interactions and it spanned across all regions running the affected software version. Orchestrator application programming interface access was unaffected, and exporting queue data to CSV was available as a workaround. The total duration was approximately 7 hours.
## Root cause
The incident was caused by a regression in the Orchestrator user interface for linked queues accessed from other folders. The regression did not correctly handle the linked-queue details flow when the requesting user lacked access to the original folder, which caused the queue item details panel to route to a 404 page instead of displaying the expected information.
## Detection
The issue was detected through an automated incident alert for the Orchestrator front end on August 28, 2026 at 3:23 pm UTC.
## Response
At 3:31 pm UTC on August 28, 2026, our engineering team described the issue as customers being unable to access queue item details panels due to a regression. At 3:41 pm UTC, a public status update confirmed that the root cause had been identified, that a fix was being deployed, and that application programming interface access was unaffected.
At 5:27 pm UTC, the scope was narrowed to linked queues where the user did not have access to the original folder. At 5:29 pm UTC, the team determined that the initial fix did not fully address the affected scenario, and a corrected fix was developed. At 5:30 pm UTC, the affected scenario was reproduced and the corrected fix was validated locally. At 5:39 pm UTC, a status page update documented CSV export as a workaround.
Deployment of the corrected fix continued across affected regions. At 10:50 pm UTC, the fix was confirmed deployed everywhere. At 10:57 pm UTC, the incident was marked resolved and the status page was updated to confirm the service was operating as expected.
## Follow up
Formal follow-up actions are being tracked through the post-incident review process initiated at resolution.
## Action items
- Expanding our automated test cases to include this scenario, and other similar scenarios related to linked objects or cross-folder scenarios.
- Ensuring we can safely turn off any changes via a flag for a faster response time.
Orchestrator logs are not visible in the US regions
Started August 27, 2026 at 5:16 PM UTC · 4h 41m
Pending
Affected components
Orchestrator
investigating
We are investigating reports of missing robot logs in the US region. Our teams are working to identify the cause and will share more details as the investigation progresses.
identified
We identified that robot logs ingestion was delayed. Robot logs are catching up to live data and we are monitoring recovery.
monitoring
Logs are now populating as expected in real time and we will continue monitoring the application.
resolved
The issue has been resolved and logs are populating as expected.
Solutions Management- Japan - Partial Outage
Started August 25, 2026 at 9:41 AM UTC · 31m
OutageMajor incident
Affected components
Solutions Management
identified
We have identified an issue affecting the Solution Deployment feature in Studio Web in the Japan region, where solution deployments may fail. A fix has been prepared and will be rolled out shortly. We will provide further updates as more information becomes available.
monitoring
The fix has been successfully deployed in the Japan region, and the issue has been mitigated. We are closely monitoring the service to ensure Solution Deployment continues to operate as expected and will provide further updates as needed.
resolved
The issue has been resolved, and Solution Deployment in Studio Web is operating as expected in the Japan region. No further impact has been observed.
Singapore - Insights - Partial Outage
Started August 25, 2026 at 3:58 AM UTC · 56m
OutageMajor incident
Affected components
Insights
investigating
We are investigating an issue affecting customers using Insights in the Singapore region, where dashboards may fail to load and display timeout errors. Our teams are working to identify the root cause and will provide further updates as more information becomes available.
monitoring
The issue has been mitigated, and we are closely monitoring the Insights service in the Singapore region to ensure dashboards continue to load as expected.
resolved
The issue has been resolved, and the Insights service in the Singapore region is operating as expected. No further impact has been observed.
postmortem
## Customer impact
Between August 25, 2026 at 3:42 am UTC and August 25, 2026 at 4:53 am UTC, customers using Insights in the Singapore region experienced failures loading Insights dashboards. Affected users saw dashboard pages time out or fail to load, and server errors were observed for related service requests. The primary impact was to Insights dashboard access, including charts and alerts. A related backend service also returned server errors during the incident, but dashboard access had recovered before final resolution.
## Root cause
The incident was attributed to a regional Microsoft service disruption in Singapore affecting infrastructure used by the Insights service. During the disruption, the Insights service was unable to reliably serve dashboard requests, resulting in request timeouts and server errors. No changes were made on our side. The service recovered as the Microsoft regional disruption was resolved, and Microsoft subsequently reported a Singapore regional service issue on its status page. A formal root cause analysis has been requested from Microsoft to confirm the specific failure mechanism and identify any additional preventive or mitigation measures.
## Detection
Automated health monitoring for the Insights service detected the issue at August 25, 2026 at 3:42 am UTC. Customer impact was confirmed on the response bridge shortly after detection.
## Response
At August 25, 2026 at 3:58 am UTC, we posted a public status update noting that Insights dashboards in the Singapore region might fail to load and display timeout errors. During the response, our engineering team verified server errors in automated monitoring, attempted to access service diagnostics, and began a recovery procedure on the underlying infrastructure.
At August 25, 2026 at 4:01 am UTC, the Insights service recovered while the recovery procedure was in progress, and dashboard loading was validated across multiple test accounts. At August 25, 2026 at 4:37 am UTC, we moved the incident to monitoring after confirming dashboards loaded successfully. A related backend service recovered at August 25, 2026 at 4:49 am UTC, and the incident was marked resolved at August 25, 2026 at 4:53 am UTC.
## Follow up
1. Request a root cause analysis from Microsoft for the Singapore regional disruption affecting the insights dashboards loading issue.
We are investigating an issue affecting customers using the Extended OCR functionality in the Document Understanding service in the Singapore region. Our teams are working to identify the root cause and will provide further updates as more information becomes available.
monitoring
The issue has been mitigated, and we are closely monitoring the service to ensure it continues to operate as expected. We will provide further updates as more information becomes available.
resolved
The issue has been resolved, and the service is operating as expected. No further impact has been observed.
postmortem
## Customer impact
On August 25, 2026, Extended OCR requests in the Document Understanding service failed with 500 status code in the Singapore region for approximately 33 minutes, between 03:08 and 03:41 UTC. The cause was a Microsoft regional service disruption in Singapore. All other Document Understanding functionality was unaffected, and no other region was impacted.
## Root cause
The failure was attributed to a regional Microsoft service disruption in Singapore affecting resources used by the Extended OCR capability. No changes had been made on our side, and Microsoft subsequently updated their own status page to reflect a Singapore regional issue. A formal root cause analysis has been requested from Microsoft.
## Detection
An automated alert for the Document Understanding service fired on August 25, 2026 at 3:13 am UTC. The alert was acknowledged promptly, and the incident was declared customer-impacting within minutes.
## Response
The on-call engineer scoped the impact to the Singapore region. Responders ruled out a recent service update as the cause, as the same update had been deployed to other regions with no comparable impact, which pointed to a regional dependency failure outside our infrastructure.
Because the failure originated in an upstream Microsoft service, no mitigating action was available or required on the UiPath side. Requests began succeeding again at 03:41 UTC as the Microsoft dependency recovered. Responders held the incident open to verify sustained recovery: ten minutes of clean traffic were confirmed at 03:51 UTC, the incident moved to Monitoring at 04:04 UTC, and it was resolved at 04:59 UTC after continued stability with no further failures.
## Follow up
1. Requested a root cause analysis from Microsoft for the Singapore regional disruption affecting Extended OCR dependencies, including how recurrence can be prevented.
Multiple Regions - UiPath Apps - Partial Outage
Started August 24, 2026 at 10:43 AM UTC · 1h 12m
OutageMajor incident
Affected components
AppsAppsAppsAppsAppsAppsApps
identified
We have identified the cause of an issue impacting a small number of customers using UiPath Apps authored from the Studio Web service across multiple regions. Our teams are ready with a fix, and deployment is about to begin across the affected regions. We will continue to monitor the deployment and provide further updates as the fix takes effect.
monitoring
The fix has been successfully deployed across all affected regions, and the issue has been mitigated. We are closely monitoring the service to ensure the fix continues to perform as expected and will provide further updates as needed.
resolved
The issue has been resolved, and the service is operating as expected. No further impact has been observed.
postmortem
## Customer impact
Between August 19, 2026 at 2:33 pm UTC and August 24, 2026 at 11:30 am UTC, a subset of customers could not load UiPath Apps projects authored from Studio Web. Affected customers experienced complete unavailability of Apps projects rather than slowness or degraded performance.
The impact was limited to a small set of customers, whose Apps service was hosted in the Japan region. The total customer-impacting duration was approximately 4 days and 21 hours.
## Root cause
The root cause was a deployment mismatch between Studio Web and UiPath Apps. On August 19, 2026, Studio Web received a scheduled update in the EU region that included a framework upgrade. The corresponding UiPath Apps update containing the framework upgrade had not yet been deployed to Japan scale unit.
An earlier Apps update that was already compatible with the new framework version had been deployed to all regions except Japan, where deployment had been deferred due to unrelated reasons. As a result, the Japan region was still running an older Apps version that was not compatible with the updated Studio Web.
## Detection
The issue was reported by a customer through their UiPath account team on August 24, 2026 at 9:36 am UTC. An incident was opened at 10:21 am UTC, and a public status page incident was declared within a minute. Existing automated monitoring did not detect the issue because it validates matched deployment combinations; improving detection for mixed configurations is part of our follow-up plan.
## Response
Within minutes of opening the incident, we identified the deployment mismatch as the root cause and decided to accelerate deployment of the corresponding UiPath Apps update to all affected regions, aligning Apps with the Studio Web version already being served to impacted customers.
At 10:43 am UTC, we posted a public status update confirming the cause had been identified and that the fix was being deployed. At 10:52 am UTC, deployment was in progress for the remaining regions, with several regions already complete. At 11:30 am UTC, we confirmed the fix had been deployed across all affected regions and marked the incident mitigated. At 11:55 am UTC, after monitoring showed no further impact, the incident was marked resolved.
## Follow up
1. Add automated alerting on production telemetry to detect Apps load failures correlated with the Studio Web version being served.
2. Implement synthetic monitoring that regularly loads an Apps project authored from Studio Web across representative customer configurations and alerts on failure.
3. Adjust release sequencing for tightly coupled Studio Web and Apps changes so that dependent Apps updates are fully deployed to all regions before the newer Studio Web experience reaches customer traffic.
US - Document Understanding - Partial Outage
Started August 21, 2026 at 12:37 PM UTC · 59m
Pending
Affected components
Document Understanding
monitoring
A fix has been implemented for the issue impacting document classification and extraction for Document Understanding in US, and we are currently monitoring the results.
resolved
The issue impacting document classification and extraction for Document Understanding in the US region has been resolved. Following a period of monitoring, service is confirmed healthy and operating normally.
postmortem
## Customer impact
Between August 21, 2026 at 11:05 am UTC and August 21, 2026 at 12:18 pm UTC, a subset of customers experienced failed Document Understanding operations, including document classification, extraction, and digitization. The estimated duration of the partial outage was 48 minutes. The impact was limited to customers using Document Understanding in the US region.
## Root cause
The incident was caused by the Document Understanding storage database failover configuration entering a broken state during a preventive database scale-up operation. The scale-up was initiated after the database approached its storage limit. During the operation, the secondary database could not be scaled, the attempted removal from the failover configuration failed, and the database platform provider had to break the replication link. This left the failover configuration in a temporary unavailable state, causing the storage and runtime services that depend on that database to fail requests.
## Detection
The incident was detected through an automated alert for the Document Understanding services, which was acknowledged on August 21, 2026 at 12:10 pm UTC.
## Response
Before the customer-impacting incident was declared, the primary database scale-up was completed and support from the database platform provider was engaged for the secondary database issue. After the failover configuration became broken, we explored redirecting the service connection but could not identify a safe, immediate path to do so given the service's current configuration.
Service was restored by deleting the unhealthy secondary database and recreating the failover configuration. By August 21, 2026 at 12:37 pm UTC, the fix had been implemented and monitoring was in progress. At 1:36 pm UTC, monitoring confirmed the service was healthy, and the incident was marked resolved.
## Follow up
1. Request a root cause analysis from the database platform provider to determine why the secondary database could not be scaled and why the failover configuration remediation required breaking replication.
2. Update database storage alerting thresholds and routing so alerts are assigned and acted on earlier, including a lower-severity warning at 75% usage and a higher-severity alert at 85% usage.
US - Agents - Some Customers May Experience Errors When Using Claude Sonnet 4.6
Started August 19, 2026 at 3:08 PM UTC · 1h 47m
OutageMajor incident
Affected components
Agents
investigating
We are investigating an issue that may affect some customers using Claude Sonnet 4.6 in Agents in the US region. Our engineering team is actively working to understand the issue and will share further updates as more information becomes available.
monitoring
We have mitigated the issue. Our engineering team is actively morning and will share further updates as more information becomes available.
monitoring
We have mitigated the issue. Our engineering team is actively monitoring and will share further updates as more information becomes available.
resolved
The issue has been resolved.
postmortem
## Customer Impact
Between August 19, 2026 at 1:36 pm UTC and August 19, 2026 at 3:53 pm UTC, a subset of customers received errors when using Claude Sonnet 4.6 in Agents. The impact lasted approximately 2 hours and 17 minutes.
The impact was to customers using Agents in the US region. Errors were also observed for Claude Opus 4.6 and Claude Opus 4.5, which are used at lower volume.
---
## Root Cause
As part of a planned infrastructure migration, we moved the platform service that routes model requests for Agents onto a new configuration delivery system, region by region.
The new configuration source was missing the routing entries for three Claude models — Claude Sonnet 4.6, Claude Opus 4.6, and Claude Opus 4.5. Without those entries, the service could not resolve a valid destination for requests to those models, and rejected them with errors. Other models were unaffected and continued to serve normally throughout.
---
## Detection
The issue was detected through customer escalations on August 19, 2026 at 2:55 pm UTC — approximately 1 hour and 19 minutes after the first affected request. Our engineering team scoped the issue to specific Claude models and began investigation. Public status communication began at 3:08 pm UTC.
Our alert monitoring is based on aggregate error rates. Although nearly all requests to the three affected models were failing, those models represented a small share of overall traffic in the region, so the aggregate signal did not cross our alerting thresholds and the issue was not raised automatically. This is the detection gap addressed in the follow-up below.
---
## Response
At 3:25 pm UTC, the incomplete configuration source was identified as the cause, and a fix was started. At 3:47 pm UTC, deployment of the fix was in progress, and the service was actively monitored as the change rolled out.
By 3:59 pm UTC, service logs confirmed the issue was mitigated, and at 4:00 pm UTC the failure rate was confirmed at 0%. The incident was marked mitigated at 4:40 pm UTC, and full resolution was declared at 4:55 pm UTC after continued monitoring and customer confirmation that the service was working as expected.
---
## Follow-Up
The infrastructure migration has been completed across all regions and routing configuration now draws from a single source, removing the mismatch that caused this incident so it cannot recur.
Automated checks are being introduced to continuously verify every supported model in every region, so an unavailable model is detected and alerted immediately — including in lower-traffic regions.
[Community] - [Agentic Orchestration] - Reports of failures in expression evaluations related to the HITL task output parameters
Started August 18, 2026 at 5:54 PM UTC · 5h 31m
OutageMajor incident
Affected components
Agentic Orchestration
investigating
We are investigating reports of an outage impacting failures in expression evaluations related to the HITL task output parameters for Asgentic Orchestratron in community users in Europe.
Impact: Users may be unable to complete HITL tasks
Next update: Our teams are working to understand the cause and scope and will share updates as available.
identified
We have identified the cause for an outage impacting expression evaluations related to the HITL task output parameters for Agentic Orchestration in community users in Europe.
Impact: Users will experience failures when they have an expressions using HITL task output parameter.
Next update: Our teams are working to understand the cause and scope and will share updates as available.
identified
We have identified the fix and the resolution is in progress.
Next update: Our teams are working on the fix and will share updates as available.
monitoring
We have deployed the fix and are monitoring the resolution.
Next update: Our teams are monitoring the resolution and will share updates as available.
resolved
The outage has been resolved and Agentic Orchestration is fully operational.
Impact: No ongoing user impact.
Multiple Regions - Studio Web & Solutions Mgmt - Resource Configuration Screen Not Loading
We have identified the root cause of an issue in Studio Web where the resource configuration screen for changing resource attributes is not loading. We are deploying a fix.
identified
Deployment of the fix is in progress. We will provide further updates as the deployment progresses.
identified
The fix has been verified and is being rolled out across all remaining regions. We are monitoring the deployment and recovery. Thanknyou for your patience.
identified
The rollout is progressing as expected across the remaining regions. We continue to monitor the deployment. Thank you for your patience.
monitoring
The rollout is progressing as expected across the remaining regions. We continue to monitor the deployment. Thank you for your patience.
resolved
The rollout is completed and the issue should be resolved.
Multiple Regions - Studio Web - New entities showing up with a delay
Started August 18, 2026 at 5:32 AM UTC · 6h 25m
Pending
Affected components
Studio WebStudio WebStudio Web
investigating
We are investigating an issue affecting Community accounts where newly created entities may take approximately one hour to appear in Studio Web. No data is lost, and existing entities are unaffected.
investigating
We continue to investigate the issue and are working to identify the root cause and restore normal processing times.
investigating
We continue to investigate the issue affecting a subset of tenants in Europe, the United States, and Japan, where newly created entities may take longer than expected to appear in Studio Web. No data is lost, and existing entities are unaffected.
identified
We have identified the cause and are working on a resolution for the issue affecting a subset of tenants in Europe, the United States, and Japan. Thank you for your patience.
monitoring
The issue has been mitigated, and we expect processing times to return to normal shortly. We are monitoring the recovery closely. Thank you for your patience.
resolved
The issue has been resolved, and processing time has returned to normal. Thank you for your patience.
postmortem
## Customer impact
Between August 18, 2026 at 5:11 am UTC and August 18, 2026 at 11:57 am UTC, newly created entities in a subset of UiPath Cloud tenants could take longer than expected to appear in Studio Web. At the time of initial assessment, newly created entities were appearing with an approximately one-hour delay. Customers in Europe, US, and Japan were affected. Entity creation itself continued successfully, no data was lost, and existing entities were unaffected.
## Root cause
The issue was caused by an unusually high volume of asset-creation requests from one tenant. Those requests generated more events than our backend entity-indexing service could process at the same rate, creating a backlog in the event-processing queue. Because Studio Web relies on this service to display newly created entities, new entities appeared only after the backlog was processed.
## Detection
The issue was identified by our engineering team by entity processing latency alert, and an incident was declared at 5:11 am UTC on August 18, 2026. By 5:23 am UTC, analysis confirmed the last processed entity was approximately one hour behind.
## Response
At 5:32 am UTC, we posted an initial customer update noting delayed visibility for newly created entities in Studio Web. By 6:07 am UTC, investigation identified unusually high asset-creation traffic from one tenant, and at 7:03 am UTC, we adjusted database resources for the affected service to help processing recover.
At 7:59 am UTC, the source of the high request volume had stopped sending requests, and the queue began draining. At 9:59 am UTC, we started a data synchronization process for the affected tenant, and at 10:31 am UTC, we removed the problematic queued events so normal processing could catch up faster. Queue depth decreased from 95,000 items at 8:34 am UTC to 1,000 items at 11:44 am UTC.
The incident was marked mitigated at 11:19 am UTC after processing recovered substantially, and resolved at 11:57 am UTC after processing time returned to normal.
## Follow up
Re-trigger data synchronization for the affected tenant and monitor ingestion until the tenant's data is confirmed consistent.
Improve entity handling in the indexing backend so the backlog will not pile up at this rate.
Orchestrator Robot Logs - US
Started August 14, 2026 at 9:24 PM UTC · 1h 59m
OutageMajor incident
Affected components
Orchestrator
identified
We have identified the cause of the degraded performance impacting Orchestrator in US region and are working on mitigation.
Impact: Users may experience delayed loads and views on Orchestrator Robot logs. Additional updates will be provided as we move toward resolution.
resolved
The issue has been resolved and Orchestrator Robot logs performance has returned to expected levels after degraded performance impacted Robot logs to load in US region.
Impact: No ongoing user impact.
postmortem
## Customer impact
Between August 14, 2026 at 8:54 pm UTC and August 15, 2026 at 2:09 AM UTC, a subset of customers in the US region experienced significant slowness in the Orchestrator Jobs and Logs pages, and robot logs appeared later than expected in the logs view. Performance had substantially recovered by 11:22 PM UTC on August 14, with full recovery confirmed with affected customers at 2:09 AM UTC on August 15.
Automation execution was not affected, jobs continued to be scheduled and to run normally throughout. No log data was lost. Logs continued to be recorded and became visible once the system caught up. Requests did not fail, so no errors were surfaced, pages were slow to load and recent activity appeared missing or delayed. No other region was impacted.
## Root cause
Orchestrator stores and retrieves robot logs using a dedicated search and storage system. Routine maintenance on that system causes data to be redistributed internally across the cluster. Our analysis indicates that a redistribution larger than anticipated consumed capacity that would otherwise have served customer requests, slowing both the retrieval of existing logs and the processing of new ones.
This accounts for the majority, but not the entirety, of the slowdown observed, and analysis of the remaining contributing factor is continuing. Capacity returned to normal without intervention, at which point log visibility and page performance recovered.
## Detection
The issue was surfaced through customer reports of slow Jobs and Logs pages in the US region.
## Response
We posted a status update confirming that we were investigating degraded Orchestrator performance in the US region.
Our engineering team scoped the impact to the US region and narrowed the slowdown to the log storage and search layer. The degradation stemmed from capacity contention that eased as the redistribution completed, and responders monitored the system through recovery.
Page performance and log visibility returned to expected levels over the course of the evening, and recovery was subsequently confirmed with affected customers at 2:09 AM UTC on August 15.
## Follow up
1. We are adding monitoring and alerting on the response times customers experience and on the delay between a robot log being generated and becoming visible, so that degradation of this kind is detected proactively.
2. We are documenting an operational procedure that gives our on-call engineers defined steps to reduce customer impact during this class of degradation.
3. We are changing how routine maintenance on the log storage system is scheduled and paced in the US region so that it does not affect customer-facing performance.
4. We are increasing spare capacity in the log storage system so that internal data movement has room to complete without competing with customer requests.
IXP Communications Mining elevated error rate in US region
Started August 12, 2026 at 3:00 PM UTC · 0m
Pending
resolved
A storm of requests on a a rarely used model feature in IXP Communications Mining caused a retry loop to lock down synchronous requests in the wider IXP API. This caused requests to fail with 500's as workers were busy with long-running requests.
Automatic scaling up quickly reached it's maximum capacity and resolution was done only by applying a code fix that introduced strict timeouts to the contributing API request.
The request storm started around 15:10 UTC and detected a 15:15 UTC by automatic alarms. Resolution was confirmed around 18:15 UTC.
The incident was initially wrongly attributed to only the client making the storm of requests, but was later discovered to have impacted a wider range of users.
Total impact limited to a handful of users in the US.
postmortem
## Customer Impact
On August 12, 2026, between approximately 15:10 and 18:15 UTC, Communications Mining (IXP) users in the United States region experienced intermittent request failures.
The incident was confined to one of the region's deployment units, where users on it saw failures in bursts of five to ten minutes, with up to 5–7% of their requests failing with 5xx errors at peak.
Between bursts the service operated normally, retried requests generally succeeded, and no data was lost.
## Root Cause
A storm of requests to a rarely used API feature that computes machine-learning predictions on demand coincided with repeated retraining of the requested model. Each retraining invalidated cached predictions, turning each request into a multi-minute computation.
The API placed no time limit on how long a request could wait for this computation, so these long-running requests progressively occupied all request-processing capacity, causing unrelated requests to fail. Automatic scaling reached its maximum capacity quickly and could not compensate.
## Detection
Automated monitoring detected the failures at 15:15 UTC, about five minutes after impact began, and paged the on-call engineer.
The incident was initially attributed only to the client generating the request storm, but client reports and further investigation showed a wider set of users was affected during the failure bursts.
## Response
The on-call engineer traced the failures to the unbounded wait in the on-demand prediction path. Service capacity was repeatedly restored by automatic instance replacement while a code fix was developed.
The fix is a strict timeout on the contributing request so it fails fast without affecting other requests, and was deployed to the affected region at approximately 18:00 UTC, and resolution was confirmed at 18:15 UTC.
## Follow-Up
1. The strict timeout and load-shedding fix has been made permanent and released to all regions (completed August 13, 2026).
2. Evaluate per-client limits on on-demand prediction computation so that a single client's usage cannot degrade the shared API.
Document Understanding Outage on GXP East US Region
Started August 10, 2026 at 2:11 PM UTC · 14m
Pending
Affected components
Document UnderstandingDocument Understanding
investigating
We are investigating degraded performance impacting Front-End Services in Document Understanding across GXP East US Region.
Impact: Users may notice timeouts when accessing the UI in GXP US.
Next update: Additional updates will be provided as more information becomes available.
resolved
The issue has been resolved and the performance of the Front-End Service in Document Understanding has returned to expected levels after degraded performance impacted the UI in the GXP East US Region.
Impact: No ongoing user impact.
postmortem
## Customer impact
Between August 10, 2026 at 13:21 UTC and August 10, 2026 at 14:03 UTC, a subset of customers experienced degraded performance and timeouts when accessing the Document Understanding user interface, which provides the design-time experience, in the Delayed US region. Document processing automations were not affected.
## Root cause
During a manual deployment from a pre-existing build of the service behind the Document Understanding user interface to the Delayed US region, our deployment process reassigned the version identifier for the service being deployed. The Document Understanding interface requires that its supporting resources are available under the same version identifier as the deployed service. Because the deployment process changed this identifier, the interface could not locate the required resources, causing it to become inaccessible or time out.
## Detection
We became aware of the issue moments after the deployment was finished, through manual verification as part of the manual deployment checklist. An automated alert shortly followed, triggered at 13:28 UTC on August 10, 2026.
## Response
After identifying the reason for the failure of the manual deployment, our engineering team started to deploy a corrected update with the required resources available. At the same time, the operations team was engaged to perform a manual rollback of the deployment. The rollback was completed at 14:03 UTC, restoring access to the design-time experience.
At 14:11 UTC, we published a public status update indicating degraded performance and possible timeouts in the Document Understanding user interface for the Delayed US region. By 14:15 UTC, the corrected update had completed and the interface was confirmed as deployed with a correct new version and working.
At 14:25 UTC, the incident was marked resolved and the public status page was updated to confirm that Document Understanding interface performance had returned to expected levels.
## Follow up
1. We are reducing the time required to restore a previous version of the service, so that recovery from a failed deployment is faster.
2. We are making improvements to the manual deployment process to prevent a future situation, by validating the required resources' existence as a prerequisite.
Uipath Apps is facing outage in Delayed US region
Started August 8, 2026 at 11:54 AM UTC · 4h 6m
OutageMajor incident
Affected components
Apps
identified
We have identified the cause of the outage impacting Uipath Apps is facing outage in Delayed US region and are working on a fix.
Impact: Users may continue to be unable to access Uipath Apps and solutions dependent on Uipath Apps.
Team is working on service restoration.
identified
Team is working on service restoration. We will update the status once mitigation is completed.
identified
Team has identified an issue with an underlying resource and is actively working to restore service.
identified
Team has made progress to fix underlying resource issue and is actively working to restore service.
monitoring
Mitigation has been applied and performance is improving for the issue.
We are monitoring closely to ensure stability.
resolved
The mitigation has remained stable, and performance has returned to expected levels. We have confirmed service restoration for UiPath Apps in the Delayed US region and are marking this incident as resolved.
postmortem
## Customer impact
Between 11:20 am UTC and 2:54 pm UTC on August 8, 2026, a subset of customers in the Delayed US region experienced failures accessing UiPath Apps and solutions that depend on UiPath Apps.
Customers may have seen UiPath Apps unavailable or intermittent request failures. The impact lasted approximately 3 hours and 34 minutes.
## Root cause
The incident was caused by database connection saturation following scheduled maintenance performed by our database provider. As application services scaled up, they created additional database connections, which caused new connection attempts to fail and resulted in connection reset errors in UiPath Apps.
## Detection
Automated alerts detected the issue at 11:24 am UTC on August 8, 2026. Application telemetry showed failures beginning at approximately 11:20 am UTC.
## Response
At 11:00 am UTC, scheduled maintenance began automatically. At 11:20 am UTC, requests began failing. At 11:24 am UTC, automated alerts were triggered, and the team began investigating.
At 12:24 pm UTC, database capacity was scaled up as a mitigation. At 1:13 pm UTC, application services were restarted to reduce saturated connection usage and refresh database connections. Connection levels remained elevated, and the database automatically scaled at 1:22 pm UTC and at 2:36 pm UTC.
Following these mitigation efforts, request failures stopped at 2:54 pm UTC. At 3:33 pm UTC, the mitigation was confirmed to be stable, and performance was improving. Full recovery was confirmed at 4:01 pm UTC after performance returned to expected levels.
## Follow up
1. Obtain and review the database provider's root cause analysis explaining what caused the connection issue following their maintenance activity.
2. Implement an application-side limit on database connection creation to prevent connection saturation.
3. We are reviewing the automatic scaling behavior that amplified connection volume during the incident and address any contributing factors.
US Region Document Ingestion Degradation
Started August 6, 2026 at 10:00 AM UTC · 0m
Pending
resolved
Between 06-08-2026 10:00 UTC and 06-08-2026 13:00 UTC, some organizations in the US region were unable to complete document ingestion. A small number of search requests in the same region were also slow or timed out.
The issue was caused by a capacity constraint affecting ingestion processing in US. Normal performance was restored at 13:00 UTC.
We have monitored the affected environments since recovery and confirm the issue is fully mitigated. Ingestion requests that failed during this window were not retried automatically and will need to be re-submitted. No action is required for search.
postmortem
## Customer impact
Between August 6, 2026 at 10:00 am UTC and 1:00 pm UTC, some organizations in the US region were unable to complete document ingestion in **UiPath Context Grounding**. A small number of search requests in the same region were also slow or timed out.
**Action required:** please re-submit the affected ingestion requests. Ingestion retries a failing request automatically for a limited number of attempts. Once those attempts are exhausted the request is marked failed and is not retried again, so affected documents will not appear in your index until the request is submitted again. Failed requests are listed in the ingestion history for each index.
No action is required for search. Those requests were affected only while the issue was ongoing, and subsequent searches completed normally.
---
## Root cause
A sudden increase in concurrent document ingestion triggered a high number of simultaneous document validation steps, which created a capacity bottleneck on the underlying infrastructure resource beyond its scaling capacity.
Once that resource was saturated, ingestion operations began exceeding their time limits and failing. Automatic retries of the failed operations added further load, which sustained the condition. Search requests served by the same resource were delayed behind the same contention.
---
## Detection
Automated alerts were flagged as the condition developed, and an automated infrastructure resource capacity alert triggered at 10:37 am UTC brought it to the team's attention.
---
## Response
- **10:03 am UTC** — Automated low severity alerts started coming in.
- **10:37 am UTC** — Automated alert for resource capacity issue paged the team.
- **12:23 pm UTC** — As a mitigation step the impacted resource's capacity was increased.
- **12:57 pm UTC** — Ingestion and search operations stopped failing and response times returned to normal.
---
## Follow-up
- **The fix is deployed.** The validation step has been reimplemented to enforce the same limits at a small fraction of the previous cost, so this level of concurrent ingestion now sits well within available capacity. It was released to the affected US region on August 7, ahead of schedule, and reaches all remaining regions by early September.
- **We are improving how quickly we detect issues like this.** We are adding monitoring that tracks whether document ingestion is completing successfully for customers, so problems are identified and acted on directly rather than inferred from underlying system alerts. This will be in place across all regions by the end of August.
Uipath outage history and incident timeline | Uptimus