Degraded Performance of Project management in Phrase TMS (EU) starting on September 1, 2026 16:12 CEST
Started September 1, 2026 at 2:18 PM UTC · 2h 5m
OutageMajor incident
Affected components
Project management
investigating
We are currently investigating an issue causing slowness in the Jobs view within Phrase TMS (EU). Users may experience delayed load times or reduced performance when accessing project management features. Our team is actively working to identify the root cause and restore normal performance as quickly as possible.
investigating
We have identified the root cause of the issue. Our team is actively implementing a fix. We will provide further updates as we make progress.
monitoring
A fix has been deployed and we are monitoring the results. Customers might need to clear browser cache and close all opened tabs with TMS if still experiencing problems.
resolved
The issue has been resolved and performance is stable.
Disruption of Gengo Translation Orders - API Timeouts Affecting Order Submission (Strings EU and US DC)
Started August 31, 2026 at 3:06 PM UTC · 1d 20h
OutageMajor incident
Affected components
OrderingOrdering
investigating
We are currently experiencing intermittent errors when communicating with Gengo, one of our third-party translation vendors. This is causing timeouts and 502 errors on price calculation and order submission for translation orders routed through Gengo, affecting both EU and US data centers. Gengo has confirmed the issue on their end and is investigating.
resolved
Gengo has resolved the underlying issue on their end. Price calculations and translation order submissions to Gengo are now working again.
Degraded performance of TMS Project Management Component (EU and US DC) between August 31 2:26 PM CEST and August 31 5:57 PM CEST
Started August 31, 2026 at 1:11 PM UTC · 3h 42m
OutageCritical incident
Affected components
Project managementProject management
investigating
Users are currently unable to open existing jobs and to create new ones within their Phrase TMS projects. We are investigating the issue.
investigating
We are continuing to investigate this issue.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
The issue has been resolved.
postmortem
## Introduction
We would like to share details about an incident that affected Phrase TMS on August 31, 2026. Between 2:26 and 5:57 PM CEST, some users of the legacy Project page were unable to use Tools-menu actions such as creating or editing jobs and running analyses, and clicking to open a job did not launch the CAT web editor as expected, instead returning users to the project page. This post-mortem explains what happened, how it was resolved, and what we are doing to prevent it from happening again.
## Timeline
* **August 31, 2026 at 2:26 PM CEST** – A code change reached production containing a defect that broke script execution on the Project page for any user who loaded it.
* **August 31, 2026 at 2:49 PM CEST** – The first customer reports came in describing Tools-menu buttons as disabled and jobs failing to open in the editor.
* **August 31, 2026 at approximately 3:10 PM CEST** – Our team identified the root cause of the issue.
* **August 31, 2026 at 3:48 PM CEST** – A fix for the underlying defect was completed.
* **August 31, 2026 at 4:04 PM CEST** – The fix was verified in a pre-production environment.
* **August 31, 2026 at 5:18 PM CEST** – Deployment of the fix to production began.
* **August 31, 2026 at 5:57 PM CEST** – The fix was fully live in production and normal functionality was restored for all affected customers.
## Root Cause
The incident was caused by a code change intended to fix an unrelated, minor display issue on shared project pages. That change altered how a value was inserted into a script embedded directly in the page. The system that renders the page automatically encodes values for safety, but that encoding does not distinguish between a value being placed in regular page content versus inside a script. As a result, the embedded script's syntax was silently broken once the change reached production.
Because browsers stop executing any further code on a page once they encounter invalid script syntax, every script placed after that point on the page stopped running — not just the part related to the original change. This is why customers experienced what looked like two separate problems \(disabled menu buttons and jobs failing to open in the editor\) that were, in fact, downstream effects of the same single defect.
The issue was not caught before release because the verification performed at the time confirmed that the underlying data being inserted was correct, but did not load the actual page in a browser to confirm it rendered and executed correctly end-to-end.
## Actions to Prevent Recurrence
1. **Fix deployed** – The underlying defect was corrected and deployed to production the same day it was identified.
2. **Engineering guidance updated** – We have updated our internal engineering documentation to clearly describe this specific failure pattern and the correct, safe way to handle it, so this category of mistake is caught during code review going forward.
3. **Automated detection improvements underway** – We are working on adding automated monitoring for this class of front-end failure, so similar issues can be detected and addressed before customers are affected, rather than relying on customer reports.
Degraded Performance of Translation memory in Phrase TMS (EU)
Started August 19, 2026 at 12:16 PM UTC · 1h 53m
OutageMajor incident
Affected components
Translation memory
investigating
We are currently investigating an issue affecting Translation Memory functionality in Phrase TMS (EU). Users may experience problems with creating Translation Memories and other related operations during this time. Our team is actively working to identify the cause and restore full functionality as quickly as possible.
identified
The issue has been identified, and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the situation.
resolved
The incident has been resolved.
Degraded Performance of Branching in Phrase Strings (EU) between August 6, 2026 03:45 PM CEST and August 7, 2026 10:46 AM CEST
Started August 7, 2026 at 7:27 AM UTC · 6h 45m
OutageMajor incident
Affected components
APITranslation center
investigating
We are currently investigating an issue with branching events missing in Phrase Strings (EU), which may cause changes to not be applied when branches are merged.
investigating
We are continuing to investigate this issue.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented. The backlog of missing events is now being processed. Once the queue is fully processed, unmerged branches will automatically receive the missing changes. We are monitoring the queue and results.
resolved
Changes made to branches during the incident window are now fully applied. Merged branches that were affected have been identified and impacted customers have been contacted directly. This incident has been resolved.
Degraded Performance of Connectors in Phrase TMS (EU) on August 6, 2026
Started August 6, 2026 at 8:11 AM UTC · 5h 8m
OutageMajor incident
Affected components
ConnectorsConnectors
investigating
We are currently investigating an issue affecting Connectors in Phrase TMS. Customers may be unable to access or use Connector functionality during this time. Our engineering team is actively working to identify the cause and restore full service as quickly as possible.
investigating
We are continuing to investigate this issue.
monitoring
We have identified the root cause and the issue has now been resolved. Connectors should be appearing as active again. We will continue to monitor the situation to ensure full stability.
monitoring
We continue monitoring the situation.
resolved
The incident has been resolved and all components are back to operational.
Degraded Performance of Phrase Orchestrator (EU) Next-Gen Workflow Engine between July 28, 05:15 AM CEST and July 28, 09:56 AM CEST
Started July 28, 2026 at 7:19 AM UTC · 4h 30m
OutageMajor incident
Affected components
Next-Gen Workflow Engine
investigating
The engineering team identified an issue with Orchestrator where the new Workflow Engine is currently not executing workflows. The problem is under investigation.
investigating
We are continuing to investigate this issue.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results. The workflow engine is processing the queue of pending executions.
resolved
All executions have been processed. This incident has been resolved.
postmortem
## Introduction
We would like to share details about an incident that affected Phrase Orchestrator on July 27–28, 2026. During this period, workflow executions in the Next-Gen Workflow Engine were unable to progress and remained stuck in an "executing" state. No data was lost during the incident. This post-mortem explains what happened, when it was resolved, and the steps we have taken to prevent a recurrence.
## Timeline
* **27 July 2026 at 18:55 CEST** – The workflow engine began producing errors as database query performance degraded. Workflow executions stalled and stopped progressing.
* **27 July 2026 at 20:54 CEST** – The first customer report of executions stuck in "executing" was received.
* **27 July 2026 at 22:39 CEST** – The incident was formally declared.
* **28 July 2026 at 02:18 CEST** – A service restart provided temporary relief; workflow executions resumed.
* **28 July 2026 at 05:15 CEST** – The issue recurred as the underlying database performance problem persisted.
* **28 July 2026 at 09:56 CEST** – The root cause was identified and addressed. No executions were lost; however, due to partial service restarts, some actions within executions were retried, which may have caused a small number of executions to fail that otherwise would have succeeded.
* **28 July 2026 at 13:46 CEST** – The full backlog of stalled executions was confirmed as cleared. The system was declared stable.
* **28 July 2026 at 13:48 CEST** – Incident resolved.
## Root Cause
The incident was caused by progressive bloat in database indexes used by the workflow job scheduling system. The performance of these particular indexes gradually degraded over time as they accumulated dead index entries from prior writes and updates.
The job scheduling engine acquires database-level coordination locks while querying these indexes to determine which jobs to dispatch. As the index lookups grew slower, they began exceeding the database's configured statement timeout. When a lookup was canceled by the timeout, the scheduling process responsible for that work crashed and restarted. With no schedulers running, no workflow steps could be dispatched and all in-progress workflow executions became stuck.
The database server itself remained healthy throughout the incident, with normal CPU and connection levels. The problem was exclusively lock and latency contention within the scheduling layer. A service restart cleared the crashed processes and temporarily restored execution. However, because the index bloat was still present, the same degradation recurred once query load resumed. A manual index rebuild fully restored performance and resolved the issue.
## Actions to Prevent Recurrence
1. **Automated index maintenance added** – Scheduled automatic index maintenance has been configured for the affected indexes. This ensures bloat cannot accumulate over time and eliminates the conditions that triggered this incident.
2. **Legacy indexes removed** – Unused legacy database indexes have been identified and removed, reducing the overall maintenance surface and simplifying future index hygiene.
3. **Monitoring coverage updated** – Our monitoring landscape is being reviewed and updated to reflect the current state of the workflow engine. This work will close gaps that allowed the degradation to go undetected before the first customer report.
Degraded Performance of Phrase Orchestrator (EU) Next-Gen Workflow Engine between July 27, 06:55 PM CEST and July 28, 02:14 AM CEST
Started July 27, 2026 at 11:48 PM UTC · 52m
OutageMajor incident
Affected components
Next-Gen Workflow Engine
investigating
Engineering has identified an issue with Orchestrator where the new Workflow Engine is currently not executing workflows. The problem is under investigation.
resolved
All workflows are being executed again as expected.
postmortem
## Introduction
We would like to share details about an incident that affected Phrase Orchestrator on July 27–28, 2026. During this period, workflow executions in the Next-Gen Workflow Engine were unable to progress and remained stuck in an "executing" state. No data was lost during the incident. This post-mortem explains what happened, when it was resolved, and the steps we have taken to prevent a recurrence.
## Timeline
* **27 July 2026 at 18:55 CEST** – The workflow engine began producing errors as database query performance degraded. Workflow executions stalled and stopped progressing.
* **27 July 2026 at 20:54 CEST** – The first customer report of executions stuck in "executing" was received.
* **27 July 2026 at 22:39 CEST** – The incident was formally declared.
* **28 July 2026 at 02:18 CEST** – A service restart provided temporary relief; workflow executions resumed.
* **28 July 2026 at 05:15 CEST** – The issue recurred as the underlying database performance problem persisted.
* **28 July 2026 at 09:56 CEST** – The root cause was identified and addressed. No executions were lost; however, due to partial service restarts, some actions within executions were retried, which may have caused a small number of executions to fail that otherwise would have succeeded.
* **28 July 2026 at 13:46 CEST** – The full backlog of stalled executions was confirmed as cleared. The system was declared stable.
* **28 July 2026 at 13:48 CEST** – Incident resolved.
## Root Cause
The incident was caused by progressive bloat in database indexes used by the workflow job scheduling system. The performance of these particular indexes gradually degraded over time as they accumulated dead index entries from prior writes and updates.
The job scheduling engine acquires database-level coordination locks while querying these indexes to determine which jobs to dispatch. As the index lookups grew slower, they began exceeding the database's configured statement timeout. When a lookup was canceled by the timeout, the scheduling process responsible for that work crashed and restarted. With no schedulers running, no workflow steps could be dispatched and all in-progress workflow executions became stuck.
The database server itself remained healthy throughout the incident, with normal CPU and connection levels. The problem was exclusively lock and latency contention within the scheduling layer. A service restart cleared the crashed processes and temporarily restored execution. However, because the index bloat was still present, the same degradation recurred once query load resumed. A manual index rebuild fully restored performance and resolved the issue.
## Actions to Prevent Recurrence
1. **Automated index maintenance added** – Scheduled automatic index maintenance has been configured for the affected indexes. This ensures bloat cannot accumulate over time and eliminates the conditions that triggered this incident.
2. **Legacy indexes removed** – Unused legacy database indexes have been identified and removed, reducing the overall maintenance surface and simplifying future index hygiene.
3. **Monitoring coverage updated** – Our monitoring landscape is being reviewed and updated to reflect the current state of the workflow engine. This work will close gaps that allowed the degradation to go undetected before the first customer report.
Degraded Performance of Phrase Strings API (EU) between June 16, 2026 01:00PM CEST and June 17, 2026 11:00AM CEST
Started July 20, 2026 at 12:53 PM UTC · 0m
IssuesMinor incident
resolved
Between June 16, 2026 01:00PM CEST and June 17, 2026 11:00AM CEST, the Phrase Job Sync connector suffered degraded performance. Customers using the Job Sync with the default connection type experienced failures. The engineering team identified a root cause and provided a fix.
postmortem
### Introduction
We would like to share more details about the events that occurred with Phrase between June 16, 2026, 01:00PM CEST and June 17, 2026, 11:00AM CEST, which led to degraded performance of Phrase Strings API, causing customers using the Job Sync with the default connection type to experience failures. We apologize for the disruption and are committed to preventing similar incidents in the future.
### Timeline
**Jun 16, 2026 @ 01:00 PM CEST** – Job Sync began failing for customers using the default connection type. Connectors were unable to authenticate against the Phrase Strings API.
**Jun 17, 2026 @ 10:00 AM CEST** – The authentication failure was identified by the engineering team.
**Jun 17, 2026 @ 10:28 AM CEST** – Impact scope was confirmed: Customers using the default Job Sync connection were affected. Customers using personal access token-based connectors were not affected.
**Jun 17, 2026 @ 10:29 AM CEST** – Root cause was identified: A recent change which added multi-platform token support inadvertently introduced errors with some existing platform tokens.
**Jun 17, 2026 @ 10:40 AM CEST** – A fix was prepared and submitted for review.
**Jun 17, 2026 @ 10:53 AM CEST** – The fix was deployed and JobSync functionality was fully restored.
### Root Cause
A code change introduced to add support for multi-platform authentication tokens modified how incoming platform tokens are validated and processed in the Phrase Strings API. This change was not compatible with an existing token format used by the default connection type.
As a result, platform tokens were rejected by the Strings API with a `401 Unauthorized` response. The connector retried establishing the connection until exhausting its retry budget, causing all affected operations to fail. Customers using personal access tokens were not affected as their tokens follow a different code path.
### Actions to Prevent Recurrence
The root cause was a backwards-incompatible change to token handling that was not detected during development. The following actions are being taken:
* **Token format compatibility testing in auth changes:** When introducing a new token/authentication format alongside an existing one, tests must explicitly cover the transition case—old-format tokens being processed under the new detection logic—not just each format in isolation.
Degraded Performance of Phrase (EU & US) starting on July 16, 2026 17:24 CEST
We are currently investigating an issue affecting our Metrics service that may result in duplicated metrics consumption for some customers. Our team is actively working to identify the affected components and determine the root cause. We will provide further updates as more information becomes available.
identified
We have identified the root cause as instability on a specific broker, which is causing repeated consumer group rebalancing in our Metrics service and leading to duplicate processing of metric updates. This may result in some customers experiencing incorrect consumption calculations or unexpected account blocking. We continue working on a fix.
monitoring
We have stopped the progression of the issue and are now working to identify the full scope of impact, including which organizations are affected or blocked. Affected accounts will be prioritized for resolution through a deduplication process, though this will take some time to complete.
identified
We have successfully deduplicated metrics consumption for all blocked organizations. We are continuing work to develop per-organization deduplication tooling to ensure accurate billing cycle data, with the full team scaling these fixes tomorrow morning.
monitoring
We continue to make progress resolving the duplicated metrics consumption issue first identified on July 16. The root cause has been identified and addressed. Since our last update, our team has corrected the duplicated data for all previously blocked accounts. We are now working through a remaining set of affected accounts and expect this cleanup to continue over the coming days. No further duplication is occurring, this work addresses historical data only.
resolved
This incident has been resolved.
Performance Disruption of Phrase Orchestrator (EU) on July 16, 2026 between 4:07 PM CEST and 5:32 PM CEST
Started July 16, 2026 at 2:49 PM UTC · 1h 22m
OutageCritical incident
Affected components
Workflow Builder
investigating
We are currently investigating an issue affecting Phrase Orchestrator. We are working to identify the root cause and will provide further updates as soon as more information is available.
investigating
We are continuing to investigate this issue.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
Orchestrator (EU) is available again since 5:32 PM CEST. All pending executions are being processed. This incident has been resolved.
Degraded Performance of Translation center in Phrase Strings (EU & US) between July 10, 2026 16:21 CEST and July 10, 2026 17:13 CEST
Started July 10, 2026 at 2:24 PM UTC · 1h 1m
OutageMajor incident
Affected components
Translation centerTranslation center
investigating
We are currently investigating an issue affecting the Translation Center in Phrase Strings (EU and US regions), where strings are failing to load. Our engineering team has identified a potential cause and is actively working to resolve it. We will provide further updates as our investigation progresses.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
postmortem
## Introduction
On July 10, 2026, parts of the Phrase Strings frontend at [app.phrase.com](http://app.phrase.com/) failed to load correctly for a short period, after required stylesheets and scripts failed to load — the underlying service itself stayed available. The issue began at 16:08 CEST and was resolved by 16:27 CEST. This post-mortem explains what happened and what we're doing to prevent a recurrence.
## Timeline
* **Jul 10, 2026 at 16:08 CEST** – A deployment introduced new application assets that were never uploaded to our storage service, so the application could no longer load its stylesheets and scripts for customers.
* **Jul 10, 2026 at 16:21 CEST** – Issue detected, incident declared.
* **Jul 10, 2026 at 16:24 CEST** – Root cause identified.
* **Jul 10, 2026 at 16:26 CEST** – Fix applied.
* **Jul 10, 2026 at 16:27 CEST** – Customer-facing impact ended.
* **Jul 10, 2026 at 16:32 CEST** – Incident downgraded to monitoring while the underlying configuration was addressed.
* **Jul 10, 2026 at 17:12 CEST** – Configuration corrected, incident resolved.
## Root Cause
During a planned, routine maintenance, a configuration change unintentionally disabled the process that uploads application assets \(stylesheets, scripts\) to our storage and content-delivery service. This went unnoticed at first, since the assets already in place kept working and no new versions had been shipped yet.
On July 10, 2026, a deployment shipped new versions of these assets. Because uploading was still disabled, the new files never reached the storage service, and the application could not load its stylesheets and scripts.
## Actions to Prevent Recurrence
1. **Asset synchronization re-enabled** – The disabled process was fixed, restoring correct asset delivery.
2. **Asset availability monitoring** – Adding monitoring to detect asset load failures automatically, before customers are affected.
Degraded performance of TMS Project Management Component
Started July 8, 2026 at 12:57 PM UTC · 57m
IssuesMinor incident
Affected components
Project management
investigating
We are getting reports from users about jobs disappearing from projects and we are currently investigating the issue.
monitoring
The issue has been identified, a fix was implemented and we are monitoring the situation.
resolved
The incident has been resolved.
Performance Disruption of Phrase TMS CAT Web Editor on July 7, 2026
Started July 7, 2026 at 8:34 AM UTC · 53m
OutageMajor incident
Affected components
CAT web editor
investigating
We are currently investigating issues with CAT Web Editor.
monitoring
The issue has been identified, a fix was implemented and we are monitoring the situation.
resolved
The incident has been resolved.
TEST - Please ignore
Started June 29, 2026 at 7:30 AM UTC · 40m
OutageMajor incident
Affected components
Legacy Workflow Engine
investigating
Hi everyone, this is a test. Please ignore, thank you!
identified
Hi everyone, this is a test. Please ignore, thank you!
resolved
This incident has been resolved.
TEST - Please ignore
Started June 26, 2026 at 3:43 PM UTC · 51m
OutageMajor incident
Affected components
Legacy Workflow Engine
investigating
Hi everyone, this is a test. Please ignore. Thank you!
Test - Please ignore
Started June 26, 2026 at 7:31 AM UTC · 51m
OutageMajor incident
Affected components
Legacy Workflow Engine
investigating
Hello everyone, this is a test. Please ignore, thank you!
investigating
Hello everyone, this is a test. We're "investigating" the issue. Please ignore, thank you!
TEST - Please ignore
Started June 26, 2026 at 7:00 AM UTC · 4m
OutageMajor incident
Affected components
Legacy Workflow Engine
investigating
Hi everyone, this is a test. Please ignore, thank you!
resolved
This incident has been resolved.
TEST - Please ignore
Started June 23, 2026 at 12:53 PM UTC · 6m
OutageMajor incident
Affected components
Legacy Workflow Engine
investigating
Hello everyone, this is a test, kindly ignore.
resolved
This incident has been resolved.
TEST - Please ignore
Started June 22, 2026 at 2:16 PM UTC · 19m
OutageMajor incident
Affected components
Legacy Workflow Engine
investigating
This is a test, please ignore.
investigating
This incident has been resolved.
Memsource outage history and incident timeline | Uptimus