[Tokyo, Ap03 and US Regions] Performance degradation with Bulk Load, Bulk Import and Result Export jobs
Started July 27, 2026 at 8:29 AM UTC · 0m
Pending
Affected components
Data Connector IntegrationsData Connector IntegrationsData Connector Integrations
resolved
Between 7/25 and today, 7/27 at 11:30 AM JST, we experienced a performance degradation issue affecting BulkLoad (Data Connector Import), Result Export (Activation) and Bulk Import. Following reports received this morning, our engineering team investigated the issue and identified high system load under specific conditions (increasing delay in internal components communication) as the root cause. The issue has been resolved by full rotation of our server nodes, and full performance was restored as of 11:30 AM JST.
We sincerely apologize for any inconvenience caused. We are currently reviewing our monitoring and alerting rules to prevent similar issues in the future.
[Tokyo Region] - Treasure Work / tdx claude - Degraded Performance
Started July 7, 2026 at 5:04 PM UTC · 3h 55m
IssuesMinor incident
Affected components
Treasure Work / tdx
investigating
We are observing degraded performance in Treasure Work and tdx in the Tokyo region. This issue is limited to the Claude Haiku 4.5 model in the Tokyo region only. Users may notice that requests using Claude Haiku 4.5 return errors or time out.
At this time, we do not see any issues with other models. As a workaround, users can select a different model.
We are actively investigating this issue and will aim to provide an update in 30 minutes.
investigating
We are continuing to investigate this issue and will provide another update in 30 minutes.
identified
We are continuing to explore this issue and are currently evaluating a potential fix. We will update as soon as our testing is complete.
identified
We are continuing to work through a potential fix and working to identify any potential contributing factors. Again, as the issue is localized to Haiku 4.5 models in the Tokyo region, we recommend users temporarily use a different model if they are experiencing any elevated error rates.
We apologize for the inconvenience and will have another update soon.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This issue has been resolved. Treasure Work and tdx are operating normally, and the Claude Haiku 4.5 model in the Tokyo region is no longer experiencing errors or timeouts.
Thank you for your patience.
We are experiencing Hive query performance degradation in the US region. This issue began on June 25, 2026. Users may observe longer than expected query execution times, particularly for long-running jobs with large data volumes.
We have identified the cause and are working on a fix. We will provide an update in 30 minutes.
identified
We are experiencing Hive query performance degradation in the US region. This issue began on June 25, 2026. Users may observe longer than expected query execution times, particularly for long-running jobs with large data volumes.
We have identified the cause. To address this, we are validating a cluster configuration change prior to applying it to the production environment. We will provide an update in 30 minutes.
identified
We are continuing to work on a fix. Our team is validating a cluster configuration change prior to applying it to the production environment.
We will provide an update in 30 minutes.
monitoring
We have provisioned a new cluster with sufficient capacity and are now routing jobs to it. Query performance is expected to recover gradually as the cluster stabilizes. We are monitoring the situation closely.
We will provide an update in 1 hour.
resolved
Hive query performance degradation in the US region has been resolved. We have successfully provisioned a new cluster and routed all traffic to it. Systems are now operating normally.
Thank you for your patience.
Data Connector IntegrationsData Connector IntegrationsData Connector Integrations
resolved
Resolved — Some result export (data export) jobs in the US, Tokyo, and EU regions may have failed or been delayed between June 24, 2026 21:00 UTC and June 25, 2026 05:26 UTC.
During this period, underlying compute nodes were intermittently unstable, which could cause export jobs to terminate before completing. In some cases an affected job may have ended without a clear error message.The issue has been resolved and all systems are operating normally. Data integrity is not impacted — source data in Treasure Data was unaffected; only the delivery of export jobs to external destinations during this window could have been disrupted.If you ran a result/data export during this window, we recommend confirming it completed and re-running it if the output looks incomplete. We will also contact affected customers directly.
Thank you for your patience.
[US Region] - Delay in Email Event Data Recording
Started June 24, 2026 at 11:11 AM UTC · 0m
Pending
Affected components
Engage Studio - Delivery API
resolved
We identified a delay in recording email event data (such as open, click, bounce, delivery events, and error events) to customer event tables.
Affected period (approximate):
- June 22, 03:45 UTC – June 23, 04:15 UTC
- June 23, 07:00 UTC – June 24, 07:30 UTC
During these periods, email event data recording was delayed. Email delivery itself was not affected — emails were sent and received normally.
The issue has been resolved and all systems are operating normally. Delayed events have been fully processed.
If you ran queries on event or error_event tables during the affected periods, we suggest rerunning those queries to get accurate results.
We apologize for any inconvenience.
Realtime 2.0 Bulk Storage Outage
Started May 5, 2026 at 4:20 PM UTC · 1h 24m
IssuesMinor incident
Affected components
Realtime 2.0Realtime 2.0Realtime 2.0
identified
Incident: Real-time Profile Deletion Workflow Failures
Status: Identified
Impacted Regions: EU01, US01, Tokyo, AP03 (USA)
Update (May 5, 2026 - 11:19 AM CDT)
We have identified the root cause of the failures affecting Real-time (RT) bulk profile deletion, load, and stitch workflows in the affected regions.
Summary of the Issue:
A backend configuration issue resulted in a loss of necessary permissions for the service responsible for processing Real-time bulk profile operations. As a result, impacted workflows are failing with an "Authorization" or "400" error when attempting to invoke internal functions.
Current Status:
Our engineering team has developed a fix and is currently in the process of deploying it to all affected regions. We expect service to be restored shortly after the deployment is completed.
We will provide a further update once the fix has been fully deployed and we have confirmed that workflow processing has returned to normal.
identified
We are continuing to work on a fix for this issue.
resolved
We have applied the fix, and are observing that traffic is flowing correctly now.
Intermittent 5xx Errors for Ingest and Real-Time APIs (EU01)
Started April 27, 2026 at 5:31 PM UTC · 54m
IssuesMinor incident
Affected components
CDP Personalization - Ingest APIRealtime 2.0
monitoring
Between 8:05 AM and 8:30 AM PDT, customers operating in the EU01 (Europe) region may have experienced intermittent errors (5xx) affecting two primary areas:
Data Ingestion: Errors when attempting to send records via the Ingest API.
Real-Time Personalization: Request failures for personalization and profile lookup services.
The disruption was caused by an service interruption with an upstream provider in our data processing layer within the eu-central-1 region. This prevented our systems from successfully persisting incoming data to our streaming infrastructure and executing real-time lookups during the affected window.
The upstream provider has confirmed the issue is resolved. Our engineering team monitored the recovery, and we have confirmed that all services in the EU01 region have returned to normal operating parameters.
Data Integrity: For data ingestion, records that were rejected with a retryable error code during this window should be resubmitted. Records that have already been successfully ingested following the recovery require no further action.
Personalization: Real-time request success rates have returned to 100%.
We are continuing to monitor the stability of the region.
We apologize for any inconvenience this may have caused to your operations. For further questions, please contact our support team at [email protected].
resolved
The upstream provider has restored service.
US region: Treasure Insights unavailable
Started February 18, 2026 at 8:45 PM UTC · 1h 15m
OutageCritical incident
Affected components
Insights
investigating
We are investigating a possible problem currently affecting Treasure Insights. We will provide an update as soon as we know more.
identified
We have partially resolved the incident. We are continuously working to resolve the incident.
monitoring
The performance went back to normal, and no errors are observed now.
We will monitor at least for 15 minutes.
Starting at approximately 17:00 PST, we observed increased latency affecting all operations accessing table resources in Plazma.
At around 19:00 PST, we applied temporary tuning to the database managing table resources and confirmed recovery based on internal performance metrics.
We are currently continuing to monitor the system to ensure overall performance remains stable.
Visible impacts are:
- Streaming, Mobile, and JavaScript/Browser imports may delay
- Jobs execution may delay
- Table creation may error
- Console execution may delay
- Plazma Public API may return errors
- Treasure Workflow may delay and fail
- Workflow REST API may be unavailable
- Workflow operations in Console and CLI may become unavailable
- Presto JDBC/ODBC queries and CDP segmentation queries may fail
- Console Table preview update may delay
- The REST API to submit and cancel jobs may be unavailable
- Data Connector Integrations may become unavailable
- ADH (Ads Data Hub) and DCR (Data Clean Room) service may be unavailable
- ADL (Active Data Layer) service may be unavailable
We will send an additional update in 30 minutes
identified
We identified the same issue in the Tokyo region and applied database parameter tuning. We are currently assessing the impact of this change on system performance.
As a result of this temporary mitigation applied in both the US and Tokyo regions, the following impact is observed:
In Data Workbench, updates to table metadata for row counts and preview data are temporarily paused.
We will continue monitoring and provide further updates as we make progress.
identified
We are continuing to work on a fix for this issue.
identified
Performance in the Tokyo region improved at approximately 12:50 JST. The Tokyo region had been experiencing gradually increasing performance degradation since around 09:00 JST.
We are currently proceeding with the steps to resume updates of table metadata in Data Workbench.
Due to the temporary pause of metadata updates, application metrics that rely on this information, such as Parent Segment size, have also been temporarily not updated.
We will continue monitoring the system and share further updates as needed.
identified
In the US region, table metadata updates have been resumed. At this point, the system in the US region is expected to be fully recovered, and we are continuing to monitor it closely.
For the Tokyo region, metadata updates will be resumed after we complete confirmation of system performance. We will provide another update when this step begins.
monitoring
The system in the US region is operating normally with no observed issues.
In the Tokyo region, table metadata updates have also been resumed. At this point, the Tokyo region is expected to be fully recovered, and we will continue to monitor system performance.
Once stability is fully confirmed, we plan to close this status page.
resolved
Incident Impact Summary (Plazma / Table Resources)
We have confirmed that the system in the Tokyo region is operating normally with no remaining issues.
What happened
- All operations accessing Plazma table resources experienced increased latency.
When
- US region: 17:00 PST – 18:40 PST
- Tokyo region: 09:00 JST – 12:50 JST
Temporary impact during recovery
- Updates to table row counts and preview data were temporarily paused.
Recovery status
- US region: Table metadata updates resumed at 20:00 PST.
- Tokyo region: Table metadata updates resumed at 13:40 JST.
Both regions are now fully recovered.
Data integrity
- No data loss or data corruption occurred.
We will prepare and publish a post-mortem document to explain the circumstances under which this incident occurred one day after the recent system maintenance.
With this confirmation, the incident is now closed.
postmortem
## Post-Mortem Summary
This incident was related to a database maintenance performed on January 19, during which management of Plazma table resources was migrated from the existing RDBMS to a dedicated Aurora database. The purpose of this change was to better isolate table resource workloads and improve long-term scalability.
The migration itself completed successfully, and both the original database and the newly introduced Aurora database were operating normally immediately after the maintenance.
On the following day, when internal batch processes began running, the newly introduced database experienced unexpected load. Investigation determined that the internal batch system was configured to connect to the master database endpoint instead of a reader endpoint. As a result, internal batch processing competed with user-facing requests, causing increased latency when accessing table resources.
To mitigate the impact and enable rapid recovery, temporary database tuning was applied. This tuning was intended solely as a short-term measure to stabilize the system and has since been fully reverted. The permanent fix consisted of correcting the configuration of the internal batch system so that it accesses the appropriate database endpoint.
Following these actions, system performance recovered in both the US and Tokyo regions, and all services returned to normal operation. No data loss or data corruption occurred.
To prevent similar issues in the future, we are reviewing our configuration and deployment practices, including reducing configuration differences between staging and production environments and strengthening validation of database endpoints used by internal workloads.
[All Regions] REST API sometimes returns only partial job result
Started December 10, 2025 at 6:03 AM UTC · 0m
Pending
Affected components
REST APIREST APIREST APIREST APIREST API
resolved
An unexpected behavior occurred in the REST API "GET /v3/job/result/{job_id}" which caused it to return only a partial set of results. This issue stemmed from an update to the job result download mechanism implemented around 2025/12/09 01:00 (UTC).
The issue was resolved after a fix was applied around 2025/12/10 04:36(UTC).
The affected API is used through the following methods:
- Directly calling the REST API
- Using functions in libraries (e.g., pytd) that reference job results
- Toolbelt commands that refers to job results
- Setting store_last_results:true in the td>:/td_run> operator in Treasure Workflows
- Using the td_for_each>, td_wait> and td_wait_table>: operators in Treasure Workflows
Unfortunately, we are unable to identify which API requests were affected. Please compare the number of results retrieved by your job with the download results to see if you were affected.
We sincerely apologize for the inconvenience this has caused.
[All Region] Service Disruption on Treasure Insights
Started December 5, 2025 at 9:18 AM UTC · 12m
OutageCritical incident
Affected components
InsightsInsightsInsightsInsights
investigating
We are currently experiencing issues accessing Treasure Insights. Our investigation has identified that this is caused by an outage with a third-party service provider utilized by Treasure Insights.
We are monitoring the situation and working with the provider to resolve the issue. We will provide updates as soon as more information becomes available.
monitoring
The third-party service provider has implemented a fix. We are monitoring the results.
resolved
We confirmed the third-party service is now operational. The incident has been resolved.
[All Regions] Temporary Interruption in Premium AI Audit Logs
Started November 5, 2025 at 3:49 AM UTC · 4h 10m
Pending
Affected components
REST APIREST APIREST APIREST APIREST API
investigating
We have encountered a temporary issue with the availability of Premium Audit Logs for Artificial Intelligence services (AI Agent Foundry and Audience Agent) during the period of November 4th 23:19:26 UTC to November 5th 00:57:55 UTC.
While some data during this period may not have been captured, our engineering team is actively investigating the situation. We are taking all necessary steps to restore full service and ensure that all future logs are recorded without issue.
We apologize for any inconvenience this may have caused and will keep you updated on our progress.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
Services have been restored to normal operations, and all of our systems are reporting healthy once again.
If you observe any lingering issues, please connect with our support team, and we would be happy to help: https://docs.treasuredata.com/articles/#!pd/contacting-treasure-data-support
[All Regions] Profile API Errors and SSO Logins
Started October 20, 2025 at 8:53 AM UTC · 12h 39m
OutageMajor incident
investigating
We are aware of an issue affecting log writing and reading for the Profile API in all regions. This is due to an incident with an upstream service. Our team is monitoring the provider's status. We will provide an update shortly.
investigating
We are also investigating reports indicating that this incident might be causing login errors for customers using SSO (Single Sign-On) to log into our service. Customers attempting a new login via SSO may be unable to access the service.
We believe users who are already logged in (have an active session) and users logging in with methods other than SSO (e.g., username/password) are not affected by this login issue. We are working to confirm this.
investigating
Our upstream service provider is gradually recovering its services, and our services are gradually recovering too.
SSO login is working fine, however profile API 1.0 is still recovering its service with latency.
As the upstream service provider incident is still ongoing, you might still encounter unexpected errors or issues. We will closely monitoring our services and work with our upstream service provider.
identified
We are continuing to work with our service provider to track this issue. Customers may continue to observe higher latency than normal and may notice delays in processing. We are monitoring the situation closely and are adding additional capacity in anticipation of a recovery.
As this situation develops, we will keep you updated on any changes in expected behavior.
resolved
Delays should be improving for all customers, and services should be returning to normal. Given the similarity of the two issues, we'll continue tracking any concerns related to this interruption as part of the "[US Region] Issue Affecting Treasure Data Services" issue. For further updates, please track here: https://status.treasuredata.com/incidents/w7v5wtmkdllj
If you experience continued issues specifically relating to our Profiles API or SSO Logins, please contact our support team: https://docs.treasuredata.com/articles/#!pd/contacting-treasure-data-support
[US Region] Issue Affecting Treasure Data Services
We are currently investigating an issue originating from an upstream service provider used by our infrastructure. This may be affecting the normal operation of some Treasure Data services.
We are working to determine the scope of the impact. Observed symptoms may include, but are not limited to, the failure of various jobs (such as Trino, Hive, Result Export, Import, and Custom Script execution) to start.
Our team is actively monitoring the situation and working with the provider. We will provide further updates as more information becomes available.
investigating
Our upstream service provider is gradually recovering its services, and our services are gradually recovering too.
Most of our Treasure Data services are back to normal except for Custom Script, autoML and realtime services such as profile API 1.0, and realtime 2.0.
As the upstream service provider incident is still ongoing, you might still encounter unexpected errors or issues. We will closely monitoring our services and work with our upstream service provider.
investigating
We are continuing to monitor ongoing recovery efforts, and most of our systems remain operating. As the situation with our upstream provider develops we will react if we discover any new issues.
At this time, we are observing delays in our systems, including Profile API 1.0 and Realtime 2.0. Requests are being processed as they have been received, but users may see a slowdown in results, and may notice delays when updating configurations or journeys.
You may notice continued slower performance or unexpected errors. We are working with our service provider to address these issues and will provide an update with new information soon.
monitoring
We are seeing continued recovery across our systems, and are currently working to process any buffered or delayed messages. Customers may continue to observe delays in data processing and may notice slower-than-usual processing of configuration changes. We expect these delays to decrease as our systems catch up.
Our team is continuing to track recovery closely, and we'll inform you of any new developments.
monitoring
Service levels continue to improve as upstream services become more stable. Users may still see processing delays, though those should be decreasing as we process the backlog of messages from this event.
We are continuing to monitor our systems and deploy capacity as needed. We will update you with additional information as it develops.
resolved
Services have been restored to normal operations, and all of our systems are reporting healthy once again.
If you observe any lingering issues, please connect with our support team, and we would be happy to help: https://docs.treasuredata.com/articles/#!pd/contacting-treasure-data-support
[ US region ] Audience Studio – Intermittent Query Expired Errors When Retrieving Segment Counts
Started October 15, 2025 at 6:16 PM UTC · 1d 7h
OutageMajor incident
Affected components
Web Interface
investigating
We are currently investigating an issue where some customers are experiencing intermittent “query expired” errors when retrieving segment counts in Audience Studio. In many cases, queries may succeed after a retry, but this behavior can impact the ability to view or use segment results reliably.
Our engineering teams are actively investigating the root cause and working to mitigate the issue. We will provide updates here as soon as more information becomes available.
In the meantime, affected users may be able to work around the issue by retrying the query after a short interval.
monitoring
We've identified that an internal component between our query engine which extracts the number of records that match Segment rules and Audience Studio ran into a capacity issue during a specific time. We have increased the capacity and will continue to monitor the situation. If the issue does not reoccur by approximately 2025/10/17 01:00 (+00:00), we will consider this incident resolved and close this.
resolved
After a successful monitoring period, we can confirm that our fix is working and the issue is now resolved. Thank you for your patience.
We are currently experiencing an issue with email delivery for Treasure Data system emails, such as password resets.
The affected email deliveries are:
* Default email notification of Treasure Workflow
* Password reset
* New user invitation
* Password change
* Email address change
* User lock notifications
The following emails are NOT affected:
* Email delivery with the mail>: operator of Treasure Workflow
* Email delivery with Engage Studio
Our team is actively working to resolve this issue.
We sincerely apologize for any inconvenience or disruption this may cause.
identified
The issue has been identified and a fix is being implemented.
monitoring
The issue has been addressed with an applied fix. We are currently monitoring the system to ensure the services are fully restored and stable. Thank you for your patience while we addressed this.
resolved
The issue has been fully resolved. Thank you for your patience.
[All Regions] AI Agent Foundry May Provide Incomplete Responses
Started September 30, 2025 at 3:37 AM UTC · 3h 45m
OutageMajor incident
Affected components
Web InterfaceWeb InterfaceWeb InterfaceWeb InterfaceWeb Interface
investigating
We are currently investigating an issue where AI agents in the AI Agent Foundry may prematurely stop their response without providing a final conclusion, especially during their thinking process.
This behavior does not generate an error message. Our initial findings show it is more likely to occur in agents handling long or extensive conversations, though we have not yet identified a specific threshold.
Our engineering team is actively investigating the root cause. We will provide another update as soon as more information is available. We apologize for any inconvenience this may cause.
identified
We have an update on the root cause of the issue causing AI agents to provide incomplete responses.
Our investigation has found that when an agent uses its tools, the user's original prompt was being sent multiple times to the underlying AI model. This duplication caused the model to incorrectly process the tool results, leading it to return an empty response instead of a complete answer.
The hotfix we are preparing will correct this behavior by ensuring the user's prompt is sent only once for each task. This will allow the AI to properly process all tool results and generate a complete, final response.
We are working to deploy this fix as quickly as possible and will provide another update once it is ready. Thank you for your continued patience.
monitoring
The hotfix for the incomplete agent responses has been deployed. This should prevent the duplicate prompt issue during tool use. We are now monitoring the platform to ensure everything is working as expected. We will resolve this incident once we've confirmed stability.
resolved
After a successful 30-minute monitoring period, we can confirm that our fix is working and the issue is now resolved. All AI agents are functioning as expected. Thank you for your patience.
Profile API Data Delays - Tokyo Region
Started September 26, 2025 at 7:12 PM UTC · 4h 12m
Pending
Affected components
CDP Personalization - Lookup API
identified
We are currently experiencing delays in our Profile API replication pipeline in the Tokyo region. This issue began at approximately 13:00 UTC and is affecting all customers in this region.
Impact:
- Profile API responses may return stale data (up to several hours behind)
- Lookup functionality remains available but data may be inconsistent
Our engineering team is actively working on recovery. We estimate normal service will be restored within 8-10 hours.
We will continue to provide updates as we make progress toward resolution.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved.
[US Region] Partial Issue with Table API in Hive and Trino
Started September 24, 2025 at 6:53 AM UTC · 0m
Pending
Affected components
Presto Query EngineHadoop / Hive Query Engine
resolved
We experienced a partial issue on September 24th, between 01:55 and 02:40 UTC, where some users may have encountered incorrect permission errors when accessing their tables and databases in Hive and Trino.
The issue was caused by a high load on our table API, which temporarily struggled to retrieve correct permission data.
All underlying data remained secure and unaffected. The API load has returned to normal, and all services are operating as expected. If you had any jobs that failed with a permission error during this period, please try running them again.
We apologize for any confusion this may have caused.
Treasure Data Insights unavailable
Started August 18, 2025 at 9:03 PM UTC · 45m
IssuesMinor incident
Affected components
InsightsInsightsInsightsInsights
investigating
Treasure Data Insights is currently unavailable; we are investigating this issue.
resolved
The issue was resolved; it was linked to a vendor issue. The vendor has resolved the problem that was impacting TD Insights.
Treasure Data outage history and incident timeline | Uptimus