编码空间会话验证失败在 Generative AI 游戏场
- identified
Generative AI Playground中的代码空间会话由于验证错误而失败. 工程正致力于建立固定装置.
- identified
工程已实施一个固定装置,目前正在努力将其部署在生产集群.
- resolved
修复装置现在部署在普罗德的所有环境中,Generative AI Playground现在正在按预期工作.
自动翻译自官方事件更新。
52 Datarobot incidents · 2024年10月 — official updates, affected components, duration and resolution details.
Generative AI Playground中的代码空间会话由于验证错误而失败. 工程正致力于建立固定装置.
工程已实施一个固定装置,目前正在努力将其部署在生产集群.
修复装置现在部署在普罗德的所有环境中,Generative AI Playground现在正在按预期工作.
自动翻译自官方事件更新。
DataRobot工程组在创建DataRobot托管矢量数据库的同时观察到了问题. 目前,这正在影响所有DataRobot MTS环境,工程正在调查其根源.
工程组已应用了固定器,矢量数据库的创建现已投入使用. 该小组正在监测局势.
工程小组观察到,在DataRobot MTS的所有环境中,病媒数据库的创建都能够运行。 此事已决.
自动翻译自官方事件更新。
We are investigating issues in MTSaaS within EU region impacting prediction requests against Custom Models deployments.
The team root caused the issue causing HTTP 403 responses for Custom Model deployments in EU region. Mitigations were applied and the team is verifying the fix stabilized the service.
The issue is contained and the team has verified that predictions against Custom Models in all MTSaaS environments are healthy.
Batch jobs that are currently in the queue are completed successfully; however, their completion status is not being updated correctly. As a temporary workaround, we are manually marking these jobs as completed while we work on implementing a solution.
The affected jobs have been resolved and the issue is mitigated. Engineering is actively working on a permanent fix and testing is currently underway
The remediation script has been deployed and Engineering is actively monitoring the situation. Batch jobs may take longer than usual to show as completed until a permanent fix is rolled out in the next production deployment.
We are continuing to monitor for any further issues.
A fix has been deployed to production and the issue is now resolved. All systems are operating normally.
Feature drift statistics are currently experiencing a processing delay of approximately 1 hour. No data has been lost, and all metrics will reflect accurate values once processing catches up. Our team is actively working to resolve this.
Feature drift statistics processing has been restored to normal. All metrics are now up to date. No data was lost.
We are experiencing a service interruption with Custom Models functionality in US SAAS environment. Predictions to existing deployments are working fine, but users cannot create new custom models. Engineering is investigating the issue and will provide updates as we make further progress.
Engineering has applied a fix in the US SAAS environment which resolved the issue. At the time of issue, some users might have experienced issues with Custom Apps, Data upload and custom model creation. The issue is contained.
We are continuing to monitor for any further issues.
Custom Models, Custom Applications & Data upload services are back to Operational state in US SAAS Environment. Issue is Resolved.
We are currently experiencing intermittent service issues in US Production, which are primarily affecting the launch of new workloads for Notebooks, Custom models, and Custom Applications. This issue does not impact existing workloads. This disruption is strongly correlated with an ongoing AWS Availability Zone outage (https://health.aws.amazon.com/health/status), causing resource allocation failures. The team is actively monitoring the situation and tracking updates from AWS.
We are continuing to experience issues launching new workloads for Custom Models and Custom Applications in US Production. This is connected to an ongoing AWS outage. Our team is exploring multiple mitigation options.
Engineering resolved the underlying issue with workload scheduling and is monitoring the cluster.
Engineering confirmed the issue is resolved and all services are restored.
Processing actual messages on JP MTS is delayed due to autoscaling malfunction. Engineering scaled up the deployment to alleviate the issue. Root cause mitigation in progress
Engineering has applied the required infrastructure configuration changes. The service is operating normally and no further user impact is observed. Engineering will continue monitoring cluster health to ensure stability. The incident is now marked as Contained.
We're experiencing an elevated level of errors and are currently looking into the issue.
A fix has been implemented and we are monitoring the results.
Engineering has applied changes to mitigate the elevated error rates. Services are now operating normally. We are continuing to monitor the system while investigating the cause of the issue.
Engineering has implemented the required fixes to resolve the elevated error rates. Services are now operating normally, and no further user impact has been observed. The team will continue to monitor the system to ensure stability. The incident is now considered contained.
Our engineering team has found the the Quay outage currently happening is causing degraded performance across the DataRobot platform. Engineering is currently monitoring the situation.
We are continuing to work on a fix for this issue.
Quay.io functionality has been restored and DataRobot environments are fully stabilized.
We are experiencing performance degradation on Managed AI Cloud.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently investigating this issue.
A fix has been implemented and we are monitoring the results.
We are continuing to monitor for any further issues.
This incident has been resolved.
DataRobot is experiencing network issue related to Kubernetes in US Cluster. This will have impact on model deployment and predictions. Engineering is investigating the root cause.
Engineering has identified the root cause of the problem and a mitigation is put in place.
The mitigation implemented by Engineering has improved the network issue. The team is continuing to monitor the environment to ensure full recovery.
The mitigation implemented by Engineering has resolved the Kubernetes network issue, and the incident is now contained.
Our engineering team has found the the Quay outage currently happening is causing degraded performance across the DataRobot platform.
This incident is now resolved.
LLM blueprints deployments can not be created in JP MTS environment. Engineering is rolling back JP cluster to previous version to mitigate the issue.
Rollback of the JP cluster to the previous version is complete and the problem has been mitigated.
Agent application template is affected with the recent moderations library upgrade, fix is identified and mitigation is in progress.
New version of Agentic application template is released, the issue is resolved
We are observing issues on DataRobot US MTS environment. Users may experience degraded performance using APIs and data ingest services. The engineering team is currently investigating the root cause.
The incident has now been resolved. All services are now operational.
Some customers have reported issue connecting to DataRobot. Please do a hard refresh of your browser by clearing the cache and this should fix the problem. As always let us know if it continue to have issue connecting to DataRobot after clearing the cache.
We are continuing to monitor for any further issues.
This incident has been resolved.
Between 16:50 UTC and 17:09 UTC, one of external providers(DockerHub) had an outage that might have caused temporary delays in starting platform workloads. The issue has been resolved and normal operations have resumed. Engineering is continuing to monitor the system.
Between 09:04 UTC and 09:17 UTC, one of external providers(DockerHub) had an outage that might have caused temporary delays in starting platform workloads. The issue has been resolved and normal operations have resumed. Engineering is continuing to monitor the system.
The engineering team has confirmed recovery across all MTS and STS environments. No further updates are expected.