Vendeur Copilot connaît des taux d'erreur élevés en raison d'une panne en aval dans le centre de données Azure supportant ce cas d'utilisation. En outre, le support de chat Mindtickle peut éprouver des échecs dans la génération de réponses à vos requêtes.
identified
Le problème a été identifié et une solution est en cours de mise en oeuvre.
monitoring
un correctif a été mis en œuvre et nous suivons la situation.
resolved
L'incident a été résolu et tous les systèmes sont maintenant opérationnels.
Traduit automatiquement depuis la mise à jour officielle de l'incident.
Certains rapports d'analyse peuvent servir de données inexistantes.
Début 19 août 2026 à 17:16 UTC · 8h 50m
IssuesIncident mineur
Composants affectés
In-Platform Analytics
investigating
Nous enquêtons sur un problème où un sous-ensemble de clients peuvent voir des données statiques dans certains rapports d'analyse, en particulier: SSR, tableaux de bord personnalisés, rapports de gestionnaire et de relation, et rapports d'aperçu / adoption / performance. D'autres rapports et zones de produits ne sont pas touchés. Si votre analyse reflète les données actuelles, aucune action n'est nécessaire. Nous publierons bientôt une mise à jour.
resolved
Cette question est résolue. Tous les rapports touchés montrent maintenant les données actuelles. Merci de votre patience.
Traduit automatiquement depuis la mise à jour officielle de l'incident.
We are currently experiencing a partial service disruption on Mindtickle platform. Users may encounter errors, intermittent failures, or degraded performance across the platform. Our engineering team is actively investigating and mitigating the issue while closely monitoring recovery efforts.
identified
We have identified the cause of the elevated error rates and are deploying mitigation measures. We are monitoring the platform closely and working to restore normal service.
monitoring
Mitigation has been deployed and services are operating normally. We are continuing to monitor the platform closely to ensure stability.
resolved
The incident has been resolved and services are operating normally. We will continue to monitor the platform to ensure stability.
A detailed Root Cause Analysis will be shared shortly.
postmortem
# **Incident Summary**
On June 12, 2026, Mindtickle experienced a service disruption caused by an unexpected surge in processing activity within one of our backend workflows. The resulting increase in downstream traffic exceeded the capacity of a critical platform dependency, leading to service degradation and eventual unavailability.
The issue was mitigated by reducing processing throughput and allowing affected systems to recover.
**Impact Duration:** **3 hours and 15 minutes**.
## **Impact**
During the incident window, customers may have experienced:
* Delays in invitation-related workflows
* Elevated latency and intermittent errors across affected services
* Temporary degradation of platform functionality dependent on the impacted backend systems
## **Incident Timeline**
* **Jun 12, 03:00 PM**: A critical backend dependency becomes saturated, resulting in service degradation and elevated error rates.
* **Jun 12, 03:15 PM**: The issue is detected through automated monitoring and customer reports. Investigation begins.
* **Jun 12, 04:50 PM**: Mitigation measures are implemented to reduce processing throughput and stabilize affected systems.
* **Jun 12, 05:46 PM**: Platform functionality is restored fully
## **Root Cause**
A backend workflow generated a significantly higher volume of processing activity than historical norms. As the system worked through the resulting backlog, requests propagated through multiple downstream services, creating sustained load across the platform.
While each component behaved as designed, the combined volume exposed capacity and protection gaps within the processing chain. The resulting amplification effect overwhelmed a critical backend dependency, causing service degradation and eventual unavailability.
## **Resolution**
The team mitigated the incident by reducing the rate at which backlog processing occurred, allowing downstream systems to stabilize and recover.
Once traffic levels returned to normal operating ranges, affected services resumed normal operation and platform functionality was fully restored.
## **Corrective and Preventive Actions**
* Introduce rate limiting and additional safeguards for high-volume processing workflows to prevent a single workflow from generating excessive downstream load.
* Implement backpressure, circuit-breaking, and load-shedding mechanisms across critical processing paths to better protect downstream dependencies during traffic spikes.
* Review platform capacity limits, concurrency settings, and monitoring to improve resilience and provide earlier visibility into abnormal processing patterns.
Call AI recording disruption for select customers - June 2, 2026 (7:19 AM – 9:21 AM PDT)
Début 2 juin 2026 à 14:19 UTC · 0m
OutageIncident majeur
resolved
On June 2nd, between 7:19 AM PDT and 9:21 AM PDT, our third-party recording provider experienced a disruption that prevented the Call AI bot from joining meetings. A limited number of customers were impacted, and unfortunately, the affected recordings are non-recoverable. We sincerely apologize for the disruption and any impact to your workflows.
The issue is fully resolved, and recordings are functioning normally. If your account was affected, your account team will reach out with specific details.
Two-Way Roleplay Service Disruption
Début 21 mai 2026 à 15:29 UTC · 3h 0m
OutageIncident majeur
Composants affectés
Mission
investigating
Users may experience issues starting Two-Way Roleplay sessions. In addition admins may be unable to preview roleplays.
The team is actively investigating the issue and working toward resolution.
identified
We have identified the root cause as a service disruption at a third-party provider that powers parts of the Two-Way Roleplay workflow.
We are actively working with the vendor to restore the service and recover impacted functionality as quickly as possible.
identified
Recovery efforts are still ongoing, and we continue to work closely with our third-party partners to restore the impacted service.
Our teams remain actively engaged on the issue.
identified
We are seeing partial recovery in the impacted service, and some roleplay attempts are now completing successfully.
We continue to work with the third-party provider toward full restoration.
monitoring
The impacted systems are now operational again, and roleplay workflows are functioning successfully.
We are continuing to monitor the systems closely to ensure stability.
resolved
All impacted systems are now operating normally, and Two-Way Roleplay functionality has been fully restored.
We will continue to monitor the systems closely. The incident is now resolved.
'Featured Hubs' widget is not loading across the Mindtickle platform
Début 18 mai 2026 à 13:17 UTC · 39m
IssuesIncident mineur
Composants affectés
Asset HubMindtickle Platform
investigating
The 'Featured Hubs' widget, wherever used on the Mindtickle platform, is not loading. The Hubs and individual media/assets are working properly. We are currently investigating this.
All other capabilities of the Mindtickle platform are operating normally.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
This incident has been resolved. The 'Featured Hubs' widget is now loading as expected, and the Mindtickle platform is fully operational.
Call AI Service Outage
Début 11 mai 2026 à 14:25 UTC · 2h 10m
IssuesIncident mineur
Composants affectés
Call AI
investigating
We are currently investigating an issue impacting call recordings in Call AI. Users may experience failures while recording calls.
The team is actively working on identifying the root cause and restoring normal functionality.
Further updates will be shared shortly.
resolved
The issue has now been identified and resolved, and Call AI services are operating normally again.
We will share a detailed Root Cause Analysis (RCA) shortly.
Call AI bot unable to join calls
Début 24 mars 2026 à 15:00 UTC · 0m
IssuesIncident mineur
resolved
Between 07:50 AM and 11:25 AM PT on March 24, 2026, the Mindtickle platform experienced a service disruption affecting Call Recordings. Full service was restored by 11:25 AM PT. We are currently conducting a thorough investigation and will publish a detailed Root Cause Analysis (RCA) shortly.
postmortem
We experienced a service disruption impacting **Call AI bot functionality**. During this time, **select scheduled meetings were affected**, where Call AI bots were unable to join and record.
The issue has been fully resolved, and services have returned to normal operation.
**Impact**
* Call AI bots were unable to join **select scheduled meetings**
* Call recordings and transcriptions were not generated for impacted meetings
**Timeline**
* **07:50 AM PT:** Incident began; bots unable to join select meetings
* **07:52 AM PT:** Alerts triggered; issue detected and investigation initiated
* **09:00 AM PT:** Root cause identified within Call AI bot infrastructure
* **10:15 AM PT:** Mitigation measures applied; service recovery in progress
* **11:25 AM PT:** Service fully restored; bots resumed normal operation
**Root Cause**
This incident was caused by a **transient issue within the Call AI bot infrastructure**, which impacted service availability during the affected period.
**Next Steps**
* Exploring & implement enhanced failover mechanisms to improve system resilience
* Introducing more granular monitoring and alerting for faster detection
We apologize for the inconvenience caused and appreciate your patience.
Service Disruption: Analytics - Custom Reporting
Début 19 février 2026 à 14:33 UTC · 18m
IssuesIncident mineur
Composants affectés
In-Platform Analytics
investigating
we are currently investigating an issue affecting the Analytics Platform. While out of the box dashboards remain functional, users may be unable to open existing custom reports or create new ones.
Our engineering team is identifying the root cause, and we will provide an update as soon as more information is available. We apologize for the interruption to your workflow.
investigating
We are continuing to investigate this issue.
resolved
We have successfully resolved the issue affecting the Analytics Platform. All custom reports should now load correctly, and the ability to create new reports has been fully restored.
Our team has verified the fix, and we are monitoring the system to ensure continued stability.
Thank you for your patience while we worked to get this back online.
We are observing high error rates for both admins and end users for select customer instances. We are investigating the issue and will share an update shortly.
identified
The issue has been identified and a fix is being implemented.
monitoring
We have implemented a fix and are monitoring the progress. We are seeing the error rates coming down.
resolved
The incident has been resolved. All admins and end users should be able to access the platform without any concerns.
[Resolved] Intermittent Content Upload Failures Impacting Admins for a Subset of Customers
Début 6 février 2026 à 07:15 UTC · 0m
IssuesIncident mineur
resolved
[Resolved]
A subset of customers (admins only) experienced intermittent failures when admins attempted to upload content to the platform. End users were not impacted and were able to view existing content without issues. The issue has been resolved, and all admin content upload workflows are operating as expected.
Incident Timeline
Start Time: 05 Feb 2026, 11:15 PM PT
Resolved Time: 06 Feb 2026, 05:12 AM PT
We will share a detailed RCA once the post-incident review is complete.
Two-Way Roleplay Missions Inaccessible due to Cloudflare outage
Début 18 novembre 2025 à 14:45 UTC · 2h 32m
IssuesIncident mineur
Composants affectés
Mission
identified
The Two-Way Roleplay Missions feature is currently experiencing an outage and is inaccessible. This is a direct consequence of the ongoing, widespread Cloudflare network outage. We are closely monitoring Cloudflare's recovery efforts and expect the service to be restored as soon as their network disruption is resolved.
monitoring
Cloudflare has confirmed the deployment of a fix for the underlying network issue. We are now observing a significant decrease in error rates on our platform, with Two-Way Roleplay Missions starting to become accessible again. We will continue to monitor the platform closely as services stabilize, and will post a final resolution notice once all operations are confirmed to be fully back to normal.
monitoring
Two-Way Roleplay Missions are operational.
We will continue to monitor the platform closely to ensure stability before marking this incident as fully resolved.
resolved
All services are normal. A detailed Root Cause Analysis (RCA) will be published shortly.
Increased error rates for select customers on Mindtickle platform due to AWS Outage in US-East-1
Début 20 octobre 2025 à 16:30 UTC · 4h 25m
OutageIncident majeur
Composants affectés
Mindtickle Platform
identified
There has been a second surge of error rates and latencies reported due to an AWS outage in the US-East-1 region. The AWS team is actively working on this and we are monitoring the status closely.
Though the platform is up and some features are working, select customers and users may face intermittent failues in accessing the platform. We will keep updating the status page as soon as we receive updates from AWS.
We apologize for the inconvenience caused here.
AWS Outage Status: https://health.aws.amazon.com/health/status
identified
As per AWS, their internal subsystems are now showing early signs of recovering in a few Availability Zones (AZs) in the US-EAST-1 Region. They are now applying mitigations to the remaining AZs at which point we expect launch errors and network connectivity issues to subside. We will share an update once the AWS team has applied the mitigations and we have verified at our end.
monitoring
Post the recovery and fix from the AWS team, we are observing a significant improvement in the Mindtickle platform performance. Most Mindtickle workflows are now functioning normally. Our teams continue to monitor the situation closely and will provide an update once the issue is fully resolved.
AWS Outage Status: https://health.aws.amazon.com/health/status
resolved
The incident has been resolved, and the Mindtickle platform is now operating normally. We continue to closely monitor performance to ensure everything remains stable.
High Error Rates for Select Customers Due to AWS Outage in US-East-1
Début 20 octobre 2025 à 10:07 UTC · 55m
OutageIncident majeur
monitoring
Earlier today, some Mindtickle customers experienced high error rates and intermittent service disruptions due to an AWS outage in the US-East-1 region.
AWS has since deployed a fix, and we’re observing steady recovery across our systems. Error rates are declining, and most services are returning to normal.
Our teams continue to monitor the platform closely to ensure full stability. We’ll share a final update once all systems are fully restored.
Thank you for your patience and understanding.
resolved
The incident has been resolved, and the Mindtickle platform is operating normally.
Stale or Incorrect Data Displayed on Select Analytics Pages
Début 16 octobre 2025 à 03:54 UTC · 56m
OutageIncident majeur
Composants affectés
In-Platform AnalyticsReporting API
investigating
We’ve identified an issue where specific pages within the Analytics section are displaying stale or incorrect data for select customers. Our team is actively investigating the issue and will share further updates as soon as possible.
Impacted Areas:
1. OData Routes (Partial Routes) – Stale data visible
2. Analytics Platform Pages (Series, Modules, Learner & Learner Groups) – Incorrect data displayed
3. Custom Reports shared with customers via Email (in the last 12 hours) – Incorrect data displayed
monitoring
A fix has been implemented and we are running a sanity check at our end.
resolved
This incident has been resolved. We will share an RCA once the postmortem is complete.
Delay in Processing of Call Recordings
Début 14 octobre 2025 à 07:10 UTC · 0m
IssuesIncident mineur
resolved
Between October 14, 2025, 12:10 AM PT and October 15, 2025, 3:00 AM PT, select Call AI customers experienced a delay in the processing of call recordings.
Impact:
1. Call recordings generated during this period were temporarily not visible on the Call Listing page.
What Continued to Work:
1. All recordings captured before the incident window remained accessible on the Call Listing page.
2. For calls recorded during the incident, Recording Ready emails were successfully sent to end users, allowing access to the recordings via the email links.
We are conducting a post-incident analysis to identify the root cause. A detailed Root Cause Analysis (RCA) report will be shared once the investigation is complete.
High Error Rate in In-Platform Analytics
Début 24 septembre 2025 à 09:01 UTC · 1h 9m
IssuesIncident mineur
Composants affectés
In-Platform Analytics
investigating
We're currently experiencing an high error rate in our in-platform analytics and are actively investigating the root cause to restore full functionality.
identified
The issue has been identified and a fix is being implemented.
monitoring
A fix has been implemented and we are monitoring the results.
resolved
A fix has been applied, and the incident has been resolved. We are continuing to monitor this closely.
Mindtickle Admin and Learning Site Unavailable for Select Users
Début 20 septembre 2025 à 06:40 UTC · 0m
IssuesIncident mineur
resolved
Between 19 Sep 2025, 23:40 to 10 Sep 2025, 00:50 PT, the Mindtickle admin and learning site were unavailable for select users in our US Production region. The platform has recovered and is fully functional now. We will share an RCA post a detailed postportem.
Apologies for the inconvenience caused.
postmortem
## Incident Summary
On September 19, 2025, during a periodic platform upgrade, the Mindtickle platform experienced an outage lasting approximately 1 hour and 43 minutes.
The disruption was caused by a configuration issue on upgraded servers, which led to resource constraints. As a result, application services could not run as expected, causing downtime for a set of customers \(in the US region\).
Our engineering team identified the issue, corrected the configuration, and rotated the affected servers. Services were fully restored and stabilized thereafter.
* Start time: September 19, 2025, 11:06 PM PT
* End time: September 20, 2025, 00:49 AM PT
## Impact
* Workflows Impacted: All workflows on the platform
* Customers Impacted: A select set of customers \(in the US region\)
## Incident Timeline \(PT\)
* September 19, 2025, 11:06 PM: Users began experiencing downtime and errors
* September 19, 2025, 11:10 PM: Engineering team detected the issue and initiated an investigation
* September 19, 2025, 11:45 PM: Root cause identified \(server resource constraints\); remediation started
* September 20, 2025, 00:30 AM: Services began recovering as corrected configurations were applied
* September 20, 2025, 00:49 AM: All services confirmed healthy; incident resolved
## Root Cause
The outage was caused by misconfigured disk sizes in newly upgraded servers. This resulted in resource shortages that prevented application services from running.
This misconfiguration was not detected during pre-upgrade validation because upgrade scripts did not fully account for updated server requirements.
## Preventive Actions
To prevent recurrence, we are implementing the following measures:
1. Configuration Management: Standardize and validate server configurations across environments.
2. Upgrade Safeguards: Introduce a staggered approach with cooldown periods between server pool rotations.
3. Runbook Enhancements: Update documentation with environment-specific requirements and lessons learned.
4. Proactive Monitoring: Enhance alerts to detect early signs of resource constraints.
We sincerely apologize for the disruption this outage caused. We are committed to learning from this incident and strengthening our upgrade and validation processes to ensure greater reliability and resilience of the Mindtickle platform.
Security Update: Not Impacted by Salesloft Drift Data Breach
Début 12 septembre 2025 à 14:30 UTC · 0m
Pending
resolved
Mindtickle was not impacted by the Salesloft Drift data breach, and no customer data was compromised.
Mindtickle does not use or integrate with the Drift platform or any other Salesloft products. Our systems, including those that integrate with Salesforce, have not been compromised or pose a similar risk to the Salesloft Drift incident.
We are proactively reviewing our third-party vendors and sub-processors. To date, we have not received any notification of incidents or data breaches from any of these third parties. We'll continue to actively monitor for any potential impact and provide updates as appropriate.
For questions regarding this update, please contact [email protected].
High error rates observed on the Mindtickle platform
Début 29 août 2025 à 19:49 UTC · 0m
IssuesIncident mineur
resolved
On August 29, 2025, between 12:49 PM and 1:17 PM PT, the Mindtickle platform experienced high error rates. During this period, multiple users may have encountered issues logging into the platform or accessing programs.
The incident was resolved at 1:17 PM PT, and services have since been fully restored.
We are conducting a detailed analysis of the incident and will share a Root Cause Analysis (RCA) once it is complete. We apologize for any disruption this may have caused and thank you for your patience.
postmortem
# **Incident Summary**
On August 29, 2025, the Mindtickle platform experienced a temporary disruption where some users were unable to log in or access programs. The issue was identified and resolved within 28 minutes, restoring the platform to normal operation.
* Start time: August 29, 2025, 12:49 PM PT
* End time: August 29, 2025, 01:17 PM PT
# **Impact Area**
The following functionality was impacted during the incident:
* User logins
* Access to programs/assets \(assigned series, modules, and assets were impacted\)
# **Incident Timeline**
* **August 29, 2025, 12:49 PM PT:** Users began experiencing login and program access errors.
* **August 29, 2025, 12:55 PM PT:** The Engineering team detected elevated error rates and initiated an investigation.
* **August 29, 2025, 01:17 PM PT:** Corrective actions applied; services restored to a stable state.
# **Root Cause Analysis**
The disruption was caused by system resource exhaustion in one database cluster, which led to request timeouts and high error rates for affected services. Once identified, the engineering team stabilized services by resetting resource pools and prioritizing critical traffic.
# **Next Steps and Preventive Actions**
* **System Safeguards:** Implement circuit breakers to isolate and recover from failures faster.
* **Resiliency Improvements:** Maintain priority channels for critical operations to reduce customer impact in similar scenarios.