kustomer.help Default Domain – Sichtbarkeitsproblem – ALLE PRODs
Beginn 8. September 2026 um 13:23 UTC · 16m
IssuesGeringfügiger Vorfall
Betroffene Komponenten
Knowledge baseKnowledge baseCSATCSATCSATKnowledge base
investigating
Kustomer ist sich eines Ereignisses bewusst, das Links und Inhalte beeinflusst, die die standardmäßige kustomer.help-Domäne (nicht benutzerdefinierte Client-Domänen) verwenden - einschließlich Help Center-URLs und CSAT-Umfragen -, die zu Sichtbarkeitsproblemen führen können.
Unser Team arbeitet derzeit daran, die Ursache dieses Problems zu identifizieren, um eine Lösung zu implementieren. Bitte erwarten Sie innerhalb der nächsten 30 Minuten zusätzliche Updates, wenden Sie sich bitte an den Kustomer Support unter [email protected] für weitere Fragen oder Updates.
investigating
Wir werden dieses Problem weiter untersuchen.
resolved
Kustomer hat ein Ereignis behoben, das Links und Inhalte unter Verwendung der standardmäßigen kustomer.help-Domäne (nicht benutzerdefinierte Client-Domänen) betrifft, die Sichtbarkeitsprobleme verursacht haben.
Nach sorgfältiger Überwachung hat unser Team festgestellt, dass alle betroffenen Gebiete nun vollständig wiederhergestellt sind. Bitte wenden Sie sich an den Kustomer Support unter [email protected], wenn Sie zusätzliche Fragen oder Bedenken haben.
Automatisch aus der offiziellen Störungsmeldung übersetzt.
Kustomer ist bekannt, dass ein von einem unserer Drittanbieter (PubNub) gemeldetes Ereignis Echtzeit-Messaging für Orgs in Prod1 betrifft, das dazu führen kann, dass sich Chat-Nachrichten, Agentenbenachrichtigungen und Live-Updates verzögern oder nicht in Echtzeit auf der Plattform erscheinen.
Unser Team überwacht den Vorfall aktiv und arbeitet nach Möglichkeit mit dem Anbieter zusammen, um das Problem zu lösen.
Sie können den Status von PubNub auch hier verfolgen: https://status.pubnub.com/incidents/cz6q37z3s1d5.
Bitte erwarten Sie innerhalb der nächsten 3 Stunden weitere Updates und wenden Sie sich an den Kustomer-Support unter [email protected], wenn Sie zusätzliche Fragen oder Bedenken haben.
resolved
Das PubNub-bezogene Ereignis, von dem zuvor berichtet wurde und das Echtzeit-Messaging für Orgs in Prod1 betrifft, wurde behoben. PubNub hat bestätigt, dass das Problem am Ende gemildert wurde - keine Nachrichten wurden verpasst und während der Veranstaltung traten keine Kundenauswirkungen auf.
Alle betroffenen Gebiete bleiben voll funktionsfähig. Bitte wenden Sie sich an den Kustomer-Support unter [email protected], wenn Sie zusätzliche Fragen oder Bedenken haben.
Automatisch aus der offiziellen Störungsmeldung übersetzt.
Probleme mit Text Editor beim Tippen, Einfügen und Hinzufügen von Verknüpfungen PROD1, PROD2 und PROD4
Beginn 30. Juli 2026 um 17:43 UTC · 1h 31m
IssuesGeringfügiger Vorfall
Betroffene Komponenten
Web ClientWeb ClientWeb Client
investigating
Kustomer ist sich eines Ereignisses bewusst, das den Texteditor beeinflusst und Probleme beim Eingeben und Hinzufügen von Inhalten wie Shortcuts verursachen kann.
Unser Team arbeitet derzeit daran, die Ursache dieses Problems zu identifizieren, um eine Lösung zu implementieren. Bitte erwarten Sie innerhalb der nächsten 30 Minuten zusätzliche Updates, wenden Sie sich bitte an den Kustomer Support unter [email protected] für weitere Fragen oder Updates.
monitoring
Kustomer hat ein Update implementiert, um ein Ereignis mit Auswirkungen auf den Texteditor zu beheben, das Probleme beim Eingeben und Hinzufügen von Inhalten wie Shortcuts verursachen kann. Wenn Sie Ihren Browser aktualisieren, wird das Problem vollständig behoben.
Unser Team überwacht derzeit dieses Update, um sicherzustellen, dass das Problem vollständig gelöst ist. Bitte erwarten Sie innerhalb der nächsten 30 Minuten weitere Updates und wenden Sie sich an den Kustomer-Support unter [email protected], wenn Sie zusätzliche Fragen oder Bedenken haben.
resolved
Kustomer hat ein Ereignis mit Auswirkungen auf den Texteditor behoben, das Probleme beim Eingeben und Hinzufügen von Inhalten wie Shortcuts verursachen kann. Um dieses Problem zu beheben, hat unser Team auf eine frühere Version zurückgegriffen.
Nach sorgfältiger Überwachung hat unser Team festgestellt, dass alle betroffenen Gebiete nun vollständig wiederhergestellt sind. Bitte wenden Sie sich an den Kustomer-Support unter [email protected], wenn Sie zusätzliche Fragen oder Bedenken haben.
postmortem
# **Summary**
On July 30, 2026, the draft text editor on Kustomer’s Timeline product experienced degraded functionality in certain user workflows. Impact included sporadic cursor behavior, text erasure, and issues with copying and pasting text and shortcuts.
# **Root Cause**
This issue was introduced by a change to the editor that was not fully caught before release. Due to misalignment in our QA process our pre-release validation did not adequately cover the real-world editing patterns affected by this change. Furthermore, this misalignment contributed to a delay in Kustomer’s understanding of the incident's resolution status.
# **Timeline**
## **Jul 30, 2026**
**11:12 AM ET:** The editor change was deployed
**12:27 PM ET:** Kustomer Technical Support escalated the incident to Kustomer’s OnCall process, notifying Engineering immediately.
**12:34 PM ET:** Kustomer Engineering identified the issue and completed a rollback
**12:51 PM ET:** Internal testing incorrectly confirms that the issue has been fully resolved, due to some customers reporting the issue no longer presented itself
**1:29 PM ET:** Continued reports of issue are received from customers who did not receive the full resolution rollout
**1:34 PM ET:** Discovery made that the rollback did not fully deploy to resolve the issue. Kustomer Engineering makes additional change to ensure resolution
**2:01 PM ET:** Corrective change deployed, and fix is confirmed
# **Lessons/Improvements**
Kustomer Engineering maintains a robust CI/CD process, and a multi-environment release process designed to prevent issues of this nature. As part of our continual investment in these areas, we have identified the following action items:
* Ensure that our internal pre-production environments align with our customer-facing experience, including but not limited to the editor experience, so that issues of this nature are caught earlier in development
* Complete an audit of our automated CI/CD process and test coverage of major features, removing any gaps that are identified
Automatisch aus der offiziellen Störungsmeldung übersetzt.
Kustomer hat ein Ereignis identifiziert, das sich auf Plattformereignisse in PROD1-Orgs auswirkt und Verzögerungen bei ereignisbasierten Daten verursachen kann.
Unser Team arbeitet derzeit an der Umsetzung einer Resolution. Bitte erwarten Sie innerhalb der nächsten 30 Minuten zusätzliche Updates und wenden Sie sich an den Kustomer Support unter [email protected] für weitere Fragen oder Updates.
identified
Kustomer hat ein Ereignis identifiziert, das sich auf Plattformereignisse in PROD1-Orgs auswirkt und Verzögerungen bei ereignisbasierten Daten verursachen kann.
Unser Team arbeitet derzeit aktiv an der Umsetzung einer Resolution. Bitte erwarten Sie innerhalb der nächsten 30 Minuten zusätzliche Updates und wenden Sie sich an den Kustomer Support unter [email protected] für weitere Fragen oder Updates.
identified
Kustomer arbeitet weiterhin an dem Problem, das Plattformereignisse in PROD1-Orgs betrifft, die zu Verzögerungen bei ereignisbasierten Daten führen können.
Unser Team arbeitet aktiv an der Umsetzung einer Resolution. Bitte erwarten Sie innerhalb der nächsten 30 Minuten zusätzliche Updates und wenden Sie sich an den Kustomer Support unter [email protected] für weitere Fragen oder Updates.
monitoring
Kustomer hat ein Update implementiert, um ein Ereignis zu adressieren, das sich auf Plattformereignisse in PROD1-Orgs auswirkt und Verzögerungen bei ereignisbasierten Daten verursacht hat.
Unser Team überwacht derzeit dieses Update, um sicherzustellen, dass das Problem vollständig gelöst ist.
Bitte erwarten Sie innerhalb der nächsten 30 Minuten weitere Updates und wenden Sie sich an den Kustomer Support unter [email protected], wenn Sie zusätzliche Fragen oder Bedenken haben.
resolved
Kustomer hat ein Ereignis behoben, das sich auf Plattformereignisse in PROD1-Orgs auswirkt und Verzögerungen bei ereignisbasierten Daten verursacht hat.
Nach sorgfältiger Überwachung hat unser Team festgestellt, dass alle betroffenen Gebiete nun vollständig wiederhergestellt sind. Bitte wenden Sie sich an den Kustomer-Support unter [email protected], wenn Sie zusätzliche Fragen oder Bedenken haben.
postmortem
## Zusammenfassung
Am 25. Juli 2026 erlebten einige Kunden eine erhöhte Latenz der Plattform, die sich auf Messaging, Konversationsupdates, Routing, Sprach- und andere Kundenservice-Workflows auswirkte. Kunden haben möglicherweise gesehen, dass Nachrichten in einem Sendezustand bleiben, sich verzögernde eingehende oder ausgehende Nachrichten, langsamere Seiten- und API-Antworten, verzögertes Routing oder Zuweisung und intermittierende Sprachanrufverzögerungen.
Der Vorfall wurde durch einen ungewöhnlich großen Ausbruch von Hintergrunddatenverarbeitungsaktivitäten verursacht. Dies führte zu einem plötzlichen Anstieg der Nachfrage nach Shared-Plattform-Infrastruktur. Die automatische Skalierung der zusätzlichen Kapazität konnte den Burst jedoch nicht schnell genug absorbieren, um eine Verschlechterung in abhängigen Workflows zu verhindern.
Wir stellten den Service wieder her, indem wir die verfügbare Kapazität erhöhten, die Rate der einleitenden Arbeitsbelastung reduzierten und verspätete Arbeiten sorgfältig verarbeiteten, während wir den Zustand der Plattform überwachten.
## Auswirkungen
* **Kundeneffekt:** Intermittierende Latenz der Plattform; verzögertes eingehendes und ausgehendes Messaging; langsamere Konversations- und API-Updates; verzögertes Routing oder Zuweisung; und intermittierende Sprachanrufverzögerungen
* **Dauer:** Kundensichtbare Degradation begann um ca. 5:22 Uhr ET. Der Service erholte sich zunächst um etwa 7:31 Uhr ET, aber die Degradation trat später wieder auf. Die Leistung der breiten Plattform wurde wiederhergestellt und der Vorfall wurde um 10:30 Uhr ET behoben.
***Anwendungsbereich:** Der Vorfall betraf eine Teilmenge von Kunden innerhalb einer Produktionsumgebung. Wir haben keine entsprechende kundenorientierte Verschlechterung in anderen Produktionsumgebungen festgestellt.
## Timeline
* ** Ungefähr 5:22 Uhr ET:** Verarbeitungslatenz und Nachrichten-Backlogs begannen zuzunehmen.
* **6:51 PM ET:** Das Incident Response Team begann mit der koordinierten Untersuchung und Wiederherstellung.
* ** Ungefähr 7:00–7:23 Uhr ET:** Wir identifizierten betroffene Verarbeitungspfade und begannen, verspätete Arbeiten sorgfältig zu kontrollierten Raten wiederherzustellen.
* **7:31 PM ET:** Zusätzliche Kapazitäten waren online gegangen, die Leistung des Kunden hatte sich wesentlich verbessert, und der Vorfall wurde zunächst behoben, während die Überwachung fortgesetzt wurde.
* **8:28 PM ET:** Der Vorfall wurde nach erneuten Berichten über intermittierende Verschlechterung wieder eröffnet.
* ** Ungefähr 8:38 Uhr ET:** Die Untersuchung bestätigte, dass Messaging-, Routing-, Kanal- und Sprachsymptome dieselbe zugrunde liegende Plattformabhängigkeit aufwiesen.
* ** Ungefähr 9:52 Uhr ET:** Wir haben die Rate der initiierenden Hintergrundarbeitslast reduziert, um den Kundenverkehr zu schützen.
* **10:08 Uhr ET:** Die Überwachung bestätigte, dass der Arbeitsbelastungsdruck erheblich gesunken war und sich die Plattformgesundheit weiter verbesserte.
* **10:30 Uhr ET:** Nachdem die fortgesetzte Überwachung die Genesung bestätigt hatte, wurde der Vorfall behoben.
***Nach Auflösung:** Die Erholung der verbleibenden verspäteten Arbeiten wurde unter Kontrolle fortgesetzt.
## Wurzelursache
Ein Ausbruch von umfangreichen Hintergrunddatenverarbeitungsaktivitäten führte zu mehr nachgelagerten Arbeiten, als die gemeinsam genutzte Plattforminfrastruktur über einen kurzen Zeitraum sicher aufnehmen könnte.
Die automatische Skalierung reagierte und fügte Kapazität hinzu, aber die Kapazität wurde schrittweise hinzugefügt und kam nicht schnell genug für die Größe und Geschwindigkeit des Bursts online. Während die Plattform aufholte, wirkten sich erhöhte Latenzzeiten und Timeouts auf kundenorientierte Workflows aus, die von der gleichen Infrastruktur abhängig waren.
## Resolution
Wir haben die Plattformleistung wiederhergestellt durch:
* Erhöhung der verfügbaren Kapazität und Verbesserung der Rate, mit der zusätzliche Kapazität hinzugefügt werden könnte.
* Reduzierung der Rate der initiierenden Hintergrundarbeitslast.
* Wiederherstellung verspäteter Arbeit zu kontrollierten Raten, um eine weitere Verkehrsspitze zu vermeiden.
* Die kontinuierliche Überwachung nach der kundenorientierten Leistung kehrte auf das erwartete Niveau zurück.
## Vorbeugende Maßnahmen
Wir ergreifen die folgenden Maßnahmen, um die Wahrscheinlichkeit und die Auswirkungen eines erneuten Auftretens zu reduzieren:
* Fügen Sie stärkere Grenzwerte und Gegendruckkontrollen für hochvolumige Hintergrundoperationen hinzu.
* Verbessern Sie die Geschwindigkeit, mit der sich die gemeinsame Infrastruktur während des plötzlichen Datenverkehrs skaliert.
* Verbesserung der Isolation zwischen Hintergrundverarbeitung und latenzsensitiven Kundenworkflows.
* Erweitern Sie die Alarmierung auf schnelles Backlog-Wachstum, Ressourcendruck und ungewöhnliche Workload-Muster.
* Verstärkte kontrollierte Wiederherstellungsverfahren für verspätete Arbeiten.
* Erweitern Sie Burst-Load- und Fehlerwiederherstellungstests für Shared Platform Services.
## Aktueller Status
Die Plattformleistung kehrte auf das erwartete Niveau zurück, und die initiierende Arbeitslast blieb kontrolliert. Wir überwachten weiterhin die Gesundheit des Dienstes und die Wiederherstellung der verspäteten Arbeit nach der Auflösung.
Automatisch aus der offiziellen Störungsmeldung übersetzt.
MessageBird - unable to send outbound replies PROD 1
Beginn 25. Juni 2026 um 22:24 UTC · 1h 7m
IssuesGeringfügiger Vorfall
Betroffene Komponenten
Channel - Chat
investigating
Kustomer is aware of an event affecting PROD 1 that may affect the ability to send outbound messages for MessageBird.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support at [email protected] for any further questions or updates.
identified
Kustomer has identified the cause affecting PROD 1 that is impacting the ability to send outbound messages through MessageBird.
Please expect another update within the next 30 minutes, as we reach a resolution. If you have any additional questions or concerns, please reply to this conversation or contact Kustomer Support at [email protected].
Thank you for your patience while we work to resolve this issue.
identified
Kustomer has identified and implemented a fix for the issue affecting PROD 1 that impacted the ability to send outbound messages through MessageBird.
Our team is actively monitoring the platform to ensure the fix remains effective and that service has fully recovered. If you continue to experience issues sending outbound MessageBird messages, please contact Kustomer Support at [email protected].
We will provide another update once monitoring is complete or if there are any significant developments.
Thank you for your patience.
monitoring
Kustomer has identified and implemented a fix for the issue affecting PROD 1 that impacted the ability to send outbound messages through MessageBird.
Our team is actively monitoring the platform to ensure the fix remains effective and that service has fully recovered. If you continue to experience issues sending outbound MessageBird messages, please contact Kustomer Support at [email protected].
We will provide another update once monitoring is complete or if there are any significant developments.
Thank you for your patience.
resolved
Kustomer has resolved the issue affecting PROD 1 that impacted the ability to send outbound messages through MessageBird.
Our team has verified that the fix has been successfully deployed and service has been restored. If you continue to experience issues sending outbound MessageBird messages, please contact Kustomer Support at [email protected]
We apologize for the disruption and appreciate your patience while we worked to resolve the issue.
postmortem
## **Summary**
On June 25, 2026, some customers were unable to send outbound WhatsApp replies from within Kustomer. The issue affected reply sending for a subset of WhatsApp channel configurations, which disrupted agent workflows and prevented some automated outbound messages from being sent through the same path.
The issue was identified and resolved the same day. Service was fully restored, and the platform is operating normally.
## **Root cause**
A change released earlier that day introduced stricter validation in the outbound WhatsApp reply flow. For a subset of supported channel configurations, valid sender values were incorrectly rejected before messages were sent. This caused reply attempts in those configurations to fail.
The issue was limited to specific WhatsApp channel setups and did not affect all WhatsApp traffic equally.
## **Timeline**
* **June 25, 2026, approximately 22:12 UTC** — Reports began coming in that outbound WhatsApp replies were failing for some customers.
* **Shortly after detection** — Investigation confirmed the issue was tied to a recently released validation change in the outbound reply flow.
* **June 26, 2026, approximately 01:07 UTC** — A fix was deployed and reply sending was restored.
## **Lessons and improvements**
* Validation changes for messaging flows now require broader test coverage across supported channel configuration variants before release.
* Additional safeguards are being added to reduce the risk of valid outbound requests being rejected.
* Monitoring and regression checks around outbound messaging paths are being strengthened to detect similar issues more quickly.
Kustomer is aware of an ongoing WhatsApp issue that may cause outbound messages to not be delivered to end users.
While the issue appears to originate with WhatsApp, our team is actively monitoring the situation and working closely to assess the impact on the platform. We will continue to provide updates as more information becomes available.
Please expect an additional update within the next 30 minutes. If you have any questions, please contact Kustomer Support at [email protected].
identified
Kustomer has identified an ongoing issue with WhatsApp that may cause outbound messages to not be delivered to end users.
Our team is closely monitoring the situation and actively working to assess impact. Clients can refer to https://metastatus.com/whatsapp-business-api for the latest WhatsApp service status.
Please expect further updates within the next 30 minutes, and reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
identified
Kustomer continues to monitor the ongoing issue on Whatsapp that may cause outbound Whatsapp messages to not be delivered to end users.
Clients can also refer to https://metastatus.com/whatsapp-business-api for the latest WhatsApp service status.
Please expect further updates within the next 30 minutes, and reach out to Kustomer Support at [email protected] if you have additional questions or concerns
monitoring
Kustomer is aware of an event reported by one of our third party vendors affecting ALL PRODS that may cause Issues with WhatsApp messages not being delivered to end users.
Our team is monitoring the incident, and working with the vendor where possible to resolve the issue. We are working on redriving previously unsent messages to ensure all messages are sent successfully. Please expect further updates within the next 3 hours, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
WhatsApp has confirmed that all services are fully recovered. After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On June 12, 2026, customers using WhatsApp through Kustomer experienced a service event that prevented some outbound messages from being delivered to end users. During the event, messages could appear as sent in Kustomer before final delivery was confirmed by WhatsApp. Service recovered the same day, and affected messages from the incident window were successfully reprocessed.
## Root cause
This event was caused by a disruption in WhatsApp Business Platform services operated by Meta. While that disruption was active, Kustomer was unable to complete delivery for some outbound WhatsApp messages and queued affected messages for retry after the upstream service recovered.
## Timeline
* **10:14 AM EDT** — Customer reports led to investigation of outbound messaging behavior.
* **10:34 AM EDT** — Meta reported high disruptions affecting WhatsApp Business Platform services.
* **1:45 PM EDT** — Reprocessing of queued messages began as upstream recovery progressed.
* **2:32 PM EDT** — Meta reported recovery for the remaining affected WhatsApp services.
* **4:07 PM EDT** — Active message sending was healthy again.
* **4:34 PM EDT** — Messages from the incident window had been successfully reprocessed.
## Lessons and improvements
* We are improving monitoring and alerting for queued WhatsApp messages so degraded delivery is identified earlier.
* We are documenting a safer reprocessing procedure for queued WhatsApp traffic to reduce the chance of retry-related rate limiting during recovery.
* We are reviewing backlog-handling procedures to make recovery more predictable when an upstream provider disruption occurs.
At this time, system health for WhatsApp message delivery through Kustomer is stable.
Kustomer is aware of an event affecting Chat and other channels that may cause delayed message delivery and slower conversation load times.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes and please reach out to Kustomer Support at [email protected] for any further questions or updates.
monitoring
Kustomer has identified the cause of the event affecting Chat and other channels that resulted in delayed message delivery and slower conversation load times. The situation has stabilized and our team is continuing to monitor to ensure full resolution. Please reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Chat and other channels that caused delayed message delivery and slower conversation load times.
After careful monitoring, our team has determined that all affected areas are now fully restored.
Please reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
postmortem
**Summary**
On June 9, 2026, customers experienced delays in chat message delivery and related real-time updates for a portion of the morning. The issue was caused by a third-party service event affecting the infrastructure used to propagate real-time messaging updates. Service recovered the same morning, and current system health is stable.
**Root cause**
A third-party real-time messaging provider experienced elevated latency and intermittent publish failures in one of its regions. That disruption delayed delivery acknowledgements and other real-time updates in Kustomer. Messages continued to be stored successfully, but some updates were not reflected immediately in the client until the external service recovered or the client refreshed.
**Timeline**
* **9:05 AM EDT** — Customer impact began, including delayed chat message delivery and stale real-time updates.
* **9:45 AM EDT** — The third-party provider reported active latency affecting its service.
* **9:48 AM EDT** — The provider reported mitigation and recovery.
* **Shortly after recovery** — Real-time behavior returned to normal and monitoring confirmed stability.
**Lesson/improvements**
* We added additional monitoring tied to the third-party provider's public service health signals so similar service events can be identified faster.
* We are improving alerting around real-time messaging errors to reduce time to diagnosis.
* We are continuing to review operational signals to better distinguish external service issues from internal platform issues.
Chats, Emails, Routing latency PROD 1
Beginn 2. Juni 2026 um 14:47 UTC · 1h 5m
IssuesGeringfügiger Vorfall
Betroffene Komponenten
Channel - ChatChannel - Email
investigating
Kustomer is aware of an event affecting chats, emails, and routing that may cause latency with sending, receiving, and routing messages.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support at [email protected] for any further questions or updates.
identified
Kustomer has identified an event affecting chats, emails, and routing that may cause latency with sending, receiving, and routing messages.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting chat, email, and routing that caused latency in sending and delivering messages.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Prod 1 that caused chat, email, and routing latency. To resolve this issue, our team reverted an underlying infrastructure change as a precaution, which helped stabilize the service.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
**Summary**
On June 2, 2026, customers in our prod1 environment experienced elevated latency affecting chat creation, email sending, and conversation routing. The event caused delayed processing for a subset of customer interactions during the incident window. Service performance recovered the same day after mitigation steps were applied, and the environment is currently operating normally.
**Impact**
* Scope: prod1 only
* Customer effect: delayed chat creation, delayed email sends, and slower conversation routing
* Duration: customer-visible impact began in the morning ET on June 2 and materially improved after mitigation later that afternoon
* Other environments: prod2 and prod4 were not impacted by the customer-facing degradation
**Timeline**
* Early June 2: We detected elevated processing latency in services supporting chat, email, and routing in prod1.
* Midday to afternoon ET: Customer reports confirmed delayed system behavior in prod1.
* Afternoon ET: We increased service capacity and adjusted resource limits to reduce backlog and restore throughput.
* Later that day: We completed an infrastructure rollback on affected worker hosts and confirmed stable recovery.
* June 3: Monitoring confirmed the environment remained healthy.
**Root cause**
The incident was caused by infrastructure-level resource exhaustion on a set of worker hosts in prod1. That reduced the availability of a metadata-dependent service and created message backlog, which in turn increased latency for customer-facing workflows such as chat creation, email sending, and routing. The issue was isolated to prod1.
**Resolution**
We restored service by increasing available capacity, raising resource limits for the affected service, and rolling impacted worker infrastructure back to a stable configuration. After those changes were applied, backlog cleared and latency returned to expected levels. Current system health is stable.
**Preventative actions**
* Strengthen host-level capacity and disk safeguards before future infrastructure rollouts
* Add earlier alerting for infrastructure resource pressure to reduce time to detection
* Expand validation for production-scale logging and resource usage prior to promotion
* Continue tuning service capacity thresholds and recovery procedures for metadata-dependent workloads
WhatsApp messages failing Prods 1 + 2
Beginn 26. Mai 2026 um 22:14 UTC · 2h 7m
IssuesGeringfügiger Vorfall
Betroffene Komponenten
Channel - ChatChannel - Chat
monitoring
Kustomer has implemented an update to address an event affecting WhatsApp that caused messaging failures.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
This incident has been resolved. Our teams have confirmed recovery and platform stability following mitigation efforts.
We will be reaching out independently to affected orgs to provide additional context and follow-up as needed. Thank you for your patience while we worked through this issue.
postmortem
## Summary
On May 26, 2026, some customers experienced failures when sending WhatsApp messages from existing conversations. The issue affected outbound message delivery for a subset of WhatsApp configurations. Service was restored after the responsible change was rolled back, and the issue is now resolved.
## Impact
During the incident window, outbound WhatsApp messages failed for some existing conversations. Newly created conversations were less consistently affected, and impact varied by account configuration.
Customers may have seen message-send failures or provider errors while attempting to send WhatsApp messages.
## Timeline
All times EDT.
* **4:05 PM:** The incident window is believed to have begun after a recent WhatsApp-related change.
* **5:54 PM:** Active investigation began.
* **5:56 PM:** The team identified a likely connection to recent WhatsApp sender-selection behavior and started a rollback.
* **6:05 PM:** Rollback completed.
* **6:17 PM – 7:47 PM:** The team validated recovery across affected examples and narrowed the underlying cause.
* **8:21 PM:** The incident was declared resolved.
## Root cause
The incident was caused by an issue in WhatsApp phone-number lookup and sender selection. A recent change expanded support for multiple valid phone-number formats in certain countries. In some cases, that allowed the system to resolve the wrong sender configuration for an existing conversation.
When that happened, outbound sends could target an invalid or unavailable sender identity, causing WhatsApp message delivery to fail.
## Resolution
We mitigated the incident by rolling back the relevant WhatsApp change. After the rollback, we verified recovery against affected examples and confirmed that the incident was resolved.
## Preventative actions
* Move WhatsApp sender resolution toward more stable identifier-based matching rather than relying on ambiguous phone-number formatting.
* Expand regression coverage for country-specific phone-number formatting edge cases.
* Improve monitoring and alerting for message delivery failures and queue growth.
* Improve testing workflows for WhatsApp changes before production rollout.
Kustomer is aware of an event affecting Chats that may cause messages to fail and assistants to not follow the configured flow.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support for any further questions or updates.
monitoring
Kustomer has implemented an update to address an event affecting Chats on Prod 1 that caused messages to fail and assistants to not follow the configured flow.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Chat on Prod 1 that caused messages to fail and assistants to not follow the configured flow.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support if you have additional questions or concerns.
postmortem
## **Summary**
Between May 13 and May 14, 2026, Kustomer Chat experienced a service event that affected chat availability and performance for a subset of customers. During this period, some customers may have seen intermittent failures, elevated error rates, or degraded behavior in chat-related settings and runtime flows.
The event was mitigated through service rollback, additional capacity, and targeted configuration changes. Service health returned to normal after these actions were completed.
## **Root cause**
The event was caused by a combination of reduced service capacity during a deployment rollback and higher-than-expected traffic through chat-related request paths. Under those conditions, the affected chat service became unstable and restarted repeatedly, which reduced available capacity further and increased customer-facing errors.
Our investigation also identified specific high-volume request patterns that increased memory pressure during the event. We addressed those patterns with caching, traffic protections, and capacity changes.
## **Timeline**
* May 13, 2026, early afternoon ET: We detected elevated instability in the chat service and began incident response.
* May 13, 2026, afternoon ET: We rolled back affected changes, increased infrastructure capacity, and stabilized dependent services.
* May 13, 2026, evening ET: We continued monitoring after initial mitigation and investigated recurring memory pressure.
* May 14, 2026: We deployed additional mitigations, including higher minimum service capacity and request-path protections.
* Following the mitigation deployments, service health returned to normal and remained stable.
## **Lessons and improvements**
We completed several improvements to reduce the likelihood of recurrence:
* Increased minimum service capacity to provide more headroom during deployments and recovery.
* Added caching for high-volume chat settings requests.
* Hardened URL processing behavior with stricter filtering, failure caching, and concurrency limits.
* Added improved memory telemetry to speed up detection and diagnosis of similar issues.
* Continued follow-up work on deployment recovery procedures and service safeguards for cross-service rollbacks.
## **Current status**
The mitigations above have been deployed, and the affected chat service is operating normally.
Kustomer is aware of an event affecting Gmail Connectivity and Send & Receive that may cause emails to not route into your Kustomer Platform and impact your ability to send messages on conversations.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat for any further questions or updates.
monitoring
Kustomer has released a fix for the event affecting Gmail Connectivity and Send & Receive functionality that was causing emails to not forward as expected into your Kustomer Platform and impacted sending functionality on conversations.
Our team is currently monitoring the released fix to ensure the issue is resolved. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat for any further questions or updates.
resolved
Kustomer has resolved an event affecting Gmail Connectivity and Send & Receive functionality in Prod 1 and Prod 2.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On May 1, 2026, customers using the Gmail integration experienced a service event that caused Gmail connections to disappear and the email channel to stop functioning.
The issue affected customers across different regions, more specifically:
* EU-based clients, from 2:40 AM ET.
* US-based clients, from 5 AM ET.
The issue was fully resolved by 9:20 AM ET. No customer data was lost during this event. Messages that were delayed during the incident were fully recovered once service was restored. Current system health is stable.
## Root cause
A configuration issue in the infrastructure used by the Gmail integration prevented replacement service tasks from starting correctly during routine task rotation. As running capacity declined over several hours, the Gmail integration became unavailable.
## Timeline
| Time \(ET\) | Event |
| --- | --- |
| May 1, ~12:30 AM | Service task replacement failures began in production. |
| May 1, ~2:40 AM | Engineering was alerted after service capacity dropped to zero in one production environment. |
| May 1, 2:40–9:20 AM | Teams investigated the failure, identified the configuration problem, prepared a fix, and deployed it. |
| May 1, 9:20 AM | Fix deployed and service restored. |
## Lessons and improvements
* We are auditing related infrastructure configurations to identify and correct similar patterns in other services.
* We are standardizing how service roles are managed so this class of configuration issue is less likely to recur.
* We are improving alerting so teams are notified earlier when running service capacity drops below expected levels, before a full service interruption occurs.
* We are adding additional checks to catch configuration drift and invalid service role references earlier in the deployment lifecycle.
[DRAFTS] Internal API Errors (Prod1)
Beginn 24. April 2026 um 20:03 UTC · 1h 45m
IssuesGeringfügiger Vorfall
Betroffene Komponenten
Channel - ChatChannel - Email
monitoring
Kustomer has identified an event affecting Prod1 that may cause internal API errors and latency.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting Prod1 that caused internal API errors and latency.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Prod1 that caused API errors and latency. To resolve this issue, our team has released an update.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On April 24, 2026, customers in our prod1 environment experienced a service event that caused elevated errors and latency in messaging-related workflows. The broad cross-customer impact was limited to approximately 41 minutes, from 3:26 PM ET to 4:07 PM ET. During that window, some customers saw failed or delayed messaging operations.
The immediate platform impact was resolved the same day, and overall system health returned to normal. We then completed follow-up mitigation to stop the underlying event source and reduce the risk of recurrence.
## Impact
* Customers in prod1 experienced elevated API errors and latency in messaging-related workflows.
* The broad cross-customer impact lasted about 41 minutes.
* A subset of messaging workflows failed or were delayed during that period.
## Timeline
* **~3:25 PM ET** — A newly enabled automation began processing a large backlog of historical conversations for one tenant.
* **3:26 PM ET** — Elevated errors and latency began affecting shared messaging workflows in prod1.
* **4:01 PM ET** — We published a status update for the production issue.
* **4:07 PM ET** — Broad cross-customer impact ended as the affected services stabilized.
* **~5:17 PM ET** — We disabled the triggering automation configuration for the affected tenant.
* **~5:23 PM ET** — The remaining retry activity stopped.
## Root cause
The event was triggered when a newly enabled automation for one tenant processed a much larger set of eligible conversations than intended. That sudden volume overloaded a shared downstream service and caused elevated errors and timeouts in dependent workflows.
The incident was amplified by missing safeguards in how this automation handled backlog volume and retries. In particular, the system did not sufficiently limit the number of conversations processed at once or prevent the same failed work from being retried too aggressively.
## Resolution
We restored platform stability during the incident by allowing the affected services to recover under increased capacity, then disabled the triggering automation configuration and cleared the remaining retry backlog. System health is currently normal.
## Preventative actions
We are treating the following preventative actions as a priority bug effort. These actions are expected to be resolved by the end of May in accordance with our SLOs:
* Prevent newly enabled automation settings from processing large historical backlogs unintentionally.
* Add stronger batch limits and tenant-level throttling for this workflow.
* Reduce retry amplification by improving how failed work is tracked and re-queued.
* Improve error handling so rate-limit conditions are classified correctly and handled with the right retry behavior.
Kustomer is aware of an event affecting Knowledge bases and forms that may be displaying a 500 error code and preventing access to these urls.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via kustomer.com for any further questions or updates.
resolved
Kustomer has resolved an event affecting Knowledge Bases and forms that caused a 500 error and preventing access to these pages. To resolve this issue, our team has performed a roll back in this area.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer Support via kustomer.com if you have additional questions or concerns.
postmortem
## **Summary**
On April 14, 2026, customers experienced errors accessing the Kustomer Knowledge Base, with all KB pages returning 500 errors. The issue was caused by a dependency upgrade in a recent deployment that introduced a TLS certificate mismatch in internal service communication. Customer impact began at 9:51 AM ET when the deployment reached the first production environment, and was fully resolved by 10:38 AM ET — a ~47 minute impact window. Engineers identified the root cause and completed a full rollback across all production environments within 12 minutes of the initial alert.
## **Root Cause**
A recent deployment to the KB service included an upgrade to an internal library that contained a known issue with TLS hostname verification. This caused internal service-to-service requests to fail, resulting in 500 errors for all KB requests. The affected library version had previously been identified as problematic in a staging environment, but the fix had not been fully applied across all services before this deployment reached production.
## **Timeline**
**Apr 14, 2026**
9:51 AM ET — Deployment reached the first production environment; customers began experiencing 500 errors when accessing the Knowledge Base
10:26 AM ET — Automated alerting fired; incident response began
10:31 AM ET — Engineers identified a recent deployment as the likely cause and began investigating rollback options
10:35 AM ET — Root cause confirmed as a problematic internal library version; rollback initiated across all production environments
10:37 AM ET — Rollback completed on prod1; KB restored for affected customers
10:38–10:39 AM ET — Rollback completed across remaining production environments
10:45 AM ET — Full KB functionality confirmed restored for all customers
12:09 PM ET — Corrected fix deployed to the Knowledge Base service; remediation completed across all other affected services
## **Lessons/Improvements**
* Implementing stricter controls to prevent pre-release or beta library versions from being deployed to production
* Improving the process for tracking and completing cross-service remediation work when an issue is identified in one service, to ensure all affected services are addressed
* Enhancing our pre-production validation process to improve detection of this class of issue before it reaches production environments
[AIC/AIR] AIC/AIR May Not Respond - Prod 1
Beginn 19. März 2026 um 14:42 UTC · 3h 46m
IssuesGeringfügiger Vorfall
Betroffene Komponenten
Channel - Chat
investigating
Kustomer is aware of an event affecting AI for Customers and Reps that have resulted in them not responding.
Our team is currently working to identify the cause for the issue in order to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via kustomer.com for any further questions or updates.
identified
Kustomer has identified an event in AIC/AIR that may cause unresponsive functionality.
Our team is still continuing to work on implementing a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Kustomer.com for any further questions or updates.
identified
Kustomer has identified an event that is causing unresponsiveness in AIC/AIR.
Our team is working directly with our database provider in order to implement a resolution for this event. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Kustomer.com for any further questions or updates.
identified
Kustomer has identified an event affecting PROD 1 that may cause failing Kustomer AI-Agent features.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting PROD 1 that may result in failures with Kustomer AI Agent features.
Our team is still actively working to implement a resolution. Please expect further updates within the next 30 minutes, and feel free to reach out to Kustomer Support at [email protected] if you have any additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting Prod 1 that caused failures with Kustomer's AI features.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Prod 1 that caused failures with Kustomer's AI features.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On March 19, 2026, some Kustomer AI features were unavailable for customers hosted in one production environment \(Prod1\).
The disruption affected AI-powered responses and related AI workflows for 41 customers in our Prod 1 instance. Other production environments were not affected.
Service was fully restored within 4 hours, and the platform is operating normally.
## Root cause
The incident was caused by a production database index change that was introduced during a service deployment.
That change caused database performance to degrade significantly for a high-traffic dataset used by our AI services. As performance deteriorated, dependent services were unable to start or process requests normally, which led to broad disruption across affected AI features.
Recovery was prolonged by duplicate records created during the incident window, which complicated restoration of normal database constraints and required additional remediation before services could be brought back cleanly.
## Timeline
All times below are in EDT \(UTC-4\) on March 19, 2026.
* **~10:27 AM** — A production deployment introduced a database index change that degraded performance for AI-related services.
* **~10:37 AM** — Customer impact was confirmed and incident response began.
* **~10:42 AM** — We published a public status update and began active mitigation.
* **~11:37 AM** — We identified the primary cause and focused recovery on database stability and service restoration.
* **~12:10 PM** — After scaling database capacity and reducing load, service recovery began.
* **~1:43 PM** — AI services began recovering and queued work started draining.
* **~2:19 PM** — Backlogged work had been processed and core functionality was restored.
* **Later that afternoon** — Follow-up cleanup was completed and normal safeguards were re-applied.
## Lessons and improvements
We take this incident seriously and are making changes to reduce the likelihood of recurrence.
* We are tightening how database schema and index changes are deployed in production, including stronger pre-deployment validation and safer rollout sequencing.
* We are removing index-management behavior from service startup paths where it can create unnecessary risk during deployment.
* We are improving monitoring and alerting so service health failures and database stress are detected earlier.
* We are refining incident response procedures, access readiness, and operational runbooks to speed mitigation during future incidents.
* We are reviewing capacity and resilience safeguards for this part of the platform to better handle abnormal database load.
Kustomer has implemented an update to address an event affecting inbound Voice calls in PROD 1 & 2 that caused inbound calls to drop before getting answered by an agent.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at Email or Chat if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Voice conversations that may cause inbound calls to not be accepted properly in the system.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat or Email for any further questions or updates.
investigating
Kustomer is continuing to investigate an event impacting Voice conversations that may cause inbound calls to not be accepted properly within the system.
Our engineering team is actively working to identify the root cause and implement a resolution as quickly as possible. We will provide another update within the next 30 minutes or sooner as more information becomes available.
If you have any urgent questions or need assistance, please contact Kustomer Support via Chat or Email.
identified
Kustomer is aware of an issue affecting inbound call acceptance in our PROD1 environment where inbound calls were not properly being accepted in the system.
Our team has identified the cause of the issue within the code and are actively working on a fix to restore system behavior back to normal.
Please expect additional updates within the next 30 minutes, and reach out to Kustomer Support at Email or Chat Channels if you have additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting Voice conversations in PROD 1 that caused inbound calls to be dropped when trying to accept calls.
Our team is currently monitoring this update to ensure the issue is fully resolved. If your agents are continuing to experience this issue please have them refresh their browser to apply the fix.
Please expect further updates within the next 30 minutes, and reach out to Kustomer support at Email or Channel if you have additional questions or concerns.
resolved
Kustomer has resolved an incident affecting Voice in PROD 1 that caused inbound calls to drop after being initially accepted. To remediate the issue, our team rolled back a recent code change to a previously stable version, which successfully resolved the error.
After thorough monitoring, we have confirmed that all impacted services are fully restored. If any agents continue to experience issues, please have them refresh their browser to ensure the fix is applied.
If you have any additional questions or concerns, please reach out to Kustomer Support via Email or Chat.
postmortem
## Summary
On February 20, 2026, a frontend change intended to improve performance unintentionally disrupted call setup in our Voice experience. As a result, agents could see an incoming call and click accept/decline, but audio would not connect and calls would time out. We rolled back the change and then deployed an additional fix to make startup more reliable and prevent recurrence.
## What happened
When the Voice widget starts, it needs to receive a small set of configuration settings \(including feature-flag values\) from the main app before it can route and connect calls correctly.
A pre-existing timing edge case meant that, in some cases, that initial configuration could be sent before the Voice widget was fully ready to receive it. Historically, the widget would receive the same configuration again shortly afterward, which masked a state issue.
The performance change reduced those repeated configuration sends \(which is normally a good optimization\). But because the widget sometimes missed the first configuration during startup, it could start with incomplete settings and fall back to legacy call-routing behavior. In that fallback mode, calls could appear answerable in the UI but fail to fully connect audio.
## Root cause
A performance optimization changed how/when configuration settings were delivered during Voice widget startup. Combined with a pre-existing startup timing edge case, some sessions did not receive the required configuration in time and fell back to legacy call-routing behavior, preventing audio from being bridged correctly.
## Timeline
* 3:04 PM ET: Frontend change deployed
* 4:38 PM ET: Reports received that calls could not be answered \(accept/decline visible, no audio\)
* 4:44 PM ET: Initial rollback attempt of an unrelated change did not resolve the issue
* 4:50 PM ET: Incident communication initiated via status page
* 6:24 PM ET: Frontend change reverted/rolled back
* 8:47 PM ET: Reports continued for some agents
* 10:31 PM ET: Identified that affected sessions could persist without a browser refresh due to cached UI state; refresh/new session restored correct behavior
* Saturday 9:00 AM ET: Confirmed reports that calls were successfully connecting
## Lessons / improvements
* Harden Voice widget startup so required configuration can’t be missed during initialization.
* Add automated smoke/E2E coverage for core Voice call flows \(answer/connect audio\) to detect regressions before production.
* Improve safeguards and monitoring around call-connect failures to reduce time-to-detect and time-to-mitigate.
Reported Third Party Event - OpenAI
Beginn 12. Februar 2026 um 15:11 UTC · 2h 57m
IssuesGeringfügiger Vorfall
Betroffene Komponenten
OpenAI
monitoring
Kustomer is aware of an event reported by one of our third party vendors affecting Prod1 that may cause AIC and AIR failures within the platform.
Our team is monitoring the incident, and working with the vendor where possible to resolve the issue. OpenAI's status page for this incident can be found here: https://status.openai.com/incidents/01KH94NGSXNH9H4WBPXB3RFZWX
Please expect further updates within the next 3 hours, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
OpenAI has confirmed that all services are fully recovered. After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
[Satisfaction Surveys] CSATs may not send for voice, email and sms channels - Prod 1
Beginn 27. Januar 2026 um 19:16 UTC · 1d 3h
IssuesGeringfügiger Vorfall
Betroffene Komponenten
CSAT
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing to work to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing investigate the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing to work to investigate the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing to work to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team continues to work to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
monitoring
The team has identified the problem and mitigations have been applied. Jobs are gradually catching up and the team continues to monitor.
investigating
Kustomer has resolved an event affecting Satisfaction Surveys that may cause Surveys to not be sent. To resolve this issue, our team has released an update.
After careful monitoring, our team has determined that our systems are now fully restored, but our engineering team is still redriving surveys that did not originally send. During this redriving period lingering issues may still be present. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Satisfaction Surveys that may cause Surveys to not be sent. To resolve this issue, our team has released an update.
After careful monitoring, our team has determined that our systems are now fully restored, but our engineering team is still redriving surveys that did not originally send. During this redriving period lingering issues may still be present. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
# Post Mortem: Chat, Voice Routing, CSAT, Oauth, Scheduled Send Issues
# **Summary**
On January 22, 2024, customers experienced chat and voice conversations failing to route to agents due to a recent change in the assistant service. This triggered cascading failures across multiple backend services, degrading platform performance for several orgs.
On January 27th, a subsequent incident occurred due to our scheduled jobs queue being flooded by assistant service jobs. This caused some CSAT Surveys to not send, Oauth connections to fail to refresh, and scheduled messages to not send.
**Root Cause**
A recent change to the assistant backend service increased the maximum workflow loops before transferring a conversation to an available agent. This allowed a single rate-limited WhatsApp conversation to become stuck in an infinite retry loop while attempting to transfer. The transfer requests themselves were also rate-limited, overwhelming shared infrastructure \(e.g. “job” engine which is shared by CSAT\) and causing service degradation across multiple orgs.
# **Timeline**
**Jan 22, 2026**
1:33 PM – Users began experiencing errors with chat and voice conversations failing to route to agents
1:44 PM – Engineers identified the problematic deployment and initiated rollback across all environments
1:54 PM – Rollback completed across all environments; on-call engineers continued monitoring system status
2:02 PM – Full assistant functionality restored for all customers
6:07 PM - Customers begin to report Oauth connections that failed to refresh
**Jan 24, 2026**
6:24 PM - Customers begin to report that scheduled messages were not sending
**Jan 27, 2026**
1:00 PM - Customers begin to report that CSAT surveys are not being sent
4:00 PM - Identified cause of CSAT scheduled job processing delays
6:00 PM - Started a script to manually increase processing throughput of scheduled jobs
**Jan 28, 2026**
12:08 PM - Deployed code change to programmatically increase processing throughput of scheduled jobs
3:49 PM - Backlog of all delayed jobs processed and system restored
**Lessons/Improvements**
* Implementing new alerts to detect when scheduled job processing falls behind, enabling faster identification of similar issues
* Improving alert prioritization to reduce noise and ensure critical alerts are acted upon immediately
* Enhancing monitoring for downstream service dependencies
* Evaluating queue architecture changes to prevent a single conversation from impacting other customers \("noisy neighbor" isolation\)
* Investigating improvements to make our job scheduling service more resilient to backlogs
* Creating documentation of all services that depend on scheduled jobs to better understand incident ripple effects
[ROUTING] Chat and Voice conversations not routing [PROD 1 && PROD 2]
Beginn 22. Januar 2026 um 18:44 UTC · 46m
IssuesGeringfügiger Vorfall
Betroffene Komponenten
Channel - Chat
investigating
Kustomer is aware of an event affecting Chat and Voice conversations that may cause the conversation to not be routed to an available agent.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Email or Chat for any further questions or updates.
monitoring
Kustomer has implemented an update to address an event affecting Chats and Voice calls in PROD 1 & 2 that caused conversations to not be routed to available agents.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at Chat and Email if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Conversational Assistants in PROD1 and 2 that caused conversations to not route to available agents. To resolve this issue, our team has completed a rollback our codebase to address the failures in the assistants.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at Chat or Email if you have additional questions or concerns.
postmortem
# Post Mortem: Chat, Voice Routing, CSAT, Oauth, Scheduled Send Issues
# **Summary**
On January 22, 2024, customers experienced chat and voice conversations failing to route to agents due to a recent change in the assistant service. This triggered cascading failures across multiple backend services, degrading platform performance for several orgs.
On January 27th, a subsequent incident occurred due to our scheduled jobs queue being flooded by assistant service jobs. This caused some CSAT Surveys to not send, Oauth connections to fail to refresh, and scheduled messages to not send.
**Root Cause**
A recent change to the assistant backend service increased the maximum workflow loops before transferring a conversation to an available agent. This allowed a single rate-limited WhatsApp conversation to become stuck in an infinite retry loop while attempting to transfer. The transfer requests themselves were also rate-limited, overwhelming shared infrastructure \(e.g. “job” engine which is shared by CSAT\) and causing service degradation across multiple orgs.
# **Timeline**
**Jan 22, 2026**
1:33 PM – Users began experiencing errors with chat and voice conversations failing to route to agents
1:44 PM – Engineers identified the problematic deployment and initiated rollback across all environments
1:54 PM – Rollback completed across all environments; on-call engineers continued monitoring system status
2:02 PM – Full assistant functionality restored for all customers
6:07 PM - Customers begin to report Oauth connections that failed to refresh
**Jan 24, 2026**
6:24 PM - Customers begin to report that scheduled messages were not sending
**Jan 27, 2026**
1:00 PM - Customers begin to report that CSAT surveys are not being sent
4:00 PM - Identified cause of CSAT scheduled job processing delays
6:00 PM - Started a script to manually increase processing throughput of scheduled jobs
**Jan 28, 2026**
12:08 PM - Deployed code change to programmatically increase processing throughput of scheduled jobs
3:49 PM - Backlog of all delayed jobs processed and system restored
**Lessons/Improvements**
* Implementing new alerts to detect when scheduled job processing falls behind, enabling faster identification of similar issues
* Improving alert prioritization to reduce noise and ensure critical alerts are acted upon immediately
* Enhancing monitoring for downstream service dependencies
* Evaluating queue architecture changes to prevent a single conversation from impacting other customers \("noisy neighbor" isolation\)
* Investigating improvements to make our job scheduling service more resilient to backlogs
* Creating documentation of all services that depend on scheduled jobs to better understand incident ripple effects
[OUTBOUND WEBHOOKS] Outbound Webhooks API errors in Prod2 and Prod4
Beginn 10. Dezember 2025 um 03:38 UTC · 1h 14m
Pending
Betroffene Komponenten
Web/Email/Form HooksWeb/Email/Form Hooks
identified
Kustomer is aware of an event affecting outbound webhook delivery in our Prod2 environment. Customers may experience failures due to 404 and 503 errors when the platform attempts to send outbound webhooks.
Our team is actively investigating the cause and working toward a resolution.
Please expect additional updates within the next 30 minutes, and reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
identified
We are continuing to work on a fix for this issue.
monitoring
Kustomer has implemented a fix for the event impacting outbound webhooks in the Prod2 and Prod4 environments. We are backporting this fix across environments, and our team is monitoring to ensure service stability and continued recovery.
Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved the event that impacted outbound webhooks in the Prod2 and Prod4 environments. The backported fix has been fully deployed, and webhook delivery is functioning as expected across all affected environments. Our team has verified stability, and no further impact is anticipated.
If you have additional questions or concerns, please reach out to Kustomer support at [email protected]
postmortem
### **Summary**
On December 9, 2025, Kustomer experienced an interruption to outbound webhook delivery affecting customers in our **prod-2 \(EU\)** and **prod-4 \(IN\)** production regions. During this time, some outbound webhook requests were unable to reach customer endpoints, resulting in delivery failures and error responses.
The disruption occurred during a routine platform update related to our deployment systems. An underlying inconsistency from a previous infrastructure change caused a required backend service to become temporarily unavailable while updates were in progress. As a result, webhook endpoints could not accept or process delivery attempts for a limited period.
Engineering teams identified the issue quickly and restored full service shortly thereafter. Once the affected systems were brought back online, normal webhook delivery resumed and any queued events were successfully processed. There was no permanent data loss, and webhook functionality returned to normal operation across all impacted regions.
### **Impact**
* **Duration:** Approximately 1 hour and 40 minutes
* **Scope:** Production regions prod-2 and prod-4
* **Customer Impact:**
* Outbound webhook delivery failures during the incident window
* Webhook requests may have returned HTTP 503 or 404 responses
* Webhook events were delayed but not lost
### **Next Steps**
While safeguards already exist to protect core platform services, we are implementing additional improvements to reduce the likelihood of similar disruptions in the future. These include strengthening validation during platform updates, improving detection of incomplete infrastructure changes, and enhancing internal visibility when critical services are modified.
Reported Third Party Event affecting AIC and AIR Observability (PROD 1)
Beginn 2. Dezember 2025 um 22:44 UTC · 2h 5m
Pending
monitoring
Kustomer is aware of an event reported by one of our third party vendors affecting the observability of AI Agent for Reps and AI Agent for Customers that may cause trace logs to not show up in AI Agent logs within conversations.
Our team is monitoring the incident, and working with the vendor where possible to resolve the issue. Please expect further updates within the next 3 hours, and reach out to Kustomer support over Chat or Email if you have additional questions or concerns.
resolved
A third party incident has been resolved affecting PROD 1 AI conversations (both AIC and AIR) to not include trace logs. Functionality has returned in this area
Please reach out to Kustomer support through Chat or Email if you have additional questions or concerns.