[Kustomer Voice] Unable to Dial Outbound Calls - All PRODS
Początek 9 września 2026 19:25 UTC · 27m
IssuesDrobny incydent
Dotknięte komponenty
Kustomer Voice
investigating
Kustomer has identified an event affecting Kustomer Voice across all PRODS that may cause issues when attempting to make outbound calls.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes.
monitoring
Kustomer has implemented an update to address an event affecting Kustomer Voice across all PRODs.
Our team is currently monitoring this update to ensure the issue is fully resolved. In order to fully resolve the issue on your end please have all agents refresh their Kustomer Instance.
resolved
Kustomer has resolved an event affecting Kustomer Voice across all PRODs. To resolve this issue, our team has rolled back an update.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
kustomer.help Domena domyślna - Wyświetlanie widoczności - WSZYSTKIE PRODy
Początek 8 września 2026 13:23 UTC · 16m
IssuesDrobny incydent
Dotknięte komponenty
Knowledge baseKnowledge baseCSATCSATCSATKnowledge base
investigating
Kustomer jest świadomy zdarzenia wpływającego na linki i treści, wykorzystującego domyślną domenę kustomer.help (nie niestandardowe domeny klientów) - w tym adresy URL Centrum Pomocy i badania CSAT - które mogą powodować problemy z widocznością.
Nasz zespół pracuje obecnie nad określeniem przyczyny tej kwestii w celu wdrożenia rezolucji. Oczekuj dodatkowych aktualizacji w ciągu najbliższych 30 minut, skontaktuj się z Kustomer Support na support @ kustomer.com, aby uzyskać dalsze pytania lub aktualizacje.
investigating
Nadal badamy tę kwestię.
resolved
Kustomer rozwiązał zdarzenie wpływające na linki i zawartość przy użyciu domyślnej domeny kustomer.help (nie niestandardowe domeny klientów), która spowodowała problemy z widocznością.
Po uważnym monitorowaniu, nasz zespół ustalił, że wszystkie dotknięte obszary zostały w pełni przywrócone. Jeśli masz dodatkowe pytania lub wątpliwości, skontaktuj się z firmą Kustomer Support na adres support @ kustomer.com.
Kustomer jest świadomy zdarzenia zgłoszonego przez jednego z naszych sprzedawców firm trzecich (PubNub) wpływającego na komunikaty w czasie rzeczywistym dla orgów w Prod1, które mogą powodować wiadomości czatu, powiadomienia agentów i aktualizacje na żywo, które mają być opóźnione lub nie pojawiają się w czasie rzeczywistym na platformie.
Nasz zespół aktywnie monitoruje incydent i współpracuje ze sprzedawcą tam, gdzie to możliwe, aby rozwiązać problem.
Można również śledzić status PubNub tutaj: https: / / status.pubnub.com / invents / cz6q37z3s1d5.
Proszę się spodziewać dalszych aktualizacji w ciągu najbliższych 3 godzin i skontaktować się z obsługą Kustomer na support @ kustomer.com, jeśli masz dodatkowe pytania lub obawy.
resolved
Zdarzenie związane z PubNub1 wcześniej zgłoszone, mające wpływ na komunikaty w czasie rzeczywistym dla orgów w Produd1 zostało rozwiązane. PubNub potwierdził, że problem został złagodzony pod ich koniec - nie pominięto żadnych wiadomości i nie doszło do wpływu na klienta w trakcie wydarzenia.
Wszystkie dotknięte obszary pozostają w pełni sprawne. Prosimy o kontakt z obsługą Kustomer pod adresem support @ kustomer.com, jeśli mają Państwo dodatkowe pytania lub obawy.
Problemy z edytorem tekstu z wpisem, wklejeniem i dodawaniem skrótów PROD1, PROD2 i PROD4
Początek 30 lipca 2026 17:43 UTC · 1h 31m
IssuesDrobny incydent
Dotknięte komponenty
Web ClientWeb ClientWeb Client
investigating
Kustomer jest świadomy zdarzenia wpływającego na edytor tekstu, który może powodować problemy podczas pisania i dodawania treści, takich jak skróty.
Nasz zespół pracuje obecnie nad określeniem przyczyny tej kwestii w celu wdrożenia rezolucji. Proszę się spodziewać dodatkowych aktualizacji w ciągu najbliższych 30 minut, prosimy skontaktować się z firmą Kustomer Support poprzez support @ kustomer.com w celu uzyskania dalszych pytań lub aktualizacji.
monitoring
Kustomer wdrożył aktualizację, aby zająć się zdarzeniem wpływającym na edytor tekstu, który może powodować problemy podczas pisania i dodawania treści, takich jak skróty. Odświeżenie przeglądarki w pełni rozwiązuje problem.
Nasz zespół obecnie monitoruje tę aktualizację, aby zapewnić pełne rozwiązanie problemu. Proszę się spodziewać dalszych aktualizacji w ciągu najbliższych 30 minut i skontaktować się z obsługą Kustomer na support @ kustomer.com, jeśli masz dodatkowe pytania lub obawy.
resolved
Kustomer rozwiązał zdarzenie wpływające na edytor tekstu, które może powodować problemy podczas pisania i dodawania treści, takich jak skróty. Aby rozwiązać ten problem, nasz zespół wrócił do poprzedniej wersji.
Po uważnym monitorowaniu, nasz zespół ustalił, że wszystkie dotknięte obszary zostały w pełni przywrócone. Prosimy o kontakt z obsługą Kustomer pod adresem support @ kustomer.com, jeśli mają Państwo dodatkowe pytania lub obawy.
postmortem
# **Summary**
On July 30, 2026, the draft text editor on Kustomer’s Timeline product experienced degraded functionality in certain user workflows. Impact included sporadic cursor behavior, text erasure, and issues with copying and pasting text and shortcuts.
# **Root Cause**
This issue was introduced by a change to the editor that was not fully caught before release. Due to misalignment in our QA process our pre-release validation did not adequately cover the real-world editing patterns affected by this change. Furthermore, this misalignment contributed to a delay in Kustomer’s understanding of the incident's resolution status.
# **Timeline**
## **Jul 30, 2026**
**11:12 AM ET:** The editor change was deployed
**12:27 PM ET:** Kustomer Technical Support escalated the incident to Kustomer’s OnCall process, notifying Engineering immediately.
**12:34 PM ET:** Kustomer Engineering identified the issue and completed a rollback
**12:51 PM ET:** Internal testing incorrectly confirms that the issue has been fully resolved, due to some customers reporting the issue no longer presented itself
**1:29 PM ET:** Continued reports of issue are received from customers who did not receive the full resolution rollout
**1:34 PM ET:** Discovery made that the rollback did not fully deploy to resolve the issue. Kustomer Engineering makes additional change to ensure resolution
**2:01 PM ET:** Corrective change deployed, and fix is confirmed
# **Lessons/Improvements**
Kustomer Engineering maintains a robust CI/CD process, and a multi-environment release process designed to prevent issues of this nature. As part of our continual investment in these areas, we have identified the following action items:
* Ensure that our internal pre-production environments align with our customer-facing experience, including but not limited to the editor experience, so that issues of this nature are caught earlier in development
* Complete an audit of our automated CI/CD process and test coverage of major features, removing any gaps that are identified
Kustomer zidentyfikował zdarzenie wpływające na zdarzenia na platformie w organach PROD1, które może powodować opóźnienia w danych dotyczących zdarzeń.
Nasz zespół pracuje obecnie nad wdrożeniem rezolucji. Proszę się spodziewać dodatkowych aktualizacji w ciągu najbliższych 30 minut i skontaktować się z Kustomer Support na support @ kustomer.com na dalsze pytania lub aktualizacje.
identified
Kustomer zidentyfikował zdarzenie wpływające na zdarzenia na platformie w organach PROD1, które może powodować opóźnienia w danych dotyczących zdarzeń.
Nasz zespół aktywnie pracuje nad wdrożeniem rezolucji. Proszę się spodziewać dodatkowych aktualizacji w ciągu najbliższych 30 minut i skontaktować się z Kustomer Support na support @ kustomer.com na dalsze pytania lub aktualizacje.
identified
Kustomer nadal pracuje nad kwestią wpływającą na wydarzenia na platformie w organach PROD1, które mogą powodować opóźnienia w danych dotyczących zdarzeń.
Nasz zespół aktywnie pracuje nad wdrożeniem rezolucji. Proszę się spodziewać dodatkowych aktualizacji w ciągu najbliższych 30 minut i skontaktować się z Kustomer Support na support @ kustomer.com na dalsze pytania lub aktualizacje.
monitoring
Kustomer wdrożył aktualizację, aby zająć się zdarzeniem wpływającym na zdarzenia platform w organach PROD1, które spowodowały opóźnienia w danych opartych na zdarzeniach.
Nasz zespół obecnie monitoruje tę aktualizację, aby zapewnić pełne rozwiązanie problemu.
Proszę się spodziewać dalszych aktualizacji w ciągu najbliższych 30 minut i skontaktować się z Kustomer Support na support @ kustomer.com, jeśli masz dodatkowe pytania lub obawy.
resolved
Kustomer rozwiązał zdarzenie wpływające na zdarzenia na platformie w organach PROD1, które spowodowało opóźnienia w danych dotyczących zdarzeń.
Po uważnym monitorowaniu, nasz zespół ustalił, że wszystkie dotknięte obszary zostały w pełni przywrócone. Prosimy o kontakt z obsługą Kustomer pod adresem support @ kustomer.com, jeśli mają Państwo dodatkowe pytania lub obawy.
postmortem
Podsumowanie
W dniu 25 lipca 2026 r. niektórzy klienci doświadczyli zwiększonej opóźnienia platformy wpływającej na komunikaty, aktualizacje rozmów, routing, głos i inne usługi obsługi klienta. Klienci mogli widzieć wiadomości pozostające w stanie wysyłania, opóźnione przychodzące lub wychodzące wiadomości, wolniejsze odpowiedzi na strony i API, opóźnione routing lub przypisanie, i przerywane opóźnienia połączeń głosowych.
Incydent został spowodowany niezwykle dużym wybuchem czynności przetwarzania danych w tle. Spowodowało to nagły wzrost popytu na infrastrukturę wspólnej platformy. Automatyczne skalowanie zwiększonej pojemności, ale nie mógł wchłonąć pęknięcia wystarczająco szybko, aby zapobiec degradacji w zależnych przepływów pracy.
Przywróciliśmy obsługę poprzez zwiększenie dostępnej przepustowości, zmniejszenie tempa inicjowania pracy i staranne przetwarzanie opóźnionej pracy podczas monitorowania zdrowia platformy.
# # Impact
* * * Efekt Klienta: * * Przerwane opóźnienie platformy; opóźnione przychodzące i wychodzące wiadomości; wolniejsze rozmowy i aktualizacje API; opóźnione routing lub przypisanie; i okresowe opóźnienia połączeń głosowych
* * * Czas trwania: * * Widoczna na zamówienie degradacja rozpoczęła się około 5: 22 PM ET. Serwis początkowo odzyskał około 7: 31 PM ET, ale degradacja później się powtórzyła. Odnowiono działanie szerokiej platformy, a incydent został rozwiązany o 10: 30 PM ET.
* * * Scope * * Incydent miał wpływ na podzbiór klientów w jednym środowisku produkcyjnym. Nie zidentyfikowaliśmy odpowiedniej degradacji w innych środowiskach produkcyjnych.
# # Timeline
* * * Około 5: 22 PM ET: * * Przetwarzanie opóźnień i zaległości wiadomości zaczęło rosnąć.
* * * 6: 51 PM ET: * * Zespół reagowania na incydenty rozpoczął skoordynowane śledztwo i powrót do zdrowia.
* * * Około 7: 00- 7: 23 PM ET: * * Zidentyfikowaliśmy uszkodzone ścieżki przetwarzania i zaczęliśmy starannie odzyskiwać opóźnioną pracę w kontrolowanym tempie.
* * * 7: 31 PM ET: * * Dodatkowe moce produkcyjne pojawiły się online, wydajność klienta uległa znacznej poprawie, a incydent został początkowo rozwiązany podczas monitorowania kontynuowanego.
* * * 8: 28 PM ET: * * Incydent został ponownie otwarty po wznowieniu raportów o przerywanej degradacji.
* * * Około 8: 38 PM ET: * * Śledztwo potwierdziło, że objawy związane z komunikacją, routowaniem, kanałem i głosem są współzależne od platformy.
* * * Około 9: 52 PM ET: * * Zredukowaliśmy tempo początkowych prac w tle, aby chronić ruch w obliczu klienta.
* * * 10: 08 PM ET: * * Monitorowanie potwierdziło, że ciśnienie pracy znacznie spadło, a zdrowie platformy nadal się poprawia.
* * * 10: 30 ET: * * Po ciągłym monitorowaniu potwierdzono powrót do zdrowia, incydent został rozwiązany.
* * * Po rozdzielczości: * * Odzyskanie pozostających opóźnionych prac było kontynuowane w kontrolowanym tempie w trakcie monitorowania.
# # Root cause
Eksploatacja dużej ilości działań związanych z przetwarzaniem danych w tle spowodowała, że w ciągu krótkiego okresu prace niższego szczebla były bardziej skuteczne niż wspólna infrastruktura platformy.
Automatyczne skalowanie odpowiadało i zwiększało pojemność, ale pojemność była dodawana stopniowo i nie była wystarczająco szybko dostępna dla wielkości i prędkości wybuchu. Podczas gdy platforma nadrabiała zaległości, zwiększono opóźnienie w żądaniu i skróty czasowe wpłynęły na przepływy pracy klientów, które zależały od tej samej infrastruktury.
# # Resolution
Przywróciliśmy wydajność platformy przez:
* Zwiększenie dostępnej przepustowości i poprawa tempa, w jakim można byłoby dodać dodatkowe zdolności.
* Zmniejszenie tempa początkowego obciążenia tła.
* Odzyskanie opóźnionej pracy w kontrolowanych stawkach, aby uniknąć tworzenia kolejnego skoku ruchu.
* Kontynuowanie monitorowania po osiąganiu przez klienta oczekiwanych poziomów.
# # Działania zapobiegawcze
Podejmujemy następujące działania w celu zmniejszenia prawdopodobieństwa i wpływu nawrotu:
* Dodawanie mocniejszych wartości granicznych i kontroli ciśnienia wstecznego dla operacji w tle o dużej objętości.
* Poprawa prędkości, z jaką współdzielona infrastruktura skale podczas nagłego ruchu wzrasta.
* Poprawa izolacji między przetwarzaniem tła a przepływami pracy klienta wrażliwego na latencję.
* Rozszerzenie ostrzeżenia o gwałtowny wzrost zaległości, ciśnienie zasobów i nietypowe wzory obciążenia pracą.
* Wzmocnienie kontrolowanych procedur odzyskiwania dla opóźnionej pracy.
* Expand burst- obciążenia i niesprawności - badania odzysku dla wspólnych usług platformy.
/ Bieżący status
Wydajność platformy powróciła do oczekiwanych poziomów, a początkowe obciążenie pracą pozostawało kontrolowane. Kontynuowaliśmy monitorowanie zdrowia służb i odzyskiwanie opóźnionej pracy po przyjęciu rezolucji.
MessageBird - unable to send outbound replies PROD 1
Początek 25 czerwca 2026 22:24 UTC · 1h 7m
IssuesDrobny incydent
Dotknięte komponenty
Channel - Chat
investigating
Kustomer is aware of an event affecting PROD 1 that may affect the ability to send outbound messages for MessageBird.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support at [email protected] for any further questions or updates.
identified
Kustomer has identified the cause affecting PROD 1 that is impacting the ability to send outbound messages through MessageBird.
Please expect another update within the next 30 minutes, as we reach a resolution. If you have any additional questions or concerns, please reply to this conversation or contact Kustomer Support at [email protected].
Thank you for your patience while we work to resolve this issue.
identified
Kustomer has identified and implemented a fix for the issue affecting PROD 1 that impacted the ability to send outbound messages through MessageBird.
Our team is actively monitoring the platform to ensure the fix remains effective and that service has fully recovered. If you continue to experience issues sending outbound MessageBird messages, please contact Kustomer Support at [email protected].
We will provide another update once monitoring is complete or if there are any significant developments.
Thank you for your patience.
monitoring
Kustomer has identified and implemented a fix for the issue affecting PROD 1 that impacted the ability to send outbound messages through MessageBird.
Our team is actively monitoring the platform to ensure the fix remains effective and that service has fully recovered. If you continue to experience issues sending outbound MessageBird messages, please contact Kustomer Support at [email protected].
We will provide another update once monitoring is complete or if there are any significant developments.
Thank you for your patience.
resolved
Kustomer has resolved the issue affecting PROD 1 that impacted the ability to send outbound messages through MessageBird.
Our team has verified that the fix has been successfully deployed and service has been restored. If you continue to experience issues sending outbound MessageBird messages, please contact Kustomer Support at [email protected]
We apologize for the disruption and appreciate your patience while we worked to resolve the issue.
postmortem
## **Summary**
On June 25, 2026, some customers were unable to send outbound WhatsApp replies from within Kustomer. The issue affected reply sending for a subset of WhatsApp channel configurations, which disrupted agent workflows and prevented some automated outbound messages from being sent through the same path.
The issue was identified and resolved the same day. Service was fully restored, and the platform is operating normally.
## **Root cause**
A change released earlier that day introduced stricter validation in the outbound WhatsApp reply flow. For a subset of supported channel configurations, valid sender values were incorrectly rejected before messages were sent. This caused reply attempts in those configurations to fail.
The issue was limited to specific WhatsApp channel setups and did not affect all WhatsApp traffic equally.
## **Timeline**
* **June 25, 2026, approximately 22:12 UTC** — Reports began coming in that outbound WhatsApp replies were failing for some customers.
* **Shortly after detection** — Investigation confirmed the issue was tied to a recently released validation change in the outbound reply flow.
* **June 26, 2026, approximately 01:07 UTC** — A fix was deployed and reply sending was restored.
## **Lessons and improvements**
* Validation changes for messaging flows now require broader test coverage across supported channel configuration variants before release.
* Additional safeguards are being added to reduce the risk of valid outbound requests being rejected.
* Monitoring and regression checks around outbound messaging paths are being strengthened to detect similar issues more quickly.
Kustomer is aware of an ongoing WhatsApp issue that may cause outbound messages to not be delivered to end users.
While the issue appears to originate with WhatsApp, our team is actively monitoring the situation and working closely to assess the impact on the platform. We will continue to provide updates as more information becomes available.
Please expect an additional update within the next 30 minutes. If you have any questions, please contact Kustomer Support at [email protected].
identified
Kustomer has identified an ongoing issue with WhatsApp that may cause outbound messages to not be delivered to end users.
Our team is closely monitoring the situation and actively working to assess impact. Clients can refer to https://metastatus.com/whatsapp-business-api for the latest WhatsApp service status.
Please expect further updates within the next 30 minutes, and reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
identified
Kustomer continues to monitor the ongoing issue on Whatsapp that may cause outbound Whatsapp messages to not be delivered to end users.
Clients can also refer to https://metastatus.com/whatsapp-business-api for the latest WhatsApp service status.
Please expect further updates within the next 30 minutes, and reach out to Kustomer Support at [email protected] if you have additional questions or concerns
monitoring
Kustomer is aware of an event reported by one of our third party vendors affecting ALL PRODS that may cause Issues with WhatsApp messages not being delivered to end users.
Our team is monitoring the incident, and working with the vendor where possible to resolve the issue. We are working on redriving previously unsent messages to ensure all messages are sent successfully. Please expect further updates within the next 3 hours, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
WhatsApp has confirmed that all services are fully recovered. After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On June 12, 2026, customers using WhatsApp through Kustomer experienced a service event that prevented some outbound messages from being delivered to end users. During the event, messages could appear as sent in Kustomer before final delivery was confirmed by WhatsApp. Service recovered the same day, and affected messages from the incident window were successfully reprocessed.
## Root cause
This event was caused by a disruption in WhatsApp Business Platform services operated by Meta. While that disruption was active, Kustomer was unable to complete delivery for some outbound WhatsApp messages and queued affected messages for retry after the upstream service recovered.
## Timeline
* **10:14 AM EDT** — Customer reports led to investigation of outbound messaging behavior.
* **10:34 AM EDT** — Meta reported high disruptions affecting WhatsApp Business Platform services.
* **1:45 PM EDT** — Reprocessing of queued messages began as upstream recovery progressed.
* **2:32 PM EDT** — Meta reported recovery for the remaining affected WhatsApp services.
* **4:07 PM EDT** — Active message sending was healthy again.
* **4:34 PM EDT** — Messages from the incident window had been successfully reprocessed.
## Lessons and improvements
* We are improving monitoring and alerting for queued WhatsApp messages so degraded delivery is identified earlier.
* We are documenting a safer reprocessing procedure for queued WhatsApp traffic to reduce the chance of retry-related rate limiting during recovery.
* We are reviewing backlog-handling procedures to make recovery more predictable when an upstream provider disruption occurs.
At this time, system health for WhatsApp message delivery through Kustomer is stable.
Kustomer is aware of an event affecting Chat and other channels that may cause delayed message delivery and slower conversation load times.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes and please reach out to Kustomer Support at [email protected] for any further questions or updates.
monitoring
Kustomer has identified the cause of the event affecting Chat and other channels that resulted in delayed message delivery and slower conversation load times. The situation has stabilized and our team is continuing to monitor to ensure full resolution. Please reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Chat and other channels that caused delayed message delivery and slower conversation load times.
After careful monitoring, our team has determined that all affected areas are now fully restored.
Please reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
postmortem
**Summary**
On June 9, 2026, customers experienced delays in chat message delivery and related real-time updates for a portion of the morning. The issue was caused by a third-party service event affecting the infrastructure used to propagate real-time messaging updates. Service recovered the same morning, and current system health is stable.
**Root cause**
A third-party real-time messaging provider experienced elevated latency and intermittent publish failures in one of its regions. That disruption delayed delivery acknowledgements and other real-time updates in Kustomer. Messages continued to be stored successfully, but some updates were not reflected immediately in the client until the external service recovered or the client refreshed.
**Timeline**
* **9:05 AM EDT** — Customer impact began, including delayed chat message delivery and stale real-time updates.
* **9:45 AM EDT** — The third-party provider reported active latency affecting its service.
* **9:48 AM EDT** — The provider reported mitigation and recovery.
* **Shortly after recovery** — Real-time behavior returned to normal and monitoring confirmed stability.
**Lesson/improvements**
* We added additional monitoring tied to the third-party provider's public service health signals so similar service events can be identified faster.
* We are improving alerting around real-time messaging errors to reduce time to diagnosis.
* We are continuing to review operational signals to better distinguish external service issues from internal platform issues.
Chats, Emails, Routing latency PROD 1
Początek 2 czerwca 2026 14:47 UTC · 1h 5m
IssuesDrobny incydent
Dotknięte komponenty
Channel - ChatChannel - Email
investigating
Kustomer is aware of an event affecting chats, emails, and routing that may cause latency with sending, receiving, and routing messages.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support at [email protected] for any further questions or updates.
identified
Kustomer has identified an event affecting chats, emails, and routing that may cause latency with sending, receiving, and routing messages.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting chat, email, and routing that caused latency in sending and delivering messages.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Prod 1 that caused chat, email, and routing latency. To resolve this issue, our team reverted an underlying infrastructure change as a precaution, which helped stabilize the service.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
**Summary**
On June 2, 2026, customers in our prod1 environment experienced elevated latency affecting chat creation, email sending, and conversation routing. The event caused delayed processing for a subset of customer interactions during the incident window. Service performance recovered the same day after mitigation steps were applied, and the environment is currently operating normally.
**Impact**
* Scope: prod1 only
* Customer effect: delayed chat creation, delayed email sends, and slower conversation routing
* Duration: customer-visible impact began in the morning ET on June 2 and materially improved after mitigation later that afternoon
* Other environments: prod2 and prod4 were not impacted by the customer-facing degradation
**Timeline**
* Early June 2: We detected elevated processing latency in services supporting chat, email, and routing in prod1.
* Midday to afternoon ET: Customer reports confirmed delayed system behavior in prod1.
* Afternoon ET: We increased service capacity and adjusted resource limits to reduce backlog and restore throughput.
* Later that day: We completed an infrastructure rollback on affected worker hosts and confirmed stable recovery.
* June 3: Monitoring confirmed the environment remained healthy.
**Root cause**
The incident was caused by infrastructure-level resource exhaustion on a set of worker hosts in prod1. That reduced the availability of a metadata-dependent service and created message backlog, which in turn increased latency for customer-facing workflows such as chat creation, email sending, and routing. The issue was isolated to prod1.
**Resolution**
We restored service by increasing available capacity, raising resource limits for the affected service, and rolling impacted worker infrastructure back to a stable configuration. After those changes were applied, backlog cleared and latency returned to expected levels. Current system health is stable.
**Preventative actions**
* Strengthen host-level capacity and disk safeguards before future infrastructure rollouts
* Add earlier alerting for infrastructure resource pressure to reduce time to detection
* Expand validation for production-scale logging and resource usage prior to promotion
* Continue tuning service capacity thresholds and recovery procedures for metadata-dependent workloads
WhatsApp messages failing Prods 1 + 2
Początek 26 maja 2026 22:14 UTC · 2h 7m
IssuesDrobny incydent
Dotknięte komponenty
Channel - ChatChannel - Chat
monitoring
Kustomer has implemented an update to address an event affecting WhatsApp that caused messaging failures.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
This incident has been resolved. Our teams have confirmed recovery and platform stability following mitigation efforts.
We will be reaching out independently to affected orgs to provide additional context and follow-up as needed. Thank you for your patience while we worked through this issue.
postmortem
## Summary
On May 26, 2026, some customers experienced failures when sending WhatsApp messages from existing conversations. The issue affected outbound message delivery for a subset of WhatsApp configurations. Service was restored after the responsible change was rolled back, and the issue is now resolved.
## Impact
During the incident window, outbound WhatsApp messages failed for some existing conversations. Newly created conversations were less consistently affected, and impact varied by account configuration.
Customers may have seen message-send failures or provider errors while attempting to send WhatsApp messages.
## Timeline
All times EDT.
* **4:05 PM:** The incident window is believed to have begun after a recent WhatsApp-related change.
* **5:54 PM:** Active investigation began.
* **5:56 PM:** The team identified a likely connection to recent WhatsApp sender-selection behavior and started a rollback.
* **6:05 PM:** Rollback completed.
* **6:17 PM – 7:47 PM:** The team validated recovery across affected examples and narrowed the underlying cause.
* **8:21 PM:** The incident was declared resolved.
## Root cause
The incident was caused by an issue in WhatsApp phone-number lookup and sender selection. A recent change expanded support for multiple valid phone-number formats in certain countries. In some cases, that allowed the system to resolve the wrong sender configuration for an existing conversation.
When that happened, outbound sends could target an invalid or unavailable sender identity, causing WhatsApp message delivery to fail.
## Resolution
We mitigated the incident by rolling back the relevant WhatsApp change. After the rollback, we verified recovery against affected examples and confirmed that the incident was resolved.
## Preventative actions
* Move WhatsApp sender resolution toward more stable identifier-based matching rather than relying on ambiguous phone-number formatting.
* Expand regression coverage for country-specific phone-number formatting edge cases.
* Improve monitoring and alerting for message delivery failures and queue growth.
* Improve testing workflows for WhatsApp changes before production rollout.
Kustomer is aware of an event affecting Chats that may cause messages to fail and assistants to not follow the configured flow.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support for any further questions or updates.
monitoring
Kustomer has implemented an update to address an event affecting Chats on Prod 1 that caused messages to fail and assistants to not follow the configured flow.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Chat on Prod 1 that caused messages to fail and assistants to not follow the configured flow.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support if you have additional questions or concerns.
postmortem
## **Summary**
Between May 13 and May 14, 2026, Kustomer Chat experienced a service event that affected chat availability and performance for a subset of customers. During this period, some customers may have seen intermittent failures, elevated error rates, or degraded behavior in chat-related settings and runtime flows.
The event was mitigated through service rollback, additional capacity, and targeted configuration changes. Service health returned to normal after these actions were completed.
## **Root cause**
The event was caused by a combination of reduced service capacity during a deployment rollback and higher-than-expected traffic through chat-related request paths. Under those conditions, the affected chat service became unstable and restarted repeatedly, which reduced available capacity further and increased customer-facing errors.
Our investigation also identified specific high-volume request patterns that increased memory pressure during the event. We addressed those patterns with caching, traffic protections, and capacity changes.
## **Timeline**
* May 13, 2026, early afternoon ET: We detected elevated instability in the chat service and began incident response.
* May 13, 2026, afternoon ET: We rolled back affected changes, increased infrastructure capacity, and stabilized dependent services.
* May 13, 2026, evening ET: We continued monitoring after initial mitigation and investigated recurring memory pressure.
* May 14, 2026: We deployed additional mitigations, including higher minimum service capacity and request-path protections.
* Following the mitigation deployments, service health returned to normal and remained stable.
## **Lessons and improvements**
We completed several improvements to reduce the likelihood of recurrence:
* Increased minimum service capacity to provide more headroom during deployments and recovery.
* Added caching for high-volume chat settings requests.
* Hardened URL processing behavior with stricter filtering, failure caching, and concurrency limits.
* Added improved memory telemetry to speed up detection and diagnosis of similar issues.
* Continued follow-up work on deployment recovery procedures and service safeguards for cross-service rollbacks.
## **Current status**
The mitigations above have been deployed, and the affected chat service is operating normally.
Kustomer is aware of an event affecting Gmail Connectivity and Send & Receive that may cause emails to not route into your Kustomer Platform and impact your ability to send messages on conversations.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat for any further questions or updates.
monitoring
Kustomer has released a fix for the event affecting Gmail Connectivity and Send & Receive functionality that was causing emails to not forward as expected into your Kustomer Platform and impacted sending functionality on conversations.
Our team is currently monitoring the released fix to ensure the issue is resolved. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat for any further questions or updates.
resolved
Kustomer has resolved an event affecting Gmail Connectivity and Send & Receive functionality in Prod 1 and Prod 2.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On May 1, 2026, customers using the Gmail integration experienced a service event that caused Gmail connections to disappear and the email channel to stop functioning.
The issue affected customers across different regions, more specifically:
* EU-based clients, from 2:40 AM ET.
* US-based clients, from 5 AM ET.
The issue was fully resolved by 9:20 AM ET. No customer data was lost during this event. Messages that were delayed during the incident were fully recovered once service was restored. Current system health is stable.
## Root cause
A configuration issue in the infrastructure used by the Gmail integration prevented replacement service tasks from starting correctly during routine task rotation. As running capacity declined over several hours, the Gmail integration became unavailable.
## Timeline
| Time \(ET\) | Event |
| --- | --- |
| May 1, ~12:30 AM | Service task replacement failures began in production. |
| May 1, ~2:40 AM | Engineering was alerted after service capacity dropped to zero in one production environment. |
| May 1, 2:40–9:20 AM | Teams investigated the failure, identified the configuration problem, prepared a fix, and deployed it. |
| May 1, 9:20 AM | Fix deployed and service restored. |
## Lessons and improvements
* We are auditing related infrastructure configurations to identify and correct similar patterns in other services.
* We are standardizing how service roles are managed so this class of configuration issue is less likely to recur.
* We are improving alerting so teams are notified earlier when running service capacity drops below expected levels, before a full service interruption occurs.
* We are adding additional checks to catch configuration drift and invalid service role references earlier in the deployment lifecycle.
[DRAFTS] Internal API Errors (Prod1)
Początek 24 kwietnia 2026 20:03 UTC · 1h 45m
IssuesDrobny incydent
Dotknięte komponenty
Channel - ChatChannel - Email
monitoring
Kustomer has identified an event affecting Prod1 that may cause internal API errors and latency.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting Prod1 that caused internal API errors and latency.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Prod1 that caused API errors and latency. To resolve this issue, our team has released an update.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On April 24, 2026, customers in our prod1 environment experienced a service event that caused elevated errors and latency in messaging-related workflows. The broad cross-customer impact was limited to approximately 41 minutes, from 3:26 PM ET to 4:07 PM ET. During that window, some customers saw failed or delayed messaging operations.
The immediate platform impact was resolved the same day, and overall system health returned to normal. We then completed follow-up mitigation to stop the underlying event source and reduce the risk of recurrence.
## Impact
* Customers in prod1 experienced elevated API errors and latency in messaging-related workflows.
* The broad cross-customer impact lasted about 41 minutes.
* A subset of messaging workflows failed or were delayed during that period.
## Timeline
* **~3:25 PM ET** — A newly enabled automation began processing a large backlog of historical conversations for one tenant.
* **3:26 PM ET** — Elevated errors and latency began affecting shared messaging workflows in prod1.
* **4:01 PM ET** — We published a status update for the production issue.
* **4:07 PM ET** — Broad cross-customer impact ended as the affected services stabilized.
* **~5:17 PM ET** — We disabled the triggering automation configuration for the affected tenant.
* **~5:23 PM ET** — The remaining retry activity stopped.
## Root cause
The event was triggered when a newly enabled automation for one tenant processed a much larger set of eligible conversations than intended. That sudden volume overloaded a shared downstream service and caused elevated errors and timeouts in dependent workflows.
The incident was amplified by missing safeguards in how this automation handled backlog volume and retries. In particular, the system did not sufficiently limit the number of conversations processed at once or prevent the same failed work from being retried too aggressively.
## Resolution
We restored platform stability during the incident by allowing the affected services to recover under increased capacity, then disabled the triggering automation configuration and cleared the remaining retry backlog. System health is currently normal.
## Preventative actions
We are treating the following preventative actions as a priority bug effort. These actions are expected to be resolved by the end of May in accordance with our SLOs:
* Prevent newly enabled automation settings from processing large historical backlogs unintentionally.
* Add stronger batch limits and tenant-level throttling for this workflow.
* Reduce retry amplification by improving how failed work is tracked and re-queued.
* Improve error handling so rate-limit conditions are classified correctly and handled with the right retry behavior.
Kustomer is aware of an event affecting Knowledge bases and forms that may be displaying a 500 error code and preventing access to these urls.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via kustomer.com for any further questions or updates.
resolved
Kustomer has resolved an event affecting Knowledge Bases and forms that caused a 500 error and preventing access to these pages. To resolve this issue, our team has performed a roll back in this area.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer Support via kustomer.com if you have additional questions or concerns.
postmortem
## **Summary**
On April 14, 2026, customers experienced errors accessing the Kustomer Knowledge Base, with all KB pages returning 500 errors. The issue was caused by a dependency upgrade in a recent deployment that introduced a TLS certificate mismatch in internal service communication. Customer impact began at 9:51 AM ET when the deployment reached the first production environment, and was fully resolved by 10:38 AM ET — a ~47 minute impact window. Engineers identified the root cause and completed a full rollback across all production environments within 12 minutes of the initial alert.
## **Root Cause**
A recent deployment to the KB service included an upgrade to an internal library that contained a known issue with TLS hostname verification. This caused internal service-to-service requests to fail, resulting in 500 errors for all KB requests. The affected library version had previously been identified as problematic in a staging environment, but the fix had not been fully applied across all services before this deployment reached production.
## **Timeline**
**Apr 14, 2026**
9:51 AM ET — Deployment reached the first production environment; customers began experiencing 500 errors when accessing the Knowledge Base
10:26 AM ET — Automated alerting fired; incident response began
10:31 AM ET — Engineers identified a recent deployment as the likely cause and began investigating rollback options
10:35 AM ET — Root cause confirmed as a problematic internal library version; rollback initiated across all production environments
10:37 AM ET — Rollback completed on prod1; KB restored for affected customers
10:38–10:39 AM ET — Rollback completed across remaining production environments
10:45 AM ET — Full KB functionality confirmed restored for all customers
12:09 PM ET — Corrected fix deployed to the Knowledge Base service; remediation completed across all other affected services
## **Lessons/Improvements**
* Implementing stricter controls to prevent pre-release or beta library versions from being deployed to production
* Improving the process for tracking and completing cross-service remediation work when an issue is identified in one service, to ensure all affected services are addressed
* Enhancing our pre-production validation process to improve detection of this class of issue before it reaches production environments
[AIC/AIR] AIC/AIR May Not Respond - Prod 1
Początek 19 marca 2026 14:42 UTC · 3h 46m
IssuesDrobny incydent
Dotknięte komponenty
Channel - Chat
investigating
Kustomer is aware of an event affecting AI for Customers and Reps that have resulted in them not responding.
Our team is currently working to identify the cause for the issue in order to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via kustomer.com for any further questions or updates.
identified
Kustomer has identified an event in AIC/AIR that may cause unresponsive functionality.
Our team is still continuing to work on implementing a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Kustomer.com for any further questions or updates.
identified
Kustomer has identified an event that is causing unresponsiveness in AIC/AIR.
Our team is working directly with our database provider in order to implement a resolution for this event. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Kustomer.com for any further questions or updates.
identified
Kustomer has identified an event affecting PROD 1 that may cause failing Kustomer AI-Agent features.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting PROD 1 that may result in failures with Kustomer AI Agent features.
Our team is still actively working to implement a resolution. Please expect further updates within the next 30 minutes, and feel free to reach out to Kustomer Support at [email protected] if you have any additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting Prod 1 that caused failures with Kustomer's AI features.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Prod 1 that caused failures with Kustomer's AI features.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On March 19, 2026, some Kustomer AI features were unavailable for customers hosted in one production environment \(Prod1\).
The disruption affected AI-powered responses and related AI workflows for 41 customers in our Prod 1 instance. Other production environments were not affected.
Service was fully restored within 4 hours, and the platform is operating normally.
## Root cause
The incident was caused by a production database index change that was introduced during a service deployment.
That change caused database performance to degrade significantly for a high-traffic dataset used by our AI services. As performance deteriorated, dependent services were unable to start or process requests normally, which led to broad disruption across affected AI features.
Recovery was prolonged by duplicate records created during the incident window, which complicated restoration of normal database constraints and required additional remediation before services could be brought back cleanly.
## Timeline
All times below are in EDT \(UTC-4\) on March 19, 2026.
* **~10:27 AM** — A production deployment introduced a database index change that degraded performance for AI-related services.
* **~10:37 AM** — Customer impact was confirmed and incident response began.
* **~10:42 AM** — We published a public status update and began active mitigation.
* **~11:37 AM** — We identified the primary cause and focused recovery on database stability and service restoration.
* **~12:10 PM** — After scaling database capacity and reducing load, service recovery began.
* **~1:43 PM** — AI services began recovering and queued work started draining.
* **~2:19 PM** — Backlogged work had been processed and core functionality was restored.
* **Later that afternoon** — Follow-up cleanup was completed and normal safeguards were re-applied.
## Lessons and improvements
We take this incident seriously and are making changes to reduce the likelihood of recurrence.
* We are tightening how database schema and index changes are deployed in production, including stronger pre-deployment validation and safer rollout sequencing.
* We are removing index-management behavior from service startup paths where it can create unnecessary risk during deployment.
* We are improving monitoring and alerting so service health failures and database stress are detected earlier.
* We are refining incident response procedures, access readiness, and operational runbooks to speed mitigation during future incidents.
* We are reviewing capacity and resilience safeguards for this part of the platform to better handle abnormal database load.
Kustomer has implemented an update to address an event affecting inbound Voice calls in PROD 1 & 2 that caused inbound calls to drop before getting answered by an agent.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at Email or Chat if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Voice conversations that may cause inbound calls to not be accepted properly in the system.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat or Email for any further questions or updates.
investigating
Kustomer is continuing to investigate an event impacting Voice conversations that may cause inbound calls to not be accepted properly within the system.
Our engineering team is actively working to identify the root cause and implement a resolution as quickly as possible. We will provide another update within the next 30 minutes or sooner as more information becomes available.
If you have any urgent questions or need assistance, please contact Kustomer Support via Chat or Email.
identified
Kustomer is aware of an issue affecting inbound call acceptance in our PROD1 environment where inbound calls were not properly being accepted in the system.
Our team has identified the cause of the issue within the code and are actively working on a fix to restore system behavior back to normal.
Please expect additional updates within the next 30 minutes, and reach out to Kustomer Support at Email or Chat Channels if you have additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting Voice conversations in PROD 1 that caused inbound calls to be dropped when trying to accept calls.
Our team is currently monitoring this update to ensure the issue is fully resolved. If your agents are continuing to experience this issue please have them refresh their browser to apply the fix.
Please expect further updates within the next 30 minutes, and reach out to Kustomer support at Email or Channel if you have additional questions or concerns.
resolved
Kustomer has resolved an incident affecting Voice in PROD 1 that caused inbound calls to drop after being initially accepted. To remediate the issue, our team rolled back a recent code change to a previously stable version, which successfully resolved the error.
After thorough monitoring, we have confirmed that all impacted services are fully restored. If any agents continue to experience issues, please have them refresh their browser to ensure the fix is applied.
If you have any additional questions or concerns, please reach out to Kustomer Support via Email or Chat.
postmortem
## Summary
On February 20, 2026, a frontend change intended to improve performance unintentionally disrupted call setup in our Voice experience. As a result, agents could see an incoming call and click accept/decline, but audio would not connect and calls would time out. We rolled back the change and then deployed an additional fix to make startup more reliable and prevent recurrence.
## What happened
When the Voice widget starts, it needs to receive a small set of configuration settings \(including feature-flag values\) from the main app before it can route and connect calls correctly.
A pre-existing timing edge case meant that, in some cases, that initial configuration could be sent before the Voice widget was fully ready to receive it. Historically, the widget would receive the same configuration again shortly afterward, which masked a state issue.
The performance change reduced those repeated configuration sends \(which is normally a good optimization\). But because the widget sometimes missed the first configuration during startup, it could start with incomplete settings and fall back to legacy call-routing behavior. In that fallback mode, calls could appear answerable in the UI but fail to fully connect audio.
## Root cause
A performance optimization changed how/when configuration settings were delivered during Voice widget startup. Combined with a pre-existing startup timing edge case, some sessions did not receive the required configuration in time and fell back to legacy call-routing behavior, preventing audio from being bridged correctly.
## Timeline
* 3:04 PM ET: Frontend change deployed
* 4:38 PM ET: Reports received that calls could not be answered \(accept/decline visible, no audio\)
* 4:44 PM ET: Initial rollback attempt of an unrelated change did not resolve the issue
* 4:50 PM ET: Incident communication initiated via status page
* 6:24 PM ET: Frontend change reverted/rolled back
* 8:47 PM ET: Reports continued for some agents
* 10:31 PM ET: Identified that affected sessions could persist without a browser refresh due to cached UI state; refresh/new session restored correct behavior
* Saturday 9:00 AM ET: Confirmed reports that calls were successfully connecting
## Lessons / improvements
* Harden Voice widget startup so required configuration can’t be missed during initialization.
* Add automated smoke/E2E coverage for core Voice call flows \(answer/connect audio\) to detect regressions before production.
* Improve safeguards and monitoring around call-connect failures to reduce time-to-detect and time-to-mitigate.
Reported Third Party Event - OpenAI
Początek 12 lutego 2026 15:11 UTC · 2h 57m
IssuesDrobny incydent
Dotknięte komponenty
OpenAI
monitoring
Kustomer is aware of an event reported by one of our third party vendors affecting Prod1 that may cause AIC and AIR failures within the platform.
Our team is monitoring the incident, and working with the vendor where possible to resolve the issue. OpenAI's status page for this incident can be found here: https://status.openai.com/incidents/01KH94NGSXNH9H4WBPXB3RFZWX
Please expect further updates within the next 3 hours, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
OpenAI has confirmed that all services are fully recovered. After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
[Satisfaction Surveys] CSATs may not send for voice, email and sms channels - Prod 1
Początek 27 stycznia 2026 19:16 UTC · 1d 3h
IssuesDrobny incydent
Dotknięte komponenty
CSAT
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing to work to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing investigate the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing to work to investigate the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing to work to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team continues to work to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
monitoring
The team has identified the problem and mitigations have been applied. Jobs are gradually catching up and the team continues to monitor.
investigating
Kustomer has resolved an event affecting Satisfaction Surveys that may cause Surveys to not be sent. To resolve this issue, our team has released an update.
After careful monitoring, our team has determined that our systems are now fully restored, but our engineering team is still redriving surveys that did not originally send. During this redriving period lingering issues may still be present. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Satisfaction Surveys that may cause Surveys to not be sent. To resolve this issue, our team has released an update.
After careful monitoring, our team has determined that our systems are now fully restored, but our engineering team is still redriving surveys that did not originally send. During this redriving period lingering issues may still be present. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
# Post Mortem: Chat, Voice Routing, CSAT, Oauth, Scheduled Send Issues
# **Summary**
On January 22, 2024, customers experienced chat and voice conversations failing to route to agents due to a recent change in the assistant service. This triggered cascading failures across multiple backend services, degrading platform performance for several orgs.
On January 27th, a subsequent incident occurred due to our scheduled jobs queue being flooded by assistant service jobs. This caused some CSAT Surveys to not send, Oauth connections to fail to refresh, and scheduled messages to not send.
**Root Cause**
A recent change to the assistant backend service increased the maximum workflow loops before transferring a conversation to an available agent. This allowed a single rate-limited WhatsApp conversation to become stuck in an infinite retry loop while attempting to transfer. The transfer requests themselves were also rate-limited, overwhelming shared infrastructure \(e.g. “job” engine which is shared by CSAT\) and causing service degradation across multiple orgs.
# **Timeline**
**Jan 22, 2026**
1:33 PM – Users began experiencing errors with chat and voice conversations failing to route to agents
1:44 PM – Engineers identified the problematic deployment and initiated rollback across all environments
1:54 PM – Rollback completed across all environments; on-call engineers continued monitoring system status
2:02 PM – Full assistant functionality restored for all customers
6:07 PM - Customers begin to report Oauth connections that failed to refresh
**Jan 24, 2026**
6:24 PM - Customers begin to report that scheduled messages were not sending
**Jan 27, 2026**
1:00 PM - Customers begin to report that CSAT surveys are not being sent
4:00 PM - Identified cause of CSAT scheduled job processing delays
6:00 PM - Started a script to manually increase processing throughput of scheduled jobs
**Jan 28, 2026**
12:08 PM - Deployed code change to programmatically increase processing throughput of scheduled jobs
3:49 PM - Backlog of all delayed jobs processed and system restored
**Lessons/Improvements**
* Implementing new alerts to detect when scheduled job processing falls behind, enabling faster identification of similar issues
* Improving alert prioritization to reduce noise and ensure critical alerts are acted upon immediately
* Enhancing monitoring for downstream service dependencies
* Evaluating queue architecture changes to prevent a single conversation from impacting other customers \("noisy neighbor" isolation\)
* Investigating improvements to make our job scheduling service more resilient to backlogs
* Creating documentation of all services that depend on scheduled jobs to better understand incident ripple effects
[ROUTING] Chat and Voice conversations not routing [PROD 1 && PROD 2]
Początek 22 stycznia 2026 18:44 UTC · 46m
IssuesDrobny incydent
Dotknięte komponenty
Channel - Chat
investigating
Kustomer is aware of an event affecting Chat and Voice conversations that may cause the conversation to not be routed to an available agent.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Email or Chat for any further questions or updates.
monitoring
Kustomer has implemented an update to address an event affecting Chats and Voice calls in PROD 1 & 2 that caused conversations to not be routed to available agents.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at Chat and Email if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Conversational Assistants in PROD1 and 2 that caused conversations to not route to available agents. To resolve this issue, our team has completed a rollback our codebase to address the failures in the assistants.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at Chat or Email if you have additional questions or concerns.
postmortem
# Post Mortem: Chat, Voice Routing, CSAT, Oauth, Scheduled Send Issues
# **Summary**
On January 22, 2024, customers experienced chat and voice conversations failing to route to agents due to a recent change in the assistant service. This triggered cascading failures across multiple backend services, degrading platform performance for several orgs.
On January 27th, a subsequent incident occurred due to our scheduled jobs queue being flooded by assistant service jobs. This caused some CSAT Surveys to not send, Oauth connections to fail to refresh, and scheduled messages to not send.
**Root Cause**
A recent change to the assistant backend service increased the maximum workflow loops before transferring a conversation to an available agent. This allowed a single rate-limited WhatsApp conversation to become stuck in an infinite retry loop while attempting to transfer. The transfer requests themselves were also rate-limited, overwhelming shared infrastructure \(e.g. “job” engine which is shared by CSAT\) and causing service degradation across multiple orgs.
# **Timeline**
**Jan 22, 2026**
1:33 PM – Users began experiencing errors with chat and voice conversations failing to route to agents
1:44 PM – Engineers identified the problematic deployment and initiated rollback across all environments
1:54 PM – Rollback completed across all environments; on-call engineers continued monitoring system status
2:02 PM – Full assistant functionality restored for all customers
6:07 PM - Customers begin to report Oauth connections that failed to refresh
**Jan 24, 2026**
6:24 PM - Customers begin to report that scheduled messages were not sending
**Jan 27, 2026**
1:00 PM - Customers begin to report that CSAT surveys are not being sent
4:00 PM - Identified cause of CSAT scheduled job processing delays
6:00 PM - Started a script to manually increase processing throughput of scheduled jobs
**Jan 28, 2026**
12:08 PM - Deployed code change to programmatically increase processing throughput of scheduled jobs
3:49 PM - Backlog of all delayed jobs processed and system restored
**Lessons/Improvements**
* Implementing new alerts to detect when scheduled job processing falls behind, enabling faster identification of similar issues
* Improving alert prioritization to reduce noise and ensure critical alerts are acted upon immediately
* Enhancing monitoring for downstream service dependencies
* Evaluating queue architecture changes to prevent a single conversation from impacting other customers \("noisy neighbor" isolation\)
* Investigating improvements to make our job scheduling service more resilient to backlogs
* Creating documentation of all services that depend on scheduled jobs to better understand incident ripple effects
[OUTBOUND WEBHOOKS] Outbound Webhooks API errors in Prod2 and Prod4
Początek 10 grudnia 2025 03:38 UTC · 1h 14m
Pending
Dotknięte komponenty
Web/Email/Form HooksWeb/Email/Form Hooks
identified
Kustomer is aware of an event affecting outbound webhook delivery in our Prod2 environment. Customers may experience failures due to 404 and 503 errors when the platform attempts to send outbound webhooks.
Our team is actively investigating the cause and working toward a resolution.
Please expect additional updates within the next 30 minutes, and reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
identified
We are continuing to work on a fix for this issue.
monitoring
Kustomer has implemented a fix for the event impacting outbound webhooks in the Prod2 and Prod4 environments. We are backporting this fix across environments, and our team is monitoring to ensure service stability and continued recovery.
Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved the event that impacted outbound webhooks in the Prod2 and Prod4 environments. The backported fix has been fully deployed, and webhook delivery is functioning as expected across all affected environments. Our team has verified stability, and no further impact is anticipated.
If you have additional questions or concerns, please reach out to Kustomer support at [email protected]
postmortem
### **Summary**
On December 9, 2025, Kustomer experienced an interruption to outbound webhook delivery affecting customers in our **prod-2 \(EU\)** and **prod-4 \(IN\)** production regions. During this time, some outbound webhook requests were unable to reach customer endpoints, resulting in delivery failures and error responses.
The disruption occurred during a routine platform update related to our deployment systems. An underlying inconsistency from a previous infrastructure change caused a required backend service to become temporarily unavailable while updates were in progress. As a result, webhook endpoints could not accept or process delivery attempts for a limited period.
Engineering teams identified the issue quickly and restored full service shortly thereafter. Once the affected systems were brought back online, normal webhook delivery resumed and any queued events were successfully processed. There was no permanent data loss, and webhook functionality returned to normal operation across all impacted regions.
### **Impact**
* **Duration:** Approximately 1 hour and 40 minutes
* **Scope:** Production regions prod-2 and prod-4
* **Customer Impact:**
* Outbound webhook delivery failures during the incident window
* Webhook requests may have returned HTTP 503 or 404 responses
* Webhook events were delayed but not lost
### **Next Steps**
While safeguards already exist to protect core platform services, we are implementing additional improvements to reduce the likelihood of similar disruptions in the future. These include strengthening validation during platform updates, improving detection of incomplete infrastructure changes, and enhancing internal visibility when critical services are modified.