Kustomer, üçüncü taraf satıcılarımızdan biri tarafından bildirilen bir olayın farkındadır (PubNub) gerçek zamanlı mesajlaşmayı Prod1'de veya sohbet mesajlarına, acente bildirimlerine ve gerçek zamanlı olarak platformda görünmemelerine neden olabilir.
Ekibimiz olayı aktif olarak izliyor ve sorunu çözmek için mümkün olan satıcı ile çalışıyor.
Ayrıca PubNub'un durumunu burada takip edebilirsiniz: https://status.pubnub.com/incidents/cz6q37z3s1d5.
Lütfen önümüzdeki 3 saat içinde daha fazla güncelleme bekleyin ve [email protected]'da Kustomer desteğine ulaşırsanız ek soruları veya endişeleriniz varsa.
resolved
PubNub ile ilgili olay daha önce Prod1'deki orgs için gerçek zamanlı mesajlaşmayı etkilediğini bildirdi. PubNub, sorunu sona erdiğini doğruladı - hiçbir mesaj kaçırıldı ve olay sırasında müşteri etkisi olmadı.
Tüm etkilenen bölgeler tamamen operasyonel kalır. Lütfen [email protected]'da Kustomer desteğe daha fazla soru veya endişeniz varsa ulaşabilirsiniz.
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
Text Editor ile tipleme, geçmişleme ve kısayollar PROD1, PROD2 ve PROD4
Başlangıç 30 Temmuz 2026 17:43 UTC · 1h 31m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Web ClientWeb ClientWeb Client
investigating
Kustomer, kısa kesimler gibi içerik yazmak ve eklemek için sorunlara neden olabilecek bir metin editörü etkileyen bir olayın farkındadır.
Ekibimiz şu anda bu konunun nedenini bir karar uygulamak için bir çaba içinde tanımlamak için çalışıyor. Lütfen önümüzdeki 30 dakika içinde ek güncellemeler bekleyin, lütfen [email protected] ile daha fazla soru veya güncellemeler için Kustomer Desteğine geçin.
monitoring
Kustomer, kısa kesimler gibi içerik yazarken ve eklemek için sorunlara neden olabilecek bir metin editörüne hitap etmek için bir güncelleme uyguladı. Tarayıcınızı yenileme sorunu tamamen çözecektir.
Ekibimiz şu anda bu güncellemeyi konunun tamamen çözülmesini sağlamak için izliyor. Lütfen önümüzdeki 30 dakika içinde daha fazla güncellemeyi bekleyin ve [email protected]'da Kustomer desteğine ulaşırsanız ilave sorularınız veya endişeleriniz varsa.
resolved
Kustomer, kısa kesimler gibi içerik yazmak ve eklemek için sorunlara neden olabilecek bir metin editörü etkileyen bir olay çözdü. Bu sorunu çözmek için, ekibimiz önceki bir sürüme geri döndü.
Dikkatli izlemeden sonra, ekibimiz tüm etkilenen alanların şimdi tamamen restore edildiğini belirledi. Lütfen [email protected]'da Kustomer desteğe daha fazla soru veya endişeniz varsa ulaşabilirsiniz.
postmortem
# **Summary**
On July 30, 2026, the draft text editor on Kustomer’s Timeline product experienced degraded functionality in certain user workflows. Impact included sporadic cursor behavior, text erasure, and issues with copying and pasting text and shortcuts.
# **Root Cause**
This issue was introduced by a change to the editor that was not fully caught before release. Due to misalignment in our QA process our pre-release validation did not adequately cover the real-world editing patterns affected by this change. Furthermore, this misalignment contributed to a delay in Kustomer’s understanding of the incident's resolution status.
# **Timeline**
## **Jul 30, 2026**
**11:12 AM ET:** The editor change was deployed
**12:27 PM ET:** Kustomer Technical Support escalated the incident to Kustomer’s OnCall process, notifying Engineering immediately.
**12:34 PM ET:** Kustomer Engineering identified the issue and completed a rollback
**12:51 PM ET:** Internal testing incorrectly confirms that the issue has been fully resolved, due to some customers reporting the issue no longer presented itself
**1:29 PM ET:** Continued reports of issue are received from customers who did not receive the full resolution rollout
**1:34 PM ET:** Discovery made that the rollback did not fully deploy to resolve the issue. Kustomer Engineering makes additional change to ensure resolution
**2:01 PM ET:** Corrective change deployed, and fix is confirmed
# **Lessons/Improvements**
Kustomer Engineering maintains a robust CI/CD process, and a multi-environment release process designed to prevent issues of this nature. As part of our continual investment in these areas, we have identified the following action items:
* Ensure that our internal pre-production environments align with our customer-facing experience, including but not limited to the editor experience, so that issues of this nature are caught earlier in development
* Complete an audit of our automated CI/CD process and test coverage of major features, removing any gaps that are identified
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
Kustomer, PROD1'de platform olayları etkileyen bir olay tespit etti veya olay temelli verilerde gecikmelere neden olabilir.
Ekibimiz şu anda bir karar uygulamak için çalışıyor. Lütfen önümüzdeki 30 dakika içinde ek güncellemeler bekleyin ve daha fazla soru veya güncelleştirme için Kustomer Desteğine ulaşırsınız.
identified
Kustomer, PROD1'de platform olayları etkileyen bir olay tespit etti veya olay temelli verilerde gecikmelere neden olabilir.
Ekibimiz şu anda bir karar uygulamak için aktif olarak çalışıyor. Lütfen önümüzdeki 30 dakika içinde ek güncellemeler bekleyin ve daha fazla soru veya güncelleştirme için Kustomer Desteğine ulaşırsınız.
identified
Kustomer, PROD1'de platform olayları etkileyen konu üzerinde çalışmaya devam ediyor ve olay temelli verilerde gecikmelere neden olabilir.
Ekibimiz bir karar uygulamak için aktif olarak çalışıyor. Lütfen önümüzdeki 30 dakika içinde ek güncellemeler bekleyin ve daha fazla soru veya güncelleştirme için Kustomer Desteğine ulaşırsınız.
monitoring
Kustomer, etkinlik temelli veriler üzerinde gecikmelere neden olan PROD1 veyags'te platform olayları etkileyen bir olay ele almak için bir güncelleme uyguladı.
Ekibimiz şu anda bu güncellemeyi konunun tamamen çözülmesini sağlamak için izliyor.
Lütfen önümüzdeki 30 dakika içinde daha fazla güncellemeyi bekleyin ve [email protected]'da Kustomer Desteğine ulaşırsanız ilave sorularınız veya endişeleriniz varsa.
resolved
Kustomer, PROD1'de platform olayları etkileyen bir olayı çözdü ve olay temelli verilerde gecikmelere neden oldu.
Dikkatli izlemeden sonra, ekibimiz tüm etkilenen alanların şimdi tamamen restore edildiğini belirledi. Lütfen [email protected]'da Kustomer desteğe daha fazla soru veya endişeniz varsa ulaşabilirsiniz.
postmortem
# # # Özet
25 Temmuz 2026'da, bazı müşteriler mesajlaşma, konuşma güncellemelerini, routing, ses ve diğer müşteri hizmetleri iş akışlarını etkileyen yüksek platformu yaşadı. Müşteriler mesaj gönderme durumunda kalabilir, gecikmiş veya dışa dönük mesajlar, daha yavaş sayfa ve API yanıtları, gecikmiş routing veya atama ve geçici ses arama gecikmeleri.
Olay, alışılmadık derecede büyük bir arka plan veri işleme faaliyeti patlamasına neden oldu. Bu, paylaşılan platform altyapısına talep etme konusunda aniden bir artış yarattı. Otomatik ölçeklendirme kapasitesi eklendi, ancak bağımlı iş akışlarında bozulmayı önlemek için yeterince hızlı patlamayı absorbe edemezdi.
Mevcut kapasiteyi artırarak hizmeti restore ettik, iş yüklerinin oranını azaltın ve platform sağlığını takip ederken gecikmiş işleri dikkatlice işlememiz.
## Effects
* **Müşteri etkisi:** Intermittent platform latency; inbound ve outbound mesajlaşma; daha yavaş konuşma ve API güncellemeleri; gecikmiş routing veya atama; ve aralıklı ses-call gecikmeleri
* **Durasyon: ** Müşteriye yönelik bozulma yaklaşık 5:22 PM ET'de başladı. Hizmet başlangıçta yaklaşık 7:31 PM ET'de geri alındı, ancak daha sonra tekrar alındı. Geniş platform performansı restore edildi ve olay 10:30 PM ET'de çözüldü.
* **Scope:** Olay, bir üretim ortamında müşterilerin alt setini etkiledi. Diğer üretim ortamlarında ilgili müşteri odaklı bozulma tespit etmedik.
## Timeline
* **Approximately 5:22 PM ET:** latency ve mesaj geriloglar büyümeye başladı.
* **6:51 PM ET:** Olay yanıt ekibi koordineli soruşturma ve kurtarma başlattı.
* **Approximately 7:00-7:23 PM ET:** Etkilenen işleme yollarını tespit ettik ve kontrol oranlarında gecikmiş işleri dikkatle geri almaya başladık.
* ** 7:31 PM ET:** Ek kapasite online gelmişti, müşteri odaklı performans maddi olarak gelişmişti ve olay başlangıçta izleme devam ederken çözüldü.
* **8:28 PM ET:** Olay yenilenen geçici bozulma raporlarından sonra yeniden açıldı.
* **Approximately 8:38 PM ET:** İnceleme, mesajlaşma, routing, kanal ve ses belirtileri aynı temel platform bağımlılığını paylaştı.
* **Approximately 9:52 PM ET:** Müşteri odaklı trafiği korumak için arka plan iş yükünü azalttık.
* **10:08 PM ET:** İzleme, iş yükü basıncının önemli ölçüde düştüğünü ve platformun sağlığını geliştirmeye devam ettiğini doğruladı.
* **10:30 PM ET:** İzlemeye devam ettikten sonra, olay çözüldü.
* ** Karardan Sonra:** Kalan gecikmiş çalışmaların kurtarılması, izleme altındaki kontrol oranlarında devam etti.
## Root neden
Yüksek hacimli arka plan veri işleme faaliyetlerinin patlaması, paylaşılan platform altyapısından daha düşük bir çalışma yarattı.
Otomatik ölçeklendirme yanıt verdi ve ilave kapasite ekledi, ancak kapasite artırıldı ve patlamanın büyüklüğü ve hızı için online olarak yeterince hızlı gelmedi. Platform yetişirken, gecikmiş istek ve zaman aralıkları aynı altyapıya bağlı olan müşteri odaklı akışları etkiledi.
## Karar
Platform performansını geri aldık:
* Mevcut kapasiteyi artırmak ve ek kapasitenin ek olabileceği oranı geliştirmek.
* Arka plan iş yükünün oranını azaltın.
* Başka bir trafik artışı yaratmaktan kaçınmak için gecikmiş işleri yeniden keşfedin.
* Müşteri odaklı performanstan sonra sürekli izleme beklenen seviyelere geri döndü.
## Önleme eylemleri
Recurrence olasılığını ve etkisini azaltmak için aşağıdaki eylemleri alıyoruz:
* Yüksek hacimli arka plan operasyonları için daha güçlü sınırları ve baskı kontrollerini ekleyin.
* Otomatik trafik artışları sırasında paylaşılan altyapı ölçeklerini geliştirmek.
* Arka işleme ve geciken müşteri akışları arasındaki izolasyonu geliştirin.
* Hızlı backlog büyümesi, kaynak basıncı ve alışılmadık iş yük modelleri için uyarı genişletmek.
* gecikmiş iş için kontrollü kurtarma prosedürlerini güçlendirin.
* Paylaşılan platform hizmetleri için patlama-yükleme ve başarısızlık testi genişletin.
## Current status
Platform performansı beklenen seviyelere geri döndü ve girişimci iş yükü kontrol edildi. Servis sağlığını ve gecikmiş çalışmaların düzeltilmesini karar vermeden sürdürdük.
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
MessageBird - unable to send outbound replies PROD 1
Başlangıç 25 Haziran 2026 22:24 UTC · 1h 7m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Channel - Chat
investigating
Kustomer is aware of an event affecting PROD 1 that may affect the ability to send outbound messages for MessageBird.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support at [email protected] for any further questions or updates.
identified
Kustomer has identified the cause affecting PROD 1 that is impacting the ability to send outbound messages through MessageBird.
Please expect another update within the next 30 minutes, as we reach a resolution. If you have any additional questions or concerns, please reply to this conversation or contact Kustomer Support at [email protected].
Thank you for your patience while we work to resolve this issue.
identified
Kustomer has identified and implemented a fix for the issue affecting PROD 1 that impacted the ability to send outbound messages through MessageBird.
Our team is actively monitoring the platform to ensure the fix remains effective and that service has fully recovered. If you continue to experience issues sending outbound MessageBird messages, please contact Kustomer Support at [email protected].
We will provide another update once monitoring is complete or if there are any significant developments.
Thank you for your patience.
monitoring
Kustomer has identified and implemented a fix for the issue affecting PROD 1 that impacted the ability to send outbound messages through MessageBird.
Our team is actively monitoring the platform to ensure the fix remains effective and that service has fully recovered. If you continue to experience issues sending outbound MessageBird messages, please contact Kustomer Support at [email protected].
We will provide another update once monitoring is complete or if there are any significant developments.
Thank you for your patience.
resolved
Kustomer has resolved the issue affecting PROD 1 that impacted the ability to send outbound messages through MessageBird.
Our team has verified that the fix has been successfully deployed and service has been restored. If you continue to experience issues sending outbound MessageBird messages, please contact Kustomer Support at [email protected]
We apologize for the disruption and appreciate your patience while we worked to resolve the issue.
postmortem
## **Summary**
On June 25, 2026, some customers were unable to send outbound WhatsApp replies from within Kustomer. The issue affected reply sending for a subset of WhatsApp channel configurations, which disrupted agent workflows and prevented some automated outbound messages from being sent through the same path.
The issue was identified and resolved the same day. Service was fully restored, and the platform is operating normally.
## **Root cause**
A change released earlier that day introduced stricter validation in the outbound WhatsApp reply flow. For a subset of supported channel configurations, valid sender values were incorrectly rejected before messages were sent. This caused reply attempts in those configurations to fail.
The issue was limited to specific WhatsApp channel setups and did not affect all WhatsApp traffic equally.
## **Timeline**
* **June 25, 2026, approximately 22:12 UTC** — Reports began coming in that outbound WhatsApp replies were failing for some customers.
* **Shortly after detection** — Investigation confirmed the issue was tied to a recently released validation change in the outbound reply flow.
* **June 26, 2026, approximately 01:07 UTC** — A fix was deployed and reply sending was restored.
## **Lessons and improvements**
* Validation changes for messaging flows now require broader test coverage across supported channel configuration variants before release.
* Additional safeguards are being added to reduce the risk of valid outbound requests being rejected.
* Monitoring and regression checks around outbound messaging paths are being strengthened to detect similar issues more quickly.
Kustomer is aware of an ongoing WhatsApp issue that may cause outbound messages to not be delivered to end users.
While the issue appears to originate with WhatsApp, our team is actively monitoring the situation and working closely to assess the impact on the platform. We will continue to provide updates as more information becomes available.
Please expect an additional update within the next 30 minutes. If you have any questions, please contact Kustomer Support at [email protected].
identified
Kustomer has identified an ongoing issue with WhatsApp that may cause outbound messages to not be delivered to end users.
Our team is closely monitoring the situation and actively working to assess impact. Clients can refer to https://metastatus.com/whatsapp-business-api for the latest WhatsApp service status.
Please expect further updates within the next 30 minutes, and reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
identified
Kustomer continues to monitor the ongoing issue on Whatsapp that may cause outbound Whatsapp messages to not be delivered to end users.
Clients can also refer to https://metastatus.com/whatsapp-business-api for the latest WhatsApp service status.
Please expect further updates within the next 30 minutes, and reach out to Kustomer Support at [email protected] if you have additional questions or concerns
monitoring
Kustomer is aware of an event reported by one of our third party vendors affecting ALL PRODS that may cause Issues with WhatsApp messages not being delivered to end users.
Our team is monitoring the incident, and working with the vendor where possible to resolve the issue. We are working on redriving previously unsent messages to ensure all messages are sent successfully. Please expect further updates within the next 3 hours, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
WhatsApp has confirmed that all services are fully recovered. After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On June 12, 2026, customers using WhatsApp through Kustomer experienced a service event that prevented some outbound messages from being delivered to end users. During the event, messages could appear as sent in Kustomer before final delivery was confirmed by WhatsApp. Service recovered the same day, and affected messages from the incident window were successfully reprocessed.
## Root cause
This event was caused by a disruption in WhatsApp Business Platform services operated by Meta. While that disruption was active, Kustomer was unable to complete delivery for some outbound WhatsApp messages and queued affected messages for retry after the upstream service recovered.
## Timeline
* **10:14 AM EDT** — Customer reports led to investigation of outbound messaging behavior.
* **10:34 AM EDT** — Meta reported high disruptions affecting WhatsApp Business Platform services.
* **1:45 PM EDT** — Reprocessing of queued messages began as upstream recovery progressed.
* **2:32 PM EDT** — Meta reported recovery for the remaining affected WhatsApp services.
* **4:07 PM EDT** — Active message sending was healthy again.
* **4:34 PM EDT** — Messages from the incident window had been successfully reprocessed.
## Lessons and improvements
* We are improving monitoring and alerting for queued WhatsApp messages so degraded delivery is identified earlier.
* We are documenting a safer reprocessing procedure for queued WhatsApp traffic to reduce the chance of retry-related rate limiting during recovery.
* We are reviewing backlog-handling procedures to make recovery more predictable when an upstream provider disruption occurs.
At this time, system health for WhatsApp message delivery through Kustomer is stable.
Kustomer is aware of an event affecting Chat and other channels that may cause delayed message delivery and slower conversation load times.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes and please reach out to Kustomer Support at [email protected] for any further questions or updates.
monitoring
Kustomer has identified the cause of the event affecting Chat and other channels that resulted in delayed message delivery and slower conversation load times. The situation has stabilized and our team is continuing to monitor to ensure full resolution. Please reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Chat and other channels that caused delayed message delivery and slower conversation load times.
After careful monitoring, our team has determined that all affected areas are now fully restored.
Please reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
postmortem
**Summary**
On June 9, 2026, customers experienced delays in chat message delivery and related real-time updates for a portion of the morning. The issue was caused by a third-party service event affecting the infrastructure used to propagate real-time messaging updates. Service recovered the same morning, and current system health is stable.
**Root cause**
A third-party real-time messaging provider experienced elevated latency and intermittent publish failures in one of its regions. That disruption delayed delivery acknowledgements and other real-time updates in Kustomer. Messages continued to be stored successfully, but some updates were not reflected immediately in the client until the external service recovered or the client refreshed.
**Timeline**
* **9:05 AM EDT** — Customer impact began, including delayed chat message delivery and stale real-time updates.
* **9:45 AM EDT** — The third-party provider reported active latency affecting its service.
* **9:48 AM EDT** — The provider reported mitigation and recovery.
* **Shortly after recovery** — Real-time behavior returned to normal and monitoring confirmed stability.
**Lesson/improvements**
* We added additional monitoring tied to the third-party provider's public service health signals so similar service events can be identified faster.
* We are improving alerting around real-time messaging errors to reduce time to diagnosis.
* We are continuing to review operational signals to better distinguish external service issues from internal platform issues.
Chats, Emails, Routing latency PROD 1
Başlangıç 2 Haziran 2026 14:47 UTC · 1h 5m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Channel - ChatChannel - Email
investigating
Kustomer is aware of an event affecting chats, emails, and routing that may cause latency with sending, receiving, and routing messages.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support at [email protected] for any further questions or updates.
identified
Kustomer has identified an event affecting chats, emails, and routing that may cause latency with sending, receiving, and routing messages.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting chat, email, and routing that caused latency in sending and delivering messages.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Prod 1 that caused chat, email, and routing latency. To resolve this issue, our team reverted an underlying infrastructure change as a precaution, which helped stabilize the service.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
**Summary**
On June 2, 2026, customers in our prod1 environment experienced elevated latency affecting chat creation, email sending, and conversation routing. The event caused delayed processing for a subset of customer interactions during the incident window. Service performance recovered the same day after mitigation steps were applied, and the environment is currently operating normally.
**Impact**
* Scope: prod1 only
* Customer effect: delayed chat creation, delayed email sends, and slower conversation routing
* Duration: customer-visible impact began in the morning ET on June 2 and materially improved after mitigation later that afternoon
* Other environments: prod2 and prod4 were not impacted by the customer-facing degradation
**Timeline**
* Early June 2: We detected elevated processing latency in services supporting chat, email, and routing in prod1.
* Midday to afternoon ET: Customer reports confirmed delayed system behavior in prod1.
* Afternoon ET: We increased service capacity and adjusted resource limits to reduce backlog and restore throughput.
* Later that day: We completed an infrastructure rollback on affected worker hosts and confirmed stable recovery.
* June 3: Monitoring confirmed the environment remained healthy.
**Root cause**
The incident was caused by infrastructure-level resource exhaustion on a set of worker hosts in prod1. That reduced the availability of a metadata-dependent service and created message backlog, which in turn increased latency for customer-facing workflows such as chat creation, email sending, and routing. The issue was isolated to prod1.
**Resolution**
We restored service by increasing available capacity, raising resource limits for the affected service, and rolling impacted worker infrastructure back to a stable configuration. After those changes were applied, backlog cleared and latency returned to expected levels. Current system health is stable.
**Preventative actions**
* Strengthen host-level capacity and disk safeguards before future infrastructure rollouts
* Add earlier alerting for infrastructure resource pressure to reduce time to detection
* Expand validation for production-scale logging and resource usage prior to promotion
* Continue tuning service capacity thresholds and recovery procedures for metadata-dependent workloads
WhatsApp messages failing Prods 1 + 2
Başlangıç 26 Mayıs 2026 22:14 UTC · 2h 7m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Channel - ChatChannel - Chat
monitoring
Kustomer has implemented an update to address an event affecting WhatsApp that caused messaging failures.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
This incident has been resolved. Our teams have confirmed recovery and platform stability following mitigation efforts.
We will be reaching out independently to affected orgs to provide additional context and follow-up as needed. Thank you for your patience while we worked through this issue.
postmortem
## Summary
On May 26, 2026, some customers experienced failures when sending WhatsApp messages from existing conversations. The issue affected outbound message delivery for a subset of WhatsApp configurations. Service was restored after the responsible change was rolled back, and the issue is now resolved.
## Impact
During the incident window, outbound WhatsApp messages failed for some existing conversations. Newly created conversations were less consistently affected, and impact varied by account configuration.
Customers may have seen message-send failures or provider errors while attempting to send WhatsApp messages.
## Timeline
All times EDT.
* **4:05 PM:** The incident window is believed to have begun after a recent WhatsApp-related change.
* **5:54 PM:** Active investigation began.
* **5:56 PM:** The team identified a likely connection to recent WhatsApp sender-selection behavior and started a rollback.
* **6:05 PM:** Rollback completed.
* **6:17 PM – 7:47 PM:** The team validated recovery across affected examples and narrowed the underlying cause.
* **8:21 PM:** The incident was declared resolved.
## Root cause
The incident was caused by an issue in WhatsApp phone-number lookup and sender selection. A recent change expanded support for multiple valid phone-number formats in certain countries. In some cases, that allowed the system to resolve the wrong sender configuration for an existing conversation.
When that happened, outbound sends could target an invalid or unavailable sender identity, causing WhatsApp message delivery to fail.
## Resolution
We mitigated the incident by rolling back the relevant WhatsApp change. After the rollback, we verified recovery against affected examples and confirmed that the incident was resolved.
## Preventative actions
* Move WhatsApp sender resolution toward more stable identifier-based matching rather than relying on ambiguous phone-number formatting.
* Expand regression coverage for country-specific phone-number formatting edge cases.
* Improve monitoring and alerting for message delivery failures and queue growth.
* Improve testing workflows for WhatsApp changes before production rollout.
Kustomer is aware of an event affecting Chats that may cause messages to fail and assistants to not follow the configured flow.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support for any further questions or updates.
monitoring
Kustomer has implemented an update to address an event affecting Chats on Prod 1 that caused messages to fail and assistants to not follow the configured flow.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Chat on Prod 1 that caused messages to fail and assistants to not follow the configured flow.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support if you have additional questions or concerns.
postmortem
## **Summary**
Between May 13 and May 14, 2026, Kustomer Chat experienced a service event that affected chat availability and performance for a subset of customers. During this period, some customers may have seen intermittent failures, elevated error rates, or degraded behavior in chat-related settings and runtime flows.
The event was mitigated through service rollback, additional capacity, and targeted configuration changes. Service health returned to normal after these actions were completed.
## **Root cause**
The event was caused by a combination of reduced service capacity during a deployment rollback and higher-than-expected traffic through chat-related request paths. Under those conditions, the affected chat service became unstable and restarted repeatedly, which reduced available capacity further and increased customer-facing errors.
Our investigation also identified specific high-volume request patterns that increased memory pressure during the event. We addressed those patterns with caching, traffic protections, and capacity changes.
## **Timeline**
* May 13, 2026, early afternoon ET: We detected elevated instability in the chat service and began incident response.
* May 13, 2026, afternoon ET: We rolled back affected changes, increased infrastructure capacity, and stabilized dependent services.
* May 13, 2026, evening ET: We continued monitoring after initial mitigation and investigated recurring memory pressure.
* May 14, 2026: We deployed additional mitigations, including higher minimum service capacity and request-path protections.
* Following the mitigation deployments, service health returned to normal and remained stable.
## **Lessons and improvements**
We completed several improvements to reduce the likelihood of recurrence:
* Increased minimum service capacity to provide more headroom during deployments and recovery.
* Added caching for high-volume chat settings requests.
* Hardened URL processing behavior with stricter filtering, failure caching, and concurrency limits.
* Added improved memory telemetry to speed up detection and diagnosis of similar issues.
* Continued follow-up work on deployment recovery procedures and service safeguards for cross-service rollbacks.
## **Current status**
The mitigations above have been deployed, and the affected chat service is operating normally.
Kustomer is aware of an event affecting Gmail Connectivity and Send & Receive that may cause emails to not route into your Kustomer Platform and impact your ability to send messages on conversations.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat for any further questions or updates.
monitoring
Kustomer has released a fix for the event affecting Gmail Connectivity and Send & Receive functionality that was causing emails to not forward as expected into your Kustomer Platform and impacted sending functionality on conversations.
Our team is currently monitoring the released fix to ensure the issue is resolved. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat for any further questions or updates.
resolved
Kustomer has resolved an event affecting Gmail Connectivity and Send & Receive functionality in Prod 1 and Prod 2.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On May 1, 2026, customers using the Gmail integration experienced a service event that caused Gmail connections to disappear and the email channel to stop functioning.
The issue affected customers across different regions, more specifically:
* EU-based clients, from 2:40 AM ET.
* US-based clients, from 5 AM ET.
The issue was fully resolved by 9:20 AM ET. No customer data was lost during this event. Messages that were delayed during the incident were fully recovered once service was restored. Current system health is stable.
## Root cause
A configuration issue in the infrastructure used by the Gmail integration prevented replacement service tasks from starting correctly during routine task rotation. As running capacity declined over several hours, the Gmail integration became unavailable.
## Timeline
| Time \(ET\) | Event |
| --- | --- |
| May 1, ~12:30 AM | Service task replacement failures began in production. |
| May 1, ~2:40 AM | Engineering was alerted after service capacity dropped to zero in one production environment. |
| May 1, 2:40–9:20 AM | Teams investigated the failure, identified the configuration problem, prepared a fix, and deployed it. |
| May 1, 9:20 AM | Fix deployed and service restored. |
## Lessons and improvements
* We are auditing related infrastructure configurations to identify and correct similar patterns in other services.
* We are standardizing how service roles are managed so this class of configuration issue is less likely to recur.
* We are improving alerting so teams are notified earlier when running service capacity drops below expected levels, before a full service interruption occurs.
* We are adding additional checks to catch configuration drift and invalid service role references earlier in the deployment lifecycle.
[DRAFTS] Internal API Errors (Prod1)
Başlangıç 24 Nisan 2026 20:03 UTC · 1h 45m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Channel - ChatChannel - Email
monitoring
Kustomer has identified an event affecting Prod1 that may cause internal API errors and latency.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting Prod1 that caused internal API errors and latency.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Prod1 that caused API errors and latency. To resolve this issue, our team has released an update.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On April 24, 2026, customers in our prod1 environment experienced a service event that caused elevated errors and latency in messaging-related workflows. The broad cross-customer impact was limited to approximately 41 minutes, from 3:26 PM ET to 4:07 PM ET. During that window, some customers saw failed or delayed messaging operations.
The immediate platform impact was resolved the same day, and overall system health returned to normal. We then completed follow-up mitigation to stop the underlying event source and reduce the risk of recurrence.
## Impact
* Customers in prod1 experienced elevated API errors and latency in messaging-related workflows.
* The broad cross-customer impact lasted about 41 minutes.
* A subset of messaging workflows failed or were delayed during that period.
## Timeline
* **~3:25 PM ET** — A newly enabled automation began processing a large backlog of historical conversations for one tenant.
* **3:26 PM ET** — Elevated errors and latency began affecting shared messaging workflows in prod1.
* **4:01 PM ET** — We published a status update for the production issue.
* **4:07 PM ET** — Broad cross-customer impact ended as the affected services stabilized.
* **~5:17 PM ET** — We disabled the triggering automation configuration for the affected tenant.
* **~5:23 PM ET** — The remaining retry activity stopped.
## Root cause
The event was triggered when a newly enabled automation for one tenant processed a much larger set of eligible conversations than intended. That sudden volume overloaded a shared downstream service and caused elevated errors and timeouts in dependent workflows.
The incident was amplified by missing safeguards in how this automation handled backlog volume and retries. In particular, the system did not sufficiently limit the number of conversations processed at once or prevent the same failed work from being retried too aggressively.
## Resolution
We restored platform stability during the incident by allowing the affected services to recover under increased capacity, then disabled the triggering automation configuration and cleared the remaining retry backlog. System health is currently normal.
## Preventative actions
We are treating the following preventative actions as a priority bug effort. These actions are expected to be resolved by the end of May in accordance with our SLOs:
* Prevent newly enabled automation settings from processing large historical backlogs unintentionally.
* Add stronger batch limits and tenant-level throttling for this workflow.
* Reduce retry amplification by improving how failed work is tracked and re-queued.
* Improve error handling so rate-limit conditions are classified correctly and handled with the right retry behavior.
Kustomer is aware of an event affecting Knowledge bases and forms that may be displaying a 500 error code and preventing access to these urls.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via kustomer.com for any further questions or updates.
resolved
Kustomer has resolved an event affecting Knowledge Bases and forms that caused a 500 error and preventing access to these pages. To resolve this issue, our team has performed a roll back in this area.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer Support via kustomer.com if you have additional questions or concerns.
postmortem
## **Summary**
On April 14, 2026, customers experienced errors accessing the Kustomer Knowledge Base, with all KB pages returning 500 errors. The issue was caused by a dependency upgrade in a recent deployment that introduced a TLS certificate mismatch in internal service communication. Customer impact began at 9:51 AM ET when the deployment reached the first production environment, and was fully resolved by 10:38 AM ET — a ~47 minute impact window. Engineers identified the root cause and completed a full rollback across all production environments within 12 minutes of the initial alert.
## **Root Cause**
A recent deployment to the KB service included an upgrade to an internal library that contained a known issue with TLS hostname verification. This caused internal service-to-service requests to fail, resulting in 500 errors for all KB requests. The affected library version had previously been identified as problematic in a staging environment, but the fix had not been fully applied across all services before this deployment reached production.
## **Timeline**
**Apr 14, 2026**
9:51 AM ET — Deployment reached the first production environment; customers began experiencing 500 errors when accessing the Knowledge Base
10:26 AM ET — Automated alerting fired; incident response began
10:31 AM ET — Engineers identified a recent deployment as the likely cause and began investigating rollback options
10:35 AM ET — Root cause confirmed as a problematic internal library version; rollback initiated across all production environments
10:37 AM ET — Rollback completed on prod1; KB restored for affected customers
10:38–10:39 AM ET — Rollback completed across remaining production environments
10:45 AM ET — Full KB functionality confirmed restored for all customers
12:09 PM ET — Corrected fix deployed to the Knowledge Base service; remediation completed across all other affected services
## **Lessons/Improvements**
* Implementing stricter controls to prevent pre-release or beta library versions from being deployed to production
* Improving the process for tracking and completing cross-service remediation work when an issue is identified in one service, to ensure all affected services are addressed
* Enhancing our pre-production validation process to improve detection of this class of issue before it reaches production environments
[AIC/AIR] AIC/AIR May Not Respond - Prod 1
Başlangıç 19 Mart 2026 14:42 UTC · 3h 46m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Channel - Chat
investigating
Kustomer is aware of an event affecting AI for Customers and Reps that have resulted in them not responding.
Our team is currently working to identify the cause for the issue in order to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via kustomer.com for any further questions or updates.
identified
Kustomer has identified an event in AIC/AIR that may cause unresponsive functionality.
Our team is still continuing to work on implementing a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Kustomer.com for any further questions or updates.
identified
Kustomer has identified an event that is causing unresponsiveness in AIC/AIR.
Our team is working directly with our database provider in order to implement a resolution for this event. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Kustomer.com for any further questions or updates.
identified
Kustomer has identified an event affecting PROD 1 that may cause failing Kustomer AI-Agent features.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting PROD 1 that may result in failures with Kustomer AI Agent features.
Our team is still actively working to implement a resolution. Please expect further updates within the next 30 minutes, and feel free to reach out to Kustomer Support at [email protected] if you have any additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting Prod 1 that caused failures with Kustomer's AI features.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Prod 1 that caused failures with Kustomer's AI features.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
## Summary
On March 19, 2026, some Kustomer AI features were unavailable for customers hosted in one production environment \(Prod1\).
The disruption affected AI-powered responses and related AI workflows for 41 customers in our Prod 1 instance. Other production environments were not affected.
Service was fully restored within 4 hours, and the platform is operating normally.
## Root cause
The incident was caused by a production database index change that was introduced during a service deployment.
That change caused database performance to degrade significantly for a high-traffic dataset used by our AI services. As performance deteriorated, dependent services were unable to start or process requests normally, which led to broad disruption across affected AI features.
Recovery was prolonged by duplicate records created during the incident window, which complicated restoration of normal database constraints and required additional remediation before services could be brought back cleanly.
## Timeline
All times below are in EDT \(UTC-4\) on March 19, 2026.
* **~10:27 AM** — A production deployment introduced a database index change that degraded performance for AI-related services.
* **~10:37 AM** — Customer impact was confirmed and incident response began.
* **~10:42 AM** — We published a public status update and began active mitigation.
* **~11:37 AM** — We identified the primary cause and focused recovery on database stability and service restoration.
* **~12:10 PM** — After scaling database capacity and reducing load, service recovery began.
* **~1:43 PM** — AI services began recovering and queued work started draining.
* **~2:19 PM** — Backlogged work had been processed and core functionality was restored.
* **Later that afternoon** — Follow-up cleanup was completed and normal safeguards were re-applied.
## Lessons and improvements
We take this incident seriously and are making changes to reduce the likelihood of recurrence.
* We are tightening how database schema and index changes are deployed in production, including stronger pre-deployment validation and safer rollout sequencing.
* We are removing index-management behavior from service startup paths where it can create unnecessary risk during deployment.
* We are improving monitoring and alerting so service health failures and database stress are detected earlier.
* We are refining incident response procedures, access readiness, and operational runbooks to speed mitigation during future incidents.
* We are reviewing capacity and resilience safeguards for this part of the platform to better handle abnormal database load.
Kustomer has implemented an update to address an event affecting inbound Voice calls in PROD 1 & 2 that caused inbound calls to drop before getting answered by an agent.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at Email or Chat if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Voice conversations that may cause inbound calls to not be accepted properly in the system.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat or Email for any further questions or updates.
investigating
Kustomer is continuing to investigate an event impacting Voice conversations that may cause inbound calls to not be accepted properly within the system.
Our engineering team is actively working to identify the root cause and implement a resolution as quickly as possible. We will provide another update within the next 30 minutes or sooner as more information becomes available.
If you have any urgent questions or need assistance, please contact Kustomer Support via Chat or Email.
identified
Kustomer is aware of an issue affecting inbound call acceptance in our PROD1 environment where inbound calls were not properly being accepted in the system.
Our team has identified the cause of the issue within the code and are actively working on a fix to restore system behavior back to normal.
Please expect additional updates within the next 30 minutes, and reach out to Kustomer Support at Email or Chat Channels if you have additional questions or concerns.
monitoring
Kustomer has implemented an update to address an event affecting Voice conversations in PROD 1 that caused inbound calls to be dropped when trying to accept calls.
Our team is currently monitoring this update to ensure the issue is fully resolved. If your agents are continuing to experience this issue please have them refresh their browser to apply the fix.
Please expect further updates within the next 30 minutes, and reach out to Kustomer support at Email or Channel if you have additional questions or concerns.
resolved
Kustomer has resolved an incident affecting Voice in PROD 1 that caused inbound calls to drop after being initially accepted. To remediate the issue, our team rolled back a recent code change to a previously stable version, which successfully resolved the error.
After thorough monitoring, we have confirmed that all impacted services are fully restored. If any agents continue to experience issues, please have them refresh their browser to ensure the fix is applied.
If you have any additional questions or concerns, please reach out to Kustomer Support via Email or Chat.
postmortem
## Summary
On February 20, 2026, a frontend change intended to improve performance unintentionally disrupted call setup in our Voice experience. As a result, agents could see an incoming call and click accept/decline, but audio would not connect and calls would time out. We rolled back the change and then deployed an additional fix to make startup more reliable and prevent recurrence.
## What happened
When the Voice widget starts, it needs to receive a small set of configuration settings \(including feature-flag values\) from the main app before it can route and connect calls correctly.
A pre-existing timing edge case meant that, in some cases, that initial configuration could be sent before the Voice widget was fully ready to receive it. Historically, the widget would receive the same configuration again shortly afterward, which masked a state issue.
The performance change reduced those repeated configuration sends \(which is normally a good optimization\). But because the widget sometimes missed the first configuration during startup, it could start with incomplete settings and fall back to legacy call-routing behavior. In that fallback mode, calls could appear answerable in the UI but fail to fully connect audio.
## Root cause
A performance optimization changed how/when configuration settings were delivered during Voice widget startup. Combined with a pre-existing startup timing edge case, some sessions did not receive the required configuration in time and fell back to legacy call-routing behavior, preventing audio from being bridged correctly.
## Timeline
* 3:04 PM ET: Frontend change deployed
* 4:38 PM ET: Reports received that calls could not be answered \(accept/decline visible, no audio\)
* 4:44 PM ET: Initial rollback attempt of an unrelated change did not resolve the issue
* 4:50 PM ET: Incident communication initiated via status page
* 6:24 PM ET: Frontend change reverted/rolled back
* 8:47 PM ET: Reports continued for some agents
* 10:31 PM ET: Identified that affected sessions could persist without a browser refresh due to cached UI state; refresh/new session restored correct behavior
* Saturday 9:00 AM ET: Confirmed reports that calls were successfully connecting
## Lessons / improvements
* Harden Voice widget startup so required configuration can’t be missed during initialization.
* Add automated smoke/E2E coverage for core Voice call flows \(answer/connect audio\) to detect regressions before production.
* Improve safeguards and monitoring around call-connect failures to reduce time-to-detect and time-to-mitigate.
Reported Third Party Event - OpenAI
Başlangıç 12 Şubat 2026 15:11 UTC · 2h 57m
IssuesKüçük çaplı olay
Etkilenen bileşenler
OpenAI
monitoring
Kustomer is aware of an event reported by one of our third party vendors affecting Prod1 that may cause AIC and AIR failures within the platform.
Our team is monitoring the incident, and working with the vendor where possible to resolve the issue. OpenAI's status page for this incident can be found here: https://status.openai.com/incidents/01KH94NGSXNH9H4WBPXB3RFZWX
Please expect further updates within the next 3 hours, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
OpenAI has confirmed that all services are fully recovered. After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
[Satisfaction Surveys] CSATs may not send for voice, email and sms channels - Prod 1
Başlangıç 27 Ocak 2026 19:16 UTC · 1d 3h
IssuesKüçük çaplı olay
Etkilenen bileşenler
CSAT
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing to work to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing investigate the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
investigating
Kustomer is aware of an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing to work to investigate the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is currently working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is continuing to work to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team continues to work to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
identified
Kustomer has identified an event affecting Satisfaction surveys that may cause surveys to not be sent, once conversations are marked done within the platform.
Our team is working to implement a resolution. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
monitoring
The team has identified the problem and mitigations have been applied. Jobs are gradually catching up and the team continues to monitor.
investigating
Kustomer has resolved an event affecting Satisfaction Surveys that may cause Surveys to not be sent. To resolve this issue, our team has released an update.
After careful monitoring, our team has determined that our systems are now fully restored, but our engineering team is still redriving surveys that did not originally send. During this redriving period lingering issues may still be present. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Satisfaction Surveys that may cause Surveys to not be sent. To resolve this issue, our team has released an update.
After careful monitoring, our team has determined that our systems are now fully restored, but our engineering team is still redriving surveys that did not originally send. During this redriving period lingering issues may still be present. Please reach out to Kustomer support at [email protected] if you have additional questions or concerns.
postmortem
# Post Mortem: Chat, Voice Routing, CSAT, Oauth, Scheduled Send Issues
# **Summary**
On January 22, 2024, customers experienced chat and voice conversations failing to route to agents due to a recent change in the assistant service. This triggered cascading failures across multiple backend services, degrading platform performance for several orgs.
On January 27th, a subsequent incident occurred due to our scheduled jobs queue being flooded by assistant service jobs. This caused some CSAT Surveys to not send, Oauth connections to fail to refresh, and scheduled messages to not send.
**Root Cause**
A recent change to the assistant backend service increased the maximum workflow loops before transferring a conversation to an available agent. This allowed a single rate-limited WhatsApp conversation to become stuck in an infinite retry loop while attempting to transfer. The transfer requests themselves were also rate-limited, overwhelming shared infrastructure \(e.g. “job” engine which is shared by CSAT\) and causing service degradation across multiple orgs.
# **Timeline**
**Jan 22, 2026**
1:33 PM – Users began experiencing errors with chat and voice conversations failing to route to agents
1:44 PM – Engineers identified the problematic deployment and initiated rollback across all environments
1:54 PM – Rollback completed across all environments; on-call engineers continued monitoring system status
2:02 PM – Full assistant functionality restored for all customers
6:07 PM - Customers begin to report Oauth connections that failed to refresh
**Jan 24, 2026**
6:24 PM - Customers begin to report that scheduled messages were not sending
**Jan 27, 2026**
1:00 PM - Customers begin to report that CSAT surveys are not being sent
4:00 PM - Identified cause of CSAT scheduled job processing delays
6:00 PM - Started a script to manually increase processing throughput of scheduled jobs
**Jan 28, 2026**
12:08 PM - Deployed code change to programmatically increase processing throughput of scheduled jobs
3:49 PM - Backlog of all delayed jobs processed and system restored
**Lessons/Improvements**
* Implementing new alerts to detect when scheduled job processing falls behind, enabling faster identification of similar issues
* Improving alert prioritization to reduce noise and ensure critical alerts are acted upon immediately
* Enhancing monitoring for downstream service dependencies
* Evaluating queue architecture changes to prevent a single conversation from impacting other customers \("noisy neighbor" isolation\)
* Investigating improvements to make our job scheduling service more resilient to backlogs
* Creating documentation of all services that depend on scheduled jobs to better understand incident ripple effects
[ROUTING] Chat and Voice conversations not routing [PROD 1 && PROD 2]
Başlangıç 22 Ocak 2026 18:44 UTC · 46m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Channel - Chat
investigating
Kustomer is aware of an event affecting Chat and Voice conversations that may cause the conversation to not be routed to an available agent.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Email or Chat for any further questions or updates.
monitoring
Kustomer has implemented an update to address an event affecting Chats and Voice calls in PROD 1 & 2 that caused conversations to not be routed to available agents.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support at Chat and Email if you have additional questions or concerns.
resolved
Kustomer has resolved an event affecting Conversational Assistants in PROD1 and 2 that caused conversations to not route to available agents. To resolve this issue, our team has completed a rollback our codebase to address the failures in the assistants.
After careful monitoring, our team has determined that all affected areas are now fully restored. Please reach out to Kustomer support at Chat or Email if you have additional questions or concerns.
postmortem
# Post Mortem: Chat, Voice Routing, CSAT, Oauth, Scheduled Send Issues
# **Summary**
On January 22, 2024, customers experienced chat and voice conversations failing to route to agents due to a recent change in the assistant service. This triggered cascading failures across multiple backend services, degrading platform performance for several orgs.
On January 27th, a subsequent incident occurred due to our scheduled jobs queue being flooded by assistant service jobs. This caused some CSAT Surveys to not send, Oauth connections to fail to refresh, and scheduled messages to not send.
**Root Cause**
A recent change to the assistant backend service increased the maximum workflow loops before transferring a conversation to an available agent. This allowed a single rate-limited WhatsApp conversation to become stuck in an infinite retry loop while attempting to transfer. The transfer requests themselves were also rate-limited, overwhelming shared infrastructure \(e.g. “job” engine which is shared by CSAT\) and causing service degradation across multiple orgs.
# **Timeline**
**Jan 22, 2026**
1:33 PM – Users began experiencing errors with chat and voice conversations failing to route to agents
1:44 PM – Engineers identified the problematic deployment and initiated rollback across all environments
1:54 PM – Rollback completed across all environments; on-call engineers continued monitoring system status
2:02 PM – Full assistant functionality restored for all customers
6:07 PM - Customers begin to report Oauth connections that failed to refresh
**Jan 24, 2026**
6:24 PM - Customers begin to report that scheduled messages were not sending
**Jan 27, 2026**
1:00 PM - Customers begin to report that CSAT surveys are not being sent
4:00 PM - Identified cause of CSAT scheduled job processing delays
6:00 PM - Started a script to manually increase processing throughput of scheduled jobs
**Jan 28, 2026**
12:08 PM - Deployed code change to programmatically increase processing throughput of scheduled jobs
3:49 PM - Backlog of all delayed jobs processed and system restored
**Lessons/Improvements**
* Implementing new alerts to detect when scheduled job processing falls behind, enabling faster identification of similar issues
* Improving alert prioritization to reduce noise and ensure critical alerts are acted upon immediately
* Enhancing monitoring for downstream service dependencies
* Evaluating queue architecture changes to prevent a single conversation from impacting other customers \("noisy neighbor" isolation\)
* Investigating improvements to make our job scheduling service more resilient to backlogs
* Creating documentation of all services that depend on scheduled jobs to better understand incident ripple effects
[OUTBOUND WEBHOOKS] Outbound Webhooks API errors in Prod2 and Prod4
Başlangıç 10 Aralık 2025 03:38 UTC · 1h 14m
Pending
Etkilenen bileşenler
Web/Email/Form HooksWeb/Email/Form Hooks
identified
Kustomer is aware of an event affecting outbound webhook delivery in our Prod2 environment. Customers may experience failures due to 404 and 503 errors when the platform attempts to send outbound webhooks.
Our team is actively investigating the cause and working toward a resolution.
Please expect additional updates within the next 30 minutes, and reach out to Kustomer Support at [email protected] if you have additional questions or concerns.
identified
We are continuing to work on a fix for this issue.
monitoring
Kustomer has implemented a fix for the event impacting outbound webhooks in the Prod2 and Prod4 environments. We are backporting this fix across environments, and our team is monitoring to ensure service stability and continued recovery.
Please expect additional updates within the next 30 minutes, and reach out to Kustomer support at [email protected] if you have additional questions or concerns.
resolved
Kustomer has resolved the event that impacted outbound webhooks in the Prod2 and Prod4 environments. The backported fix has been fully deployed, and webhook delivery is functioning as expected across all affected environments. Our team has verified stability, and no further impact is anticipated.
If you have additional questions or concerns, please reach out to Kustomer support at [email protected]
postmortem
### **Summary**
On December 9, 2025, Kustomer experienced an interruption to outbound webhook delivery affecting customers in our **prod-2 \(EU\)** and **prod-4 \(IN\)** production regions. During this time, some outbound webhook requests were unable to reach customer endpoints, resulting in delivery failures and error responses.
The disruption occurred during a routine platform update related to our deployment systems. An underlying inconsistency from a previous infrastructure change caused a required backend service to become temporarily unavailable while updates were in progress. As a result, webhook endpoints could not accept or process delivery attempts for a limited period.
Engineering teams identified the issue quickly and restored full service shortly thereafter. Once the affected systems were brought back online, normal webhook delivery resumed and any queued events were successfully processed. There was no permanent data loss, and webhook functionality returned to normal operation across all impacted regions.
### **Impact**
* **Duration:** Approximately 1 hour and 40 minutes
* **Scope:** Production regions prod-2 and prod-4
* **Customer Impact:**
* Outbound webhook delivery failures during the incident window
* Webhook requests may have returned HTTP 503 or 404 responses
* Webhook events were delayed but not lost
### **Next Steps**
While safeguards already exist to protect core platform services, we are implementing additional improvements to reduce the likelihood of similar disruptions in the future. These include strengthening validation during platform updates, improving detection of incomplete infrastructure changes, and enhancing internal visibility when critical services are modified.
Reported Third Party Event affecting AIC and AIR Observability (PROD 1)
Başlangıç 2 Aralık 2025 22:44 UTC · 2h 5m
Pending
monitoring
Kustomer is aware of an event reported by one of our third party vendors affecting the observability of AI Agent for Reps and AI Agent for Customers that may cause trace logs to not show up in AI Agent logs within conversations.
Our team is monitoring the incident, and working with the vendor where possible to resolve the issue. Please expect further updates within the next 3 hours, and reach out to Kustomer support over Chat or Email if you have additional questions or concerns.
resolved
A third party incident has been resolved affecting PROD 1 AI conversations (both AIC and AIR) to not include trace logs. Functionality has returned in this area
Please reach out to Kustomer support through Chat or Email if you have additional questions or concerns.
[POSTMARK] ISP errors sending to Gmail addresses - all pods
Başlangıç 13 Kasım 2025 20:56 UTC · 7h 5m
IssuesKüçük çaplı olay
Etkilenen bileşenler
Channel - EmailChannel - EmailChannel - Email
investigating
Kustomer is aware of an event affecting kustomerapp.com email addresses that may cause messages sent to gmail.com addresses to bounce.
Our team is currently working to identify the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support [email protected] for any further questions or updates.
investigating
Kustomer is aware of an event affecting kustomerapp.com email addresses that may cause messages sent to gmail.com addresses to bounce.
Our team is continuing to work toward identifying the cause of this issue in an effort to implement a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support [email protected] for any further questions or updates.
investigating
Kustomer is aware of an issue impacting emails sent from @kustomerapp.com addresses, which may result in delivery failures to Gmail recipients.
Our team is actively investigating the cause and working to implement a resolution. We will provide another update within 30 minutes.
For further questions or updates, please contact Kustomer Support via Chat or Email.
investigating
Kustomer continues to investigate an issue affecting emails sent from @kustomerapp.com addresses, which may cause delivery failures to Gmail (@gmail.com) recipients.
Our team is working diligently to identify the root cause and implement a fix. We will share another update within 30 minutes.
For any questions or further assistance, please contact Kustomer Support via Chat or Email.
investigating
Kustomer is still investigating an ongoing issue impacting emails sent from @kustomerapp.com addresses, which may result in delivery issues when reaching Gmail (@gmail.com) recipients.
Our team continues to work toward identifying the root cause and implementing a resolution. We will provide another update within 30 minutes.
For any questions or additional support, please reach out to Kustomer Support via Chat or Email.
identified
Kustomer has identified an event affecting mail.kustomerapp.com email addresses that may cause messages sent to gmail.com addresses to bounce.
Our team is still continuing to work on implementing a resolution. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat or Email for any further questions or updates.
identified
Kustomer has identified an event affecting mail.kustomerapp.com email addresses that may cause messages sent to gmail.com addresses to bounce.
A fix is being rolled out incrementally to customers and we continue to monitor the situation as it improves. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat or Email for any further questions or updates.
identified
A fix is being rolled out incrementally to customers and we continue to monitor the situation as it improves. Please expect additional updates within the next 30 minutes, please reach out to Kustomer Support via Chat or Email for any further questions or updates.
monitoring
Kustomer has implemented a fix, which we are currently monitoring, to address the event affecting [POSTMARK] ISP errors.
Our team is currently monitoring this update to ensure the issue is fully resolved. Please expect further updates within the next 30 minutes, and reach out to Kustomer support via Chat or Email if you have additional questions or concerns.
resolved
Issues encountered during sending emails to gmail addresses have been resolved. Within the hour, the issue should be fully resolved as changes propagate throughout DNS records. If you experience any further problems, please reach out to support and we'll look into it with high priority.
postmortem
# **Summary**
New, tightened bulk sender rules by Gmail went into effect in November 2025, which affected Kustomer’s ability to deliver transactional emails to Gmail mailboxes. Specifically, DMARC alignment failures due to messages not being aligned to the from domain resulted in emails sent from [mail.kustomerapp.com](http://mail.kustomerapp.com) domain to [gmail.com](http://gmail.com) addresses to be blocked.
**Root Cause**
Beginning in early November 2025, Gmail rolled out stricter requirements for organizations that send large volumes of email. Because Kustomer sends more than 5,000 messages per day to Gmail addresses, Gmail classifies our traffic as “bulk sender” traffic and applies stronger authentication checks.
Many of the emails sent from Kustomer customers through Postmark were not fully meeting Gmail’s updated authentication requirements. Although the messages were being sent from addresses that looked like they belonged to each customer \(eg: [[email protected]_](mailto:[email protected])\), the DMARC policy check was now failing. As a result, Gmail could not verify that these messages truly came from the domain shown in the “From” address.
When Gmail applied its new enforcement rules, any email that did not clearly authenticate as coming from the correct domain was blocked before delivery. This caused a spike in rejected emails for customers sending to Gmail inboxes.
# **Timeline**
**Early November 2025**
Gmail begins enforcing newly updated bulk-sender requirements for high-volume senders. These rules require stricter domain authentication for all messages sent to Gmail inboxes.
**Nov 11, 2025**
**2:18 PM EST — Initial Reports**Customer support teams begin receiving reports that emails sent to Gmail addresses are bouncing or failing to deliver.
**2:46 PM EST — Issue Confirmed**
Responding teams identify that the issue is widespread and related to Gmail’s new enforcement behavior.
**3:11 PM EST — Root Cause Narrowed Down**Misalignment in required email authentication settings is identified as the likely cause of Gmail rejecting messages.
**4:00 PM EST - 7:45 PM EST – Fix Validation**
Responding engineers work on implementing and validating the fix.
**8:01 PM EST — Fix deployment**The necessary configuration updates are confirmed, including correcting DNS records so Gmail can properly verify the sender.
**10:00 PM–11:00 PM EST — Remediation & Stabilization**Updated DNS records are applied across affected customer domains. Gmail begins accepting previously rejected messages, and email delivery to Gmail addresses returns to normal.
**Lessons/Improvements**
**Improved Email Authentication Setup Across All Customers:** All customer organizations now have the correct DNS records and Postmark sender signatures in place to ensure their messages meet Gmail’s authentication requirements.
**Enhanced Monitoring for Email Delivery Issues:** We are implementing new logging and alerting to detect spikes in email bounces earlier, allowing us to respond more quickly if delivery issues arise in the future.
**Stronger Setup Process for New Integrations:** Going forward, all new Postmark email integrations will be configured with the correct authentication settings from the start to prevent this issue from occurring.