Batch API outage
- identified
Sorunlar tespit edildi ve bir düzeltme üzerinde çalışıyoruz
- identified
Bu sorun için bir düzeltme üzerinde çalışmaya devam ediyoruz.
- resolved
Bu olay çözüldü.
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
50 Kickbox incidents · Nisan 2015 — official updates, affected components, duration and resolution details.
Sorunlar tespit edildi ve bir düzeltme üzerinde çalışıyoruz
Bu sorun için bir düzeltme üzerinde çalışmaya devam ediyoruz.
Bu olay çözüldü.
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
This incident has been resolved.
Currently investigating issue with web application and batch API.
The issue has been identified and a fix is being implemented.
We are continuing to work on a fix for this issue.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
Currently investigating issue.
We are still investigating an intermittent issue with list downloads.
Incident resolved. Monitoring will take place for a time to ensure functionality.
We are actively addressing an issue causing inconsistencies in the verification results for Yahoo and AOL emails. This problem has been traced to a breaking change in a recent update from Yahoo/AOL that affects our results. We have identified the issue and we are working towards a resolution.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
Currently investigating an issue that is causing an outage on our EU instances.
We are continuing to investigate this issue.
This incident has been resolved.
Issue currently under investigation.
Fix implemented and results being monitored.
Incident resolved.
We are currently investigating and issue causing slower than normal API responses for a fraction of requests
Issue identified. Preparing deployment for fix.
A fix has been deployed. We are now monitoring services.
This incident has been resolved.
Partial outage. Investigating
We are continuing to investigate this issue.
This incident has been resolved.
Team is investigating the cause of a minor service degradation occurring between 2:30pm and 3:00pm (central time). Details to come.
# August 21, 2024 Service Degradation Investigation **Date**: 08/21/2024 **Status**: Mitigated ### Summary A service degradation was detected by our monitoring services between 12:34 PST and 12:48 PST. We removed the impacted servers from rotation restoring full-service. We investigated and solved the issue before bringing the impacted servers back into production rotation. ### Impact API service degradation in 45-second bursts for a 12 minute period, beginning at 12:34 PST causing some requests to respond with a 502 ‘Bad Gateway’. The issue was limited to a small subset of our production servers, so fewer than 3% of global API requests during this time were affected. ### Root Cause\(s\) During a routine canary upstream service update, a small number of canary servers became unable to connect to an upstream internal service. Due to this problem, a fraction of requests to these servers caused a connection timeout, resulting in 502 Bad Gateway errors from our proxies.This problem was sporadic, so the proxies did not automatically remove these servers from rotation. ### Resolution Our internal monitoring systems detected this degradation and we worked to manually remove these servers from production rotation. This restored the API to full health. After our investigation, we resolved the communication issues with a full code re-deploy and service restart. We then ran a full suite of API health checks, confirmed the problem was resolved, and re-added the servers into production. ### Action Items We will perform a review of our internal monitoring systems to decrease the time to alert, and review our proxy health-check endpoints to see if they can be improved for partial service degradation. ## Timeline
We’re currently experiencing reduced performance with several of our integrations, including Iterable, HubSpot, and Klaviyo within Kickbox, specifically affecting functionality; our team is actively working to resolve these issues to restore full service. Users can still actively leverage our email verification services through list verification or API. If you have any additional questions, please reach out to [email protected].
This incident has been resolved.
A 7 minute unexpected partial outage occurred. Root cause still under investigation.
We are currently investigating an issue with the EU verification instance
We are continuing to investigate this issue.
A fix has been implemented and we are monitoring the results.
A networking error caused a severed connection to the EU verification engine. Manual reconnection has restored services.
A DNS change with an Amazon service temporarily broke sign-in for some users Sunday morning (API unaffected.
There was a small outage 10 minute outage due to an unforeseen network issue during a routine deployment.
A minor configuration change \(adding an alias to a Redis instance\) caused an unexpected error and prevented the Kickbox app from restarting after deployment. The issue was quickly detected and resolved by the engineering team. The total outage was 11 minutes. The misconfiguration was identified quickly and a rollback tag was deployed to restore service within minutes.
We're investigating a period of degraded service beginning at 5:54AM CST and lasting until 7:30AM CST.
We had degraded service due to a partial datastore service degradation. We repaired and restarted the datastore. No data was lost.
An issue with a upstream service provider created a window of unplanned downtime for a small fraction of EU users (smaller than 1%). The partial service outage lasted from 3pm September 19th until about 3pm the next day.
In the deliverablity tools suite, DMARC data has stopped populating as of September 1st. Engineering is investigating the issue and we hope to have more information and a resolution to this issue soon.
Still investigating the case.
This incident has been resolved.
Our system failed to properly log and summarized inbound aggregate DMARC reports. Our engineers are aware of the issue and are actively working to address this issue as soon as possible.
We are continuing to work on a fix for this issue.
We’ve corrected the issue with DMARC that was preventing ongoing data collection. Ongoing, inbound DMARC reports should be properly received, parsed, logged and displayed in reporting. If you have any questions, please don’t hesitate to reach out to us at [email protected]
Just after 11:00 am (central) a routine TLS certificate update caused a misconfiguration. Some users experienced a partial outage.
The new SSL cert that was implemented did not support `TLS1.2` . This resulted in issues for only some users where that support was required.