The latency and error rate issues that AWS is currently experiencing is now affecting the performance of the Mezmo platform. We will do everything we can to mitigate the impact and recover as quickly as possible once AWS has resolved their issues. See here for more detail on what is happening with AWS: https://health.aws.amazon.com/health/status
monitoring
AWS is reporting that services on their end are recovering, and we are seeing evidence to support that on our end.
monitoring
We are continuing to monitor for any further issues.
monitoring
We are continuing to monitor our platform for any further issues during AWS' recovery.
resolved
Mezmo is no longer seeing impacts from the AWS events of today. We will continue to monitor our services and respond if we notice any abnormalities.
Search performance issues
์์ 2025๋ 9์ 15์ผ PM 8:18 UTC ยท 58m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
LivetailSearch
identified
We have identified that there are some performance issues related to search in Mezmo Log Analysis. We are currently working to resolve the problem. Your data is still be ingested, indexed, and retained but are slow to display in the live tail because of the issues with Search. Please subscribe to the page to stay updated.
monitoring
Search performance should be returning to normal. You may notice some degraded performance.
resolved
The issues with Search in Live Tail have been resolved, and Search performance is back to normal.
Problems with Search
์์ 2025๋ 9์ 9์ผ PM 8:49 UTC ยท 3h 3m
Pending
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
LivetailSearch
identified
Apologies Mezmo customers as we are again experiencing problems with our Search Engine and live tail. We have identified the problem and are actively working to fix it now.
identified
We are continuing to work on a fix for this issue.
monitoring
At this time services are restored. Mezmo continues to perform diagnosis to identify and address the root causes or triggers for degradations in service, with engineering staying on high alert for any recurrences. We sincerely apologize for the impacts this incident may have had to your services.
resolved
This incident has been resolved.
Issue with Search
์์ 2025๋ 9์ 8์ผ PM 7:55 UTC ยท 1h 58m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Search
investigating
Mezmo has identified an issue with our search engine which is preventing events from being displayed in live tail. We are actively working on the problem and expect to have it resolved shortly. Please subscribe to this page for updates on this issue.
monitoring
Impacts from this issue are currently mitigated. Mezmo engineering continues to investigate the underlying causes of this issue and will leave this ticket in a monitoring state for now.
resolved
Mezmo has made a backend change that we believe should mitigate this issue in the future. Mezmo will continue to proactively monitor the platform for any signs of recurrence. Please reach out through our normal support channels if you are still experiencing issues at this time or need further information about this incident. We apologize for the impacts this incident may have had to your services.
Intermittent issues
์์ 2025๋ 6์ 9์ผ AM 10:02 UTC ยท 42m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Web App
investigating
We are experiencing intermittent issues with the Web App
identified
The problem has been identified and our infrastructure team is working on the resolution
monitoring
A fix has been implemented. The backlog is being processed.
resolved
The issue has been resolved. The processing of the backlog is progressing, and all delayed logs will be available shortly
Log availability and search performance temporarily degraded
์์ 2025๋ 4์ 14์ผ PM 8:12 UTC ยท 2h 17m
Pending
investigating
The cause is known and is being actively mitigated.
resolved
This incident has been resolved.
Issues accessing logs in boards and screens
์์ 2025๋ 4์ 11์ผ PM 7:17 UTC ยท 1h 47m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Web App
investigating
We are currently investigating this issue.
resolved
This incident has been resolved. All services are fully operational.
Intermittent Pipeline Service Degradation
์์ 2025๋ 4์ 2์ผ PM 8:08 UTC ยท 2h 9m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Ingestion / Sources
investigating
We are currently experiencing intermittent issues with our Pipeline service and are actively investigating the matter.
Some customers may experience delays in the Pipeline data.
resolved
The Pipeline service has resumed. All services are fully operational.
Unable to view the logs
์์ 2025๋ 3์ 18์ผ PM 1:33 UTC ยท 56m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Web AppSearch
investigating
We are currently investigating the issue.
identified
The issue has been identified and the fix implemented.
monitoring
We are monitoring the performance. There is some lag and once it's been fully processed all the recent logs will be available.
resolved
This incident has been resolved.
New logs are not available for some accounts
์์ 2024๋ 12์ 12์ผ PM 9:51 UTC ยท 24m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Search
investigating
New logs being submitted are not available for some accounts. We are investigating.
monitoring
A fix has been implemented and new logs are now available for the accounts. We are monitoring the results.
resolved
New logs are available for all accounts. This incident has been resolved.
postmortem
**Dates:**
Start Time: Thursday, December 12, 2024 at 19:24 UTC
End Time: Thursday, December 12, 2024 at 20:54 UTC
Duration: 1 hour and 30 minutes
โ
**What happened:**
Log lines submitted to Mezmo for ingestion into Log Analysis were never made available in our WebUI for Searching, Graphing, and Timelines โ neither during the incident, nor afterwards. This affected a significant number of accounts. Log lines were still passed through Pipeline and made available at all times within Live Tail. The log data still triggered both Telemetry Pipeline and Log Analysis based alerts, and were still archived in both places.
โ
**Why it happened:**
A single pod within our indexing service ran out of disk space. This was due to a sudden increase in the volume of log lines sent to Mezmo by accounts that happened to be assigned to this pod.
Our pods are configured to limit how much disk space is available for writing new data; these limits should have prevented any pod from running out of disk space, in any scenario. After the incident, we discovered that the limits had been configured incorrectly, which explains why it was possible for this pod to run out of disk space.
Our service is designed to tolerate the loss of a single pod within its indexing service without any widespread impact to customers. Instead, for reasons still under investigation, the impact expanded to many other pods.
Our service is also designed to retain submitted log lines, even if the indexing portion of our service cannot process them immediately; these log lines can be indexed later, when the service is functional again. In this incident, however, the log lines were never indexed. The reason for this failure is also under investigation.
โ
**How we fixed it:**
We restarted the pod that had run out of disk space. It immediately had enough free disk space to accept new log lines for indexing. All other indexing pods also returned to a normal operational state.
โ
**What we are doing to prevent it from happening again:**
We have properly configured our pods to prevent them from running out of disk space. This step alone should prevent any recurrence of the same problem.
We have updated our monitoring to send high priority alerts when any indexing pod is in danger of running out of disk space. These alerts were in place before, but set to โlowโ priority; they did not come to our attention in time to prevent the incident.
We will continue to actively investigate why the impact spread to other indexing pods and why log lines were not retained for indexing in the future.
Intermittent Pipeline Service Degradation
์์ 2024๋ 12์ 9์ผ PM 5:32 UTC ยท 2h 56m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Destinations
investigating
We are currently experiencing intermittent issues with our Pipeline service and are actively investigating the matter.
resolved
The Pipeline service has resumed. All services are fully operational.
postmortem
**Dates:**
Start Time: Monday, December 9, 2024 at 17:32 UTC
End Time: Monday, December 9, 2024 at 20:28 UTC
Duration: 2 hours and 56 minutes
โ
**What happened:**
For some accounts, data sent to Mezmo Pipelines was slow to be processed and sent onwards to their destination, or was not processed at all for the duration of the incident.
โ
**Why it happened:**
We released a new version of the agent \(3.10.1\) that improves how logs lines are sent to the Mezmo service for ingestion. Most applications add newly written log lines by a process known as โappendingโ; by contrast, a small number use a process called โtruncationโ. The new agent version has improved its ability to handle log lines added to logs using truncation, particularly when log lines are written frequently.
Many accounts do not monitor any logs that use truncation. A few accounts do, but they write new log lines infrequently. However, a handful of customer accounts have applications that write log lines using truncation at a very high frequency. When these accounts upgraded to agent 3.10.1, there was a very large increase in the volume of data sent to the Mezmo Pipeline service.
The increase was detected by our monitoring. It also caused some pods in some Vector partitions to crash, which affected all accounts on the same partition. Data was still ingested and cached within our service, but it was not being quickly processed or sent on to destinations for all accounts.
โ
**How we fixed it:**
We temporarily paused the processing of newly ingested data, which stopped pods from crashing. We then identified the accounts \(approximately 20\) that were sending us an increased volume of data and moved them to a newly created Vector partition on a newly commissioned node. This allowed the accounts on all other Vector partitions to function normally again; we restarted Pipeline processing on their pods and data began to flow to destinations again.
For the remaining affected accounts, we applied exclusion rules within our service to remove redundant data and thereby reduce the overall volume of data. We also contacted the owners of the handful of accounts sending log lines at very high volumes and helped them apply exclusion rules within their locally running Mezmo agents. With these changes, the overall volume of ingested data was reduced and again able to be processed and sent to destinations. After monitoring these accounts and seeing no ill effects, we moved them back to the general pool of Vector partitions.
โ
**What we are doing to prevent it from happening again:**
We discovered our Vector partitions are configured to use more CPU cores than necessary; under high load, the increased CPU usage causes pods to crash. Accordingly, we will limit the number of cores available to Vector per node.
We will rebalance the number of accounts assigned to Vector partitions, aiming to assign less accounts to each one. This will reduce the impact of any similar incidents in the future.
We will explore how to rate limit data sent to Pipelines by individual agents in the future. Rate-limiting will prevent any impact in a similar incident in the future.
Pipeline Web UI is Unresponsive
์์ 2024๋ 10์ 26์ผ AM 1:00 UTC ยท 19h 23m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Web App
investigating
Our Pipeline WebUI is not loading pages. We are investigating.
investigating
The Pipeline WebUI is still unavailable. Ingress and egress are unaffected. Our engineers are investigating.
monitoring
The Pipeline WebUI is working now. We are monitoring.
monitoring
We are continuing to monitor for any further issues.
monitoring
The Pipeline WebUI is available, but at times pages are slow to load and metrics may be unavailable. We are taking remedial action and continuing to monitor.
monitoring
The Pipeline WebUI is available and loading pages normally. We're working to resolve the root cause permanently and continuing to monitor.
resolved
The Pipeline UI is now fully functional.
postmortem
**Dates:**
Start Time: Saturday, October 26, 2024 at 01:00 UTC
End Time: Saturday, October 26, 2024 at 20:23 UTC
Duration: 19 hours and 23 minutes
\(Note that customer impact was limited to the first four hours of the incident.\)
โ
**What happened:**
For approximately two hours, pages on the Web UI for Pipeline did not load. Afterwards, pages were able to load, but often only after delays of 15 seconds or longer. After another two hours, the Web UI returned to normal usage. The incident was kept open until all remediation was completed.
The Web UI for Log Analysis and the ingress and egress of Pipeline data were unaffected.
โ
**Why it happened:**
On Wednesday, October 23, 2024, we deployed a change to our Pipeline service, by which metrics on pipeline usage began to be saved to a postgres database. The same database also stores account configuration information; this information must be accessed to display pages in the Web UI.
The deployment on Wednesday changed the performance profile of the database, most notably in the number and frequency of writes. There was no immediate customer impact, but we noted that backups of the database were unable to run successfully because of the increased load.
Customer impact only began on Saturday, October 26th when an unrelated user action โ running a Profiler from the Web UI โ placed even more demands on the postgres database. It was unable to process queries and, according to its design, moved into โread-onlyโ mode to prevent any loss of data. At times the database was entirely inaccessible or very slow to respond. This made our Web UI unable to display pages, as it relies on configuration information stored in the database.
โ
**How we fixed it:**
We first enabled two replicas of the database, which allowed the Web UI to load pages again, albeit slowly.
We then deployed a new code change that removed all superfluous writes about metric usage to the postgres database. This reduced the number of queries and Web UI usage returned to normal. We kept the incident open while working on further remediation.
Finally we took steps to bring the postgres database back to a normal state. This remediation phase was complicated by the fact that a full backup had not completed successfully in the last two days. By the end of the incident, the database was operating normally and a full backup had completed.
โ
**What we are doing to prevent it from happening again:**
The problems with the deployment on Wednesday, October 23, 2024 only revealed themselves under the load of our production environment. To better simulate production, we will update our testing environment and processes to use higher data loads.
We will closely evaluate the current workload of the postgres database and consider making a separate database to store just metrics, rather than combining metrics and configurations in one location. Separate databases would have prevented Web UI pages from not loading, thus avoiding any customer impact.
We will re-evaluate our backup strategy for the postgres database, since this slowed down the remediation phase.
Intermittent user session timeouts, requiring periodic re-authentication
์์ 2023๋ 12์ 4์ผ PM 12:06 UTC ยท 1h 13m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Web App
investigating
The Web UI is currently encountering user session timeouts, prompting customers to log in every 1-2 minutes. Our team is actively investigating the root cause of this issue, while the remaining aspects of the service remain fully functional.
monitoring
We have implemented a fix for the user session timeouts on the Web UI, but will continue to monitor the situation closely.
resolved
The issue has been resolved, and no further issues have been observed with user sessions.
postmortem
**Dates:**ย
Start Time: Monday, December 4, 2023, at 10:29 UTC
End Time: Monday, December 4, 2023, at 12:01 UTC
Duration: 92 minutes
โ
**What happened:**
Web UI users were logged out frequently โ usually within 1-2 minutes of logging in. Users could successfully login again without any issues, but the session would expire shortly afterwards.
โ
**Why it happened:**
It was identified that both Web UI pods and the Redis database pods, which are responsible for storing user sessions, experienced a critical memory shortage, leading to uncontrolled data purging. When this same issue happened in July 2023, our engineering team deployed a fix that enhanced how Redis stores the user session keys. This fix successfully prevented any recurrence of the problem until today. The team is still determining what made it exceed the memory limit this time.
โ
**How we fixed it:**
Initially, the Web UI pods were restarted, but that did not resolve the problem permanently. The engineering team then restarted the Redis database pods and the session stopped expiring.
โ
**What we are doing to prevent it from happening again:**
The team will revise the previous fix, including implementing a mechanism for the pod to automatically restart upon reaching its limit and setting up alerts to notify an engineer when it's approaching that threshold.
Web UI is unresponsive and ingestion of log lines halted
์์ 2023๋ 8์ 29์ผ PM 9:01 UTC ยท 3h 16m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Web App
investigating
Our WebUI is not loading pages consistently. We are investigating. [Reference #3204]
investigating
The webUI is loading consistently now, but we are still investigating.
resolved
This incident has been resolved.
postmortem
**Dates:**
Start Time: 8:32 pm UTC, Tuesday August 29th, 2023
End Time: 10:04 pm UTC, Tuesday August 29th, 2023
Duration: 92 minutes
โ
**What happened:**
Our Kong Gateway service stopped functioning and all connection requests to our ingestion service and web service failed. The Web UI did not load and log lines could not be sent by either our agent or API. Log lines sent using syslog were unaffected.
Kong was unavailable for two periods of time: one lasting 27 minutes \(8:32 pm UTC to 8:59 pm UTC\) and another lasting 9 minutes \(9:43 pm UTC to 9:52 pm UTC\). Once Kong became available, the Web UI was immediately accessible again. Agents resent locally cached log lines \(as did any APIs implemented with retry strategies\). Our service then processed the backlog of log lines, passing them to downstream services such as alerting, live tail, archiving, and indexing \(which makes lines visible in the Web UI for searching, graphing, and timelines\). The extra processing was completed ~20 minutes after Kong returned to normal usage the first time, and ~10 minutes after the second time.
โ
**Why it happened:**
The pods running our Kong Gateway were overwhelmed with connection requests. CPU increased to a point that health checks started to fail and the pods were shut down. Weโve determined through research and experimentation that the cause was a sudden, brief increase in the volume of traffic directed to our service. Our service is designed to handle increases in traffic, but these were approximately 100 times above normal usage. The source\(s\) of the traffic are unknown. The increase came in two spikes, which correspond to the two periods when Kong became unavailable.
โ
**How we fixed it:**
We manually scaled up the number of pods devoted to running our Kong Gateway. During the first spike of traffic, we doubled the number of pods; during the second, we quadrupled the number. This certainly helped speed up the processing of the backlog of log lines sent by agents once Kong was again available. Itโs unclear whether the higher number of pods would have been able to process the spikes of traffic as they were happening.
โ
**What we are doing to prevent it from happening again:**
We are running our Kong service with more pods so there are more resources to handle any similar spikes in traffic. We will add auto-scaling to the Kong service so more pods are made available automatically as needed. Weโll also add metrics to identify the origin of any similar spikes in traffic.
User sessions are timing out and customers are required to login again
์์ 2023๋ 6์ 19์ผ AM 11:09 UTC ยท 2h 49m
Issues๊ฒฝ๋ฏธํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Web App
investigating
User sessions to our Web UI are timing out and customers using the UI have to log in every 1-2 minutes. We are investigating why this is happening, but the rest of the service is fully functional. No other components are affected.
identified
The issue has been identified, and a fix is being implemented.
monitoring
The fix was implemented and we are now monitoring the user login sessions.
resolved
This incident has been resolved.
postmortem
**Dates:**
Start Time: Monday, June 19, 2023, at 10:31 UTC
End Time: Monday, June 19, 2023, at 12:35 UTC
Duration: 124 minutes
**What happened:**
Users were being logged out of our WebUI frequently โ within 1-2 minutes of logging in. Users could successfully login again, but the new session would also expire quickly.
**Why it happened:**
The cache of logged in users held in our Redis database was being cleared every 1-2 minutes. This caused all user sessions to expire and new logins to be required. We have yet to ascertain why the cache was being periodically cleared at frequent intervals.
**How we fixed it:**
We restarted the pods running the Redis database and the cache behavior returned to normal.
**What we are doing to prevent it from happening again:**
We will investigate further to learn why the Redis cache was being frequently cleared.
The Web UI is not accessible
์์ 2023๋ 5์ 1์ผ PM 8:18 UTC ยท 9m
Outage์ฌ๊ฐํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Web App
identified
The Web UI is not accessible.
resolved
This incident has been resolved.
postmortem
**Dates:**
Start Time: Monday, May 1, 2023, at 19:55 UTC
End Time: Monday, May 1, 2023, at 20:11 UTC
Duration: 16 minutes
โ
**What happened:**
The WebUI was unresponsive, returning an error of โfailure to get a peer from the ring-balancer.โ
**Why it happened:**
All Mezmo services run within a service mesh. The portion of the mesh dedicated to the pods running our Mongo database began receiving many connection requests, more than its allocated CPU and memory could handle at once. This portion of the mesh \(which itself runs on pods\) quickly ran out of memory. This made the Mongo database unavailable to other services. The WebUI relies entirely on Mongo for account information and therefore became unresponsive, returning an error of โfailure to get a peer from the ring-balancer.โ
While the immediate reason for the incident is clear, the root cause is still unknown. We suspect there was a change in user usage patterns \(e.g. increased traffic, login attempts, etc\) which triggered the incident.
**How we fixed it:**
We removed the WebUI from the service mesh. The Mongo service has more CPU and memory resources allocated to it and was able to accept the high level of connection requests successfully. WebUI usage immediately returned to normal.
**What we are doing to prevent it from happening again:**
We will change the default settings for the service mesh to allocate more CPU and memory resources, permanently. Afterwards, we will add the Mongo service back to the service mesh.
Searches are running slowly
์์ 2023๋ 2์ 10์ผ PM 6:11 UTC ยท 3m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
Search
identified
Searches are running slowly. We have identified the cause and are implementing a fix.
(Reference # 3018)
resolved
Searches are running at normal speeds again. All services are fully operational.
postmortem
**Dates:**
Start Time: Friday, February 10, 2023 at 16:45 UTC
End Time: Friday, February 10, 2023 at 18:14 UTC
Duration: 89 minutes
**What happened:**
Searches returned results slowly or not at all. No data was lost and ingestion was not halted.
โ
**Why it happened:**
In a previous incident on February 6, 2023 \(more details at [https://status.mezmo.com/incidents/3yl9x1t7qcw5\),](https://status.mezmo.com/incidents/3yl9x1t7qcw5),) two pods storing logs were temporarily removed from the pool of pods available for receiving and inserting new batches of logs into our data store. The pods continued to return results for previously processed logs. We took this step because the pods had fallen behind on their tasks, which we believe was a consequence of an ungraceful shutdown during the incident. We gave the pods several days to catch up on tasks and then made them available for insertion of new logs into the data store again, a change we expected to have no impact.
One of the pods immediately began integrity checks to confirm the same data existed on its local disk and on our S3 storage. As a side effect of the previous incident, the pod incorrectly determined that data was missing from the local disk and began sending http requests to our S3 storage to locate the missing data. In fact, the data in question is designed to only reside on local disk and was not supposed to be stored on S3.
The requests failed with 404 errors when the data was not found on S3 \(as expected\). Every new attempt to retrieve search results generated another request. The rate of requests was high enough to slow down all requests related to search results within the podโs zone \(one out of three total\). This led to search results being returned slowly or not at all.
โ
**How we fixed it:**
We removed the pod from the pool available for receiving and inserting new batches of logs into our data store. The pod continued to return results for previously processed logs.
โ
**What we are doing to prevent it from happening again:**
We marked this pod to remain unavailable for new logs until all previously processed logs on the pod have passed their retention period, whose maximum is 30 days. At that time, the pod will be rebuilt and begin accepting newly submitted logs again.
Weโll fix the logic of our search engine so it doesnโt request data from S3 that is intentionally not stored there. This will prevent the widespread 404 errors that slowed down all searching, should a pod again incorrectly determine it is missing data from its local disk.
We have added alerting and monitoring to detect high latency in search speeds and the average time to compact newly inserted logs.
Intermittent delays loading Web UI and running searches
์์ 2023๋ 2์ 6์ผ PM 8:48 UTC ยท 5h 19m
Outage์ค๋ํ ์ธ์๋ํธ
์ํฅ์ ๋ฐ์ ๊ตฌ์ฑ ์์
LivetailWeb AppSearch
identified
We are seeing intermittent delays loading the Web UI and running searches. We are taking remedial action.
identified
We are continuing to work on a fix for this incident.
identified
Web UI pages are loading at normal speeds and searches are returning quickly again. A small amount of data (<2%) is not being returned in search results. We are working to restore access to the results.
resolved
All data is being returned in search results. All services are fully operational.
postmortem
**Dates:**
Start Time: Monday, February 6, 2023, at 20:05 UTC
End Time: Tuesday, February 7, 2023, at 00:30 UTC
Duration: 4 hours and 25 minutes
**What happened:**
Searches returned results slowly or not at all. Our Web UI was intermittently unresponsive, particularly for pages like Live Tail, Graphing, and Timelines. No data was lost and ingestion was not halted.
**Why it happened:**
We initiated an upgrade of all nodes in our service, including the nodes that store logs. Pods were gradually moved to other nodes and restarted, so as to prevent any interruption in service.
A single pod that stores logs did not restart normally. Upon investigation, we found that it had not shut down cleanly and some files essential to a normal startup had not been written to disk. More significantly, we discovered that all nodes that store logs were using a podManagementPolicy of โorderedReadyโ \(the default setting\). This forced pods to restart in an ordered sequence. The single pod that would not restart was in the middle of the sequence; all the pods later in the sequence followed the policy and did not start either. In effect, about 25% of the pods within one zone \(out of the three zones devoted to storing logs\) were unable to start.
The remaining pods in the zone were forced to take on extra work, such as accepting new logs, compacting data, and answering queries from our internal APIs. This led to slow searches and slow load times for any part of the Web UI that displays data about logs.
**How we fixed it:**
We temporarily added more pods to run API calls to increase the odds of them succeeding. We changed the podManagementPolicy to โParallelโ to allow all pods to restart, regardless of their position in the ordered sequence for starting up. We made manual edits to the pod that had not restarted cleanly so it could start again. These steps brought search latency back to normal speeds and made API calls work again.
We cordoned off two pods that had fallen far behind in processing to allow them to recover without taking on new tasks. This temporarily removed ~2% of logs from all search results. When these pods were caught up with all pending tasks, we made them available again for search queries.
โ
**What we are doing to prevent it from happening again:**
We have changed the podManagementPolicy to โParallelโ for all nodes that store logs.
We will review the podManagementPolicy of all other areas of our service and make changes where appropriate.
We will add alerting and monitoring to detect high latency in search speeds and the average time to compact newly inserted logs.
Weโll explore options for adding more resources to each zone of pods, so they are less likely to fall behind on processing tasks when some pods are unavailable.
Weโll explore ways to prevent unclean shutdowns of pods when nodes are upgraded.
Degraded performance for WebUI, Ingestion, Alerting, Searching, Live Tail, Graphing, and Timelines
This incident has been resolved. All services are fully operational.
postmortem
**Dates:**
Start Time: Wednesday, October 5, 2022, at 14:27 UTC
End Time: Wednesday, October 5, 2022, at 14:45 UTC
Duration: 00:18
**What happened:**
The ingestion of logs was partially halted. The WebUI was mostly unresponsive and most API calls failed. Because many newly submitted logs were not being ingested, new logs were not immediately available for Alerting, Searching, Live Tail, Graphing, Timelines, and Archiving.
**Why it happened:**
We recently added a new API gateway - Kong - to our service, that acts as a proxy for all other services. We had gradually increased the amount of traffic directed through the API gateway over several weeks and seen no ill effects. Prior to the incident, only some of the traffic for ingestion wen through the gateway.
Kong was restarted after a routine configuration change. After the restart, all traffic for our ingestion service began to go through Kong. Our monitoring quickly revealed the Kong service did not have enough pods to keep up with the increased workload, causing many requests to fail.
**How we fixed it:**
We manually added more pods to the Kong service. Ingestion, the WebUI, and API calls began to work normally again. Once ingestion had resumed, LogDNA agents running on customer environments resent all locally cached logs to our service for ingestion. No data was lost.
**What we are doing to prevent it from happening again:**
We updated Kubernetes to always assign enough pods for the Kong API gateway service to be able to handle all traffic.
Weโll update the Kong gateway to more evenly distribute ingestion traffic across available pods.
We will adjust our deployment processes so pods are restarted more slowly, which will reduce the impact in a similar scenario.
Weโll explore autoscaling policies so more pods could be added automatically in a similar situation.
Some customers' logs are not currently being processed
Some logs submitted to our service in the last 1.5 hours have not been processed. We are taking remedial action now.
resolved
Newly submitted logs are now being processed and retained. Some logs submitted by some customers during the incident were discarded and not successfully retained. [Reference #2792]
postmortem
**Dates:**
Start Time: Thursday, June 30, 21:40 UTC
End Time: Thursday, June 30, 23:32 UTC
Duration: 1 hour and 52 minutes
**What happened:**
Some log lines for some customers were discarded by our service. The log lines were successfully accepted by our ingestion service, but a downstream service โ the parser โ removed some of them. All further downstream services, such as Alerting, Live Tail, Searching, and Archiving never received these logs. In some cases, lines were received by Live Tail and were appended with the phrase โ\(not retained\)โ.
The great majority of customers โ 94.2% โ were unaffected and had no log lines discarded. Approximately 3.5% had a relatively small number of log lines discarded. Approximately 2.3% had most or all of the log lines submitted during the incident discarded.
**Why it happened:**
We inadvertently released code into production that contained a bug in the parser service. This bug was known to us and in the process of being fixed in our development environment, but was not yet ready for release to production.
The parser service is where exclusion rules are applied to recently submitted log lines that have been ingested but not yet passed to downstream services \(e.g. Alerting, Live Tail, Searching, and Archiving\). The bug made the parser exclude log lines that matched rules for inactive exclusion rules.
This included exclusion rules made by customers in the past and then disabled. Customers with such rules had some log lines excluded: whichever lines matched the inactive rules. If those rules had the โPreserve these lines for live-tail and alertingโ option enabled, then the excluded lines would still be processed for alerts and appear in Live Tail with the phrase โ\(not retained\)โ appended. This affected 3.5% of our customer accounts.
The usage quota feature is implemented as a particular type of exclusion rule even though it is not presented in the UI as an exclusion rule. The bug made the parser exclude all log lines if the usage quota feature was enabled for an account. This affected 2.3% of our customer accounts.
Our monitoring did not detect the decrease in lines being passed from the parser to downstream services because the change was within the range of normal fluctuation rates. These rates vary significantly as traffic changes and as customers choose to enable/disable exclusion rules.
**How we fixed it:**
We reverted the last release of parser code to the previous version. Once the previous version was deployed to all pods running the parser service, log lines stopped being discarded.
**What we are doing to prevent it from happening again:**
We added a code level test to ensure inactive exclusion rules are never applied by the parser \(such tests are part of our standard operating procedure\).
We will review our release process to understand how the code containing the bug was moved into production and improve our processes to prevent a similar event in the future.