Canvas is not loading for customers in the US and EU regions and we have identified what we believe to be the cause and are rolling out a fix.
monitoring
We have deployed the fix and are seeing recovery of normal Canvas functionality.
resolved
This incident has been resolved.
US1 cold query outage
Έναρξη 10 Αυγούστου 2026 στις 1:01 μ.μ. UTC · 7h 34m
IssuesΜικρό περιστατικό
Επηρεαζόμενα στοιχεία
ui.honeycomb.io - US1 Querying
monitoring
From 11:36UTC to 12:43UTC, queries in US1 that read older data returned errors or timed out. Queries over recent data were un-affected, and no telemetry is being lost.
The issue has been identified, and a fix is in place. We are currently monitoring the situation to confirm resolution.
resolved
The fix is in place and we've confirmed resolution. Queries and SLOs are processing normally.
Querying issues
Έναρξη 31 Ιουλίου 2026 στις 2:30 μ.μ. UTC · 0m
IssuesΜικρό περιστατικό
resolved
Between 10:30 and 11:00 AM EDT, we experienced elevated querying failures and slowness. The impact has subsided, and we are currently investigating and monitoring the situation.
postmortem
Starting at 9:50 AM Central Time on July 31st we received alerts around querying health and latency. The cause was eventually tracked down to a clash of two different factors. First, we have been introducing Lambda Managed Instances \(LMI\) into our query architecture, which behave similar to classic lambdas in many ways and had been introduced as part of work around querying performance improvements with the goal of a better and faster user experience. Unfortunately, a difference between classic lambdas and LMIs we discovered is the 15 minute hard stop of classic lambdas does not exist on LMIs, which means that certain assumptions built into our querying system no longer held. At the time of the incident, our newer Canvas feature was running batches of queries for a customer which should have a self imposed one minute time out. This timeout was not properly communicated back to the lambda, and while with classic lambdas the task would automatically cut off at 15 minutes, the new LMIs were not enforcing this sort of safety mechanism. As a result, some large queries were running for up to 30 minutes, driving up query latency across the board and resulting in errors for users trying to run fresh queries. The system recovered by 10:10 AM CT, 20 minutes later, when the large queries cleared out, but we continued to investigate the cause of the incident.
Once we identified this weakness in our querying architecture we began work to fix it and prevent the same incident from occurring again. We have shipped all associated incident follow ups.
We are seeing triggers intermittently failing in the EU region. We are actively investigating the cause.
investigating
We are continuing to see issues with triggers and are also noticing issues with querying. We are investigating for the cause of both.
monitoring
We have identified the source of the issue, applied a temporary fix, and are working on a permanent fix. We continue to monitor the situation.
monitoring
We are seeing a regression in performance with partial degradation of both triggers and querying. We are working to remediate.
monitoring
The partial degradation has been resolved for all customers except for those specifically contacted. We are working on mitigation measures so that we can fully restore trigger functionality for everyone.
monitoring
We are continuing to monitor for any further issues.
resolved
The degradation in querying and triggers has been resolved and we're back to full functionality.
honeycomb.io marketing website not working
Έναρξη 2 Ιουλίου 2026 στις 6:10 μ.μ. UTC · 6h 21m
OutageΣοβαρό περιστατικό
Επηρεαζόμενα στοιχεία
www.honeycomb.io
investigating
We are currently investigating this issue.
monitoring
Clearing the CDN cache appears to have solved the issue.
resolved
This incident has been resolved.
Activity Log delayed in US
Έναρξη 29 Ιουνίου 2026 στις 10:07 μ.μ. UTC · 2h 40m
IssuesΜικρό περιστατικό
Επηρεαζόμενα στοιχεία
ui.honeycomb.io - US1 Activity Log
monitoring
Starting at 14:24 PT, there was an issue with the Activity Log causing events to stop being processed. You may notice a gap starting at that time. We are slowly backfilling these events and do not expect any data loss to occur. We expect the full data log to be restored around 17:30 PT.
resolved
Backfill of the Activity Log events has completed. All events during the affected period should be present.
Ingest outage in EU
Έναρξη 25 Ιουνίου 2026 στις 6:10 μ.μ. UTC · 2h 32m
We are investigating an issue with delayed ingest in the US and EU region. Our engineers are rolling back a recent deploy. SLOs and Trigger evaluations may also be delayed.
monitoring
We have rolled back the deploy and ingest service has recovered. There has been an ingest outage from 17:55 - 18:14 UTC. We have also observed an delay for Service Maps.
monitoring
Triggers, SLOs and Service Maps are recovered. We are continuing to monitoring the Ingest Service
resolved
All services have been fully recovered.
Impact window: 17:55–18:17 UTC
Regions affected: Primary impact in US. The EU region was affected to a much lesser degree.
During this period, the following effects may have occurred:
Ingest: Some data was not ingested, resulting in complete or partial data loss for events sent during the window.
Triggers: Triggers may have failed to fire within the impact window.
SLOs: SLI values across the affected window are skewed by the missing data and may show artificial dips or accelerated budget burn.
Service Maps: Maps covering the outage window are incomplete. Services and dependencies may be under-counted, as their traces were not ingested.
We apologize for the disruption. Please reach out if you have any questions about how this may have affected your data.
postmortem
On June 24, we experienced approximately 30 minutes of severe data ingestion degradation in Honeycomb’s US and EU instances. During this degradation, Honeycomb rejected a significant percentage of inbound telemetry, and customers would have seen up to a half-hour gap in ingested telemetry. Additionally, during this same time window, customers would have experienced data processing degradation and service disruption within Anomaly Detection and Service Maps.
A deployment containing a change to how services retrieve dataset schemas into local caches caused a cascading set of failures, resulting in failed remote cache retrievals, and subsequently, a MySQL stampede due to several services falling back to retrieving schemas from the database. This stampede resulted in significant memory utilization growth, causing some services to crash loop with out-of-memory exceptions while retrieving schemas. Customers would have seen 5xx errors from our ingestion APIs when this occurred, because ingestion services were among those that were crash looping. This also included crash looping from services responsible for Anomaly Detection and Service Maps.
Right after these service crash loops began, alerts fired for the crashing services, and procedures were run to pin all services back to the last known good build before the breaking change landed. Once services were running the previous deploy’s build, the 5xx error responses from ingestion APIs returned to nominal baseline levels, and services that were crash looping recovered and resumed normal operation.
The underlying issue centered on how dataset schema cache payloads and metadata were serialized to memcached for remote cache retrieval, and how memcached clients deserialize that data before writing it to a local cache. The code change contained a feature flag to toggle the serialization behavior, and backwards-compatibility safeguards for clients deserializing different schema versions from memcached. However, services’ schema deserialization prior to the feature flag flip on did not behave as intended, treating each backwards-compatible schema read as a hard error rather than as a fallback, which subsequently caused every remote schema cache read to fall through to the database. Subsequent analysis and review of the code identified the issue and we have re-deployed this change with a fix along with additional tests. Services now correctly and successfully deserialize schemas from memcached in a safe, backwards-compatible manner.
MCP tool access degraded
Έναρξη 19 Ιουνίου 2026 στις 4:39 μ.μ. UTC · 42m
IssuesΜικρό περιστατικό
investigating
We are aware of an issue affecting Honeycomb MCP. Write tool functionality may be degraded. Read tools are unaffected.
monitoring
We are aware of an issue affecting Honeycomb MCP. Write tool functionality may be degraded. Read tools are unaffected.
resolved
MCP Write Tools have been restored. MCP users may need to reconnect to see all available tools.
Beginning June 17th, enhance-related requests were failing to complete. We've since identified and deployed a fix, and enhance is now fully operational. Thank you for your patience.
Querying Issues
Έναρξη 17 Ιουνίου 2026 στις 2:01 μ.μ. UTC · 7h 13m
IssuesΜικρό περιστατικό
Επηρεαζόμενα στοιχεία
ui.eu1.honeycomb.io - EU1 Querying
investigating
We are continuing to investigate intermittent slowness and query failures affecting our Production EU Region. We will provide an update as soon as we have more information.
monitoring
Querying in our Production EU Region is fully operational. We will continue to monitor for any abnormal behavior.
resolved
We have identified the root cause of the intermittent slowness and query failures affecting our Production EU Region. The issue was triggered by a large volume of dataset deletions that caused elevated database load, resulting in query instability. No data was lost during this incident. Service has been restored. A fix has been applied to help prevent recurrence.
Querying Issues in EU
Έναρξη 17 Ιουνίου 2026 στις 12:01 μ.μ. UTC · 31m
IssuesΜικρό περιστατικό
Επηρεαζόμενα στοιχεία
ui.eu1.honeycomb.io - EU1 Querying
investigating
We are currently investigating the cause of slowness and query failures in our Production EU Region.
resolved
The issue is now resolved and querying is back to fully functional.
Elevated API errors in production-eu1
Έναρξη 11 Ιουνίου 2026 στις 3:30 μ.μ. UTC · 0m
IssuesΜικρό περιστατικό
resolved
Requests to Honeycomb's Query Data and Management APIs in the production-eu1 environment encountered increased error rates between approximately 15:35 and 22:24 UTC. Event ingestion was not affected. The issue has been resolved.
We are investigating an issue with delayed ingest in the EU region. Received events are still being stored, but there may be a delay in event retrieval in queries. SLOs and Trigger evaluations may also be delayed.
investigating
Queries should no longer be delayed. SLOs and Trigger evaluations remain impacted.
identified
The issue has been identified and a fix is being implemented.
We have identified and are working to resolve an issue that is causing query results to return inconsistent results.
investigating
Impact from this issue has concluded.
For a period of time (different between US and EU instances, noted below), queries spanning data older than the most recent 2 hours with certain GROUP / WHERE clauses returned inconsistent results. Data was never lost, and reruns of the same queries will show the previously missing data.
US impact times: 21:50 UTC - 23:20 UTC
EU impact times: 20:50 UTC - 23:20 UTC
resolved
This is an administrative update, marking the incident as Resolved. There has been no further impact since 23:20UTC on May 21 2026.
activity log for US instance is delayed
Έναρξη 8 Μαΐου 2026 στις 8:00 π.μ. UTC · 3d 8h
IssuesΜικρό περιστατικό
Επηρεαζόμενα στοιχεία
ui.honeycomb.io - US1 Activity Log
identified
Activity log delay is rising, and is more than an hour behind due to a MySQL replica failover. Estimated recovery by 04:00 PDT.
monitoring
Backfilling of the past 2h30m of data is now in progress and should complete shortly.
monitoring
We believe that the activity log should now be caught up.
monitoring
We are continuing to monitor for any further issues.
monitoring
All activity log streams are now caught up, except for the query runs table, which should have all data since the start of the incident backfilled by 1600 PDT.
monitoring
Activity Log recovered at 07:00 UTC-07:00 today (about 6.5h ago). All streams are caught up.
resolved
This incident has been resolved.
Gap in Activity Log data in the EU region.
Έναρξη 1 Μαΐου 2026 στις 1:00 μ.μ. UTC · 0m
Pending
resolved
The database failover that caused query issues earlier in the EU (https://status.honeycomb.io/incidents/n855d8kzp32y) has also had knock-on effects on our Activity Log beta feature. Due to recovery issues on the data replication flow used for this mechanism, activity log events will be missing from 12:45 UTC until 19:45 UTC, representing a 5 hour gap in events.
We have identified configuration parameters that will be adjusted to reduce the likelihood of such losses in the future.
Query interuption due to database failover in the EU region
Έναρξη 1 Μαΐου 2026 στις 12:30 μ.μ. UTC · 0m
OutageΣοβαρό περιστατικό
resolved
At 12:50, our main database underwent an automated failover. This failover led to issues with some of our query engine connection pools, and caused queries to fail for roughly 10 minutes before self resolving.