Canvas no está cargando para clientes en las regiones de EE.UU. y de la UE y hemos identificado lo que creemos que es la causa y estamos sacando una solución.
monitoring
Hemos implementado la solución y estamos viendo la recuperación de la funcionalidad normal de Canvas.
resolved
Este incidente ha sido resuelto.
Traducido automáticamente desde la actualización oficial del incidente.
Salario de consultas frías
Comenzó August 10, 2026 at 1:01 PM UTC · 7h 34m
IssuesMinor incident
Componentes afectados
ui.honeycomb.io - US1 Querying
monitoring
De 11:36UTC a 12:43UTC, las consultas en EE.UU.1 que leyeron datos antiguos devolvieron errores o se timed out. Las consultas sobre datos recientes no se vieron afectadas, y no se está perdiendo telemetría.
Se ha identificado la cuestión y se ha establecido una solución. Actualmente estamos monitoreando la situación para confirmar la resolución.
resolved
La solución está en su lugar y hemos confirmado la resolución. Las consultas y las SLO se procesan normalmente.
Traducido automáticamente desde la actualización oficial del incidente.
Cuestiones de consulta
Comenzó July 31, 2026 at 2:30 PM UTC · 0m
IssuesMinor incident
resolved
Entre las 10:30 y las 11:00 AM EDT, experimentamos altas fallas de búsqueda y lentitud. El impacto ha disminuido, y actualmente estamos investigando y monitoreando la situación.
postmortem
A partir de las 9:50 AM Central Hora del 31 de julio recibimos alertas sobre la búsqueda de salud y latencia. La causa fue finalmente rastreada hasta un choque de dos factores diferentes. En primer lugar, hemos estado introduciendo Lambda Managed Instances \(LMI\) en nuestra arquitectura de consulta, que se comportan de forma similar a las lambdas clásicas de muchas maneras y se habían introducido como parte del trabajo para buscar mejoras de rendimiento con el objetivo de una mejor y más rápida experiencia de usuario. Desafortunadamente, una diferencia entre las lambdas clásicas y los LMIs que descubrimos es la parada dura de 15 minutos de las lambdas clásicas no existe en los LMI, lo que significa que ciertas suposiciones construidas en nuestro sistema de consulta ya no se sostienen. En el momento del incidente, nuestra nueva característica de Canvas estaba ejecutando lotes de consultas para un cliente que debería tener un auto impuesto un minuto de tiempo fuera. Este tiempo no fue comunicado correctamente de vuelta a la lambda, y mientras que con los lambdas clásicos la tarea se cortaría automáticamente a 15 minutos, los nuevos IMC no estaban haciendo cumplir este tipo de mecanismo de seguridad. Como resultado, algunas grandes consultas se estaban ejecutando por hasta 30 minutos, conduciendo latencia de consultas a través de la junta directiva y resultando en errores para los usuarios que intentan realizar nuevas consultas. El sistema se recuperó en 10:10 AM CT, 20 minutos después, cuando las grandes consultas se retiraron, pero continuamos investigando la causa del incidente.
Una vez que identificamos esta debilidad en nuestra arquitectura de búsqueda comenzamos a trabajar para arreglarlo y evitar que el mismo incidente vuelva a ocurrir. Hemos enviado todos los seguimientos de incidentes asociados.
Traducido automáticamente desde la actualización oficial del incidente.
Estamos viendo desencadenantes que fallan intermitentemente en la región de la UE. Estamos investigando activamente la causa.
investigating
Continuamos viendo problemas con los desencadenantes y también estamos notando problemas con la búsqueda. Estamos investigando por la causa de ambos.
monitoring
Hemos identificado la fuente de la cuestión, aplicado una solución temporal y estamos trabajando en una solución permanente. Seguimos vigilando la situación.
monitoring
Estamos viendo una regresión en el rendimiento con degradación parcial de los desencadenantes y la búsqueda. Estamos trabajando para remediar.
monitoring
La degradación parcial se ha resuelto para todos los clientes excepto para aquellos específicamente contactados. Estamos trabajando en medidas de mitigación para que podamos restaurar plenamente la funcionalidad de los desencadenantes para todos.
monitoring
Continuamos monitoreando cualquier otro problema.
resolved
La degradación en la búsqueda y los disparadores se ha resuelto y estamos de vuelta a la funcionalidad completa.
Traducido automáticamente desde la actualización oficial del incidente.
honeycomb.io marketing website not working
Comenzó July 2, 2026 at 6:10 PM UTC · 6h 21m
OutageMajor incident
Componentes afectados
www.honeycomb.io
investigating
We are currently investigating this issue.
monitoring
Clearing the CDN cache appears to have solved the issue.
resolved
This incident has been resolved.
Activity Log delayed in US
Comenzó June 29, 2026 at 10:07 PM UTC · 2h 40m
IssuesMinor incident
Componentes afectados
ui.honeycomb.io - US1 Activity Log
monitoring
Starting at 14:24 PT, there was an issue with the Activity Log causing events to stop being processed. You may notice a gap starting at that time. We are slowly backfilling these events and do not expect any data loss to occur. We expect the full data log to be restored around 17:30 PT.
resolved
Backfill of the Activity Log events has completed. All events during the affected period should be present.
We are investigating an issue with delayed ingest in the US and EU region. Our engineers are rolling back a recent deploy. SLOs and Trigger evaluations may also be delayed.
monitoring
We have rolled back the deploy and ingest service has recovered. There has been an ingest outage from 17:55 - 18:14 UTC. We have also observed an delay for Service Maps.
monitoring
Triggers, SLOs and Service Maps are recovered. We are continuing to monitoring the Ingest Service
resolved
All services have been fully recovered.
Impact window: 17:55–18:17 UTC
Regions affected: Primary impact in US. The EU region was affected to a much lesser degree.
During this period, the following effects may have occurred:
Ingest: Some data was not ingested, resulting in complete or partial data loss for events sent during the window.
Triggers: Triggers may have failed to fire within the impact window.
SLOs: SLI values across the affected window are skewed by the missing data and may show artificial dips or accelerated budget burn.
Service Maps: Maps covering the outage window are incomplete. Services and dependencies may be under-counted, as their traces were not ingested.
We apologize for the disruption. Please reach out if you have any questions about how this may have affected your data.
postmortem
On June 24, we experienced approximately 30 minutes of severe data ingestion degradation in Honeycomb’s US and EU instances. During this degradation, Honeycomb rejected a significant percentage of inbound telemetry, and customers would have seen up to a half-hour gap in ingested telemetry. Additionally, during this same time window, customers would have experienced data processing degradation and service disruption within Anomaly Detection and Service Maps.
A deployment containing a change to how services retrieve dataset schemas into local caches caused a cascading set of failures, resulting in failed remote cache retrievals, and subsequently, a MySQL stampede due to several services falling back to retrieving schemas from the database. This stampede resulted in significant memory utilization growth, causing some services to crash loop with out-of-memory exceptions while retrieving schemas. Customers would have seen 5xx errors from our ingestion APIs when this occurred, because ingestion services were among those that were crash looping. This also included crash looping from services responsible for Anomaly Detection and Service Maps.
Right after these service crash loops began, alerts fired for the crashing services, and procedures were run to pin all services back to the last known good build before the breaking change landed. Once services were running the previous deploy’s build, the 5xx error responses from ingestion APIs returned to nominal baseline levels, and services that were crash looping recovered and resumed normal operation.
The underlying issue centered on how dataset schema cache payloads and metadata were serialized to memcached for remote cache retrieval, and how memcached clients deserialize that data before writing it to a local cache. The code change contained a feature flag to toggle the serialization behavior, and backwards-compatibility safeguards for clients deserializing different schema versions from memcached. However, services’ schema deserialization prior to the feature flag flip on did not behave as intended, treating each backwards-compatible schema read as a hard error rather than as a fallback, which subsequently caused every remote schema cache read to fall through to the database. Subsequent analysis and review of the code identified the issue and we have re-deployed this change with a fix along with additional tests. Services now correctly and successfully deserialize schemas from memcached in a safe, backwards-compatible manner.
MCP tool access degraded
Comenzó June 19, 2026 at 4:39 PM UTC · 42m
IssuesMinor incident
investigating
We are aware of an issue affecting Honeycomb MCP. Write tool functionality may be degraded. Read tools are unaffected.
monitoring
We are aware of an issue affecting Honeycomb MCP. Write tool functionality may be degraded. Read tools are unaffected.
resolved
MCP Write Tools have been restored. MCP users may need to reconnect to see all available tools.
Beginning June 17th, enhance-related requests were failing to complete. We've since identified and deployed a fix, and enhance is now fully operational. Thank you for your patience.
Querying Issues
Comenzó June 17, 2026 at 2:01 PM UTC · 7h 13m
IssuesMinor incident
Componentes afectados
ui.eu1.honeycomb.io - EU1 Querying
investigating
We are continuing to investigate intermittent slowness and query failures affecting our Production EU Region. We will provide an update as soon as we have more information.
monitoring
Querying in our Production EU Region is fully operational. We will continue to monitor for any abnormal behavior.
resolved
We have identified the root cause of the intermittent slowness and query failures affecting our Production EU Region. The issue was triggered by a large volume of dataset deletions that caused elevated database load, resulting in query instability. No data was lost during this incident. Service has been restored. A fix has been applied to help prevent recurrence.
Querying Issues in EU
Comenzó June 17, 2026 at 12:01 PM UTC · 31m
IssuesMinor incident
Componentes afectados
ui.eu1.honeycomb.io - EU1 Querying
investigating
We are currently investigating the cause of slowness and query failures in our Production EU Region.
resolved
The issue is now resolved and querying is back to fully functional.
Elevated API errors in production-eu1
Comenzó June 11, 2026 at 3:30 PM UTC · 0m
IssuesMinor incident
resolved
Requests to Honeycomb's Query Data and Management APIs in the production-eu1 environment encountered increased error rates between approximately 15:35 and 22:24 UTC. Event ingestion was not affected. The issue has been resolved.
We are investigating an issue with delayed ingest in the EU region. Received events are still being stored, but there may be a delay in event retrieval in queries. SLOs and Trigger evaluations may also be delayed.
investigating
Queries should no longer be delayed. SLOs and Trigger evaluations remain impacted.
identified
The issue has been identified and a fix is being implemented.
We have identified and are working to resolve an issue that is causing query results to return inconsistent results.
investigating
Impact from this issue has concluded.
For a period of time (different between US and EU instances, noted below), queries spanning data older than the most recent 2 hours with certain GROUP / WHERE clauses returned inconsistent results. Data was never lost, and reruns of the same queries will show the previously missing data.
US impact times: 21:50 UTC - 23:20 UTC
EU impact times: 20:50 UTC - 23:20 UTC
resolved
This is an administrative update, marking the incident as Resolved. There has been no further impact since 23:20UTC on May 21 2026.
activity log for US instance is delayed
Comenzó May 8, 2026 at 8:00 AM UTC · 3d 8h
IssuesMinor incident
Componentes afectados
ui.honeycomb.io - US1 Activity Log
identified
Activity log delay is rising, and is more than an hour behind due to a MySQL replica failover. Estimated recovery by 04:00 PDT.
monitoring
Backfilling of the past 2h30m of data is now in progress and should complete shortly.
monitoring
We believe that the activity log should now be caught up.
monitoring
We are continuing to monitor for any further issues.
monitoring
All activity log streams are now caught up, except for the query runs table, which should have all data since the start of the incident backfilled by 1600 PDT.
monitoring
Activity Log recovered at 07:00 UTC-07:00 today (about 6.5h ago). All streams are caught up.
resolved
This incident has been resolved.
Gap in Activity Log data in the EU region.
Comenzó May 1, 2026 at 1:00 PM UTC · 0m
Pending
resolved
The database failover that caused query issues earlier in the EU (https://status.honeycomb.io/incidents/n855d8kzp32y) has also had knock-on effects on our Activity Log beta feature. Due to recovery issues on the data replication flow used for this mechanism, activity log events will be missing from 12:45 UTC until 19:45 UTC, representing a 5 hour gap in events.
We have identified configuration parameters that will be adjusted to reduce the likelihood of such losses in the future.
Query interuption due to database failover in the EU region
Comenzó May 1, 2026 at 12:30 PM UTC · 0m
OutageMajor incident
resolved
At 12:50, our main database underwent an automated failover. This failover led to issues with some of our query engine connection pools, and caused queries to fail for roughly 10 minutes before self resolving.