La tela non sta caricando per i clienti nelle regioni degli Stati Uniti e dell'UE e abbiamo identificato che cosa crediamo essere la causa e stanno tirando fuori una correzione.
monitoring
Abbiamo implementato la correzione e stiamo vedendo il recupero della normale funzionalità Canvas.
resolved
Questo incidente è stato risolto.
Tradotto automaticamente dall'aggiornamento ufficiale dell'incidente.
U.S.1 esaurimento delle query fredde
Inizio 10 agosto 2026 alle ore 13:01 UTC · 7h 34m
IssuesIncidente minore
Componenti interessati
ui.honeycomb.io - US1 Querying
monitoring
Da 11:36UTC a 12:43UTC, domande in US1 che leggono i dati più vecchi restituiscono errori o timed out. Le domande sui dati recenti non sono state influenzate, e nessuna telemetria viene persa.
Il problema è stato identificato, e una correzione è in atto. Stiamo monitorando la situazione per confermare la risoluzione.
resolved
La soluzione è in atto e abbiamo confermato la risoluzione. Le query e gli SLO stanno elaborando normalmente.
Tradotto automaticamente dall'aggiornamento ufficiale dell'incidente.
Questioni di query
Inizio 31 luglio 2026 alle ore 14:30 UTC · 0m
IssuesIncidente minore
resolved
Tra le 10:30 e le 11:00 del mattino, abbiamo sperimentato elevati errori di querying e lentezza. L'impatto è diminuito, e stiamo attualmente indagando e monitorando la situazione.
postmortem
A partire dalle 9:50 Tempo il 31 luglio abbiamo ricevuto avvisi circa querying salute e latenza. La causa fu infine rintracciata in uno scontro di due fattori diversi. In primo luogo, abbiamo introdotto Lambda Managed instances \(LMI\) nella nostra architettura di query, che si comportano simili a a agnelli classici in molti modi e sono stati introdotti come parte del lavoro intorno a query miglioramenti delle prestazioni con l'obiettivo di un'esperienza utente migliore e più veloce. Purtroppo, una differenza tra agnello classico e LMI che abbiamo scoperto è la fermata dura 15 minuti di agnelli classici non esiste su LMI, il che significa che alcune ipotesi costruite nel nostro sistema di querying non sono più tenute. Al momento dell'incidente, la nostra nuova funzione Canvas era in esecuzione lotti di query per un cliente che dovrebbe avere un auto imposto un minuto di tempo. Questo timeout non è stato correttamente comunicato all'agnembda, e mentre con i classici agnelli l'operazione si sarebbe tagliata automaticamente a 15 minuti, i nuovi LMI non stavano rafforzando questo tipo di meccanismo di sicurezza. Di conseguenza, alcune grandi domande sono state in esecuzione per un massimo di 30 minuti, guidando la latenza di query attraverso la scheda e con conseguente errori per gli utenti che cercano di eseguire domande fresche. Il sistema è stato recuperato entro le 10:10 AM CT, 20 minuti dopo, quando le grandi domande sono state cancellate, ma abbiamo continuato a indagare la causa dell'incidente.
Una volta identificata questa debolezza nella nostra architettura di querying abbiamo iniziato a lavorare per risolvere e impedire che lo stesso incidente si verifichi di nuovo. Abbiamo spedito tutti gli incidenti associati.
Tradotto automaticamente dall'aggiornamento ufficiale dell'incidente.
Triggers intermittently fail nella regione dell'UE
Stiamo assistendo a inneschi intermittentemente inadeguati nella regione dell'UE. Stiamo indagando attivamente sulla causa.
investigating
Stiamo continuando a vedere problemi con i trigger e stanno anche notando problemi con la querying. Stiamo indagando per la causa di entrambi.
monitoring
Abbiamo identificato la fonte del problema, applicato una correzione temporanea, e stiamo lavorando su una fissazione permanente. Continuiamo a monitorare la situazione.
monitoring
Stiamo vedendo una regressione nelle prestazioni con degrado parziale di entrambi i trigger e querying. Stiamo lavorando per rimediare.
monitoring
Il degrado parziale è stato risolto per tutti i clienti tranne per quelli specificamente contattati. Stiamo lavorando a misure di mitigazione in modo da poter ripristinare pienamente la funzionalità di trigger per tutti.
monitoring
Stiamo continuando a monitorare eventuali ulteriori problemi.
resolved
Il degrado in querying e trigger è stato risolto e siamo tornati a piena funzionalità.
Tradotto automaticamente dall'aggiornamento ufficiale dell'incidente.
honeycomb.io marketing website not working
Inizio 2 luglio 2026 alle ore 18:10 UTC · 6h 21m
OutageIncidente maggiore
Componenti interessati
www.honeycomb.io
investigating
We are currently investigating this issue.
monitoring
Clearing the CDN cache appears to have solved the issue.
resolved
This incident has been resolved.
Activity Log delayed in US
Inizio 29 giugno 2026 alle ore 22:07 UTC · 2h 40m
IssuesIncidente minore
Componenti interessati
ui.honeycomb.io - US1 Activity Log
monitoring
Starting at 14:24 PT, there was an issue with the Activity Log causing events to stop being processed. You may notice a gap starting at that time. We are slowly backfilling these events and do not expect any data loss to occur. We expect the full data log to be restored around 17:30 PT.
resolved
Backfill of the Activity Log events has completed. All events during the affected period should be present.
We are investigating an issue with delayed ingest in the US and EU region. Our engineers are rolling back a recent deploy. SLOs and Trigger evaluations may also be delayed.
monitoring
We have rolled back the deploy and ingest service has recovered. There has been an ingest outage from 17:55 - 18:14 UTC. We have also observed an delay for Service Maps.
monitoring
Triggers, SLOs and Service Maps are recovered. We are continuing to monitoring the Ingest Service
resolved
All services have been fully recovered.
Impact window: 17:55–18:17 UTC
Regions affected: Primary impact in US. The EU region was affected to a much lesser degree.
During this period, the following effects may have occurred:
Ingest: Some data was not ingested, resulting in complete or partial data loss for events sent during the window.
Triggers: Triggers may have failed to fire within the impact window.
SLOs: SLI values across the affected window are skewed by the missing data and may show artificial dips or accelerated budget burn.
Service Maps: Maps covering the outage window are incomplete. Services and dependencies may be under-counted, as their traces were not ingested.
We apologize for the disruption. Please reach out if you have any questions about how this may have affected your data.
postmortem
On June 24, we experienced approximately 30 minutes of severe data ingestion degradation in Honeycomb’s US and EU instances. During this degradation, Honeycomb rejected a significant percentage of inbound telemetry, and customers would have seen up to a half-hour gap in ingested telemetry. Additionally, during this same time window, customers would have experienced data processing degradation and service disruption within Anomaly Detection and Service Maps.
A deployment containing a change to how services retrieve dataset schemas into local caches caused a cascading set of failures, resulting in failed remote cache retrievals, and subsequently, a MySQL stampede due to several services falling back to retrieving schemas from the database. This stampede resulted in significant memory utilization growth, causing some services to crash loop with out-of-memory exceptions while retrieving schemas. Customers would have seen 5xx errors from our ingestion APIs when this occurred, because ingestion services were among those that were crash looping. This also included crash looping from services responsible for Anomaly Detection and Service Maps.
Right after these service crash loops began, alerts fired for the crashing services, and procedures were run to pin all services back to the last known good build before the breaking change landed. Once services were running the previous deploy’s build, the 5xx error responses from ingestion APIs returned to nominal baseline levels, and services that were crash looping recovered and resumed normal operation.
The underlying issue centered on how dataset schema cache payloads and metadata were serialized to memcached for remote cache retrieval, and how memcached clients deserialize that data before writing it to a local cache. The code change contained a feature flag to toggle the serialization behavior, and backwards-compatibility safeguards for clients deserializing different schema versions from memcached. However, services’ schema deserialization prior to the feature flag flip on did not behave as intended, treating each backwards-compatible schema read as a hard error rather than as a fallback, which subsequently caused every remote schema cache read to fall through to the database. Subsequent analysis and review of the code identified the issue and we have re-deployed this change with a fix along with additional tests. Services now correctly and successfully deserialize schemas from memcached in a safe, backwards-compatible manner.
MCP tool access degraded
Inizio 19 giugno 2026 alle ore 16:39 UTC · 42m
IssuesIncidente minore
investigating
We are aware of an issue affecting Honeycomb MCP. Write tool functionality may be degraded. Read tools are unaffected.
monitoring
We are aware of an issue affecting Honeycomb MCP. Write tool functionality may be degraded. Read tools are unaffected.
resolved
MCP Write Tools have been restored. MCP users may need to reconnect to see all available tools.
Beginning June 17th, enhance-related requests were failing to complete. We've since identified and deployed a fix, and enhance is now fully operational. Thank you for your patience.
Querying Issues
Inizio 17 giugno 2026 alle ore 14:01 UTC · 7h 13m
IssuesIncidente minore
Componenti interessati
ui.eu1.honeycomb.io - EU1 Querying
investigating
We are continuing to investigate intermittent slowness and query failures affecting our Production EU Region. We will provide an update as soon as we have more information.
monitoring
Querying in our Production EU Region is fully operational. We will continue to monitor for any abnormal behavior.
resolved
We have identified the root cause of the intermittent slowness and query failures affecting our Production EU Region. The issue was triggered by a large volume of dataset deletions that caused elevated database load, resulting in query instability. No data was lost during this incident. Service has been restored. A fix has been applied to help prevent recurrence.
Querying Issues in EU
Inizio 17 giugno 2026 alle ore 12:01 UTC · 31m
IssuesIncidente minore
Componenti interessati
ui.eu1.honeycomb.io - EU1 Querying
investigating
We are currently investigating the cause of slowness and query failures in our Production EU Region.
resolved
The issue is now resolved and querying is back to fully functional.
Elevated API errors in production-eu1
Inizio 11 giugno 2026 alle ore 15:30 UTC · 0m
IssuesIncidente minore
resolved
Requests to Honeycomb's Query Data and Management APIs in the production-eu1 environment encountered increased error rates between approximately 15:35 and 22:24 UTC. Event ingestion was not affected. The issue has been resolved.
We are investigating an issue with delayed ingest in the EU region. Received events are still being stored, but there may be a delay in event retrieval in queries. SLOs and Trigger evaluations may also be delayed.
investigating
Queries should no longer be delayed. SLOs and Trigger evaluations remain impacted.
identified
The issue has been identified and a fix is being implemented.
We have identified and are working to resolve an issue that is causing query results to return inconsistent results.
investigating
Impact from this issue has concluded.
For a period of time (different between US and EU instances, noted below), queries spanning data older than the most recent 2 hours with certain GROUP / WHERE clauses returned inconsistent results. Data was never lost, and reruns of the same queries will show the previously missing data.
US impact times: 21:50 UTC - 23:20 UTC
EU impact times: 20:50 UTC - 23:20 UTC
resolved
This is an administrative update, marking the incident as Resolved. There has been no further impact since 23:20UTC on May 21 2026.
activity log for US instance is delayed
Inizio 8 maggio 2026 alle ore 08:00 UTC · 3d 8h
IssuesIncidente minore
Componenti interessati
ui.honeycomb.io - US1 Activity Log
identified
Activity log delay is rising, and is more than an hour behind due to a MySQL replica failover. Estimated recovery by 04:00 PDT.
monitoring
Backfilling of the past 2h30m of data is now in progress and should complete shortly.
monitoring
We believe that the activity log should now be caught up.
monitoring
We are continuing to monitor for any further issues.
monitoring
All activity log streams are now caught up, except for the query runs table, which should have all data since the start of the incident backfilled by 1600 PDT.
monitoring
Activity Log recovered at 07:00 UTC-07:00 today (about 6.5h ago). All streams are caught up.
resolved
This incident has been resolved.
Gap in Activity Log data in the EU region.
Inizio 1 maggio 2026 alle ore 13:00 UTC · 0m
Pending
resolved
The database failover that caused query issues earlier in the EU (https://status.honeycomb.io/incidents/n855d8kzp32y) has also had knock-on effects on our Activity Log beta feature. Due to recovery issues on the data replication flow used for this mechanism, activity log events will be missing from 12:45 UTC until 19:45 UTC, representing a 5 hour gap in events.
We have identified configuration parameters that will be adjusted to reduce the likelihood of such losses in the future.
Query interuption due to database failover in the EU region
Inizio 1 maggio 2026 alle ore 12:30 UTC · 0m
OutageIncidente maggiore
resolved
At 12:50, our main database underwent an automated failover. This failover led to issues with some of our query engine connection pools, and caused queries to fail for roughly 10 minutes before self resolving.
Cronologia delle interruzioni di Honeycombio | Uptimus