Connectivity Issues with our Data Center
- resolved
Ongoing incident with our Hetzner data center UPDATE: there was an outage at Hetzner causing website and ingestion downtime from about 18:45-19:35 UTC
28 Hosted Graphite incidents · Φεβρουάριος 2025 — official updates, affected components, duration and resolution details.
Ongoing incident with our Hetzner data center UPDATE: there was an outage at Hetzner causing website and ingestion downtime from about 18:45-19:35 UTC
We have noticed some ingestions problems with our Heroku integration. We are looking into restoring the service as soon as possible.
We are upgrading K8 versions in the environment that manages all Hosted Grafana instances. You may experience up to 5min of downtime with your Grafana, but please let us know if you experience any additional issues: [email protected]
We are currently experiencing connectivity issues with the AWS US-West region affection ingestion of AWS Cloudwatch metrics.
Delayed AWS Cloudwatch => HG ingestion is only occurring on certain Hosted Graphite accounts, and a fix is being implemented. UPDATE: this issue has been resolved
We're experiencing network instability with our Data Center provider, resulting in slightly delayed ingestion time. UPDATE: this issue is now marked as 'Resolved'
We are currently experiencing connections issues with AWS specifically US-EAST. There could be delays in ingestion from Cloudwatch services. We are currently investigating the incident. UPDATE: this is resolved and missing metrics will backfill
We are investigating the cause of the website outage. UPDATE: there was a brief outage at our data center causing downtime on Hosted Graphite systems from: 21:00-21:11 UTC Hetzner Status Details can be seen here: https://status.hetzner.com/incident/5e901309-0873-42af-90c5-84f38304f592
We are working on a fix to restore services as soon as possible. Update: this issue has been resolved and any dropped data points should backfill to 100%
We are working to restore any instances of Grafana that are currently down. Update: we have located the cause of the issue and are implementing a fix. Update2: This incident has been resolved. If your Grafana instance is still down for any reason, please let us know: [email protected]
We are upgrading the production Grafana env to a newer version of Kubernetes (v1.33). We expect zero to minimal downtime, but please reach out if you experience any issues with your Grafana dashboards: [email protected]
We're experiencing delays in AWS metric collection for the US-East-1 region and are actively investigating the cause.
At this moment we are seeing any more issues with the 30s resolution. If you experience any issues, please don't hesitate to reach out.
All systems are operational. If you are still experiencing issues, please don't hesitate to contact us.
We are continuing to monitor the incident. We had to remove the IP endpoint: 136.243.28.200. If you were using this endpoint please update to any of other direct IPs to avoid any other disruptions. 195.201.153.19 195.201.153.20 144.76.143.215 176.9.21.64 136.243.95.166 136.243.95.165 94.130.127.162
Our data storage clusters are experiencing longer GET times than usual, we are actively investigating the cause and solution around this issue. 10/7 6:36 UTC: We have implemented a fix and this issue is now being marked as resolved.
We are experiencing an issue that may cause delayed alerts for some users, and we are working on a fix to resolve this issue.
Connectivity with our provider has been restored and all dashboard instances should be operational. If you encounter an issue, please don't hesitate to reach out to get it fixed.
We had an incident on 8/3 with a DB server that stores data for the 3600s resolution. This is only affecting some metrics and we are working to recover the data shortly. Please reach out to us with any questions or concerns: [email protected] We have identified the issues and we have a started the data recovery process. 8/11/25 All 3600s data has been recovered, please reach out to us if you notice any metric series that are not rendering historic data from before 8/3/25. [email protected]
We expect the connectivity issues with our data center to resolve shortly, data ingestion should not be affected at all.