Retards dans les webhooks d'Opsgenie
- resolved
Statut: Resolved Opsgenie webhook livraison retourné à la normale vers 17:48 BST composants touchés Opsgenie (opérationnel)
Traduit automatiquement depuis la mise à jour officielle de l'incident.
30 incident.io Status incidents · septembre 2025 — official updates, affected components, duration and resolution details.
Statut: Resolved Opsgenie webhook livraison retourné à la normale vers 17:48 BST composants touchés Opsgenie (opérationnel)
Traduit automatiquement depuis la mise à jour officielle de l'incident.
Statut: Résolue Entre 00:00 et 02:44 UTC, certaines fonctionnalités AI ont retourné des taux d'erreur accrus. Nous assistons à un rétablissement complet. Composants touchés Jira (opérationnel) Atlassian Statuspage (opérationnel) PagerDuty (opérationnel) Opsgenie (opérationnel) Cortex (opérationnel) Linear (opérationnel) Email (opérationnel) Application mobile (opérationnel) Google Docs (opérationnel) Zapier (opérationnel) Vanta (opérationnel) Calendrier (opérationnel) Google Meet (opérationnel) Shortcut (opérationnel) Zendesk (opérationnel) API (opérationnel) OpsLevel (opérationnel) Asana (opérationnel) Okta (opérationnel) Dashboard (opérationnel) Status pages (opérationnel) Datadog (opérationnel) Confluence (opérationnel) Site web (opérationnel) Spunk (opérationnel) Téléphone (opérationnel) SMS (opérationnel) Entrée (opérationnel) GitHub (opérationnel) Slack (opérationnel) Insights (opérationnel) ClickUp (opérationnel) Slack app (opérationnel) Zoom (opérationnel) Microsoft Teams app (opérationnel) @incident (opérationnel) SAML login (opérationnel)
Traduit automatiquement depuis la mise à jour officielle de l'incident.
Statut : La connexion SAML résolue a continué à fonctionner normalement. Composants touchés Connexion SAML (opérationnelle)
Traduit automatiquement depuis la mise à jour officielle de l'incident.
Status: Resolved We observed a limited increase in latency as well as a small number of dashboard errors between 19:31:30 and 19:34:30 UTC. This affected 0.03% of calls to our dashboard. This was detected by our internal synthetic monitoring. The cause of these elevated errors was due to networking issues within GCP at the time. This is confirmed by our logs showing our GCP load balancer not forwarding success responses, as well as network timeouts when trying to reach external services. The impact was limited to a small number of users and was solved by refreshing the dashboard. Affected components Dashboard (Operational)
Traduit automatiquement depuis la mise à jour officielle de l'incident.
Status: Resolved Slack have confirmed their services have largely returned to operational levels. We've not seen any direct impact on our Slack app since ~23:20 UTC yesterday. We're resolving this on our side, and you can follow any further updates on Slack’s status page here: https://slack-status.com/2026-05/46d9d5d41fbedd9f. Affected components Slack app (Operational)
Traduit automatiquement depuis la mise à jour officielle de l'incident.
We have identified increased error rates with our Slack integration. We are currently investigating how this is impacting our platform - this seems to be an issue on the Slack side, particularly around channel creation.
We've confirmed the issue is relating to errors being returned by Slack. This is affecting a number of features of our integration, including: • Creating incident channels • Sending messages • Syncing Slack user groups Our web dashboard continues to be available and Scribe is unaffected.
Slack have acknowledged the incident on their side. You can follow their status page [here](https://slack-status.com/2026-05/fe557ca05fdb64fa).
We're continuing to monitor the situation - Slack have posted an update on their status page to confirm that they're working on it.
Le problème a été résolu du côté de Slack et nous avons confirmé que les choses nous semblent normales.
Traduit automatiquement depuis la mise à jour officielle de l'incident.
The website is now up and working as normal.
The web dashboard is unavailable due to a JavaScript deployment issue.
Traduit automatiquement depuis la mise à jour officielle de l'incident.
Nous voyons des retards allant jusqu'à 30 minutes dans la réception de webhooks d'Opsgenie, ce qui peut retarder la livraison d'alerte. Nos systèmes fonctionnent normalement et traitent les alertes dès qu'elles arrivent; le problème semble être du côté Opsgenie. Nous suivons de près et partagerons les mises à jour au fur et à mesure que nous en apprendrons davantage.
Webhooks semble arriver comme normal à nouveau, de sorte que vous ne devriez plus voir de retards.
Traduit automatiquement depuis la mise à jour officielle de l'incident.
We have identified increased error rates with our Atlassian Statuspage integration. Atlassian have published this to their status page [here](https://metastatuspage.com/incidents/94fwk40f4z15).
Error rates have returned to normal levels.
We have identified increased error rates coming from Microsoft Teams. We are currently investigating how this is impacting our platform and we will update with more information shortly as we learn more.
We have identified increased error rates coming from Microsoft Teams that are causing delays in sending messages to Teams channels and DMs. We'll continue to monitor the situation and we will update with more information shortly as we learn more. We will provide the next update by 1830 UTC.
We are continuing to see an increase in error rates from Microsoft Teams, causing a delay in sending some messages to Teams channels and DMs. We are monitoring the situation and will provide an update by 1930 UTC.
We have seen an improvement in error rates and latency for requests to Microsoft Teams, and we expect the performance of sending messages to Teams channels and DMs to have returned to normal. We are continuing to monitor to ensure performance remains as expected, and we will provide a further update by 2100 UTC.
We are no longer seeing elevated latency and error rates from Microsoft Teams.
We have identified increased error rates with an upstream call transcription provider. This is affecting Scribe for all customers. Our team are investigating and we will update with more information shortly as we learn more.
The issues were isolated to one region and we have performed a failover to a secondary region, restoring Scribe for all customers. You may need to ask Scribe to rejoin your call. We'll continue to monitor the situation and are working closely with the upstream provider while they resolve the issue.
Our failover region is continuing to work as expected, and Scribe functionality has been restored. You may need to ask Scribe to rejoin your calls. We'll continue to monitor the situation and are working closely with the upstream provider while they resolve the issue.
We continue to use the failover region, and Scribe performance has returned to expected levels.
We are currently investigating reports that users are unable to login using Slack.
Users in some Slack workspaces are unable to sign in with Slack. We have escalated this issue to Slack. Signing in with SAML and Microsoft Teams is not affected.
Slack have rolled back a change which caused this issue, and we have confirmed recovery. Please reach out to support if you're still experiencing any issues signing in with Slack.
Slack have confirmed that the issue has been fully resolved, and affected customers have verified that they can log in successfully.
We've identified a drop in inbound traffic to [incident.io](http://incident.io "incident.io") from Opsgenie. This means that alerts may be slow to show up in [incident.io](http://incident.io "incident.io"), and paging via Opsgenie may be slow to reflect updates. This appears to be an issue with outbound webhooks on Opsgenie's side, we're monitoring the situation and will update as we know more.
Traffic appears to be back to normal levels now.
We have identified increased error rates with our Atlassian Statuspage integration. Atlassian have acknowledged an issue on their side (<https://metastatuspage.com/incidents/ctw5xv6t643h>). We'll continue to monitor this until the issue is resolved.
Atlassian's status page continues to show the issue with Statuspage (<https://metastatuspage.com/incidents/ctw5xv6t643h>). To restore the integration you can follow Atlassian's recommended workaround to regenerate your API key and reconnect in incident.io. We'll continue to monitor this until the issue is resolved.
Atlassian's status page continues to show the issue with Statuspage (<https://metastatuspage.com/incidents/ctw5xv6t643h>). To restore the integration you can follow Atlassian's recommended workaround to regenerate your API key and reconnect in [incident.io.](http://incident.io./ "incident.io.") We'll continue to monitor this until the issue is resolved.
Atlassian's status page continues to show the issue with Statuspage (<https://metastatuspage.com/incidents/ctw5xv6t643h>). To restore the integration you can follow Atlassian's recommended workaround to regenerate your API key and reconnect in [incident.io.](http://incident.io./ "incident.io.") We'll continue to monitor this and give an update in the next hour.
Atlassian's status page now reports that they are rolling out a fix (<https://metastatuspage.com/incidents/ctw5xv6t643h>). In the meantime, to restore the integration you can follow Atlassian's recommended workaround to regenerate your API key and reconnect in [incident.io.](http://incident.io./ "incident.io.") We continue to monitor the issue until it is resolved.
Atlassian have confirmed that the issues with Statuspage have been resolved. You should now be able to reuse your original API key in our integration.
We have identified increased error rates with our Sentry integration. We are currently investigating how this is impacting our platform and we will update with more information shortly as we learn more.
Sentry have acknowledged an issue on their side ([https://status.sentry.io/incidents/wpg3tnwnwctj](https://status.sentry.io/incidents/wpg3tnwnwctj "https://status.sentry.io/incidents/wpg3tnwnwctj")), and we're seeing alerts come through from them ok. We'll continue to monitor things before closing this out.
We've not seen any errors from Sentry since 15:35 UTC so we're closing this out. Please see Sentry's status page for any further details.
We have identified an increased rate of Sentry integrations becoming disconnected. Alert ingestion is unaffected. Please don't uninstall our application from Sentry: we are working to restore these connections on our side, and no action is currently required. We will provide an update in 30 minutes.
We're currently testing the fix that will allow us to re-connect any disconnected Sentry integrations on behalf of customers, and will have another update on this in 30 minutes.
We've verified a fix locally, and we're starting to roll this out in production now.
We have confirmed that we are able to restore these connections, and we are now in the process of restoring these. When this happens, you'll see that the Sentry integration no longer flags up as requiring reconnection. Once we have finished restoring access, we'll post another update here.
We have managed to re-establish a connection to Sentry for almost all of the affected customers. If Sentry is still showing up as disconnected, you will need to remove the incident.io integration from the Sentry side, and then go through the reinstallation flow by opening incident.io and going to Settings → Integrations → Sentry. If you do this, please note that Sentry will remove the "Notify incident.io" action from alert rules when you uninstall the integration. If that was the only action in the alert rule, the entire rule is disabled. You will therefore need to verify any alert rules after re-installing.
We are aware of an issue that is causing delays updating data in Insights. The team are currently investigating, and we'll provide further updates soon.
We are aware of an issue that is causing delays updating data in Insights. The team are currently working on mitigating this, and we'll provide further updates soon.
The team have resolved this issue, and Insights are now syncing as expected.
We have identified increased error rates with a number of our platform integrations, these relate to an ongoing Cloudflare outage. The list of impacted integrations includes: SAML sign-on, Zoom, Zendesk, Gitlab and Linear, but others may be experiencing degradation as well. On-call notifications and our Slack and Microsoft Teams apps are unaffected.
Cloudflare has implemented a [fix for the issue](https://www.cloudflarestatus.com/incidents/lfrm31y6sw9q) and we've seen recovery. SAML sign ins are now successful, other integration health appears to be recovering. We'll have more updates as and when things change.
[Cloudflare have resolved](https://www.cloudflarestatus.com/incidents/lfrm31y6sw9q) the incident. We've seen a recovery in system health - all integrations have returned to normal.
We have identified increased error rates with our Sentry integration. This is being tracked on [their status page](https://status.sentry.io/incidents/4w7y638y0zrh). Alerts from Sentry are still being received, but may be missing some additional context.
We're continuing to see degraded performance from Sentry. Alerts from Sentry are still being received, Sentry Issue Alerts may be missing some additional context, Sentry Metric Alerts are unaffected. This incident is being tracked on [their status page](https://status.sentry.io/incidents/4w7y638y0zrh).
We've seen a recovery in error rates from Sentry. We're continuing to monitor the recovery, and will post any updates as and when we see the situation change. Sentry are tracking [this incident on their status page](https://status.sentry.io/incidents/4w7y638y0zrh).
Sentry have confirmed the recovery on [their status page](https://status.sentry.io/incidents/4w7y638y0zrh), and are investigating the root cause. We have not seen the earlier errors return, and are continuing to monitor our integration's health. We'll have more updates as and when things change.
Sentry has resolved the incident.
We are experiencing degraded performance from some of our LLM model providers, causing error responses from @incident. /inc commands remain unaffected. More info on [Anthropic's status page](https://status.claude.com/incidents/xbbf1wq2dlz7 "Anthropic's status page")
We have resolved the degraded performance affecting @incident chat. This issue was caused by an overload from one of our AI model providers, which led to intermittent errors when using AI actions such as generating summaries or drafting updates. \inc commands were unaffected during this time. We have confirmed that the provider's issue has been resolved and all @incident chat features are now functioning normally.