CLI Downloads başarısız
- investigating
Frekansı indirmeye çalışırken hata raporları aldık. Şu anda konuyu araştırıyoruz.
- resolved
Bu olay mühendislik ekibimiz tarafından çözülmüştür.
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
37 Doppler incidents · Aralık 2018 — official updates, affected components, duration and resolution details.
Frekansı indirmeye çalışırken hata raporları aldık. Şu anda konuyu araştırıyoruz.
Bu olay mühendislik ekibimiz tarafından çözülmüştür.
Resmî olay güncellemesinden otomatik olarak çevrilmiştir.
We are currently investigating an issue involving delays impacting integration sync, webhook and email delivery.
We are continuing to investigate this issue.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
Doppler is currently affected by an ongoing Cloudflare outage. This is impacting Dashboard and API access. Services accessing Doppler secrets through an integration (e.g., Kubernetes operator or syncs) should remain unaffected, but direct attempts to access the API will be unreliable until this issue resolves.
Services are still impacted. Our team is continuing to monitor the situation.
Cloudflare identified the issue and has been working on implementing a fix. Our dashboard and API are currently available and we're continuing to monitor the situation.
This incident is now resolved.
We're currently investigating an issue involving delays in background job processing that's impacting integration syncs, email delivery, activity log delivery, and webhook delivery (amongst other things). Impacted integration syncs will show as failed with "Out of sync". If expected emails aren't being received, it's likely related to this as well.
We are continuing to investigate the issue.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we're monitoring the situation. All unprocessed jobs have now been processed.
This incident has been resolved.
We are currently investigating this issue.
We believe the issue is resolved and are monitoring for any additional issues.
The issue is resolved.
We're currently investigating a system-wide outage related to disruptions with GCP and Cloudflare.
GCP and Cloudflare have been gradually resolving their outages and Doppler services have been partially restored.
This incident has been resolved.
We're currently experiencing downtime due to extended database maintenance.
This incident has been resolved.
Certain users may be unable to fetch from personal configs.
We are continuing to investigate this issue.
A fix has been identified and should be deployed soon.
Only users with access to personal configs via groups are affected. Users with direct access to personal configs are not affected.
This incident has been resolved.
Activity Logs stopped publishing to Splunk as of May 30th at 19:55 UTC. The issue has been identified and a fix is in progress. After the issue has been resolved and the connections reconnected, missed activity logs will be republished.
We are continuing to work on a fix for this issue.
A fix has been deployed and we are monitoring reconnections. We will backfill missed Activity Logs as needed after workplaces reconnect Splunk.
For a period of 18 minutes, secrets related API and dashboard actions saw increased error rates (503) and timeouts.
We are working on a fix.
We are continuing to work on a fix for this issue.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
We are currently investigating this issue.
We are continuing to investigate this issue.
This incident has been resolved.
This incident was caused by a faulty Kubernetes NetworkPolicy change. We’ll be evaluating how we can adjust our deployment procedures to catch these changes in the future.
We're currently investigating this issue. If you need immediate availability, you can install Doppler directly from our published release artifacts instead of via your OS's package registry by running: ``` (curl -Ls --tlsv1.2 --proto "=https" --retry 3 https://cli.doppler.com/install.sh || wget -t 3 -qO- https://cli.doppler.com/install.sh) | sh ```
An incident was identified at our provider for OS package registry downloads. They've fixed the issue and installs from OS package registries are available again. Impact: The incident affected all Doppler users attempting to install the Doppler CLI via their OS package registeries. This resulted in disruptions to their workflows relying on the Doppler CLI. Resolution: Cloudsmith resolved the incident after around an hour, and users were once again able to install the Doppler CLI from their OS package registries.
For several hours on February 27, 2023, users intermittently encountered errors when attempting to install the Doppler CLI. This issue was due to a related incident on GitHub, where Doppler hosts all CLI binaries and signatures. According to GitHub's incident report, GitHub experienced degraded performance and increased error rates for their packages service. Impact: The incident affected all Doppler users attempting to install the Doppler CLI. This resulted in disruptions to their workflows and GitHub actions relying on the Doppler CLI. Resolution: GitHub resolved the incident after a few hours, and users were once again able to install the Doppler CLI.
During this incident, changes to secrets could not be synced to integrations or trigger webhook updates. Additionally, activity log notifications could not be posted to Slack, Microsoft Teams, Sumo Logic, Splunk, or Datadog.
# Summary From 2023-01-16 19:11 UTC to 2023-01-16 20:04 UTC, Doppler experienced a partial outage which prevented sync integrations, webhooks, and activity log notifications from executing. The outage also prevented an internal job from firing which recomputes the version hashes for Doppler configs. This resulted in API clients \(e.g. the Doppler CLI and Kubernetes Operator\) failing to receive secrets updates which were made during this window. A recovery migration was run at 2023-01-17 00:37 UTC, re-triggering all syncs, webhooks, and activity log notifications — as well as recomputing config version hashes to restore the functionality for all clients to fetch secret updates. # Incident Details Doppler uses RabbitMQ to queue jobs which need to be executed as a result of secret updates. On 2023-01-13, Doppler’s security team rotated a RabbitMQ password, mistakenly identifying the credential as unused in production. It took several days for the RabbitMQ sessions in Doppler’s production services to expire and once they did, queue jobs could no longer be published. Once the incident was identified, Doppler’s security team created new RabbitMQ users to be used by our production services. The change was deployed and the incident was resolved at 2023-01-16 20:04 UTC. At 2023-01-17 00:37 UTC, Doppler ran a recovery migration to re-fire queue events for sync integrations, webhooks, activity log notifications, and secret version hash recomputations that were meant to fire during the incident window. # Next Steps Doppler has switched from using a single RabbitMQ credential to using one user per service. RabbitMQ users are now clearly named to mitigate the risk of accidental rotation in the future. We’ve also identified that the ability for API clients to fetch secrets should not be dependent on our application’s ability to connect to RabbitMQ. Our engineering team will move the config version hash computation to our atomic secrets write operation to ensure that the latest secrets are always fetched by clients. Lastly, our engineering team is reconfiguring the way we queue asynchronous jobs to ensure that if secrets are modified during a partial infrastructure failure, all post-update jobs will eventually be executed — without the need for manual recovery migrations.
New and existing Heroku sync integrations are currently not functioning. We have identified the issue and are working on resolving it. Secrets previously synced to Heroku apps will remain available but new secret changes in Doppler will not be synced until this is resolved.
This incident may persist for up to 24 hours. In the meantime, if you have an urgent secret change that needs to be made, please update the secret in your Doppler dashboard and then manually update the secret in Heroku directly via `heroku config:set`. Updating the secret in Doppler will ensure that the value in Heroku is not overwritten once the incident is resolved.
This issue is now resolved. Additional action is required to re-enable existing Heroku syncs. All workplaces will need to reconnect their Heroku sync integrations from the [workplace Settings page](https://dashboard.doppler.com/workplace/settings). Once integrations are reconnected, all syncs will automatically be triggered to sync any pending changes in Doppler. A postmortem will be available soon with more details regarding what happened.
**Root Cause** Doppler syncs secrets to Heroku via a Heroku OAuth application. This application is created in the Heroku dashboard and must be owned by a single Heroku user account. Doppler’s previous Heroku OAuth application was owned by a specific Doppler employee’s Heroku account without access to any additional resources. During a routine external account audit, this account was mistakenly identified as unused and manually deleted by our security team. This irrecoverably deleted Doppler’s existing Heroku OAuth application, thereby breaking any existing syncs and requiring the creation of a new OAuth application in a new account. **Resolution** Because users had authorized our previous Heroku OAuth application to their Heroku account\(s\), users need to authorize the new Heroku OAuth application. This involves reconnecting the integration from the [Doppler dashboard](https://dashboard.doppler.com/workplace/settings). Once the integration is reconnected, Doppler will re-enable all associated syncs that have been disabled and perform a fresh sync. Note that the previous OAuth application was deleted and therefore no action is required to remove its access to your Heroku account. **Next Steps** Internally, we’re reorganizing how shared accounts used for critical functionality are stored in 1Password. This new 1Password organization should help prevent this kind of accidental deletion in the future. We avoid shared accounts whenever possible, but this isn’t always feasible given third party implementations. We'll also be adding our individual integrations to our status page. This will allow customers to more easily see which integrations, if any, are currently experiencing issues.
We are currently investigating this issue.
A fix has been implemented and we are monitoring the results.
This issue was caused by a redis server running out of memory. We've increased the amount of memory allotted to this server. We'll also be setting up additional alerting on redis resource usage so that we can catch and prevent these issues in the future.
An error was discovered while running a data backfill job that resulted in users being unable to update secrets for affected environments. The data that was generated triggered an unexpected assertion in the secrets validation code. Users that attempted to save secrets in these environments would have received a 500 response from the server with the message "An error has occurred but don't fret. Our team has been notified." The data has been fixed for all environments and secrets write operations are fully functional again.
We are currently investigating this issue.
This issue is due to a widespread Cloudflare outage. https://www.cloudflarestatus.com/incidents/xvs51y9qs9dj
A fix has been implemented and we are monitoring the results.
We are continuing to monitor for any further issues.
This incident has been resolved.
We are currently investigating this issue.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.