Voice & Audio

status.daily.co

Daily

Realtime video & voice agents

All Systems Operational

Operational
Latency
27ms
Checked
just now
Active incidents
0
Components
6
Source
Statuspage.io API

Overview

Observed uptime · 1 day100%
2026-09-202026-09-20
Component health
6 up0 degraded0 down

Components6

6 operational0 degraded0 outage6 total
  • API
    Operational
  • Dashboard
    Operational
  • Core Call Experience
    Operational
  • Recording/Live Streaming
    Operational
  • SIP/PSTNDial-in, dial-out, and SIP connectivity
    Operational
  • Webhook Delivery
    Operational

Incidents50

History 50

Minor

Continued issues with joining calls from certain ISPs/Regions

Started
Wed, Jun 10, 2026, 07:20:46 PM
Updated
Thu, Jun 11, 2026, 02:47:17 AM
Resolved
Thu, Jun 11, 2026, 02:47:17 AM
Duration
7h 26m
  1. resolved

    The RegistryDNS.co servers continue to be intermittently unavailable, but AT&T's DNS servers seem to be more responsive. We encourage you to update to daily-js 0.91.0 as soon as possible to avoid future issues for your users.

  2. monitoring

    daily-js 0.91.0 is now available on npm and GitHub. If your app uses call object mode, you can update your dependency to that version and redeploy, and this issue should be resolved for you. We're publishing an update to our networking guide soon, but in the meantime, anything that mentions a `.daily.co` hostname in that doc is also available at `.dailywebrtc.com` and `.dailywebrtc.net`. We'll have updates for embedded Prebuilt and direct link customers soon. We'll leave this incident open while RegistryDNS.co remains offline.

  3. identified

    We expect to have daily-js 0.91.0 available within the next 60 to 90 minutes. The update includes automated failover to `dailywebrtc.com` or `dailywebrtc.net` if `daily.co` is unreachable. If you haven't updated daily-js in a while, you can do preliminary testing with 0.90.0 in order to be ready to release an update with 0.91 as soon as it's ready. 0.91 will only contain this feature, as well as a few dependency version updates for security. If you have customers with restrictive networks, they may need to add `*.dailywebrtc.com` and `*.dailywebrtc.net` to their firewall/security allowlists. These domains are relatively new, so they may be flagged by Cisco app firewalls and similar security appliances.

  4. identified

    We're seeing a recurrence of the problems with RegistryDNS.co and AT&T DNS from Monday. Users trying to join calls from affected regions using ISP DNS may get "Unable to join call" errors. We know specifically that AT&T DNS servers 68.94.157.1 and 68.94.156.1 are affected. We've been working hard on solving this problem since the incident on Monday. We're releasing daily-js 0.91.0 shortly to address this.

Minor

Issues with joining calls from certain ISPs/regions

Started
Mon, Jun 8, 2026, 02:32:48 PM
Updated
Tue, Jun 9, 2026, 02:57:28 PM
Resolved
Tue, Jun 9, 2026, 02:57:28 PM
Duration
1d
  1. resolved

    CentralNIC confirmed that there was indeed an issue with the .co nameservers, and they resolved it at approximately 06:30 UTC today. Our monitoring of both their DNS servers and downstream providers confirmed this. We're continuing to work urgently on updating our infrastructure to make our services available on alternate hostnames. We'll post updates to our Networking Guide when we do: https://docs.daily.co/guides/privacy-and-security/corporate-firewalls-nats-allowed-ip-list We'll also provide a root cause analysis for this incident in the next few days.

  2. identified

    We are working on implementing a fallback to a .com domain so that users on ISPs that are not fixing this will not see any issues starting calls.

  3. identified

    Users that are using their ISP's DNS service (specifically AT&T users in the southeast US and/or Texas) are experiencing intermittent problems accessing any .co domain, including daily.co. This may result in those users seeing an “Unable to join call” error in the browser when trying to join a call. This is being caused by failures from the DNS servers that serve the .co TLD itself. Several third-party DNS providers like Google, Cloudflare, and Quad9 are working around this in order to continue to make .co domains available. To work around this, you can encourage your users to use a third-party DNS provider, like Cloudflare or Quad9.

Minor

API and call connection issues

Started
Fri, Apr 17, 2026, 01:33:25 PM
Updated
Mon, Apr 20, 2026, 03:36:18 PM
Resolved
Fri, Apr 17, 2026, 11:06:00 PM
Duration
9h 32m
  1. resolved

    This incident has been resolved.

  2. identified

    Our tests just confirmed that DNS resolution is no longer returning errors for .co domains in our tests. This issue has been resolved!

  3. identified

    The rate of bundle download failures has decreased, but it's also tracking our overall usage volume through the day. The number of affected users appears to be very small, but we still have tests that can replicate the failure. Unfortunately this is completely out of our control. If you have users that are continuing to be affected, you can suggest that they use Google's or Quad9's DNS servers at 1.1.1.1 or 9.9.9.9.

  4. identified

    Cloudflare has resolved their status incident, but we're unsure if that means the underlying issue is resolved. We're continuing to monitor our own error rates and health checks.

  5. identified

    Cloudflare has posted that they've implemented a fix, and they are monitoring the results. We're still not certain that Cloudflare is the actual root cause of 100% of our affected users, but we'll be monitoring error rates.

  6. identified

    The .co registry appears to be experiencing issues. Cloudflare's recursive DNS is affected (and they've acknowledged it). Many regional ISPs, such as AT&T in the southeast US, rely on Cloudflare's DNS to power their own DNS. Unlike Cloudflare, other DNS providers, like Google and Quad9, are serving .co records from stale cache. If they stop doing this, then anyone using those DNS services will start to experience these same failures. This won't be fully resolved until the .co registry comes back online. In the meantime, different DNS providers may see problems come and go. All of Daily's services remain online and functional, so if your users can resolve .co hostnames, they can use our services without problems.

  7. identified

    Cloudflare has posted an incident concerning intermittent DNS failures for .co domains: https://www.cloudflarestatus.com/incidents/z3b5zxjtp6g1 This aligns with the troubleshooting we've done so far. Some internet discussions suggest that it can be somewhat ISP-dependent, but this is unconfirmed. If this is indeed related to the .co domain itself, it would mean that it affects participants' ability to join calls, but it could also affect the ability to make API requests to api.daily.co. Updates to follow.

  8. identified

    We've confirmed through several different end users that this is a DNS issue. Affected users from multiple regions have been able to join calls by pointing DNS to Google or CloudFlare, using IPs like 1.1.1.1 or 8.8.8.8 for DNS. Obviously, this solution doesn't scale; we're working with AWS to identify the root cause of the DNS resolution issue for c.daily.co.

  9. identified

    We're receiving reports of a few other regions experiencing similar connection issues. If you're monitoring client errors and you see messages that start with "Failed to load call object bundle https://c.daily.co/....", you're being affected by this. We're working with AWS to get to the bottom of this.

  10. identified

    We've identified an issue that's causing some users to fail to join calls. The daily-js library has to download a bundle of additional JavaScript as part of joining a call. This bundle is downloaded from c.daily.co, which is using Amazon's CloudFront CDN. They aren't reporting issues yet, but we're engaging with their support to figure out why this is happening.

  11. investigating

    We're receiving reports of some users having problems connecting to Daily calls. It seems to be localized to the Texas area. We're investigating.

Minor

Issues joining calls

Started
Fri, Apr 10, 2026, 03:36:15 PM
Updated
Fri, Apr 10, 2026, 04:30:42 PM
Resolved
Fri, Apr 10, 2026, 04:30:42 PM
Duration
54m
  1. resolved

    This incident has been resolved.

  2. monitoring

    There was a sudden increase in activity across several databases. We've addressed the cause of the issue, and platform metrics are returning to normal. We're continuing to monitor for any further issues.

  3. investigating

    We're seeing an elevated rate of errors for API requests and room joins. We're addressing it right now.

Minor

Issues with inbound phone calls

Started
Mon, Mar 23, 2026, 08:18:08 PM
Updated
Mon, Mar 23, 2026, 09:16:15 PM
Resolved
Mon, Mar 23, 2026, 09:16:15 PM
Duration
58m
  1. resolved

    This incident has been resolved. We're still collecting more information on the root cause here; we'll update this status post when we have more info.

  2. monitoring

    It seems that SignalWire has resolved their issue. We're monitoring for any other problems.

  3. investigating

    SignalWire has identified an issue causing problems with incoming and outgoing PSTN/SIP calls for Daily customers. If you're making dialout calls, you may see an error that says "DialOut stopped: Remote busy or Remote did not answer or Remote ended the call". If you're using dialin, you may hear a message that says "the number you have dialed is not configured correctly and cannot receive calls" or something similar.

  4. investigating

    We're investigating elevated error rates for incoming PSTN/SIP calls.

Minor

Webhook delivery delays

Started
Mon, Mar 16, 2026, 05:12:36 PM
Updated
Mon, Mar 16, 2026, 09:57:56 PM
Resolved
Mon, Mar 16, 2026, 09:57:56 PM
Duration
4h 45m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We've identified an issue with the webhook system that was causing brief delays when finding meeting start and end events for webhooks. We've addressed that issue. Webhook delivery time was only minimally impacted during the event, and it has returned to normal as of about 15 minutes ago.

  3. investigating

    We're investigating an issue that may be causing delayed delivery for some webhooks.

Critical

SIP / PSTN Disruption

Started
Wed, Feb 18, 2026, 11:31:42 AM
Updated
Wed, Feb 18, 2026, 03:55:29 PM
Resolved
Wed, Feb 18, 2026, 03:55:29 PM
Duration
4h 23m
  1. resolved

    This incident has been resolved.

  2. monitoring

    The fix from our upstream carrier partner has been applied, and service has been restored for SIP/PSTN dial-in (inbound) calling. We are seeing successful inbound call connections again, and our internal testing confirms recovery. We are continuing to monitor closely to ensure stability and full recovery.

  3. identified

    Our upstream carrier partner has identified the root cause of the issue impacting SIP/PSTN dial-in (inbound) calling and has implemented a fix. Service is beginning to recover, though some degradation may still be present while the fix fully propagates. We are actively monitoring and working to confirm full restoration of inbound calling. Outbound (dial-out) calling remains unaffected.

  4. investigating

    We are currently investigating an issue affecting SIP / PSTN inbound calls. Customers may encounter 480 “Temporarily Unavailable” errors when attempting to receive calls. Our engineering team is actively working to identify the root cause and restore service.

Minor

Increase in "Failed to load call object bundle" errors

Started
Wed, Feb 4, 2026, 11:59:24 PM
Updated
Thu, Feb 5, 2026, 09:51:30 AM
Resolved
Thu, Feb 5, 2026, 08:15:39 AM
Duration
8h 16m
  1. resolved

    The issue has been resolved. We’ve seen the number of “Failed to load call object bundle” errors return to normal levels, and call joins are operating as expected.

  2. monitoring

    The initial increase in error rates has subsided. Based on reports from a few customers, it seems like the impact was primarily in the Texas area. We're continuing to work with AWS to determine the root cause.

  3. identified

    We've seen an increase in "Failed to load call object bundle" errors within the past hour. This happens when the browser is unable to download our JavaScript bundle from the CDN. If this happens to you or your users, you should be able to refresh the page and join your call successfully. We're working with our CDN provider to identify the issue.

Minor

Issues with transcriptions and captions

Started
Mon, Feb 2, 2026, 03:45:47 PM
Updated
Mon, Feb 2, 2026, 10:37:23 PM
Resolved
Mon, Feb 2, 2026, 10:37:23 PM
Duration
6h 51m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We're seeing extremely low error rates, but Deepgram hasn't resolved their incident yet, so we'll keep monitoring.

  3. identified

    We're continuing to see very brief occurrences of slow or missing transcription data from the Deepgram streaming endpoint. We're talking to Deepgram engineers and continuing to monitor the situation.

  4. identified

    We're tracking a Deepgram incident that is causing problems with transcription events, as well as captions in Daily Prebuilt: https://status.deepgram.com/incidents/rp0jzltym9kw

Minor

Webhook deliveries delayed

Started
Mon, Jan 26, 2026, 04:47:23 PM
Updated
Mon, Jan 26, 2026, 09:24:40 PM
Resolved
Mon, Jan 26, 2026, 09:24:40 PM
Duration
4h 37m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We've worked through the webhook backlog, and delivery times should be normal. We'll be monitoring for any further issues.

  3. identified

    We're continuing to work through a backlog of webhook events, and the delay is starting to decrease. It will still take some time to get fully caught up.

  4. identified

    We’ve identified the cause of the issue. While a backlog of webhooks remains, we’re actively working to reduce it. Webhook deliveries will continue to be delayed for some time, but no data is being lost.

  5. investigating

    We are currently experiencing an abnormally large queue of webhooks awaiting delivery. Our team is investigating and working to restore normal processing.

Minor

Issues starting SIP/PSTN calls

Started
Thu, Jan 15, 2026, 09:32:26 PM
Updated
Thu, Jan 15, 2026, 10:11:29 PM
Resolved
Thu, Jan 15, 2026, 10:11:29 PM
Duration
39m
  1. resolved

    This incident has been resolved.

  2. identified

    We've rolled back the release that was in process when we started seeing the errors, which seems to have resolved the issue. We'll be monitoring for any regressions.

  3. identified

    We're seeing an increased rate of errors when dialing out SIP/PSTN calls.

Minor

Webhook deliveries delayed

Started
Tue, Jan 13, 2026, 10:13:02 PM
Updated
Tue, Jan 13, 2026, 11:46:06 PM
Resolved
Tue, Jan 13, 2026, 11:46:06 PM
Duration
1h 33m
  1. resolved

    Webhook delivery times have returned to normal.

  2. identified

    We're continuing to work through the backlog of webhooks.

  3. identified

    We're seeing an abnormally large queue of webhooks to be delivered, so we've scaled up to work through the backlog. We'll post another update when delivery times return to normal.

None

Issues with PSTN/SIP availability

Started
Tue, Jan 13, 2026, 12:16:48 AM
Updated
Tue, Jan 13, 2026, 12:16:48 AM
Resolved
Tue, Jan 13, 2026, 12:16:48 AM
Duration
0m
  1. resolved

    SignalWire experienced an outage earlier this morning. Their status site post is here: https://status.signalwire.com/incidents/MJVYwKKYYwbG Our first customer report of problems was at 1:51 AM PST (09:51 UTC). During the incident, many SignalWire SIP calls were returning 603 Declined. As we were investigating, SignalWire opened an incident on their status site at 2:26 AM PST (10:26 UTC). Their incident post didn't provide enough information for us to determine the overall impact to our customers. As we continued investigating our own logs and pressing SignalWire for more information, our affected customers started reporting that the issue was resolved around 2:35 AM PST. SignalWire ultimately resolved their incident at 03:27 PST.

Minor

Impact from the ongoing AWS us-east-1 outage

Started
Mon, Oct 20, 2025, 12:57:36 PM
Updated
Tue, Oct 21, 2025, 03:22:41 PM
Resolved
Tue, Oct 21, 2025, 03:22:41 PM
Duration
1d 2h
  1. resolved

    This incident has been resolved.

  2. monitoring

    AWS has resolved their incident. We're continuing to run the same call servers in us-east-1 as we were before the incident started, but we're still routing the majority of US traffic through us-west-2. We plan to return things to normal tomorrow.

  3. identified

    AWS is continuing to address issues in us-east-1. We're still seeing normal activity levels across our APIs and call servers for this time of day, so we have no reason to believe that any your customers' calls are actually being affected. Part of AWS's mitigation efforts involve limiting the rate at which we can start new call servers. Out of an abundance of caution, we're temporarily removing us-east-1 from our list of available regions for new call sessions. This means your users will likely connect to a call through us-west-2 instead. This failover happens automatically and transparently, and you don't need to do anything. We've already scaled up call servers in us-west-2 to handle the increase in sessions. We'll keep you posted as we continue to monitor the situation.

  4. identified

    AWS is addressing an issue with DynamoDB that is causing problems with many of their services in us-east-1. We have seen issues loading the dashboard and docs sites, and for a few hours we've been unable to log in to Statuspage to update this site. We're sorry about that delay. Up until a few minutes ago, our metrics suggested that core call functionality wasn't being impacted. We just saw a small number of call failures in us-east-1, though. We're continuing to monitor the situation and do what we can to minimize this outage's impact on you and your customers. We'll post more information here as we have it.

Minor

Dashboard API and webhook logs temporarily disabled

Started
Wed, Sep 24, 2025, 06:17:59 PM
Updated
Thu, Sep 25, 2025, 09:22:01 PM
Resolved
Thu, Sep 25, 2025, 09:22:01 PM
Duration
1d 3h
  1. resolved

    This incident has been resolved.

  2. investigating

    We've temporarily disabled displaying the API and webhook log views on the Dashboard Developer page. The Daily API is still fully operational, and these logs are still being collected. We expect to re-enable these log views in the next few hours. If you have an urgent need for this data, please contact help@daily.co.

Minor

Delayed data from the /logs API endpoint

Started
Thu, Sep 4, 2025, 02:48:34 PM
Updated
Thu, Sep 4, 2025, 04:45:29 PM
Resolved
Thu, Sep 4, 2025, 04:45:29 PM
Duration
1h 56m
  1. resolved

    This incident has been resolved.

  2. investigating

    We're experiencing delays in making log data available via the /logs REST API endpoint. We're not losing any log data, but the process that ingests logs for the REST API is bottlenecked. This means that the most recent data available from the /logs endpoint is about 6 hours old. We'll update this incident once the situation has been resolved.

Minor

Issues with SIP dialin/dialout

Started
Thu, Jul 17, 2025, 08:52:59 PM
Updated
Thu, Jul 17, 2025, 09:40:35 PM
Resolved
Thu, Jul 17, 2025, 09:40:35 PM
Duration
47m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We've rolled back a server release, and the SIP error messages have stopped. We're monitoring for any further problems.

  3. identified

    We're in the process of rolling back a release to address this issue.

  4. investigating

    We're investigating reports of SIP/PSTN dialin and dialout failing, or stopping immediately after starting.

Minor

Webhook delivery delays

Started
Mon, Jul 14, 2025, 04:43:47 PM
Updated
Mon, Jul 14, 2025, 05:25:36 PM
Resolved
Mon, Jul 14, 2025, 05:25:36 PM
Duration
41m
  1. resolved

    We're continuing to investigate the root cause, but we've resolved all impact to our running services.

  2. identified

    We're aware of an infrastructure issue that may be causing some intermittent failures for Daily API requests.

  3. investigating

    We're seeing alerts that webhooks are taking longer than normal to deliver.

Minor

Issues with logging and telemetry

Started
Thu, Jun 12, 2025, 06:39:33 PM
Updated
Thu, Jun 12, 2025, 09:47:56 PM
Resolved
Thu, Jun 12, 2025, 09:47:56 PM
Duration
3h 8m
  1. resolved

    This incident has been resolved.

  2. identified

    We had a brief spike in response time, but platform operations appear normal since approximately 20:05 UTC. We're continuing to monitor for any follow-on effects, and we'll leave this status post open until various upstream providers close their respective status incidents.

  3. identified

    We're seeing some increases in response time from our internal API endpoint that returns connection information to daily-js. This may be causing "room lookup timed out" errors for users attempting to join calls. We believe this is related to the ongoing GCP incident.

  4. identified

    This appears to be related to the Google Cloud outage posted here: https://status.cloud.google.com/incidents/ow5i3PPK96RduMcb1SsW We'll post here as we learn more.

  5. investigating

    Daily is being affected by widespread internet outages in authentication services. This is currently affecting our ability to collect or view some call metrics. We'll post more as soon as we have more information.

Minor

Elevated error rates for SIP/PSTN

Started
Thu, Apr 17, 2025, 07:33:17 PM
Updated
Thu, Apr 17, 2025, 08:36:46 PM
Resolved
Thu, Apr 17, 2025, 08:36:46 PM
Duration
1h 3m
  1. resolved

    This incident has been resolved.

  2. monitoring

    It appears that SignalWire rolled back the change, and the errors have stopped. We are monitoring for any further issues.

  3. identified

    SignalWire has confirmed that they made an unannounced change that's causing the problem, and they're working to revert it. We're also working on a quick production update to work around the issue if they can't revert quickly enough.

  4. investigating

    We're investigating elevated error rates from SignalWire when provisioning SIP/PSTN resources. If you're creating rooms with dial-in or dial-out enabled and it isn't absolutely necessary, you can remove those params from your room creation request to successfully create rooms.

Minor

Issues with SIP/PSTN audio quality

Started
Tue, Apr 15, 2025, 02:11:01 PM
Updated
Tue, Apr 15, 2025, 09:56:20 PM
Resolved
Tue, Apr 15, 2025, 09:56:20 PM
Duration
7h 45m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We've deployed an update to production servers to remediate the issue, and we're testing to ensure everything is fixed.

  3. identified

    We've identified an issue with how Daily and SignalWire are negotiating audio codecs for incoming SIP/PSTN calls. We're working on an update that will change the default audio codec used between Daily and SignalWire to Opus. Once this is deployed, you may still experience audio issues if you've explicitly set your SIP audio codec to PCMU as described here: https://docs.daily.co/guides/products/dial-in-dial-out/sip#sip-dial-in-audio-and-video

  4. identified

    We've identified an issue causing degraded ("broken up" or "choppy") audio between some Daily sessions and SignalWire SIP/PSTN endpoints. PSTN dialout and PIN dialin do not seem to be affected. SIP dialout seems to be moderately affected. PIN-less PSTN dialin and SIP dialin seem to be experiencing a much higher proportion of affected calls.

  5. investigating

    We're investigating reports of poor quality audio for SIP/PSTN participants in some calls.

Critical

Elevated SIP/PSTN error rates

Started
Wed, Apr 9, 2025, 12:15:21 PM
Updated
Wed, Apr 9, 2025, 09:48:01 PM
Resolved
Wed, Apr 9, 2025, 09:48:01 PM
Duration
9h 32m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We're still waiting on some additional fixes from Signalwire. Room creation with dial-in/dial-out is working, but you still may experience problems updating dial-in/dial-out settings for existing rooms. Daily Bots and Pipecat Cloud users should be unaffected.

  3. monitoring

    Signalwire has restored service, but we're still seeing some error responses from their API. We believe the majority of room creation failure issues have been resolved. We'll post here again when we've handled this last issue.

  4. identified

    We've been told we should be back online after a hotfix at approximately 18:00 UTC, or 15 minutes from now. We'll post another update as soon as we have more info.

  5. identified

    We're still waiting on Signalwire to restore service.

  6. identified

    We're still monitoring. In addition to dialout_enabled, you'll need to remove other room properties related to SIP/PSTN, dialin, and/or dialout to create rooms.

  7. identified

    We're continuing to monitor an issue from our SIP/PSTN provider. The rest of the Daily platform is unaffected. If you're getting an error when trying to create a room, you can remove the dialout_enabled property from the room creation request and try again.

  8. identified

    Our SIP/PSTN provider has identified an issue and they're deploying a fix. In the meantime, you should still be able to create Daily rooms without provisioning SIP/PSTN.

  9. investigating

    We're seeing elevated error rates from our SIP/PSTN provider. If you're creating rooms with dial-in support and getting errors, you may want to retry creating those rooms without dial-in, and then add dial-in with an update REST request.

Minor

Delayed audio for some SIP/PSTN dial-in calls

Started
Wed, Mar 12, 2025, 03:19:17 PM
Updated
Wed, Mar 12, 2025, 08:46:33 PM
Resolved
Wed, Mar 12, 2025, 08:46:33 PM
Duration
5h 27m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We've deployed a fix, and we're monitoring for any further issues.

  3. identified

    We've identified the issue, and we're testing a fix in our staging infrastructure. We'll post another update when we've deployed the fix.

  4. identified

    We're getting reports of delayed audio from some customers using SIP dialin and dialout. When the phone user joins the call, they can talk and others will hear them, but the phone user won't hear any audio from other call participants (bot or human) for the first 20-30 seconds. We're continuing to troubleshoot the issue, and we'll post here as soon as we have more info.

  5. investigating

    We're investigating an issue that's causing delays in audio connection for a some SIP/PSTN calls.

Minor

Issues connecting to rooms

Started
Tue, Jan 28, 2025, 04:21:07 PM
Updated
Tue, Jan 28, 2025, 04:58:39 PM
Resolved
Tue, Jan 28, 2025, 04:58:39 PM
Duration
37m
  1. resolved

    This issue has been resolved.

  2. monitoring

    We've resolved the issue and we're monitoring to ensure the platform is operating normally.

  3. investigating

    We're investigating an issue that may be preventing some users from joining meeting rooms.

Major

Networking issues

Started
Mon, Nov 25, 2024, 04:48:56 PM
Updated
Mon, Nov 25, 2024, 07:26:36 PM
Resolved
Mon, Nov 25, 2024, 07:26:36 PM
Duration
2h 37m
  1. resolved

    We've deployed an update that increases the throughput of the database that was the bottleneck in today's incident. We'll have more info about additional remediations and a postmortem for today's incident within the next few days.

  2. monitoring

    Our metrics have stayed at normal levels since our remediating actions about 30 minutes ago. We're continuing to monitor the platform while we discuss longer-term solutions to make absolutely sure we've addressed the root cause here.

  3. monitoring

    We've made some changes to the affected database, and our metrics and error rates have returned to normal. We've also re-enabled delivery of all webhooks, and we're monitoring for any further issues.

  4. identified

    We're addressing an issue with an internal database that's causing problems with existing meetings, as well as starting new ones. Your users are likely seeing some failures when trying to join meeting sessions, and users in ongoing sessions are seeing occasional meeting moves. We'll post more information as soon as soon as it's available.

  5. identified

    We've temporarily disabled the component that sends webhooks.

  6. investigating

    We're continuing to investigate the source of meeting disruptions. Customers may be experiencing 'meeting moves' where a call session moves from one server to another, causing a 2-3 second disruption to the call. You may also see delays in receiving meeting.started and meeting.ended webhooks.

  7. investigating

    We're investigating issues that may be causing problems with network connections between regions.

Minor

Issues starting recordings

Started
Wed, Oct 30, 2024, 06:59:54 PM
Updated
Tue, Nov 5, 2024, 12:16:34 AM
Resolved
Wed, Oct 30, 2024, 09:03:52 PM
Duration
2h 3m
  1. resolved

    We're still making a few small infrastructure changes, but our internal metrics have been back at normal levels for some time.

  2. identified

    We've identified an issue causing some recordings to fail to start, specifically in the Oracle Cloud San Jose region. We've already made some infrastructure changes that should be routing new recording requests to other regions. If you've seen a recording fail to start, you can try starting it again using daily-js or the REST API. We'll keep you posted on our progress resolving the issue.

Minor

Delayed API calls

Started
Tue, Oct 22, 2024, 06:46:26 PM
Updated
Wed, Oct 23, 2024, 10:47:48 PM
Resolved
Tue, Oct 22, 2024, 10:02:30 PM
Duration
3h 16m
  1. postmortem

    On Tuesday, October 22, around 17:15 UTC \(9:15 AM PDT\), a Daily customer started running a series of load tests. Their test involved rapidly creating and deleting a large number of rooms that used PSTN dial-out, cloud recording, and webhooks. This eventually caused several capacity threshold alerts to fire around 18:15 UTC \(10:15 AM PDT\) as our system scaled out to handle the load. We noticed that their test was running a script that created a room and started dial-out, but almost every instance of the script was exiting the room uncleanly before the outgoing call even connected to anything. This exposed an edge case that caused a ‘zombie’ PSTN participant to stay in that session and continue to try and send presence updates indefinitely. We’re already working on fixing that bug. This has probably happened before, but in much smaller quantities, since it involves a very unusual combination of events—but since this was an automated load test, it was causing too many of these ‘zombies’ to build up, all trying to write frequent presence updates to the database. Soon, the database response time began to slow under the increased load. Around that same time \(18:15 UTC, 10:15 AM PDT\), we noticed an increase in API error rates—specifically, actions that required writing to the database. Our team started to work both problems at once: safely get rid of the ‘zombie’ sessions without affecting other customers, and alleviate the load on the database to improve API response times. API error rates for POST requests spiked as high as 8%, and error rates for all requests peaked at 2-3%. We were able to return API error levels and latency back to normal by around 19:50 UTC \(12:50 PDT\) by refreshing several database instances. We contacted the customer and stopped the load tests, and then we were able to remove the ‘zombie’ sessions through our normal deploy process. We’re sorry for the disruption this caused. We’re already working on several remediations, including fixing the bug that caused the ‘zombie’ sessions, as well as adjusting platform rate limits to prevent this from happening again.

  2. resolved

    This issue has been resolved. We will post more information about this incident in the near future.

  3. monitoring

    API latency and errors have stayed at normal levels for a while now, but we're continuing to monitor for any further impact.

  4. identified

    API error levels have decreased considerably, but we're still working on full remediation. More updates to come.

  5. identified

    We've identified an issue causing some slowdowns in one of our databases, leading to some delayed or failed API responses. We've solved the root cause of the issue, but we're being cautious about restoring the database to full functionality, so we expect the delays to continue for a short time.

  6. investigating

    We're investigating an issue that's causing delays with some API operations, such as creating rooms and starting recordings. We'll post more info as soon as we have it.

None

Missing meeting webhook deliveries

Started
Mon, Oct 21, 2024, 05:50:07 PM
Updated
Tue, Oct 22, 2024, 02:38:47 AM
Resolved
Tue, Oct 22, 2024, 02:38:47 AM
Duration
8h 48m
  1. resolved

    This incident has been resolved. Customers needing assistance with missing webhook deliveries should contact support via help@daily.co.

  2. monitoring

    Between 14:01 UTC and 17:12 UTC webhooks for meeting.started and meeting.ended events were not delivered. We have applied a mitigation and are continuing to monitor. The underlying cause for the missing deliveries is still under investigation.

Minor

dashboard.daily.co availability

Started
Wed, Aug 7, 2024, 07:41:43 PM
Updated
Wed, Aug 7, 2024, 08:17:01 PM
Resolved
Wed, Aug 7, 2024, 08:17:01 PM
Duration
35m
  1. resolved

    This incident has been resolved.

  2. monitoring

    The upstream issue has been resolved, and we're monitoring for any more issues.

  3. monitoring

    The upstream issue has been resolved, and we're monitoring for any more issues.

  4. investigating

    Some customers are seeing 400 BAD_REQUEST messages when trying to load dashboard.daily.co. This is likely related to a Vercel incident: https://www.vercel-status.com/incidents/f6b2blrl5f5f

None

Elevated latency on some API endpoints

Started
Fri, May 24, 2024, 08:03:34 AM
Updated
Fri, May 24, 2024, 08:56:13 AM
Resolved
Fri, May 24, 2024, 08:56:13 AM
Duration
52m
  1. resolved

    API latency has returned to normal levels.

  2. monitoring

    We have applied a mitigation and are continuing to monitor the situation.

  3. investigating

    We are currently investigating increased latency affecting some of our APIs.

Minor

Degraded logging and metrics API performance

Started
Thu, Dec 21, 2023, 02:04:43 PM
Updated
Thu, Dec 21, 2023, 11:49:48 PM
Resolved
Thu, Dec 21, 2023, 11:49:48 PM
Duration
9h 45m
  1. resolved

    The impaired database system has fully recovered and is operating normally. API performance has returned to normal levels.

  2. monitoring

    The degraded logging and metrics API performance was the result of an impaired database system. The initial impact was resolved earlier today, but we continue to monitor the system as recovery completes.

  3. investigating

    We are currently investigating an issue with degraded performance with the logging and metrics API.

Minor

Elevated latency / intermittent failures on API endpoints

Started
Mon, Sep 18, 2023, 08:21:41 PM
Updated
Tue, Sep 19, 2023, 10:36:34 AM
Resolved
Tue, Sep 19, 2023, 10:36:34 AM
Duration
14h 14m
  1. resolved

    Network-level issues are resolved and service is operating nominally.

  2. monitoring

    Network-level mitigations have been applied and we are seeing latency back to normal levels.

  3. investigating

    We're currently investigating this issue.

Major

Issue with sessions in ap-northeast-2

Started
Thu, Aug 10, 2023, 02:13:22 PM
Updated
Thu, Aug 10, 2023, 02:42:30 PM
Resolved
Thu, Aug 10, 2023, 02:42:30 PM
Duration
29m
  1. resolved

    This issue has been resolved, and call sessions in ap-northeast-2 are working normally.

  2. investigating

    We've confirmed an issue preventing some users from joining calls hosted in the ap-northeast-2 region. Sessions in other regions are unaffected. If you've set the 'geo' property on your domain or a specific room to 'ap-northeast-2', you may want to temporarily change it to 'ap-south-1' or another nearby region.

  3. investigating

    We're investigating an issue preventing some users from joining call sessions in the ap-northeast-2 region.

Minor

Problems connecting to rooms

Started
Mon, Mar 13, 2023, 08:36:30 PM
Updated
Mon, Mar 13, 2023, 11:10:52 PM
Resolved
Mon, Mar 13, 2023, 11:10:52 PM
Duration
2h 34m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We've identified an issue that was causing some users to receive an error when trying to join a call. Affected users would see an error in the console starting with "web socket connection failed". We've rolled back a platform update from earlier today, and the errors have stopped. We're still diagnosing the problem with the platform update, but operations are back to normal.

  3. investigating

    We're investigating reports of problems when trying to join calls.

Minor

Issues connecting to calls

Started
Tue, Feb 7, 2023, 02:54:02 PM
Updated
Tue, Feb 28, 2023, 04:16:27 PM
Resolved
Thu, Feb 9, 2023, 11:37:38 PM
Duration
2d 8h
  1. postmortem

    On Tuesday, February 7 at 9:47 AM Eastern time \(14:47 UTC\), our database reported a performance issue under normal operational load. We had upgraded the database server over the weekend, but it had been operating normally since Monday. The alerts indicated a high level of lock contention on the newly upgraded database, which was causing problems for our call servers \(SFUs\). The SFUs are designed to shut themselves down if they are not able to connect to our database. When an SFU shuts down, our autoscaling will start a new SFU to replace it. With several SFUs shutting down at the same time \(and several new ones starting\), we experienced a larger than normal volume of “meeting moves”, which added additional load to a database that was already struggling. A “meeting move” occurs when an old SFU is shutting down. Our webapp automatically moves any ongoing call sessions on that SFU to a different SFU. During a meeting move, users will usually notice everyone else’s video drop out for a second or two before reappearing. The next few paragraphs shows the sequence of events between 09:51 to 10:47 which helped us identify the cause. By 9:51 \(T\+4 minutes\), engineers had found a potential culprit: a large volume of queries stuck in a deadlock. These were “meeting events” from the SFU, noting when participants joined or left meetings. This was causing the webapp API requests to time out and return 5xx errors, and ultimately causing the SFUs to drop their connections and restart. By 10:13 \(T\+26 minutes\), we had found one potential cause of the deadlocks. After our database migration from the previous weekend, we were still using MySQL binary log replication to keep our old database up to date. We disabled binlog replication and restarted the database to try and reduce the overall load on the database. This helped, but many of the SFUs retried the queries that were causing the deadlocks, so the problem persisted. We continued investigating, and also contacted AWS support to see if they had any insight on the issue. At 10:47 AM \(T\+1 hour\), engineers were working on a script that would terminate stuck queries when the database suddenly restarted itself. This restart took slightly longer than the one at 10:13, and it allowed the SFUs to discard the now-stale meeting updates without being disconnected long enough to cause them to restart. At this time, the SFUs and the platform went back to normal operation. We were ultimately able to prove that the deadlocking behavior was caused by a low-level behavior change introduced in a point release of MySQL. Our database maintenance from the previous weekend had upgraded us to that version and introduced the change. Working around that behavior change involved updating an index on one affected table. We spent the rest of the week developing and testing a plan to update the production database, and we completed that work with no user impact on Saturday evening. At 11:01, we decided we could move into a monitoring state while continuing to investigate the root cause. We left the status incident in a “monitoring” state until Friday, because we wanted to make sure we fully understood the initial cause of the deadlocks took any necessary action to avoid it in the future. One such action was the addition of rate limiting to the room creation API endpoint. The overall impact of this incident was limited to almost exactly one hour, between 14:47 and 15:47 UTC. During that time, some users in Daily calls experienced the “meeting moves” described earlier. There may have been a small number of users that weren’t able to join a room if they happened to try in the middle of a “meeting move”, which lasts a few 10s of seconds. They would have joined on reattempting a few seconds later. Similarly, some REST API requests may have returned 5xx error codes as well. We are continuing to work with AWS to make sure that the deadlocks issue we saw in production with Aurora MySQL 2.11.0 is fully documented, understood, and fixed in a future release. A more conservative approach to deadlocks was a known change in MySQL 5.7 \(which Aurora MySQL 2 is based on\). However, the severity of the deadlocks that we experienced during this incident was a surprise to us and to the AWS Aurora team. We try hard to test all infrastructure changes under production-like workloads. In this case, we failed to test with a synthetic workload that had the right “shape” to trigger these deadlocks. As a result of this incident, we have added additional API request patterns to our testing workload. We’ve also added some new production monitoring alarms that are targeted at more fine-grained database metrics.

  2. resolved

    We've identified the issue that caused the incident on Tuesday morning. While we've already deployed fixes that helped prevent the problem from reoccurring, we still need to perform one more database update that will require a short scheduled maintenance. That will likely happen this weekend. We will post a full retro after completing the final database maintenance operation.

  3. monitoring

    We've deployed a platform update with a few improvements designed to mitigate the impact of the current database performance issue. The only thing you may notice is that you'll no longer see 429 rate limit responses in your Dashboard API logs. Our database metrics have remained normal today, but we'll continue to monitor the platform to verify these fixes and watch for further issues.

  4. monitoring

    While we were able to restore platform functionality earlier today, we've continued to troubleshoot the underlying issue that caused the problem. As a precautionary measure, we've temporarily enabled rate limiting on the REST API endpoint used to create rooms. The limit for <tt>POST /rooms</tt> is now the same as the <a href="https://docs.daily.co/reference/rest-api#rate-limits">DELETE /rooms/:name endpoint</a>. You can expect about 2 requests per second, or 50 over a 30-second window.

  5. monitoring

    We’ve addressed the issue with the database, and platform operations have returned to normal. We are monitoring alerts and metrics for any further issues.

  6. identified

    We've identified an issue with one of our databases that coordinates activity between call servers. This is causing elevated rates of "meeting moves", which is when an ongoing call session has to move from one call server to a different one. If you're in a call when this happens, you'll notice everyone's video and audio drop out and come back within a few seconds. You may also need to restart recording or live streaming when this happens. You may also experience timeouts when making REST API requests. We'll post more information as soon as it's available.

  7. investigating

    We are investigating elevated platform error rates. Users may get websocket connection errors when trying to join calls.

Critical

Issues connecting to calls.

Started
Mon, Feb 6, 2023, 03:01:05 PM
Updated
Mon, Feb 6, 2023, 05:33:27 PM
Resolved
Mon, Feb 6, 2023, 05:33:27 PM
Duration
2h 32m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been applied and we are monitoring to be sure that all underlying issues are resolved.

  3. identified

    We have identified the issue and are applying a fix.

  4. identified

    The issue has been identified and a fix is being implemented.

  5. investigating

    We're investigating an issue preventing some users from connecting to calls.

Minor

Missing metrics in call participant logs

Started
Wed, Jan 25, 2023, 03:18:26 PM
Updated
Wed, Jan 25, 2023, 08:08:42 PM
Resolved
Wed, Jan 25, 2023, 08:08:42 PM
Duration
4h 50m
  1. resolved

    We've confirmed the initial report that there are a small number of recent call sessions that didn't log any metrics data. This can happen if your app has multiple call object instances running on the same page. Your app may do this if you are calling <tt>createCallObject()</tt> more often than you think; for example, in a React effect hook. Multiple call objects usually cause a variety of other errors on the page, so if you aren't already troubleshooting app issues related to this problem, you don't need to be worried about missing metrics. We are adding functionality to daily-js to help customers identify if they have multiple call objects on the same page. If you need help resolving this issue in your app, please feel free to contact support.

  2. investigating

    We're investigating reports of missing metrics data in participant logs from a small number of users. This may date back to some time around 2023-01-23 17:00 UTC (9:00 AM PST on Monday, Jan 23).

  3. investigating

    We're investigating reports of missing metrics data in participant logs.

Minor

Problems creating raw-tracks recordings

Started
Wed, Jan 18, 2023, 08:54:54 PM
Updated
Thu, Jan 19, 2023, 04:27:58 AM
Resolved
Thu, Jan 19, 2023, 04:27:58 AM
Duration
7h 33m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We've deployed new call servers and recording infrastructure to resolve the issue. You should be able to start a raw-tracks recording from any call session that started on or after approximately 02:15 UTC. Existing long-running call sessions may still be running on older call server instances. Those sessions may still experience errors with raw-tracks recordings. Those sessions will automatically move to new call server instances within the next few hours as part of our normal deploy process. We'll resolve this incident when all of the old call server instances have been retired and operations are back to normal.

  3. identified

    We're in the process of deploying updates to resolve this issue. We'll resolve the incident as soon as the fix is live in production.

  4. identified

    We've confirmed an issue preventing the creation of raw-tracks recordings. Other recording types are unaffected, including "cloud" recordings to your own S3 bucket. If you need to record an important call during this incident, you can change the <tt>enable_recording</tt> property on your domain, room, or meeting token to <tt>cloud</tt> to make a cloud recording.

  5. investigating

    We're investigating reports of errors from customers trying to create "raw-tracks" recordings.

Minor

Intermittent issues with cloud recordings

Started
Tue, Dec 6, 2022, 12:37:03 AM
Updated
Tue, Dec 6, 2022, 01:21:57 AM
Resolved
Tue, Dec 6, 2022, 01:21:57 AM
Duration
44m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We are experiencing an issue where cloud recordings are intermittently returning all black frames. We have pushed a fix to production and are currently monitoring the situation.

Major

Intermittent issues starting cloud recordings and livestreams

Started
Tue, Nov 29, 2022, 01:17:12 PM
Updated
Tue, Nov 29, 2022, 04:46:50 PM
Resolved
Tue, Nov 29, 2022, 04:46:50 PM
Duration
3h 29m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Daily’s auto scaling system experienced a failure to communicate with some internal services, preventing it from adding capacity for cloud recording and live streaming quickly enough to keep up with demand. We resolved the issue, and we're monitoring platform operations to ensure that everything has returned to normal.

  3. monitoring

    We have pushed a fix to production and are currently monitoring the situation.

  4. identified

    We have identified the issue and are currently testing a fix.

  5. investigating

    We are experiencing an issue where customers attempting to start livestreams or cloud recordings are intermittently receiving a temporarily-unavailable error.

Minor

Problems joining rooms in us-east-1

Started
Mon, Oct 17, 2022, 10:14:28 PM
Updated
Mon, Oct 17, 2022, 10:41:13 PM
Resolved
Mon, Oct 17, 2022, 10:41:13 PM
Duration
26m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We identified a brief issue with DNS while deploying a call server (sigh, it's always DNS). This would have caused intermittent join problems for some users for several minutes. We resolved the issue, and we're monitoring platform operations to ensure that everything has returned to normal.

  3. investigating

    We're investigating an issue preventing some users from joining meetings hosted in the <tt>us-east-1</tt> region.

Minor

Issues connecting to rooms

Started
Wed, Sep 28, 2022, 04:57:00 PM
Updated
Wed, Sep 28, 2022, 11:39:37 PM
Resolved
Wed, Sep 28, 2022, 11:39:37 PM
Duration
6h 42m
  1. resolved

    AWS has resolved their issue, and our operations have returned to normal.

  2. monitoring

    We're still watching the ongoing AWS status incident until it's resolved. We'll provide another update if anything changes in the meantime.

  3. monitoring

    AWS has acknowledged an issue with API Gateway in the <tt>us-west-2</tt> region. We're routing API requests to other regions for now, so everything should be operating normally for you and your users. We'll leave this issue open until AWS has resolved their underlying issue and our health checks return to normal.

  4. identified

    We're routing around a possible networking issue to our API gateways in <tt>us-west-2</tt>. This should allow your users to connect to calls, but we're still watching for other networking problems or follow-on effects<a href="https://news.ycombinator.com/item?id=33010341" target="_blank">.</a>

  5. investigating

    We're investigating an issue preventing some users from connecting to rooms in the <tt>us-west-2</tt> region.

Minor

Problems connecting to rooms

Started
Wed, Aug 24, 2022, 05:16:18 PM
Updated
Wed, Aug 24, 2022, 06:11:52 PM
Resolved
Wed, Aug 24, 2022, 06:11:52 PM
Duration
55m
  1. resolved

    We've re-enabled our API Gateways in us-west-2, and users are connecting to rooms successfully. This incident is resolved.

  2. monitoring

    We've confirmed that the issues with joining calls were a result of an AWS incident posted on their status site. AWS has resolved that incident, and we're seeing successful responses from our us-west-2 resources in our staging environment. We should be re-enabling our us-west-2 API Gateways shortly.

  3. investigating

    We've temporarily removed our affected us-west-2 API Gateways while AWS works to resolve the underlying issues. That should solve the problem that was preventing users from joining calls. Existing call sessions should be unaffected. We're closely monitoring other parts of our infrastructure, and we'll provide updates here if further issues emerge.

  4. investigating

    We are investigating an AWS issue with API Gateways in the us-west-2 region that is preventing some users from joining Daily sessions.

Minor

Users may be unable to view meeting session data in dashboard

Started
Thu, Jun 2, 2022, 11:41:55 AM
Updated
Thu, Jun 2, 2022, 05:54:20 PM
Resolved
Thu, Jun 2, 2022, 05:54:20 PM
Duration
6h 12m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been deployed, and dashboard users should now be able to access meeting session data. We're continuing to monitor the situation.

  3. identified

    We've identified an error impacting the ability for some users to view meeting session data in the Daily Dashboard. A fix is being implemented.

Minor

Customers may experience difficulty downloading meeting recordings.

Started
Wed, Apr 6, 2022, 01:45:09 PM
Updated
Wed, Apr 6, 2022, 02:31:09 PM
Resolved
Wed, Apr 6, 2022, 02:31:09 PM
Duration
45m
  1. resolved

    The issue impacting downloads of meeting recordings is resolved.

  2. monitoring

    Daily has resolved the issue impacting downloads of meeting recordings, and continues to monitor the situation.

  3. identified

    Daily has identified a problem that is impacting the ability to download meeting recordings via the dashboard and access-link APIs, and is implementing a fix.

Major

Connectivity issues

Started
Wed, Dec 15, 2021, 03:32:27 PM
Updated
Wed, Dec 15, 2021, 04:35:55 PM
Resolved
Wed, Dec 15, 2021, 04:35:55 PM
Duration
1h 3m
  1. resolved

    Error rates and network metrics have returned to normal levels, so we're considering this issue resolved.

  2. identified

    We've seen overall error rates decrease as AWS has been working to resolve the networking issue. Things are improving, but you may still experience delays and errors until this issue is fully resolved.

  3. identified

    We're continuing to see issues across our platform as a result of the ongoing AWS outage. You'll likely experience problems joining calls, accessing the Dashboard, or using the REST API. We'll continue to post more information here as we have it.

  4. identified

    We're experiencing network delays and timeouts throughout our infrastructure as a result of a larger-scale AWS incident. You may experience problems connecting to calls, viewing your Dashboard, or making REST API requests. We'll update this incident as we know more.

  5. investigating

    We're investigating reports of problems connecting to calls.

Minor

Degraded audio and video call experience

Started
Fri, Dec 10, 2021, 09:05:21 PM
Updated
Mon, Dec 13, 2021, 01:09:37 PM
Resolved
Mon, Dec 13, 2021, 01:09:37 PM
Duration
2d 16h
  1. resolved

    We have restored our service provider configuration to its nominal state after confirming that all providers are operating normally.

  2. monitoring

    We have temporarily routed traffic through another service provider, which should resolve call connection issues for most users.

  3. identified

    We are continuing to work on a fix for this issue.

  4. identified

    We are currently experiencing an issue with one of our service providers that may be affecting connection to calls (slow connections or timeouts). We are implementing a fix.

Major

Audio and video calls impacted by an ongoing incident in us-east-1

Started
Tue, Dec 7, 2021, 05:43:17 PM
Updated
Wed, Dec 8, 2021, 01:13:27 PM
Resolved
Wed, Dec 8, 2021, 01:13:27 PM
Duration
19h 30m
  1. resolved

    AWS has resolved the regional issues in us-east-1. We have re-activated the us-east-1 region for new audio and video calls.

  2. monitoring

    We have attempted to work around this AWS issue for users that would normally be routed to our us-east-1 resources by temporarily removing all of our us-east-1 DNS records from our AWS API Gateway configurations whilst we wait for AWS to recover the region. We'll continue to monitor the situation.

  3. investigating

    Users close to the us-east-1 (Northern Virginia) region of AWS may be unable to join video calls. We are investigating.

Minor

Internet connectivity - South America region

Started
Tue, Nov 9, 2021, 01:15:25 AM
Updated
Tue, Nov 9, 2021, 04:38:23 AM
Resolved
Tue, Nov 9, 2021, 04:38:23 AM
Duration
3h 22m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Connectivity issues in the South America - Brazil region have been resolved. We are continuing to monitor the situation.

  3. identified

    AWS is experiencing intermittent Internet connectivity issues in the South America - Brazil region. Users may experience a degraded experience connecting to audio and video calls in the region. Connecting to calls may take longer, and users in the region may connect to a server in another region until connectivity returns to normal.

Minor

Dashboard degraded: logging and telemetry impacted

Started
Mon, Oct 11, 2021, 06:04:25 PM
Updated
Mon, Oct 11, 2021, 09:14:07 PM
Resolved
Mon, Oct 11, 2021, 09:14:07 PM
Duration
3h 9m
  1. resolved

    This incident has been resolved.

  2. monitoring

    The impacted data repository has returned to normal operation. Audio and video call logs and metrics should now be available in the dashboard.

  3. identified

    We have identified a problem retrieving audio and video call logging and telemetry from the dashboard. Customers may experience a 'Session not found' error message when attempting to view call logs and telemetry.

Watch Daily
Email alerts on every status change — outages, degradations, new incidents, and resolutions.

Watching all 1 providers. Customize on the alerts page.

Details

Aliases
daily.co, daily bots
Indicator
none
Path
/daily

Related in Voice & Audio