Voice & Audio

status.livekit.io

LiveKit

Realtime voice & agent media

All Systems Operational

Operational
Latency
84ms
Checked
just now
Active incidents
0
Components
125
Source
Statuspage.io API

Overview

Observed uptime · 1 day100%
2026-09-202026-09-20
Component health
125 up0 degraded0 down

Components125

125 operational0 degraded0 outage125 total
  • Global Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • US West - SIPLiveKit Cloud SIP
    Operational
  • US East - Cloud AgentsLiveKit Cloud Agents
    Operational
  • US East - InferenceLiveKit Inference
    Operational
  • US Phone Numbers
    Operational
  • Cloud Dashboard (cloud.livekit.io)LiveKit Cloud Dashboard (analytics, settings, and billing)
    Operational
  • US West - Analytics IngestionLiveKit analytics ingestion
    Operational
  • US East - SIPLiveKit Cloud SIP
    Operational
  • Japan - EgressLiveKit Cloud Egress
    Operational
  • Europe Central - Cloud AgentsLiveKit Cloud Agents
    Operational
  • Australia - InferenceLiveKit Inference
    Operational
  • Japan - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • Brazil - Analytics IngestionLiveKit analytics ingestion
    Operational
  • Brazil - TURNLiveKit RTC Proxy
    Operational
  • US East - IngressLiveKit Cloud Ingress
    Operational
  • Global EgressLiveKit Cloud Egress
    Operational
  • Europe South - InferenceLiveKit Inference
    Operational
  • India - Cloud AgentsLiveKit Cloud Agents
    Operational
  • Brazil - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • India - Analytics IngestionLiveKit analytics ingestion
    Operational
  • Japan - TURNLiveKit RTC Proxy
    Operational
  • Global IngressLiveKit Cloud Ingress
    Operational
  • US West - IngressLiveKit Cloud Ingress
    Operational
  • US West - EgressLiveKit Cloud Egress
    Operational
  • India - SIPLiveKit Cloud SIP
    Operational
  • Brazil - InferenceLiveKit Inference
    Operational
  • Global SIPLiveKit Cloud SIP
    Operational
  • US East - EgressLiveKit Cloud Egress
    Operational
  • Japan - IngressLiveKit Cloud Ingress
    Operational
  • Brazil - SIPLiveKit Cloud SIP
    Operational
  • South Africa - InferenceLiveKit Inference
    Operational
  • India - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • US West - TURNLiveKit RTC Proxy
    Operational
  • Singapore - EgressLiveKit Cloud Egress
    Operational
  • Singapore - IngressLiveKit Cloud Ingress
    Operational
  • Australia - SIPLiveKit Cloud SIP
    Operational
  • Global Cloud AgentsLiveKit Cloud Agents
    Operational
  • Singapore - InferenceLiveKit Inference
    Operational
  • US West - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • Japan - Analytics IngestionLiveKit analytics ingestion
    Operational
  • India - TURNLiveKit RTC Proxy
    Operational
  • India - EgressLiveKit Cloud Egress
    Operational
  • Australia - IngressLiveKit Cloud Ingress
    Operational
  • Japan - SIPLiveKit Cloud SIP
    Operational
  • Global InferenceLiveKit Inference
    Operational
  • US West - InferenceLiveKit Inference
    Operational
  • Regional Real Time CommunicationStatus of RTC services in each data center. This is informational and does not impact end users. When individual data centers have outages, traffic is automatically rerouted to healthy data centers.
    Operational
  • Australia - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • Australia - Analytics IngestionLiveKit analytics ingestion
    Operational
  • Australia - TURNLiveKit RTC Proxy
    Operational
  • South Africa - IngressLiveKit Cloud Ingress
    Operational
  • Saudi Arabia - InferenceLiveKit Inference
    Operational
  • Regional TURNLiveKit RTC Proxy
    Operational
  • US Central - SIPLiveKit Cloud SIP
    Operational
  • Saudi Arabia - IngressLiveKit Cloud Ingress
    Operational
  • Israel - InferenceLiveKit Inference
    Operational
  • Regional Analytics IngestionLiveKit analytics ingestion service
    Operational
  • UAE - EgressLiveKit Cloud Egress
    Operational
  • United Kingdom - InferenceLiveKit Inference
    Operational
  • Singapore - TURNLiveKit RTC Proxy
    Operational
  • Singapore - Analytics IngestionLiveKit analytics ingestion
    Operational
  • Singapore - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • Brazil - EgressLiveKit Cloud Egress
    Operational
  • Israel - SIPLiveKit Cloud SIP
    Operational
  • Europe Central - InferenceLiveKit Inference
    Operational
  • Regional Phone Numbers
    Operational
  • Regional SIPLiveKit Cloud SIP
    Operational
  • US Central - TURNLiveKit RTC Proxy
    Operational
  • US Central - Analytics IngestionLiveKit analytics ingestion
    Operational
  • US Central - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • South Africa - EgressLiveKit Cloud Egress
    Operational
  • Saudi Arabia - SIPLiveKit Cloud SIP
    Operational
  • India - IngressLiveKit Cloud Ingress
    Operational
  • US Central - InferenceLiveKit Inference
    Operational
  • Regional IngressLiveKit Cloud Ingress
    Operational
  • South Africa - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • South Africa - TURNLiveKit RTC Proxy
    Operational
  • US Central - EgressLiveKit Cloud Egress
    Operational
  • Japan - InferenceLiveKit Inference
    Operational
  • Regional EgressLiveKit Cloud Egress
    Operational
  • South Africa - Analytics IngestionLiveKit analytics ingestion
    Operational
  • UAE - TURNLiveKit RTC Proxy
    Operational
  • Saudi Arabia - EgressLiveKit Cloud Egress
    Operational
  • UAE - SIPLiveKit Cloud SIP
    Operational
  • Brazil - IngressLiveKit Cloud Ingress
    Operational
  • UAE - InferenceLiveKit Inference
    Operational
  • UAE - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • UAE - Analytics IngestionLiveKit analytics ingestion
    Operational
  • South Africa - SIPLiveKit Cloud SIP
    Operational
  • Israel - IngressLiveKit Cloud Ingress
    Operational
  • Regional Cloud AgentsLiveKit Cloud Agents
    Operational
  • India - InferenceLiveKit Inference
    Operational
  • Israel - TURNLiveKit RTC Proxy
    Operational
  • Israel - Analytics IngestionLiveKit analytics ingestion
    Operational
  • Israel - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • Singapore - SIPLiveKit Cloud SIP
    Operational
  • Israel - EgressLiveKit Cloud Egress
    Operational
  • UAE - IngressLiveKit Cloud Ingress
    Operational
  • Regional InferenceLiveKit Inference
    Operational
  • Saudi Arabia - TURNLiveKit RTC Proxy
    Operational
  • Saudi Arabia - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • Saudi Arabia - Analytics IngestionLiveKit analytics ingestion
    Operational
  • Australia - EgressLiveKit Cloud Egress
    Operational
  • Europe Central - SIPLiveKit Cloud SIP
    Operational
  • Europe South - IngressLiveKit Cloud Ingress
    Operational
  • US East - TURNLiveKit RTC Proxy
    Operational
  • United Kingdom - Analytics IngestionLiveKit analytics ingestion
    Operational
  • Europe South - SIPLiveKit Cloud SIP
    Operational
  • Europe South - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • United Kingdom - EgressLiveKit Cloud Egress
    Operational
  • United Kingdom - IngressLiveKit Cloud Ingress
    Operational
  • Europe Central - TURNLiveKit RTC Proxy
    Operational
  • US East - Analytics IngestionLiveKit analytics ingestion
    Operational
  • United Kingdom - SIPLiveKit Cloud SIP
    Operational
  • United Kingdom - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • US Central - IngressLiveKit Cloud Ingress
    Operational
  • Europe Central - EgressLiveKit Cloud Egress
    Operational
  • United Kingdom - TURNLiveKit RTC Proxy
    Operational
  • Europe South - Analytics IngestionLiveKit analytics ingestion
    Operational
  • US East - Real Time CommunicationLiveKit Cloud API and media backend
    Operational
  • Europe Central - IngressLiveKit Cloud Ingress
    Operational
  • Europe South - EgressLiveKit Cloud Egress
    Operational
  • Europe Central - Analytics IngestionLiveKit analytics ingestion
    Operational
  • Europe South - TURNLiveKit RTC Proxy
    Operational
  • Europe Central - Real Time CommunicationLiveKit Cloud API and media backend
    Operational

Incidents50

History 50

Major

Cloud Dashboard, Analytics Services, and Billing API failures

Started
Tue, Sep 15, 2026, 12:08:27 PM
Updated
Thu, Sep 17, 2026, 06:21:17 PM
Resolved
Tue, Sep 15, 2026, 03:30:05 PM
Duration
3h 21m
  1. resolved

    This incident has been resolved. Between 12:08 and 14:23 UTC, customers were unable to access the Cloud Dashboard and the Analytics Services APIs. The cause was a control plane failure in our last remaining legacy Kubernetes cluster, which took out a dependency of those services. Our user-facing services (RTC, SIP, and others) were unaffected, as they run redundantly across our modern clusters and have no dependency on the legacy cluster. We have no indication of data loss. Recovery took longer than we would like, as mitigation required working within the legacy cluster. Work was already underway to make the affected dependency redundant across our modern clusters, and that cluster is scheduled for decommissioning, which removes this class of failure. We will follow up with a postmortem.

  2. monitoring

    A fix has been deployed as of 14:23 UTC, and access to the Cloud Dashboard and Analytics Services APIs is returning to normal. We are continuing to monitor before marking this resolved. Recovery may take a few additional minutes for some clients as the change fully propagates; retrying or reconnecting will restore access.

  3. identified

    We have identified the cause of the issue affecting the Cloud Dashboard and Analytics Services APIs, and recovery work is underway. Customers may continue to see errors accessing the Cloud Dashboard and Analytics Services APIs during this time. Room APIs, RTC APIs, and all other services remain healthy and unaffected. We have no indication of data loss from this incident. We will post another update as recovery progresses.

  4. investigating

    We are continuing to investigate this issue.

  5. investigating

    We are continuing to investigate this issue affecting LiveKit Cloud dashboard, Analytic Services APIs and Billing APIs. Users may be unexpectedly signed out and unable to sign back in, and API requests may return 522 errors.

  6. investigating

    We are investigating an issue where users may be unexpectedly signed out of the LiveKit Cloud dashboard and unable to sign back in.

Minor

Monitoring reports of intermittently increased API latency impacting Egress, Ingress, and SIP

Started
Tue, Sep 15, 2026, 10:10:33 PM
Updated
Thu, Sep 17, 2026, 03:33:02 PM
Resolved
Tue, Sep 15, 2026, 11:15:15 PM
Duration
1h 4m
  1. postmortem

    Please see a full postmortem [here](https://status.livekit.io/incidents/dbqb4jhxcg9h).

  2. resolved

    We have resolved the underlying issue and have not observed any further occurrences of this incident. We will follow up with a detailed postmortem in the original incident.

  3. monitoring

    From 21:11 to 21:21 UTC we observed an occurrence of the same incident experienced earlier today impacting Global SIP, Ingress, and Egress due to the same underlying issue. We are currently operating normally and are continuing to monitor.

Minor

Investigating reports of intermittently increased API latency impacting Egress, Ingress, and SIP

Started
Tue, Sep 15, 2026, 05:04:22 PM
Updated
Thu, Sep 17, 2026, 03:32:19 PM
Resolved
Tue, Sep 15, 2026, 06:16:09 PM
Duration
1h 11m
  1. postmortem

    ### Summary On 15 September, LiveKit Cloud experienced three periods of degraded performance between 13:34–13:41, 16:06–16:11 and 21:10–21:22 UTC. Ingress creation was the most affected: more than half of `CreateIngress` requests failed at the worst point. A number of in-progress recordings ended prematurely, and a small percentage of inbound SIP calls failed during a one-minute period \(during each impact window\). Realtime connections were not affected and media already flowing continued normally. The cause was a dropped connection to our primary metadata database that neither the client nor server side detected, which left a transaction open and holding row locks for up to 17 minutes. Requests queued behind those locks, and the sudden release of that backlog — not the wait itself — is what briefly degraded the wider platform. ### Root Cause Our services share a distributed global metadata database that stores ingress and egress state. At the time of the incident, a connection between one of our services and that database was dropped in a way neither end observed: the client treated the connection as closed and returned an error, while the database continued to consider the session live. Because the client had a transaction open at the time, the database kept that transaction's row locks held. The impact then came in two distinct phases. **While the lock was held**, only requests that needed the same rows were affected. Ingress creation failed, and recordings that could not report their status ended early. Most other APIs were unaffected, because they never touched the locked rows. **When the lock was released**, several hundred transactions that had queued behind it — some waiting more than seventeen minutes — all executed within a fraction of a second. That burst briefly saturated the database and degraded it for every service using it, not just those touching the original rows. This is why the broadest impact, including room management APIs and SIP participant creation, appears at the very end of each window rather than during it. The database only reclaimed the abandoned session when the operating system's TCP timeout expired, roughly 17 minutes after the connection was dropped. Two design choices amplified a single stuck row into regional impact: * **Recording workers treated "cannot report status" as "cannot accept work."** Every worker in a region reports through one shared service, so when that service became slow, all workers in the region stopped accepting new work at the same time and our router saw no available capacity. * **Several internal queries scanned the affected table without a narrowing filter.** That meant one locked row could block reads that were otherwise unrelated to it, which is what allowed the backlog to grow large enough to be disruptive on release. ### Scope of Impact Ingress creation was the most affected API: at the worst point in each window, 50.3%, 62.4% and 58.1% of `CreateIngress` requests returned server errors. Many failed requests are retried automatically, so the share of ingresses that ultimately could not be created is lower than those figures suggest. Starting a new egress was largely unaffected, staying under 2% throughout. SIP was affected at the end of each window, with `CreateSIPParticipant` reaching 7.50% and between 1.3% and 2.2% of attempted inbound calls failing during call setup; calls already connected were unaffected. Room management APIs stayed below 0.5%. Realtime connections were not affected, and media already flowing continued normally. The more consequential impact was to recordings already in progress. Up to 16.5% of egresses started during an affected window ended prematurely — 0.39% of all egresses that day — and did so without surfacing an error to indicate the recording had failed. ### Mitigations and Follow-ups Completed: * We have set a database-side idle transaction timeout so an abandoned transaction can no longer hold locks for more than 10 seconds. This caps both the wait and the size of any backlog that can accumulate behind it. We have verified this against a reproduction of the original failure. Underway: * We are removing an unnecessary transaction wrapper around single-statement writes, which shrinks the window in which a dropped connection can leave locks held. * We are adding database-side statement and lock timeouts so that no query can wait indefinitely on a lock. * We are changing recording workers so that a single status-reporting failure no longer removes an entire regional fleet from service. * We are making egress startup retry transient database errors rather than aborting the job. We appreciate your understanding and are committed to continuously improving our platform's reliability. If you have any questions, please reach out to our support team.

  2. resolved

    We are resolving this incident as the underlying issue has been resolved. Thank you for your patience and apologies for the disruption. We will follow up with a postmortem as soon as possible.

  3. monitoring

    We are continuing to monitor, but want to keep users up to date with how we believe impact may have materialized. The first impact window started at 13:40 UTC and impacted Ingress and Egress services for about 3 minutes. The second impact window started at 16:11 UTC and impacted SIP, Ingress, and Egress services for about 1 minute.

  4. monitoring

    We are still investigating, but believe that customer impact should be mostly mitigated - although some spikes may still be occurring. We are investigating reports of aborted egresses which we believe were also related to this issue. We will post another update as soon as possible.

  5. investigating

    We are investigating reports of intermittently increased API latencies starting at 13:30 UTC. Users may observe increased latency spikes across all APIs.

Minor

Investigating reports of elevated egress & connector API errors

Started
Thu, Aug 20, 2026, 09:10:42 PM
Updated
Thu, Sep 10, 2026, 07:20:46 PM
Resolved
Thu, Aug 20, 2026, 09:47:51 PM
Duration
37m
  1. postmortem

    On August 20, 2026, customers in our US East region experienced approximately 10 minutes of 503 errors on the LiveKit Ingress API and failures launching media track egress, between 20:52 and 21:01 UTC. All other regions were unaffected. ‌ **Root Cause** An internal routing service lost one of its two instances, and a concurrent infrastructure issue with our cloud provider was preventing new compute nodes from provisioning in US East, so the replacement couldn't start. At 20:51 UTC, an automated maintenance process removed that last instance, making the service fully unavailable. It recovered at 21:01 UTC once the evicted instance freed capacity for a replacement. ‌ **Timeline \(UTC\)** * Aug 20, 20:52 - Automated maintenance removes last instance in the pool; customer-facing errors begin * Aug 20, 21:01 - Service restored * Aug 21, 03:00 - Cloud provider resolves the underlying infrastructure issue; region fully healthy ‌ **Mitigations & Follow-ups** * The service is restored with two instances guaranteed to run on separate physical nodes. * We updated the automated maintenance process to avoid evicting the last healthy instance of a service. * We added disruption protection for the affected service and are continuing broader resilience improvements. * Our cloud provider resolved the infrastructure issue that had prevented new nodes from scaling up.

  2. resolved

    This incident has been resolved. Between 20:52 and 21:01 UTC, we observed an increase in failed egress and connector API requests. Errors have returned to baseline as of 21:01 UTC. We will follow up with a postmortem with further details.

  3. investigating

    We had a spike of errors between 20:52 and 21:01 UTC, and the errors have subsided to baseline as of 21:01 UTC. We are actively monitoring and will share more updates in 30 minutes.

  4. investigating

    We are currently investigating reports of elevated API errors across the Egress service. We will post another update in 5 minutes.

Minor

Identified increased agent join latencies in eu-central

Started
Sat, Sep 5, 2026, 11:36:53 PM
Updated
Tue, Sep 8, 2026, 10:01:01 PM
Resolved
Sun, Sep 6, 2026, 01:12:49 AM
Duration
1h 35m
  1. postmortem

    **Root Cause** A configuration change introduced a formatting error in the startup settings for the server pools that run hosted agents in our eu-central region. Newly provisioned servers failed to start, so the region could not add capacity. Agents that were already running were unaffected, but new agent deployments, version rollouts, and automatic scale-ups were delayed or stuck pending until the change was reverted. **Timeline \(UTC\)** `2026-09-04 23:56` - Configuration change deployed. The error only affected newly provisioned servers, so there was no immediate impact. `2026-09-05 18:07` - New servers began failing to start. Impact begins, intermittent at first. `2026-09-05 22:23` - Our monitoring alerted us to agent deployments stuck pending and we began investigating. `2026-09-06 00:50` - We identified the malformed startup configuration and reverted the change. `2026-09-06 01:17` - New capacity provisioned successfully, all pending deployments recovered, and we validated the fix. **Scope of Impact** Limited to hosted agents in our eu-central region. During the impact window, new agent deployments and rollouts were delayed or stuck pending, and automatic scale-up was blocked. Agents already running continued to serve sessions normally, and we estimate that, at the peak, roughly 1% of agent instances in the region were affected. No other regions or products were affected. **Mitigations and Follow-ups** * The faulty configuration change was reverted and provisioning has been stable since. * We are adding alerting for new servers that fail to start via this failure mode, which today fails silently. This lets us detect this class of failure directly and validate future changes to server provisioning quickly. We appreciate your understanding and are committed to continuously improving our platform's reliability. If you have any questions, please reach out to our support team.

  2. resolved

    The mitigation was successful and users shouldn't notice any issues with join latencies or creating new agents. We will follow up as soon as possible with a postmortem.

  3. monitoring

    We have applied a mitigation and believe that symptoms should be improving. We will update again with confirmation as soon as possible.

  4. identified

    We have identified an issue with LiveKit Cloud hosted agents in eu-central. Existing agents in eu-central are facing scaling issues resulting in increased join latencies. Existing agents are still functional and can serve sessions normally, but they may be slower to join new sessions. Additionally, new agents cannot be created in this region. We are actively working on a mitigation to alleviate this issue. We will provide another update as soon as possible.

None

Some sessions incorrectly shown as active

Started
Fri, Sep 4, 2026, 10:00:00 AM
Updated
Tue, Sep 8, 2026, 08:45:12 AM
Resolved
Fri, Sep 4, 2026, 10:00:00 AM
Duration
0m
  1. resolved

    A subset of room sessions that ended around 4 September 09:57 UTC were not recorded with an end time. These sessions continue to appear as active in the Cloud dashboard and in session records, and no `room_ended` end time is reflected for them. This was a reporting issue only. Rooms ended normally for participants, live traffic was not affected, and there is no impact on usage or billing. The underlying cause has been identified and fixed. Affected sessions from this window may continue to display as active; they can be safely disregarded.

Minor

Investigating reports of one-way audio issues in LiveKit Phone Numbers

Started
Wed, Sep 2, 2026, 09:03:43 PM
Updated
Wed, Sep 2, 2026, 09:47:25 PM
Resolved
Wed, Sep 2, 2026, 09:45:19 PM
Duration
41m
  1. resolved

    This incident has been resolved. We will follow up with a detailed postmortem as soon as possible.

  2. monitoring

    An upstream carrier reverted a codec change that was causing one-way audio issues on calls to LiveKit-issued phone numbers. As of 21:10 UTC calls are connecting normally and we are no longer seeing the issue. The upstream codec change has impacted calls from T-Mobile and AT&T carriers. No other SIP providers were affected. We are monitoring before marking this resolved.

  3. investigating

    We are investigating elevated issues in calls connected to LiveKit Phone numbers where callee audio is not being heard by the caller. This is currently only experienced for numbers issues by LiveKit, and no other SIP provider connectivity should be affected. Our team is actively investigating this.

Major

LiveKit dashboard failing to load

Started
Wed, Sep 2, 2026, 08:15:56 AM
Updated
Wed, Sep 2, 2026, 02:39:44 PM
Resolved
Wed, Sep 2, 2026, 11:57:01 AM
Duration
3h 41m
  1. resolved

    This incident has been resolved and dashboard load times returned to normal by 10:55 UTC. As a note, our beta agent simulations feature was also impacted from approximately 06:07 - 07:52 UTC.

  2. monitoring

    A fix has been deployed as of 10:55 UTC, and the dashboard is now loading successfully. We are continuing to monitor before marking this resolved.

  3. identified

    We have identified the cause: a problematic node in US-East and are deploying a fix. Impact continues to be limited to the dashboard.

  4. investigating

    We are currently investigating reports of our dashboard failing to load across multiple regions, beginning at 07:00 UTC. Other services remain unaffected and the impact is limited to the dashboard.

Minor

Investigating reports of failed SIP Transfer calls in US East

Started
Mon, Aug 17, 2026, 06:10:46 PM
Updated
Tue, Sep 1, 2026, 11:45:42 AM
Resolved
Mon, Aug 17, 2026, 07:22:38 PM
Duration
1h 11m
  1. postmortem

    ## Summary On 17 August, between 17:24 and 18:05 UTC, a small percentage of SIP call transfers failed in our US East region during a routine configuration rollout. An issue in the automation that manages our SIP signaling servers prevented outgoing servers from being taken out of service safely, and those servers shut down while calls were still active on them. Transfers have completed normally since 18:05 UTC. Calls themselves stayed connected, and no other region was affected. ## Root Cause The configuration change restarts the servers that handle SIP signaling one at a time. Before a server is shut down, it is removed from service so that no new calls reach it, and it is then given some time to finish the calls it is already handling. In this case that removal did not complete expectedly and the outgoing servers kept receiving new calls for the entire drain duration, then shut down on schedule with calls still active on them. ## Timeline \(UTC\) * 16:52 - A routine configuration rollout begins in US East. * 17:24 - SIP transfer requests begin to fail for a subset of active calls. * 17:36 - Automated monitoring detects the elevated failure rate and our team begins investigating. * 17:40 - A further subset of active calls is affected. * 17:41 - Replacement capacity comes online. * 18:05 - Last of the impacted calls attempts a transfer and record a failure. ## Scope of Impact Only SIP call transfers \(`TransferSIPParticipant`\) in our US East region were affected, between 17:24 and 18:05 UTC. This represented 0.04% of all active calls in that window, and 1.1% of the calls that attempted a transfer. Customers who were impacted would have seen affected transfer requests return a 408; the underlying call stayed connected and only the transfer failed. Inbound and outbound calling were unaffected, as were calls that did not attempt a transfer, and no other region was affected. ## Mitigations and Follow-ups * We have deployed an alert for calls that end unexpectedly when a server shuts down. * We are changing our rollout process so that a server which cannot be removed from service safely halts the rollout. * We are preventing new calls from being routed to servers that are shutting down. * We are improving monitoring of the automation that manages SIP server rotation. * We are returning a more specific error when a transfer request cannot be delivered.

  2. resolved

    We have identified the root cause of both spikes and have ensured safeguards to prevent a recurrence. We've been monitoring since 18:05 UTC and have seen no further transfer failures, and have confirmed that SIP transfers are operating normally. We will follow up with a detailed post-mortem.

  3. investigating

    No transfer failures has been observed since 18:05 UTC and SIP transfers are currently completing normally. Impact was limited to a small percentage of calls between 15:48-16:08 UTC and 17:24-18:05 UTC. We're actively monitoring while we investigate the cause and put safeguards in place to prevent a recurrence.

  4. investigating

    We're investigating an elevated rate of failures when transferring active SIP calls in the US East region. SIP calls themselves remain connected, and inbound and outbound calling are otherwise operating normally.

None

Elevated inference errors — US East

Started
Wed, Aug 26, 2026, 10:44:00 AM
Updated
Wed, Aug 26, 2026, 12:30:03 PM
Resolved
Wed, Aug 26, 2026, 12:22:01 PM
Duration
1h 38m
  1. resolved

    Between 08:47 and 08:54 UTC, the service that routes LLM, STT and TTS requests through LiveKit Inference in our US East region was unavailable, causing inference requests from agents in that region to fail or time out. Sessions depending on those responses may have degraded or ended prematurely. Replacement capacity came online at 08:54 UTC and request volumes returned fully to normal levels by 09:02 UTC. We have confirmed no further impact since then. The cause was a loss of capacity during automated infrastructure maintenance. We have identified fixes to prevent a recurrence.

  2. investigating

    We are investigating a period of elevated error rates and timeouts (08:47:00-08:53:35 UTC) for LiveKit Inference requests in our US East region. Service has since recovered and inference requests are completing normally. We are continuing to investigate the cause and confirm the full scope and duration of impact. A further update will follow.

Minor

Agent session analytics displaying incorrect values in EU Central and India regions

Started
Mon, Aug 24, 2026, 10:11:45 AM
Updated
Mon, Aug 24, 2026, 02:14:46 PM
Resolved
Mon, Aug 24, 2026, 02:14:46 PM
Duration
4h 3m
  1. resolved

    Agent session analytics in the Cloud dashboard have recovered and current data is reporting correctly for all affected projects. Some gaps may remain in historical data from the affected window (August 21–24 UTC); these will be backfilled.

  2. monitoring

    We are continuing to monitor while the historical data backfill completes, and will resolve this incident once all analytics for the affected window are fully restored.

  3. monitoring

    We identified an issue where agent session analytics in the Cloud dashboard displayed zero concurrent sessions, beginning late on August 21 (UTC). This was a reporting issue only — agent sessions, calls, and all real-time services in these regions operated normally throughout. Ingestion has been restored and we are backfilling historical analytics for the affected window; some charts may show incomplete history until this completes. We are monitoring while the backfill finishes.

None

Egress recordings terminated early in US East

Started
Fri, Aug 21, 2026, 09:00:00 AM
Updated
Sun, Aug 23, 2026, 09:51:15 PM
Resolved
Sun, Aug 23, 2026, 09:00:00 PM
Duration
2d 12h
  1. resolved

    From 09:18 UTC on 21 August to 21:00 UTC on 23 August, 0.008% of egress recordings processed in our US East region were terminated before completion. Affected recordings were marked EGRESS_FAILED ("egress timed out"). For file outputs, media was not written to the configured destination; for HLS and streaming destinations, media was delivered up to the point of failure. The cause was node scale-down evicting egress workers with active recordings. A fix was deployed at 21:00 UTC on 23 August.

Minor

Investigating delayed session data ingestion on Cloud Dashboard

Started
Wed, Aug 19, 2026, 07:33:25 PM
Updated
Thu, Aug 20, 2026, 12:32:36 AM
Resolved
Thu, Aug 20, 2026, 12:30:54 AM
Duration
4h 57m
  1. resolved

    This incident has been resolved. Between 19:20 and 00:10 UTC, session data and observability ingestion was delayed by 45-60 min. The root cause has been addressed and data ingestion returned to normal by 00:10 UTC. No data was lost during this window.

  2. monitoring

    Mitigations have been deployed and we are now catching up to realtime.

  3. identified

    The amount of data we are ingesting in the sessions view is outpacing our ability to ingest, thus causing a delay in ingestion. We are pursuing a few mitigations right now including moving to a larger database instance to relieve the processing bottleneck.

  4. identified

    We are continuing to work on a fix for this issue.

  5. identified

    Sessions data ingest remains delayed for about an hour. It's not impacting other dashboards. We've identified the source of the slow ingest and are working on a mitigation. We will update again in 30 mins.

  6. investigating

    We are currently investigating delayed session data ingestion on Cloud Dashboard. No services appear to be impacted and we don't expect any data to be lost.

Minor

Hosted agent builds failing in US East

Started
Fri, Aug 14, 2026, 09:10:19 PM
Updated
Mon, Aug 17, 2026, 05:15:39 PM
Resolved
Fri, Aug 14, 2026, 11:01:30 PM
Duration
1h 51m
  1. postmortem

    ### Summary An issue with the image registry which powers our hosted agents offering temporarily caused new hosted agent builds in US East to fail. Existing agent workloads were not affected, and no sessions or calls were impacted. ### Timeline \(UTC\) * 19:54 - Hosted agent builds begin failing in US East * 20:48 - Our team noticed the increased failure rate and began investigating * 21:04 - First mitigation deployed * 21:08 - Most builds succeeding * 22:14 - Second mitigation deployed, covering the remaining affected infrastructure * 22:24 - All builds succeeding ### Scope of Impact Only new hosted agent builds and deploys in US East were affected. Agents already running were unaffected, as were all sessions, calls, and other regions. Customers who were impacted would have seen their deploy fail with “unable to deploy agent: an error occurred while building your agent, please try again”. ### Mitigations and Follow-ups * We are improving alerting on build failures. * We are improving monitoring of our downstream image registries. ‌ We appreciate your understanding and are committed to continuously improving our platform's reliability. If you have any questions, please reach out to our support team.

  2. resolved

    We pushed a second mitigation at 22:14 UTC and all US East builds have been successful since 22:25 UTC. We will follow up with a postmortem describing the timeline, root cause, and follow ups as soon as possible.

  3. monitoring

    The total build error rate has decreased significantly due to the mitigation, but a small percentage of builds in US East are still failing. We are continuing to monitor and will follow up with another update as soon as possible.

  4. monitoring

    We have deployed a mitigation and builds are now successful. We will continue monitoring while we collect any further details to share.

  5. identified

    New hosted agent builds are failing in US East. Existing agent workloads are not affected. We are currently deploying a mitigation and will follow up with more information shortly.

Minor

Investigating longer than usual dashboard loading times

Started
Thu, Aug 13, 2026, 08:42:26 PM
Updated
Fri, Aug 14, 2026, 12:15:44 AM
Resolved
Fri, Aug 14, 2026, 12:15:44 AM
Duration
3h 33m
  1. resolved

    We have resolved the incident as session load times have returned to baseline.

  2. monitoring

    We are seeing a 5-10 minute delay in loading some sessions, but we are confident that no data is being lost. We are currently monitoring and will update again once ingestion times have returned to baseline.

  3. investigating

    We are currently investigating longer than usual dashboard loading times. No services appear to be impacted and we don't expect any data to be lost.

None

Investigating alerts for increased timeout error rates on ListParticipants APIs

Started
Wed, Aug 12, 2026, 06:34:59 PM
Updated
Wed, Aug 12, 2026, 08:10:17 PM
Resolved
Wed, Aug 12, 2026, 08:09:42 PM
Duration
1h 34m
  1. resolved

    Error rates on ListParticipants API requests recovered at 18:58 UTC and have remained stable since.

  2. monitoring

    Error rates on ListParticipants API requests returned to normal as of 18:58 UTC. We will follow up with more details as soon possible.

  3. investigating

    We are currently investigating alerts for elevated server error rates on ListParticipants APIs across multiple regions. The impact appears to be intermittent and limited to a subset of room-management API requests. We will follow up with more details as soon possible.

Minor

A percentage of cross-region calls between US West and US Central not completing

Started
Fri, Aug 7, 2026, 09:22:30 PM
Updated
Fri, Aug 7, 2026, 11:07:30 PM
Resolved
Fri, Aug 7, 2026, 10:27:50 PM
Duration
1h 5m
  1. resolved

    This incident has been resolved. Between 19:45 and 20:00 UTC, calls crossing the US West <-> US Central link experienced intermittent degradation in Room Service and SIP performance. The underlying network between the two regions had episodes of high packet loss during this time period. We rerouted traffic and performance has returned to baseline as of 22:00 UTC. We will follow up with a postmortem, which will include further details regarding scale and scope.

  2. monitoring

    We have applied a mitigation to the affected regions and are now monitoring to ensure that performance signals return to baseline. We are still working on determining impact (including how users can determine if they were affected) and will share those details as soon as possible.

  3. investigating

    We are currently investigating automated alerting which triggered for degraded Room service performance in the US West and US Central regions beginning around 19:45 UTC. We are working to determine if there is any customer-facing impact. If there is impact, we don't currently have reason to believe that it is widespread. We will update again within 30 minutes.

Minor

Elevated error rate on LiveKit Inference (google/gemma-4-31b-it)

Started
Thu, Aug 6, 2026, 07:47:33 PM
Updated
Fri, Aug 7, 2026, 02:51:28 AM
Resolved
Fri, Aug 7, 2026, 02:51:27 AM
Duration
7h 3m
  1. resolved

    Today there was an issue with one of our GPU providers that caused it run at about 50% capacity. Requests above capacity return 503s, which we route to another, fallback provider. This fallback provider was not able to rise to the occasion and also returned 503s. In the very near future we will be onboarding additional providers to avoid these kinds of issues.

  2. monitoring

    A burst of requests increased the failure rate. Both primary and secondary model providers failed to fulfill the requests at the time; we are root-causing the issue. We are continuing to monitor.

  3. identified

    We are actively investigating elevated error rate on LiveKit Inference (google/gemma-4-31b-it)

  4. identified

    Between approximately 18:45 and 19:15 UTC today, a subset of LiveKit Inference requests using the google/gemma-4-31b-it model returned errors. Affected agent sessions would have seen an LLM request error on those requests. Other models and other LiveKit services were not affected. No action is needed on your part. If you continue to see errors, please reach out to support. We apologize for the disruption.

Minor

Elevated timeouts on LiveKit Inference (google/gemma-4-31b-it)

Started
Tue, Aug 4, 2026, 08:45:37 PM
Updated
Tue, Aug 4, 2026, 10:11:21 PM
Resolved
Tue, Aug 4, 2026, 10:11:21 PM
Duration
1h 25m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We are monitoring elevated timeouts that affected LiveKit Inference between 14:05 and approximately 15:00 UTC today. During that window, roughly 3% of LLM requests using the google/gemma-4-31b-it model timed out. Impacted customers would have seen LLM request timeouts in their agent sessions (for example, APITimeoutError in agent logs). Other models were not affected. We have identified the cause. During a period of increased traffic, response times for this model slowed, and a defect in our health-monitoring logic incorrectly marked a backup deployment as unhealthy and removed it from rotation, preventing requests from failing over as designed. Error rates returned to baseline by approximately 15:00 UTC, and a fix for the health-monitoring defect is in progress. We are continuing to monitor before marking this resolved. No action is needed on your part. If you continue to see timeouts, please reach out to support. We apologize for the disruption.

Minor

Sessions view not populating in the Cloud Dashboard

Started
Mon, Aug 3, 2026, 03:30:50 PM
Updated
Mon, Aug 3, 2026, 07:48:48 PM
Resolved
Mon, Aug 3, 2026, 07:48:48 PM
Duration
4h 17m
  1. resolved

    We have fixed the data pipeline delay and all sessions should be loading now in the Cloud Dashboard.

  2. investigating

    We have traced the underlying issue to a backlog in our data pipeline, and we are continuing to investigate the root cause while working to clear the backlog. This data ingest delay results in the Sessions view returning no results for any time range ending at the current time, including all Quick Ranges options in the dashboard. As a temporary workaround, select "Custom range" and set the end time at least 30 minutes in the past to load your sessions. Real-time connectivity is not affected and no data has been lost.

  3. investigating

    We are currently investigating an issue where the sessions view does not populate in the Cloud Dashboard. This issue is specific to the dashboard display; all LiveKit services are operating normally and no sessions or sessions data are affected.

None

Intermittent webhook delivery failures

Started
Thu, Jul 30, 2026, 10:55:56 PM
Updated
Thu, Jul 30, 2026, 10:57:22 PM
Resolved
Thu, Jul 30, 2026, 10:55:56 PM
Duration
0m
  1. postmortem

    ## Summary Between 2026-07-16 and 2026-07-29, a caching bug caused webhooks to be intermittently not delivered for about one hour for projects whose configuration was updated during this period. ## Root Cause A bug in an internal configuration service caused webhook settings to be dropped from a project's cached configuration whenever that project's record was updated. Once triggered, our system stopped delivering webhooks until the cache refreshed from the database, up to one hour later. Because of this caching behavior, the impact was intermittent: webhooks would stop, then resume on their own, then stop again after the next project update. Additionally, the "test webhook" feature in the dashboard reads webhook configuration through a different path and was unaffected, so test webhooks succeeded even while real event delivery was failing. ## Timeline \(UTC\) * 2026-07-16 23:48 - Change containing the bug reached production; intermittent impact begins * 2026-07-24 - Some customer reports of missing webhooks received; investigation begins * 2026-07-28 16:23 - Root cause identified; fix developed and rollout begins * 2026-07-29 16:26 - Fix deployed to all regions; incident resolved ## Scope of Impact Customers using webhooks were affected across all regions between 2026-07-16 and 2026-07-29, with most impact occurring after 2026-07-22. A project was affected only if its configuration was updated during the window, and each occurrence lasted up to one hour. Affected projects saw `room_started` and `participant_*` webhooks silently not delivered, while `room_finished` webhooks and dashboard webhook tests often continued to work. If you observed missing webhooks and had recently made any settings changes to your project, you were likely impacted. ## Mitigations and Follow-ups * The bug has been fixed and a regression test has been added. * We added logging of webhook delivery response codes to improve visibility into delivery failures. * We are adding per-project webhook delivery monitoring and alerting that can detect missing deliveries independent of overall volume. * We are unifying the configuration paths used by webhook delivery and the dashboard webhook test, so a passing test reflects actual delivery behavior. * We are making the configuration cache lifetime adjustable so similar issues can be mitigated quickly in the future.

  2. resolved

    Between 2026-07-16 and 2026-07-29, a caching bug caused webhooks to be intermittently not delivered for about one hour for projects whose configuration was updated during this period.

None

Investigating longer than usual dashboard loading times

Started
Thu, Jul 30, 2026, 06:53:17 PM
Updated
Thu, Jul 30, 2026, 10:25:24 PM
Resolved
Thu, Jul 30, 2026, 10:25:24 PM
Duration
3h 32m
  1. resolved

    An internal one-off analysis job ran some expensive DB queries which slowed down load times of the dashboard's sessions page. The job was removed and load times are now back to baseline.

  2. investigating

    We are currently investigating longer than usual dashboard loading times. No services appear to be impacted and we don't expect any data to be lost.

None

Elevated SIP API error rates in the London region

Started
Tue, Jul 28, 2026, 02:23:20 PM
Updated
Tue, Jul 28, 2026, 02:46:33 PM
Resolved
Tue, Jul 28, 2026, 02:23:20 PM
Duration
0m
  1. resolved

    Between approximately 10:30 and 11:55 UTC on July 28, customers with telephony workloads in our London region experienced elevated error rates on SIP APIs. Small number of call transfers were affected. We mitigated by removing the affected infrastructure from service at 11:48 UTC, and error rates returned to normal by 11:55 UTC.

None

Elevated connectivity issues in US West

Started
Mon, Jul 27, 2026, 05:00:00 PM
Updated
Mon, Jul 27, 2026, 06:10:46 PM
Resolved
Mon, Jul 27, 2026, 05:00:00 PM
Duration
0m
  1. resolved

    On July 27, our San Jose region experienced a minor networking incident beginning at 16:14 UTC. Some users connected to this region may have experienced disconnects and increased session start times. Traffic was routed away from this region by 16:39 UTC and errors have since returned to baseline.

None

Increased egress latency and media connection interruptions in London region

Started
Thu, Jul 23, 2026, 08:00:00 AM
Updated
Sat, Jul 25, 2026, 10:47:03 AM
Resolved
Thu, Jul 23, 2026, 08:00:00 AM
Duration
0m
  1. resolved

    On July 23, between 07:50 and 08:07 UTC, media servers in our London region experienced a host-level network fault, causing them to fail health checks and restart. Users connected to these servers may have experienced brief connection interruptions, and egress operations may have seen increased latency. The affected servers were removed and replaced. The region has been operating normally since.

Minor

Elevated RoomService API error rates in London

Started
Tue, Jul 14, 2026, 09:49:13 AM
Updated
Tue, Jul 14, 2026, 02:52:33 PM
Resolved
Tue, Jul 14, 2026, 10:41:18 AM
Duration
52m
  1. resolved

    Error rates for room service api requests in our London region have remained normal since 09:06 am UTC and we are no longer seeing any issues.

  2. monitoring

    Between 08:54 am and 09:06 am UTC, 5.4% of room service api requests in our London region failed with errors. A fix was applied and error rates returned to normal by 09:06 am UTC. We are continuing to monitor.

Minor

Elevated connection failures in Japan

Started
Mon, Jul 6, 2026, 06:40:24 PM
Updated
Tue, Jul 7, 2026, 10:58:13 AM
Resolved
Tue, Jul 7, 2026, 10:30:30 AM
Duration
15h 50m
  1. resolved

    Users connecting to our Japan region experienced elevated connection failures between approximately 18:18 and 18:43 UTC, caused by a network disruption affecting connectivity between our Tokyo data centres. Connectivity was restored by 18:43 UTC. Update (July 7): After the initial recovery, one replacement media server in Tokyo came online serving only IPv6. From approximately 22:00 UTC on July 6, connections routed to this server failed for clients on networks without IPv6 connectivity. The server was removed from service at 10:30 UTC on July 7, and connection success rates in Japan have returned to normal.

  2. investigating

    We are currently investigating increased connection failures affecting our Japan region beginning at 18:18 UTC.

Minor

Elevated rate of egress ending early in India region

Started
Tue, Jun 16, 2026, 09:49:00 AM
Updated
Tue, Jun 23, 2026, 05:48:25 PM
Resolved
Tue, Jun 16, 2026, 08:42:35 PM
Duration
10h 53m
  1. postmortem

    **Summary** Between June 5 and June 16, 2026, an intermittent loss of connectivity between internal systems caused room state to be altered incorrectly in a single region in India. As a result, a small number of active sessions in that region ended earlier than expected, and affected egress recordings could stop mid-session. While the intermittent issue was undetected for some time, the impacted number of egresses was low and limited to the one affected region in India - other regions were unaffected. **Timeline \(UTC\)** * 2026-06-05 05:00 - Intermittent loss of connectivity between internal systems begins causing room state to be incorrectly altered in one India region, leading to some egresses being terminated prematurely. * 2026-06-16 06:45 - Mitigation was deployed and monitoring began. Eligible connections were moved to a different region and added redundancy was deployed to the affected internal tooling. * 2026-06-16 20:42 - No further instances were observed since the mitigation was applied. Incident was marked resolved. **Resolution** We mitigated the outage by moving eligible connections to a different region and by adding redundancy to the affected internal tooling so it could better tolerate intermittent connectivity loss. As a longer-term fix, we hardened the communication channels between the involved systems. We have not observed any further occurrences since the mitigation was applied, and we are continuing to evaluate additional improvements to our tolerance for this class of failure.

  2. resolved

    We have not seen further instances of this issue since the mitigation was applied. We will follow up with a postmortem as soon as possible.

  3. monitoring

    We identified an issue in the India region where a small number of active sessions may have ended earlier than expected. Affected sessions could experience egress recordings stopping mid-session. We've deployed a mitigation and are monitoring to confirm resolution. Other regions are unaffected.

Major

Connection failures and SIP transfer errors in US Central

Started
Mon, Jun 22, 2026, 01:05:45 PM
Updated
Tue, Jun 23, 2026, 05:21:14 PM
Resolved
Mon, Jun 22, 2026, 02:40:26 PM
Duration
1h 34m
  1. postmortem

    ## Summary LiveKit Cloud is a globally distributed system. Our data centers continuously probe one another, and our overlay network uses those probes to choose healthy paths between regions. When a brief network disruption occurs in or between data centers, the overlay is designed to mark affected paths as unhealthy, route traffic around them, and return them to "healthy" once the underlying network has cleared. On June 18, 2026 \(US East 1\) and again on June 22, 2026 \(US Central\), a short packet-loss event in the affected data center triggered a compound failure in this mechanism. Two separate defects, one in the overlay network and one in our link-monitoring service, interacted in a way that left the affected region effectively isolated from other regions long after the underlying network had cleared. For most customers and most products, the impact was bounded. Our global routing automatically moved traffic to other US regions within approximately one minute of onset. SIP transfers anchored in the affected region took longer to recover but cleared within approximately thirty minutes. A small number of LiveKit Agents running self-hosted workers experienced an extended outage in agent dispatch, up to about two hours in some cases. The majority of agent workloads continued to receive jobs normally after the routing flip. The workaround at the time was to restart the affected agent worker process so that it re-registered against a controller in a healthy region. This is a separate defect in our agent dispatch routing, and we are fixing it. This post-mortem describes what happened, why the second incident occurred so soon after the first, and the steps we are taking to prevent recurrence. ## Impact * **Realtime connections:** minimal impact. Clients automatically reconnected to alternative data centers. * **API:** approximately 1 minute of degraded availability while traffic rerouted to a nearby region. * **SIP transfers:** approximately 30 minutes for in-progress transfers anchored in the affected region. * **Agent dispatch:** approximately 1 minute for the majority of agents. A small number of agent workers saw extended dispatch outages of up to approximately 2 hours until the worker process was restarted. * **Egress**: A few number of in-progress egresses running in the affected regions were ended prematurely during the approximately 1-minute outage window, between 16:47 and 16:48 UTC on June 18, and between 12:24 and 12:25 UTC on June 22. These egresses were incorrectly marked as successful and no error was surfaced to indicate the recording had failed. ## Root Cause The trigger in both incidents was the same: a short packet-loss event in the affected data center. The expected behavior was that affected paths would be marked unhealthy briefly while the overlay rerouted traffic around them, and would return to "healthy" once the underlying network cleared. Two separate defects interacted in a way that kept routing state stuck long after the network had recovered: 1. **Inbound and outbound freshness collapsed onto a single indicator.** In the overlay network, the two directions of a connection \(inbound and outbound\) shared a single "freshness" indicator. When fresh data arrived in one direction, our code incorrectly assumed data in the other direction was also fresh. As a result, the affected region continued to treat its outbound connections to other regions as unreachable based on stale inbound data, even after outbound traffic was not impacted. 2. **Link-monitor key exchange could de-sync during a disruption.** Our link-monitoring service relies on a key exchange between regions to validate probe traffic. The packet-loss event caused this key exchange to fall temporarily out of sync between the affected region and its neighbors. With probes failing validation, the link monitors in neighboring regions marked the affected region as "to avoid," and held that state past the actual network recovery. Either defect on its own would not have produced an outage. Together, they produced a stuck state in both directions: neighboring regions believed they could not reach the affected region \(because probe validation kept failing\), and the affected region believed it could not reach the others \(because the freshness collapse kept its outbound state pinned to "unreachable"\). Traffic stayed routed away from the affected region until the state was manually cleared. Following Incident 1 on June 18, we developed and merged a fix for the overlay freshness defect on the same day. The fix shipped as a new version of the overlay software and was being validated in staging at the time of Incident 2. On June 22, the same compound failure occurred on US Central, which was still running the unpatched version of the overlay. At the time of writing, the fix is now in active rollout to production. ## Corrective Actions & Prevention **Immediate \(in active rollout\)** * **Roll the overlay freshness fix to production region-by-region.** As of this writing, the rollout is underway and we expect global production coverage in the next two days. We are giving this rollout priority given the severity. **Underway** * **Improve resilience of the link-monitor probing process to key-exchange de-sync.** This addresses the second of the two defects described above. Eliminating it is necessary to prevent the compound failure even after the overlay freshness fix is in place. _ETA: Thursday, June 25._ * **Migrate existing agent worker connections on health changes.** We correctly detected the data center as unhealthy and rerouted new traffic away quickly, but did not migrate existing agent worker connections. The change is to automatically migrate agent worker connections to healthy data centers as soon as the health indicator starts failing. _ETA: June 30._ * **Make egress resilient to region disruptions.** In-progress Egress should use the same mechanism available in our realtime SDK to reconnect to alternative regions. _ETA: June 30._ ## For customers running self-hosted agent workers Until our agent dispatch routing fix is deployed, the most reliable mitigation if you observe a dispatch outage is to restart your agent worker processes. This causes them to re-register against controllers in healthy regions. To catch the condition early and automate the recovery, consider adding a health check that: * Tracks the time since the worker last received a dispatch \(or last completed a session\). * Triggers a restart of the worker process if this exceeds an expected idle window for your workload. We recognize that this workaround places a burden on customers and is not a substitute for the underlying server-side fix. Eliminating it is a priority. We sincerely apologize for the disruption to customers whose traffic and agent workloads were affected. Thank you for your patience, and we welcome any additional feedback from customers who were impacted.

  2. resolved

    We are resolving this incident as the mitigation is in place and successful. We have received some reports of self-hosted agents which connect to US Central needing to be restarted in order to continue receiving dispatches and will investigate opportunities to improve this behavior. Apologies for the disruptions and we will follow up with a postmortem as soon as possible.

  3. monitoring

    Update: As soon as the US Central region became unresponsive around 12:25 UTC, traffic was automatically re-routed to the nearest healthy region. Customers may have noticed failed API requests for approximately 1 minute while the re-route took place. Failed SIP Transfers continued until we finished manually draining the impacted region at 13:09 UTC. We are also investigating impact to Agent Dispatches and will post another update when we know more.

  4. monitoring

    We have begun routing traffic away from US Central after our monitoring system triggered due to a high API error rate. Customer impact may include WebRTC connection failures and SIP transfer errors.

Minor

Investigating connection failure and SIP transfer error reports in US East

Started
Thu, Jun 18, 2026, 05:09:32 PM
Updated
Fri, Jun 19, 2026, 02:41:49 PM
Resolved
Thu, Jun 18, 2026, 05:40:13 PM
Duration
30m
  1. resolved

    TransferSIPParticipant for calls happening in US East required a manual intervention to redirect. That API is fully restored as of 17:18 UTC. We will follow up with a postmortem as soon as possible.

  2. monitoring

    US East has experienced a networking failure. Traffic was automatically rerouted away from US East within a minute and connection failure errors are back to baseline. The disruption started happening as of 16:47 UTC, and traffic was rerouted at 16:48. We are continuing to investigate and monitor SIP transfer errors.

  3. investigating

    We are currently investigating connection failure alerts in US East.

Minor

Partial Outage in LiveKit Cloud Dashboard

Started
Thu, Jun 18, 2026, 05:49:37 PM
Updated
Thu, Jun 18, 2026, 07:10:58 PM
Resolved
Thu, Jun 18, 2026, 07:10:58 PM
Duration
1h 21m
  1. resolved

    As of 18:45 UTC, all features of Cloud Dashboard are fully restored and functional.

  2. investigating

    Session data (except Agent Insights) on the dashboard can still be accessed by changing the timeline on the Sessions page to "Past 7 days" or greater. We're still working to restore full access.

  3. investigating

    Due to networking issues faced in our US East region, certain features of Cloud Dashboard are not accessible, including Sessions page and Agent Insights. There is no data loss or ongoing session impact associated with this. Our team is actively working on restoring the service.

None

Certain dashboard services partially unavailable

Started
Thu, Jun 11, 2026, 01:00:00 PM
Updated
Wed, Jun 17, 2026, 09:59:59 PM
Resolved
Thu, Jun 11, 2026, 01:00:00 PM
Duration
0m
  1. resolved

    On 2026-06-11 around 13:00 UTC, a serialization bug was introduced via frontend components which rendered some services unavailable via the dashboard for some users. These services included SIP trunk and dispatch rule creation/management as well as several advanced features in Agent Builder. A fix was pushed on 2026-06-17 at 17:43 UTC which resolved the issue.

Minor

Intermittent errors for Google Gemini models via Inference

Started
Thu, Jun 11, 2026, 05:52:59 PM
Updated
Thu, Jun 11, 2026, 11:47:09 PM
Resolved
Thu, Jun 11, 2026, 11:47:09 PM
Duration
5h 54m
  1. resolved

    This incident has been resolved. The errors were caused by an upstream issue at our model provider (Google) that incorrectly triggered an account-level usage cap on our Gemini API access, returning rate-limit errors for a subset of Gemini requests routed through LiveKit Inference. Google identified the cause, rolled back the change on their side, and raised our account limits to prevent recurrence. Gemini requests are now serving normally and error rates have returned to baseline. Other models and providers were unaffected throughout.

  2. identified

    We have identified the cause as an upstream limit on our Google Gemini API account. Automatic failover to an alternate Gemini deployment is in place and has reduced the impact, but a subset of requests to Google Gemini models may still intermittently return errors when failover capacity is exceeded. We are actively engaged with Google support to restore full capacity and will share further updates as we have them. Other models and providers remain unaffected.

  3. investigating

    We are currently investigating intermittent 429 errors for Google Gemini models routed through LiveKit Inference.

None

Elevated errors for Google Gemini models via Inference

Started
Thu, Jun 11, 2026, 03:30:00 PM
Updated
Thu, Jun 11, 2026, 03:51:30 PM
Resolved
Thu, Jun 11, 2026, 03:30:00 PM
Duration
0m
  1. resolved

    Between 13:00 and 15:05 UTC, a subset of requests to certain Google Gemini preview models (gemini-3.1-flash-lite and gemini-3-flash-preview) routed through LiveKit Inference returned errors due to a project-level misconfiguration. These preview models did not yet have model failover configured; all other models and providers were unaffected. This was resolved by routing the traffic via an alternate provider and fixing the misconfiguration. As resolution measures, we are extending model-level failover to cover these models and adding increased monitoring for such failures in the future.

Minor

Elevated latency in Hyderabad region

Started
Fri, Jun 5, 2026, 09:06:34 AM
Updated
Fri, Jun 5, 2026, 11:08:32 AM
Resolved
Fri, Jun 5, 2026, 11:08:32 AM
Duration
2h 1m
  1. resolved

    We identified a connectivity problem in the Hyderabad cluster. As mitigation, we've drained the cluster to route traffic away from it. Mumbai traffic is unaffected and serving normally. We'll undrain Hyderabad once we're confident it's functioning normally.

  2. investigating

    As of 09:21 UTC, we've applied a mitigation by routing traffic away from the Hyderabad region. Latency rates are returning to normal levels. We're continuing to monitor the issue.

  3. investigating

    We're investigating elevated latency affecting Ingress and SIP services in our Hyderabad region.

None

Degraded SIP Connectivity in India

Started
Wed, Jun 3, 2026, 08:00:00 PM
Updated
Thu, Jun 4, 2026, 11:53:13 PM
Resolved
Wed, Jun 3, 2026, 08:00:00 PM
Duration
0m
  1. resolved

    Summary On June 2, 2026, a subset of users placing or receiving SIP calls through LiveKit Cloud's India region experienced call failures between 12:00 and 12:40 UTC. During a routine maintenance, both of our SIP-serving regions in India were temporarily disabled in error. With no SIP capacity available in India, traffic was routed to our Dubai region, where outbound calls to Indian numbers were rejected at the destination for region-pinned projects and inbound calls targeting India-specific endpoints failed to connect. We mitigated this by re-enabling SIP in the India region and restoring normal capacity. Timeline (UTC) 12:00 — Both India SIP regions disabled during maintenance; start of measurable customer impact 12:40 — India SIP regions re-enabled; customer impact ends Root cause The incident was caused by an operational error during planned maintenance. LiveKit operates three data centers in India, two of which serve SIP traffic. While disabling regions as part of our routine maintenance, the operator disabled both SIP-serving data centers, believing the third data center was also running SIP and would continue to carry customer traffic. Because the third data center does not serve SIP, this removed all SIP capacity in the India region and forced traffic to our Dubai region, where it could not complete region-pinned calls. Resolution Once we identified that India SIP capacity had been removed, we re-enabled SIP in the affected regions, which restored normal inbound and outbound calling in the region.

Minor

Degraded Connectivity in US Central

Started
Tue, Jun 2, 2026, 02:05:52 PM
Updated
Thu, Jun 4, 2026, 02:03:13 PM
Resolved
Tue, Jun 2, 2026, 03:18:08 PM
Duration
1h 12m
  1. postmortem

    **Summary** On June 2, 2026, a small subset of users connected to LiveKit Cloud US Central region experienced elevated connection failures between 12:06 and 14:38 UTC. A single node in the region entered a degraded state in which it continued to accept new connections but could not reliably complete the real-time media connection those calls depend on. As a result, a portion of participant connections, including inbound and outbound SIP calls routed through Chicago failed to connect or dropped shortly after starting. We mitigated this by routing SIP traffic away from the affected server, suspending and removing the faulty node, and restoring normal service to the region. **Timeline \(UTC\)** * 12:06 — Start of measurable customer impact * 14:04 — Chicago SIP traffic drained as mitigation * 14:38 — Problematic server suspended; customer impact ends * 15:07 — Chicago SIP service fully restored **Root cause** The incident was caused by a "gray failure" of a single node in the Chicago region, a partial failure in which a server appears healthy to automated systems but is not actually functioning correctly. The server continued to be assigned new calls and reported itself as available, but could not reliably bring the underlying real-time media connections to a fully active state. **Resolution** We first drained SIP traffic out of the Chicago region to route the traffic to other regions, and once the failing node was identified, we suspended and removed the faulty node. After confirming error rates had returned to normal, we restored full SIP service to the region. **Corrective actions and prevention** We are actively addressing the two gaps that allowed a single failing node to impact the region: * **Error aversion in load balancing:** configuring our load balancing servers to automatically detect and steer traffic away from nodes exhibiting this class of failure. * **Monitoring:** closing the alerting gap so that "gray failure" servers which appear healthy but cannot complete connections are detected and alerted promptly.

  2. resolved

    This incident has been resolved.

  3. monitoring

    We've identified the root cause as a single problematic media node in US Central, which has been suspended at 14:38 UTC. We are continuing to monitor before marking this resolved.

  4. monitoring

    We have routed the traffic away from the US Central region at 14:00 UTC and are seeing the connection failures returning to normal levels. We are continuing to monitor the issue.

  5. investigating

    We're investigating elevated SIP call connection failures in our US Central region beginning at ~12:06 UTC. We are working to mitigate this issue.

Critical

Elevated Reports of Participant Connection Latency and Errors In US East Region

Started
Thu, May 28, 2026, 02:53:26 PM
Updated
Sat, May 30, 2026, 06:19:11 PM
Resolved
Thu, May 28, 2026, 08:13:30 PM
Duration
5h 20m
  1. postmortem

    ## Summary LiveKit Cloud validates incoming API and connection requests via an internal authentication service, accessed over gRPC. The service has been in place since 2022 and is designed for resilience: it runs as multiple redundant instances, has pod-level health monitoring that automatically reaps unhealthy pods, and is backed by in-memory caches so it can survive transient database failures. On 2026-05-28, between 13:55 UTC and 15:45 UTC, a percentage of requests and new connections in our US East region failed or timed out. The root cause was a rare failure mode on a single instance of the authentication service. The instance remained reachable and its TCP connections stayed alive, but it began responding to gRPC requests extremely slowly, in a way that did not trip our existing pod-level health checks or cause gRPC clients to fail over. We sincerely apologize to customers whose traffic was disrupted. We've let you down, and we are taking this very seriously. In addition to the corrective actions outlined below, we are performing a thorough audit of our systems for other failure modes we may not have anticipated. ## Timeline and Impact \(all times in UTC on 2026-05-28\) * **13:55**: Elevated rate of authentication failures observed. * **14:15**: On-call team notified; investigation begins. * **14:46**: Error rate continues to climb; incident escalated to Sev 1. * **15:35**: US East drained; error rate begins to subside. * **16:32**: Root cause identified; faulty instance shut down. * **16:32**: Webhooks backlogged during the incident begin to deliver. Between 2026-05-28 13:55 UTC and 15:35 UTC, a percentage of API requests and new connection attempts in our US East region failed or timed out. This included new participant connections, SIP connection requests, and agent connections originating in US East. The blast radius of this incident was substantially reduced by our local in-process auth cache. Each service maintains an in-memory cache of recently-seen authentication details, which allowed the majority of requests to continue flowing whenever they landed on a server that had already cached the relevant credentials. The failures concentrated on requests that landed on servers without those credentials cached, which disproportionately impacted customers with lower overall traffic to US East. Existing realtime sessions in US East continued to operate, and all traffic in other regions was unaffected. The impact subsided once US East was drained and traffic was rerouted to neighboring regions. There was also a secondary effect for customers who subscribe to outgoing webhooks. During the incident, webhook deliveries that depended on the affected auth path were queued internally while waiting for a valid signing token. Once the faulty instance was removed and the previously stuck operations unblocked, the queued webhooks were delivered in a burst rather than smoothly over time. Customers whose webhook endpoints have limited throughput may have observed a short-lived spike that exceeded their normal load. ## Root Cause The trigger was a rare hardware failure mode on one instance of our authentication service. Typically, when hardware fails, the underlying machine is shut down either by the hypervisor or by Kubernetes \(via failed health checks\), and gRPC clients reconnect to a healthy instance. In this case, the faulty machine did not terminate. It remained reachable and continued to respond, but extremely slowly. The TCP connection to that instance stayed up, so gRPC clients that had pinned their requests onto that connection continued to use it, and pod-level health checks did not fire. As a result, a percentage of services in US East were unable to fail over to healthy instances and saw their auth requests time out. Two additional fallback mechanisms did not engage as fully as we had designed: 1. **Client-driven cross-region failover.** LiveKit clients can automatically fail over from one region to a neighboring healthy region. However, the failover path itself depended on passing the same authentication gate, so it failed for the same reason. 2. **Local in-process auth cache refresh path.** As noted in the Impact section, the in-process cache absorbed a significant share of requests during the incident. However, services always attempt to refresh auth details against the auth service first to ensure they have the most recent data, and fall back to the cache only when that refresh fails or times out. The refresh attempt has a 5-second timeout. Many client SDKs use a default request timeout shorter than 5 seconds and gave up before the cache fallback could engage, even when the relevant credentials were present in the cache. When the impact was first observed, our initial signal pointed to database connection timeouts in US East. Because of that, we did not immediately drain the region. Two reasons informed that decision: 1. only a portion of requests were failing \(some connections were still succeeding\) 2. if the underlying database had truly been the cause, draining US East would have shifted that load onto neighboring regions and risked spreading the impact rather than containing it. After confirming that the database was healthy and that the issue was isolated to US East, we drained the region. Error rates began dropping immediately as traffic moved to neighboring regions. ## Corrective Actions & Prevention This incident exposed a machine failure mode we had not yet designed for. The following changes have been implemented and will be rolled out in the next week: * **Detect and time out slow gRPC connections from internal clients.** Internal gRPC clients will detect connections that are alive but unresponsive, and tear them down so they can re-establish against a healthy instance. * **Improved load balancing between gRPC servers.** Retries from internal clients will spread across alternative instances, so a single slow server cannot black-hole a meaningful share of traffic. * **Reduce inline auth refresh timeout to 1 second.** This allows the local in-process auth cache fallback to engage well within the timeout budget that client SDKs typically use. * **Allow client-driven cross-region failover to function when auth has failed.** Region failover will no longer depend on passing the auth gate, so a degraded auth service in one region cannot prevent clients from reaching a healthy region. We sincerely apologize for the disruption to customers whose traffic was affected. Thank you for your patience, and we welcome any additional feedback from customers who were impacted.

  2. resolved

    Connection errors have remained cleared since the US East drain at 15:44 UTC, and as of 19:30 UTC, US East is back online. We'll share a detailed RCA in the coming days.

  3. monitoring

    Service is operating normally with traffic routing through other regions. We are working on bringing US East back online and will share a detailed RCA once complete.

  4. monitoring

    Connection errors have fully cleared since the US East drain completed at 15:44 UTC. We'll continue monitoring. We will be sharing a detailed RCA.

  5. monitoring

    US East drain is complete and all new traffic is now routing to other regions. Error rate is trending down and we'll continue monitoring.

  6. identified

    We are observing database connection timeouts in the US East region and are currently seeing an impact to all services in the region. We are currently draining the region and routing away all traffic to other regions.

  7. investigating

    We are currently investigating reports of spikes in participant connection latency and errors

Major

Partial data outage on LiveKit Cloud Dashboard

Started
Tue, May 26, 2026, 07:22:21 PM
Updated
Tue, May 26, 2026, 09:32:19 PM
Resolved
Tue, May 26, 2026, 09:32:19 PM
Duration
2h 9m
  1. resolved

    All missing data has now been backfilled and the incident is resolved.

  2. monitoring

    A fix has been implemented and we are seeing no more data outage for newly created sessions. We are monitoring the fix and are currently working on backfilling the missing data.

  3. identified

    We have identified an issue causing partial data unavailability in LiveKit Cloud Dashboard: - For some users, Agent Observability is temporarily unavailable for some sessions. - Some sessions in the 24-hour filter are showing as Active even though they have ended. A fix has been identified and we are backfilling the missing data.

  4. investigating

    We are investigating reports of Agent Observability not accessible for some sessions.

None

Network instability in Frankfurt

Started
Fri, May 22, 2026, 08:26:01 PM
Updated
Sat, May 23, 2026, 04:36:55 PM
Resolved
Sat, May 23, 2026, 04:36:55 PM
Duration
20h 10m
  1. resolved

    We have been monitoring the clusters for the past 16 hours and have not found any further disruptions.

  2. monitoring

    We have re-enabled our Frankfurt cluster and are monitoring for further disruptions.

  3. investigating

    We have observed network instability impacting cross-region traffic from Frankfurt beginning at 18:15 UTC. Traffic has been redirected from Frankfurt as of 19:55 UTC. Participants originally connecting via Frankfurt may see slightly increased latency until the region is restored.

Minor

Intermittent connectivity disruptions in US regions

Started
Fri, May 15, 2026, 04:57:21 PM
Updated
Fri, May 22, 2026, 02:53:27 PM
Resolved
Sat, May 16, 2026, 02:50:40 AM
Duration
9h 53m
  1. postmortem

    ## Summary LiveKit Cloud runs as a distributed realtime network, with data centers around the world interconnected via dedicated networking. Even with dedicated fiber, momentary disruptions between any two data centers can and do occur. To handle these blips, we have built a comprehensive set of resilience mechanisms that automatically reroute and relay traffic over alternate healthy paths. For example, if data centers A and B cannot reach one another cleanly but both can reach data center C, we use C as a relay so that traffic flows A to C to B. Under normal circumstances, network blips between our data centers are handled transparently there would be no visible impact. Between 2026-05-07 and 2026-05-19, a small number of these otherwise routine network blips did become customer-visible in our US regions. Affected sessions experienced a brief interruption to media \(under 5 minutes\) before recovering on a new path. ## Impact Five short windows of disruption were observed, each tied to a brief network blip between certain clusters: * 2026-05-07 17:16 UTC * 2026-05-12 05:00 UTC * 2026-05-12 17:00 UTC * 2026-05-13 19:00 UTC * 2026-05-15 07:30 UTC * 2026-05-19 15:01 UTC The majority of sessions traversed alternate paths normally and were not affected. A subset of sessions whose traffic happened to be relayed through a region in the specific failure state described below experienced a media interruption of up to ~5 minutes before re-routing onto a healthy path. ## Root Cause On 2026-05-06, a change was deployed that altered the relay process. The change introduced a subtle bug that required two conditions to occur simultaneously to manifest: 1. The relay region was under-utilized and had not cached a particular piece of state needed by the relay process. 2. A large volume of traffic was directed to that relay region in a very short window of time. When both conditions were present, the relay process would be stuck and would take minutes to fully catch up. During that time, neither endpoint of the relayed session could continue to receive media from the other. Because both conditions are narrow \(a cold-cache relay region absorbing a sudden burst of traffic\), the bug did not surface during pre-deploy testing, and it did not trigger on every network blip. It only manifested when a real network disruption happened to redirect a sufficiently large burst of traffic to a relay region that had not warmed its cache. Once that occurred on a given relay, sessions flowing through it stalled until traffic shifted off that path. We tracked down the root cause on 2026-05-20, and a fix has been fully rolled out across the fleet. ## Corrective Actions & Prevention * **Improved integration tests for relay locks under load.** We are adding tests that exercise the relay process with cold caches and sudden traffic bursts, so this class of contention is caught before deploy. We are also auditing related lock paths to ensure they behave correctly under similar load profiles. We sincerely apologize for the disruption to customers whose sessions were affected during these windows. Thank you for your patience, and we welcome any additional feedback from customers who were impacted.

  2. resolved

    We are resolving this incident as we have not observed any further connectivity disruptions. We will follow up with a postmortem once available.

  3. monitoring

    We have not observed further instances of connectivity disruptions since 13:19 UTC. We are continuing to investigate why our reroute mechanism did not activate in these isolated instances and are monitoring for further disruptions.

  4. investigating

    We are investigating intermittent connectivity disruptions affecting a subset of sessions, caused by internal network degradation between several of our US clusters (most heavily impacting US East). Affected sessions may see brief interruptions (less than 5 minutes) to media relay. We have observed these interruptions at the following times: - 5/12/26 05:00 UTC - 5/12/26 17:00 UTC - 5/13/26 19:00 UTC - 5/15/26 07:30 UTC We are actively investigating. Further updates to follow.

Major

Outbound SRTP calls are failing

Started
Wed, May 6, 2026, 07:46:39 PM
Updated
Sun, May 10, 2026, 03:28:06 PM
Resolved
Wed, May 6, 2026, 11:04:10 PM
Duration
3h 17m
  1. postmortem

    **Summary** On 2026-05-06 between 18:07 and 19:56 UTC, a code change to the API server which handles CreateSIPParticipant calls caused outbound calls to incorrectly inherit their encryption settings from outbound trunks. This resulted in calls that didn't specify an encryption setting in the request itself to be sent without SRTP, even if the LiveKit trunk was configured to require it. If users had configured their carrier’s trunk to expect SRTP \(by enabling Twilio’s Secure Trunking, for example\) those INVITEs were rejected. Approximately 2% of global outbound SIP calls were affected during the incident window. **Impact** - Window: 2026-05-06 18:07 - 19:56 UTC - Scope: Outbound TLS calls where the encryption setting was not set in the request and the carrier was configured to expect SRTP. - Customer symptom: Outbound calls failed at the carrier with `488` rejections such as Twilio's `32208 SIP trunk or domain is required to use secure media (SRTP)`, `Not Acceptable Here`, or `Media Encryption Required`. - Identifying impact: Customers can search their call records during the window for failed outbound calls returning `488` from their carrier with one of the above reason phrases. **Root cause** A recent code change introduced a regression in how SIP trunk encryption settings were applied to outgoing calls. SIP trunks configured to require SRTP had that requirement silently dropped on outbound INVITEs when the per-call request did not also explicitly set an encryption mode. As a result, outbound calls were sent without SRTP, and carriers that were configured to expect SRTP on TLS correctly rejected them. **Timeline \(UTC\)** - 18:07 - Production deployment containing the regression went live. - 18:39 - Internal monitoring paged on a periodic end to end functionality test. - 19:42 - First customer-reported outbound call failures. - 19:56 - Fix deployed to production. Failure rate returned to baseline. **Resolution** A fix was deployed that ensures trunk-level SRTP requirements are correctly applied to outbound calls. Regression tests covering this and the adjacent code paths have been added to prevent the same class of issue. **What we are doing to prevent recurrence** - Strengthening pre-deployment validation requirements for changes that affect our SIP infrastructure. - Improving the reliability of our automated test coverage so similar regressions are caught before reaching production.

  2. resolved

    We are resolving this incident as we have not observed any additional errors since the fix was rolled out. We will follow up with a postmortem as soon as possible.

  3. monitoring

    We believe this incident was not provider-specific and as such have removed Twilio from the incident title.

  4. monitoring

    A fix has been implemented and we are monitoring the results. The issue persisted from 18:10 - 19:58 UTC.

  5. investigating

    We have noted that outbound calls with both TLS and media encryption enabled via Twilio are failing. We are currently investigating and will update here as soon as we know more.

None

Cloud Agents experiencing deployment failures in US East

Started
Wed, May 6, 2026, 04:01:53 AM
Updated
Sat, May 9, 2026, 07:37:28 AM
Resolved
Wed, May 6, 2026, 09:41:46 AM
Duration
5h 39m
  1. postmortem

    **Incident** On May 6, around 04:00 UTC we started to receive internal alarms about cluster availability from a single cluster in our us-east hosted agents region. Upon investigation we found that etcd was unresponsive for this cluster. Without etcd and control plane availability, new actions in the cluster were not processed. Anything already running previous to the start of the incident continued to run. However, builds, autoscaling and other similar events sent to this cluster failed. Around 04:30 UTC we escalated the issue to our cloud provider. Resolution took longer than expected for several reasons, including additional complexities in getting etcd and kube control plane back into a good state. We have multiple clusters in the region but we did not want to strain other clusters with the complete workload from the affected cluster. In order to ensure stability in other clusters we decided to add additional capacity to our fleet. Around 05:00 UTC we began setting up additional clusters and started making plans to migrate new deployments. Around 08:00 UTC a scale up operation began to revive the affected cluster. Around 09:15 UTC control plane and etcd resources began to recover. During the recovery process some existing workload became unstable, but our reconciler resolved the issue soon after. Around 09:40 UTC everything had recovered. **Post-incident** We've found better methods of escalation and recovery to ensure a similar delay doesn't happen in the future for cases like this. We've also been working on tuning our workloads so that a similar incident doesn't reoccur. We'll be adding additional monitoring and alarms for several specific scenarios that we've uncovered in our investigation. In addition, we're continuing our work to stand up additional compute resources so that individual cluster failures won't block things like deployments and scaling actions from completing. We've been working on a few initiatives along these lines already, but will be increasing the priority to ensure these items are completed soon.

  2. resolved

    US-East Cloud Agents is now fully functional.

  3. monitoring

    US-East recovered as of 08:55 UTC and has been serving traffic normally since. New and existing agent deployments are working as expected. We are moving to monitoring while we confirm sustained stability. A full post-incident review will follow.

  4. investigating

    We are continuing to work on restoring the Kubernetes API server in US-east. Starting at 08:15 UTC, we are observing impact to existing agent deployments in addition to new deployments. Mitigation is in progress.

  5. investigating

    etcd service in our US East Kubernetes cluster is currently down, resulting in API server unresponsiveness and failures for new deployments and redeployments. We are actively working on mitigating this with our data center provider.

  6. investigating

    Our team is actively working to resolve deployment failures in US East. Existing deployed agents should not be affected in any way.

  7. investigating

    New agent deployments are temporarily unavailable in the US East region, and our team is actively working on this. Dispatches to previously deployed agents are not affected.

  8. investigating

    We are currently investigating intermittent deployment failures on Cloud Agents in US East.

  9. investigating

    We are currently investigating degraded performance on Agents Hosted on LiveKit Cloud in US East.

None

Brief cross-region connectivity disruption in US East

Started
Thu, May 7, 2026, 05:16:00 PM
Updated
Fri, May 8, 2026, 08:28:32 PM
Resolved
Thu, May 7, 2026, 09:15:00 PM
Duration
3h 59m
  1. resolved

    Between ~17:16 and 17:21 UTC on May 7, 2026, inter-region connectivity between our US East region and other regions was briefly degraded due to a network connectivity issue in our underlying cloud provider's data center. During this window, sessions that relied on cross-region track relay through US East may have experienced brief connection failures or reconnections for some participants. Our system recovered automatically once the cloud provider's data center network connectivity was restored.

None

Investigating SIP participant timeouts and signalling connection errors

Started
Fri, Apr 24, 2026, 06:12:17 PM
Updated
Wed, Apr 29, 2026, 02:41:31 PM
Resolved
Wed, Apr 29, 2026, 02:41:31 PM
Duration
4d 20h
  1. resolved

    We are closing this incident as we have not received further reports and the signalling error rate has dropped back to baseline. We made some fine-tuning adjustments to our stack to improve performance and will continue to seek out opportunities to keep latency low.

  2. identified

    A fix is being implemented. We still have not received any further reports and believe the impact to be minor, but we're continuing to monitor for further issues.

  3. investigating

    We received a single report of CreateSIPParticipant timeouts in Singapore. While investigating this, we discovered an increased rate of signalling connection errors on a very small minority of requests, mainly in India. We haven’t received any other reports, but are proactively creating this incident in case users encounter increased connection latency or failed inbound SIP calls. We are continuing to investigate and will update here once we know more.

None

Rejected INVITEs in US East

Started
Wed, Apr 22, 2026, 01:00:00 PM
Updated
Thu, Apr 23, 2026, 04:17:57 PM
Resolved
Wed, Apr 22, 2026, 01:00:00 PM
Duration
0m
  1. resolved

    Between 13:15 and 14:25 UTC on April 22, inbound SIP calls routed through our US East region where the toUser field was left empty (approximately 5% of total inbound calls in this region) were rejected by our system due to a failing internal trunk lookup. Upstream carriers surfaced these rejections to end users as 503 "Service Unavailable" responses. A recent change to our internal service responsible for SIP trunk authorization lookups caused trunk queries to return empty results if the toUser field was empty. When our SIP service received an empty trunk lookup, it rejected the inbound INVITE. The regression was deployed to one US region as part of a staged rollout. Our routine checks during the release identified the issue. Once the offending change was rolled back, inbound call rejections returned to baseline within minutes and full service was restored. Other regions and all outbound calls were unaffected. We are introducing a dedicated monitor for this specific failure mode so that any recurrence pages our on-call engineers immediately, rather than relying on broader error-rate signals.

Major

Increased Latency in RoomService APIs, brief period of higher error rate

Started
Wed, Apr 15, 2026, 11:23:49 PM
Updated
Fri, Apr 17, 2026, 04:43:48 PM
Resolved
Thu, Apr 16, 2026, 04:09:10 AM
Duration
4h 45m
  1. postmortem

    ## Summary LiveKit's core realtime and agent services are designed to tolerate database failures. WebRTC media, SIP calls, and hosted agent sessions continue to operate even when our database backend is slow or unavailable. A subset of Room APIs, specifically `CreateRoom`, `DeleteRoom`, and `UpdateRoomMetadata`, do depend on a database for consistency and disaster recovery. That database is highly available and globally distributed, with no single point of failure. When it is under significant contention, these Room APIs can return errors or time out, while realtime traffic continues to flow normally. On 2026-04-15, database contention caused a percentage of Room API calls to fail in our US-West region. Remediation work later produced a 26-minute global outage of the Room APIs. Realtime sessions, SIP calls, and agent processes were unaffected throughout. We sincerely apologize to customers whose applications were disrupted during this incident. ## Impact The incident had two distinct phases of customer impact. ### Phase 1: Elevated Room API timeouts in US-West \(2026-04-15 22:10 UTC to 2026-04-16 03:14 UTC\) A percentage of `DeleteRoom`, `UpdateRoomMetadata`, and `ListRooms` calls timed out, primarily in our US-West region. Other regions saw limited impact during this phase. Customers with high Room API volume in US-West observed elevated error rates on their integrations; the majority of customers were not affected. ### Phase 2: Global Room API outage \(2026-04-16 03:14 to 03:40 UTC, ~26 minutes\) While we were swapping in a rebuilt `rooms` table, the table was briefly missing from the database, and the majority of Room API calls globally returned HTTP 500 with `ERROR: relation "rooms" does not exist`. WebRTC sessions, SIP calls, and agent processes continued to function, and realtime connection counts remained stable. Applications that depend on Room APIs to start or manage sessions saw visible failures during this window. ## Root Cause The sweeper is a background process that removes rows from the `rooms` table as sessions end. Earlier on 2026-04-15, its throughput dropped significantly, and over roughly 8 hours stale rows accumulated to the point where the table was many times larger than its intended steady-state size. At approximately 21:00 UTC, a routine schema migration was applied to a different, unrelated table. The migration itself did not touch `rooms`, but it raised overall database disk utilization and background load. Combined with the oversized `rooms` table, this produced enough contention to slow down reads and writes against it. The effect first appeared in US-West, where the regional mix of Room API traffic was most sensitive to the contention. Once we identified the oversized table as the underlying cause, we needed to restore it to a healthy size. Because the table was already contended, deleting rows directly would have taken additional locks and worsened the contention. We instead chose to rebuild the table: create a new table with the same schema, copy over the active rows, then atomically swap the new table into place via a pair of renames. The copy phase completed quickly. The first rename \(moving the old `rooms` table aside\) completed in about 2.5 minutes. The second rename, moving the new table into the `rooms` name, stalled on our globally distributed database for significantly longer than we anticipated. During the stall, the `rooms` table did not exist from the perspective of any region, and all Room API calls globally returned errors. After roughly 10 minutes, we aborted the stalled rename, created a fresh `rooms` table from scratch, and inserted the active rows into it. Room API traffic recovered globally shortly after. ## Corrective Actions & Prevention The following improvements have been implemented or initiated to reduce the likelihood and impact of similar incidents: * **Enhance monitoring for sweeper throughput and active room count.** We are adding and hardening alerts on sweeper throughput and active room count, so that any future divergence pages on-call well before it threatens production. * **Improve sweeper resilience and throughput.** We are investigating the cause of the sweeper's throughput drop and adding capacity headroom so a transient slowdown cannot translate into multi-hour backlog growth. * **Remove database as a dependency for Room APIs**. This incident reaffirmed our long-held design principle that realtime services should not depend on databases. We believe this is the only way to build a system that approaches 100% uptime, and we will continue the work to ensure Room APIs do not depend on a database either. The Phase 2 outage was caused by our own remediation, and we recognize how disruptive it was for applications that depend on the Room APIs. We are committed to the work above to reduce both the likelihood and the blast radius of a similar failure in the future. Thank you for your patience, and we welcome any additional feedback from customers who were affected.

  2. resolved

    This issue is now fully resolved. We will be posting a detailed RCA.

  3. monitoring

    Our fix is fully implemented and we are not seeing any more failures or high latencies of the RoomService APIs. We are continuing to monitor the issue. We did observe a period of 15 minutes with high API failures while mitigation steps were being applied.

  4. investigating

    While applying a fix for the API latencies, we are temporarily seeing increased failure rates in RoomServices APIs, including CreateRoom, UpdateRoomMetadata, and DeleteRoom. We are actively working on mitigating this. Impact has been upgraded to major.

  5. identified

    We continue to see the long Room API latencies which are now also impacting other regions. The latency increases appear to originate from a specific table in our distributed database. The issue has been escalated with the database vendor and we are working on a workaround for decreasing the API latencies. Other services are not impacted.

  6. investigating

    We believe these elevated latencies began around 22:00 UTC. We have confirmed that only API requests in US-West should be impacted. The current list of impacted APIs appears to be CreateRoom, DeleteRoom, and UpdateRoomMetadata. We are working on mitigating the issue to return latencies back to normal.

  7. investigating

    We are continuing to investigate this issue.

  8. investigating

    We are investigating reports of increased latencies in RoomService APIs in the US West region, specifically on CreateRoom, DeleteRoom, and UpdateRoomMetadata APIs.

None

LiveKit Cloud Dashboard Missing Observability Sessions

Started
Wed, Apr 8, 2026, 04:18:32 PM
Updated
Wed, Apr 8, 2026, 04:25:49 PM
Resolved
Wed, Apr 8, 2026, 04:18:32 PM
Duration
0m
  1. resolved

    We have identified a bug where LiveKit Cloud projects created after April 1 at 20:41 UTC which enabled Agent Insights encountered a bug where their observability data (recordings, traces, and agent logs) was not saved correctly. We have resolved the issue on April 8 at 03:19 UTC and Agent Insights should be fully operational for all new and existing projects. No action is required from impacted users. We will follow up with a postmortem as soon as possible.

Critical

LiveKit Cloud Dashboard Down

Started
Fri, Apr 3, 2026, 05:13:43 PM
Updated
Fri, Apr 3, 2026, 11:24:37 PM
Resolved
Fri, Apr 3, 2026, 05:23:12 PM
Duration
9m
  1. postmortem

    **Root Cause** A gap in our DNS update process resulted in a misconfiguration of the DNS for [livekit.io](http://livekit.io), causing several A records to be overwritten. This resulted in an incorrect IP address being returned for [livekit.io](http://livekit.io). All real-time services - calls, RTC, and hosted agents - were unaffected as they operate on separate domains. We do have DNS monitoring in place, but it was not configured to page on-call. As a result, the issue was identified by LiveKit engineering after a short delay. **Timeline** * **2026-04-03 17:46 UTC**: DNS configuration change applied; resolution begins failing for [livekit.io](http://livekit.io). * **2026-04-03 17:52 UTC**: Issue identified by LiveKit. * **2026-04-03 18:15 UTC**: Root cause identified and fix applied. * **2026-04-03 18:18 UTC**: Fix fully propagated; services restored. Total duration: ~32 minutes. **Mitigations** * Enforce tighter automated guardrails around DNS updates. * Enable paging on existing DNS health checks so any future issues are caught even sooner.

  2. resolved

    A fix has been implemented and we are monitoring the results. All users should now be able to access the Cloud Dashboard.

  3. investigating

    We are currently investigating this issue and will update as soon as we know more.

Minor

Degraded Connectivity Issues – EU (Frankfurt)

Started
Wed, Apr 1, 2026, 10:37:57 AM
Updated
Wed, Apr 1, 2026, 04:34:34 PM
Resolved
Wed, Apr 1, 2026, 01:01:46 PM
Duration
2h 23m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We identified degraded connection errors affecting a small percentage of requests to Real Time Communication and TURN services in our EU (Frankfurt) region. Automatic retries prevented applications from experiencing failures. A fix has been implemented and we are actively monitoring to confirm stability.

Watch LiveKit
Email alerts on every status change — outages, degradations, new incidents, and resolutions.

Watching all 1 providers. Customize on the alerts page.

Details

Aliases
livekit agents, livekit cloud
Indicator
none
Path
/livekit

Related in Voice & Audio