Google Cloud Platform Identity & Access Management
Operational
Incidents19
History 19
Minor
Studio feature regression
Started
Fri, Jul 24, 2026, 09:11:37 PM
Updated
Fri, Jul 24, 2026, 09:46:03 PM
Resolved
Fri, Jul 24, 2026, 09:46:03 PM
Duration
34m
resolved
The incident has been resolved and Studio functionality is back to normal operation.
identified
The issue has been identified and the team and working on a resolution.
investigating
We are investigating WellSaid Studio showing incorrect functionality to users.
Critical
TTS Service Disruption
Started
Wed, Apr 15, 2026, 02:20:28 PM
Updated
Wed, Apr 15, 2026, 02:23:09 PM
Resolved
Wed, Apr 15, 2026, 02:20:28 PM
Duration
0m
resolved
We experienced a TTS service disruption at approximately 9 UTC.
Our infrastructure team is worked to identify the root cause and implement a solution at approximately 5 UTC. Studio and API users were unable to generate new audio clips during that time. However Studio projects and existing audio were accessible.
Major
Standard avatar partial outage
Started
Wed, Feb 4, 2026, 04:27:00 PM
Updated
Thu, Feb 5, 2026, 05:38:36 PM
Resolved
Wed, Feb 4, 2026, 04:27:00 PM
Duration
0m
resolved
Initial assessment: Starting at 1:55 AM PT there is an ongoing partial outage for Studio and API customer on "standard"/v10 avatar voices.
Resolved 8:27AM PT.
None
GPU Node Provisioning Failure
Started
Tue, Sep 16, 2025, 03:00:06 PM
Updated
Tue, Sep 30, 2025, 11:16:06 PM
Resolved
Tue, Sep 30, 2025, 11:16:06 PM
Duration
14d 8h
resolved
This incident has been resolved.
monitoring
Provisioning of new GPU nodes in our GKE environments began returning to normal overnight. We are now seeing only minimal delays in node creation, and service performance has largely stabilized. We are working on expanding our infrastructure and are actively monitoring node availability.
identified
We are currently experiencing issues provisioning new GPU nodes in our GKE environments. This is causing a reduced capacity in the number of requests we're able to serve simultaneously, leading to longer response times and clips failing to generate.
Major
Cloud provider core infrastructure unavailability
Started
Thu, Jun 12, 2025, 06:33:14 PM
Updated
Thu, Jun 12, 2025, 08:46:49 PM
Resolved
Thu, Jun 12, 2025, 08:46:49 PM
Duration
2h 13m
resolved
Services are now fully operational, and all systems are performing as expected.
monitoring
Services are now operational, though some requests may still be slower than normal. Clip creation is functioning, but clips generated during the outage remain unavailable. Our team is continuing to monitor the system and investigating options to recover the affected clip data.
monitoring
Services are beginning to recover, and clip creation is now fully operational. However, clips that were created during the outage remain unavailable. Our engineering team is actively investigating options to recover those affected clips.
identified
Our cloud provider has reported that most locations have fully recovered. However, some features—such as clip interactions—remain unavailable. We’re actively working to restore full functionality and will provide further updates as they become available.
identified
Google Cloud has reported a major service outage across IAM affecting the entire platform https://status.cloud.google.com/incidents/ow5i3PPK96RduMcb1SsW
identified
We are continuing to work on a fix for this issue.
identified
The issue has been identified as outside our infrastructure and reported by our cloud providers.
investigating
We are currently experiencing issues with our Cloud provider's core infrastructure, which is affecting a large portion of network traffic.
We are monitoring this outage at the PaloAltoNetworks level https://status.paloaltonetworks.com/incidents/l4zv6n67r8x6
Critical
Developer API Outage
Started
Tue, May 20, 2025, 08:30:00 AM
Updated
Wed, May 21, 2025, 04:39:39 PM
Resolved
Tue, May 20, 2025, 08:30:00 AM
Duration
0m
resolved
Service failure impacting all API customer traffic
Incident Start: 2025-05-20 08:17:00 UTC
Incident End: 2025-05-20 09:24:00 UTC
Incident Duration: 67 minutes
Impact: Complete service outage resulting in 502 errors for customer traffic
Minor
Studio clip conversion/script downloading outage
Started
Tue, Apr 22, 2025, 05:06:12 PM
Updated
Tue, Apr 22, 2025, 05:17:34 PM
Resolved
Tue, Apr 22, 2025, 05:17:34 PM
Duration
11m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
Minor
Developer TTS API authorization failures
Started
Sun, Nov 10, 2024, 01:41:04 AM
Updated
Sun, Nov 10, 2024, 03:23:21 AM
Resolved
Sun, Nov 10, 2024, 02:02:13 AM
Duration
21m
postmortem
# Summary
During a migration of our internal TTS service infrastructure, our external facing TTS Developer API experienced a partial outage due to authentication issues with the new TTS service. The root cause was identified as the use of invalid authentication keys from the Developer API which had not been updated to reflect those within the new internal TTS service. This resulted in a service degradation leading to 19 failed requests.
# Timeline
* 2024-11-10 01:39 UTC : Migration of internal TTS service to new infrastructure begins
* 2024-11-10 01:41 UTC : Developer TTS API Gateway nodes begin to pick up new routing change
* 2024-11-10 01:41 UTC : The nodes routing to the new TTS service begin to fail requests
* 2024-11-10 01:42 UTC : Monitoring system shows failing requests between the two services
* 2024-11-10 01:42 UTC : Routing logic changed back to existing infrastructure
* 2024-11-10 01:43 UTC : Root cause of mismatching keys is identified
* 2024-11-10 01:43 UTC : Monitoring system alert for failed requests to Developer TTS fires
* 2024-11-10 01:44 UTC : Impact analysis begins
* 2024-11-10 01:44 UTC : Developer TTS API Gateway successfully authenticates to existing internal TTS service
* 2024-11-10 01:44 UTC : Existing authentication keys used by Developer TTS API Gateway are manually added to new TTS service
* 2024-11-10 01:51 UTC : Verification of existing keys against new TTS infrastructure is performed
* 2024-11-10 02:00 UTC : Impact analysis finds only Developer TTS API is affected
* 2024-11-10 02:01 UTC : Routing logic changed back to new infrastructure
* 2024-11-10 02:02 UTC : Developer TTS API Gateway begins to successfully authenticates to internal TTS service
* 2024-11-10 02:05 UTC : Infrastructure team continues to monitor both services
# Impact
* Some users of the Developer TTS API hitting nodes connecting to the new TTS service would have experienced server errors while attempting to generate new clips, in total 19 requests over a 3 minute time frame failed
# Action Items
* A planned extension of this infrastructure migration is to implement automatic rotation and refreshing of API keys between the two services to remove the need for manual syncing
resolved
All Developer TTS API gateway nodes are properly communicating with internal TTS service.
monitoring
Keys between the two services have been synced and we are seeing successful requests coming from the Developer API gateway. We are continuing to monitor to ensure all nodes reflect this change.
identified
Some TTS Developer API nodes have started performing requests to internal TTS service with incorrect api keys.
Major
Studio login issues
Started
Thu, Oct 10, 2024, 06:57:06 PM
Updated
Thu, Oct 10, 2024, 10:42:46 PM
Resolved
Thu, Oct 10, 2024, 10:42:46 PM
Duration
3h 45m
resolved
This incident has been resolved.
monitoring
Stripe team has implemented a fix on their side and the errors have stopped. We are still monitoring our systems.
identified
Stripe has reported API failures on their end for the components that are affecting our system. Actively monitoring updates on their side while looking for a workaround.
investigating
Requests to stripe API are failing preventing users from logging into the studio application
Critical
Developer API Outage
Started
Thu, Oct 10, 2024, 07:30:00 PM
Updated
Thu, Oct 10, 2024, 09:11:41 PM
Resolved
Thu, Oct 10, 2024, 07:30:00 PM
Duration
0m
resolved
Service failure impacting all API customer traffic.
Incident Start: 2024-10-10 19:40:00 UTC
Incident End: 2024-10-10 20:31:00 UTC
Incident Duration: 51 minutes
Impact: Complete service outage resulting in 502 errors for customer traffic
Minor
Developer API servers returning 502
Started
Tue, Oct 1, 2024, 04:00:44 PM
Updated
Mon, Oct 7, 2024, 11:27:36 PM
Resolved
Mon, Oct 7, 2024, 11:27:36 PM
Duration
6d 7h
resolved
This incident has been resolved.
monitoring
Error rate has returned to normal
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating this issue.
Major
Studio Outage
Started
Wed, Sep 18, 2024, 11:00:04 PM
Updated
Thu, Sep 19, 2024, 01:20:17 AM
Resolved
Wed, Sep 18, 2024, 11:20:42 PM
Duration
20m
postmortem
# Summary
WellSaid Labs Studio experienced an outage due to a traffic routing policy change made at the load balancer level. The change, intended to reduce the latency of clip generation and downloading, configured the routing rules for the Studio API and Studio web to services that failed to be updated leading to the load balancer being unable to route incoming requests to the backend services.
# Timeline
* 2024-09-18 22:57 UTC : Infrastructure change begins to roll out
* 2024-09-18 23:00 UTC : Routing rules for studio production services are updated
* 2024-09-18 23:00 UTC : Studio page and API requests begin failing
* 2024-09-18 23:05 UTC : Outage detected by automated systems and infrastructure team is alerted
* 2024-09-18 23:06 UTC : Cause of outage identified
* 2024-09-18 23:10 UTC : Fix identified
* 2024-09-18 23:15 UTC : Changes to fix routing rules begin to roll out
* 2024-09-18 23:18 UTC : Routing changes applied to production
* 2024-09-18 23:18 UTC : Services begin to spin back up to handle traffic
* 2024-09-18 23:19 UTC : Infrastructure team is able to access studio web and call the studio API
* 2024-09-18 23:20 UTC : System reports healthy status
# Impact
All users attempting to load pages within the WellSaid Labs studio would have experienced failing requests. Those attempting to generate clips would have been unaffected.
# Resolution
The infrastructure team removed and recrated the failing backend services using the automated deployment pipeline allowing the routing layer to properly reach them.
# Follow-Up
## Next steps
The infrastructure team is working on changes to ensure alignment between the different layers involved in routing requests to the studio services and tightening the dependency links between them such that the new backends must exist before traffic is attempted to be routed to them.
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
None
Expired SSL certificate impacting Studio users
Started
Thu, Mar 7, 2024, 06:00:00 PM
Updated
Thu, Mar 7, 2024, 07:21:38 PM
Resolved
Thu, Mar 7, 2024, 06:00:00 PM
Duration
0m
resolved
The Studio application was temporarily serving traffic under an expired SSL certificate.
Incident Start: 2024-03-07 09:56 PT
Incident End: 2024-03-07 10:22 PT
Incident Duration: 26 minutes
Impact: Studio users during the time of incident may have experienced failed requests due to an invalid (expired) SSL certificate
Major
Increased latency & error rates for API customers
Started
Sat, Mar 2, 2024, 10:00:00 PM
Updated
Tue, Mar 5, 2024, 05:35:54 PM
Resolved
Sat, Mar 2, 2024, 10:00:00 PM
Duration
0m
postmortem
## Incident Summary
Elevated latency and error rates impacting all API customers. The errors were a result of a single gateway pod serving traffic in a faulty state after it failed to initialize a system component properly.
## Impact
Roughly 20% of API customer traffic experienced increased latency and/or error rates for a period of 24 hours.
## Root Cause
The source of error rates and latency was pinned down to a single pod responsible for handling API customer traffic. This pod in particular failed to initialize a middleware component properly but continued to unsuccessfully serve traffic. Upon termination of the problematic pod, service was restored.
## Preventative Measures
* Improve gateway initialization logic and handling of failure scenarios
* Improve logging severity to ensure relevant errors trigger alerts accordingly
* Adjust alerting policies around error rates and latency for gateway components
resolved
Elevated latency and error rates impacting all API customers.
Incident Start: 2024-03-02 13:48 PT
Incident End: 2024-03-03 13:48 PT
Incident Duration: 24hrs
Impact: Roughly 20% of API Customer traffic experienced increased latency and/or error rates
Minor
Clip conversion and script downloading failing
Started
Thu, Nov 9, 2023, 02:30:00 PM
Updated
Thu, Nov 9, 2023, 05:06:39 PM
Resolved
Thu, Nov 9, 2023, 02:30:00 PM
Duration
0m
resolved
Functionality has been restored
Minor
Intermittent Studio Timeouts
Started
Sat, Aug 26, 2023, 11:35:00 AM
Updated
Tue, Sep 5, 2023, 07:05:57 PM
Resolved
Sat, Aug 26, 2023, 11:35:00 AM
Duration
0m
resolved
Summary: From 2023-08-26 04:35 PT to 2023-08-27 09:16 PT our web servers responsible for the Studio editor experienced issues resulting in intermittent delayed loading, render time, and failed logins. Due to the reduced traffic of the weekend and the intermittentness of the issues, our monitoring failed to alert our infrastructure team of the issue.
Impact: 3.34% of requests failed resulting in a small number of users being unable to access various parts of the Studio Editor intermittently.
Recovery: Once notified by our Support team the Infrastructure team was able to resolve the issue by resetting the connections to database.
Preventative Measures taken: Our Infrastructure team is working on improving alerting policies to better handle signals while at lower traffic.
Minor
Intermittent Studio Timeouts
Started
Tue, Sep 5, 2023, 12:30:00 PM
Updated
Tue, Sep 5, 2023, 07:03:34 PM
Resolved
Tue, Sep 5, 2023, 12:30:00 PM
Duration
0m
resolved
Summary: From 2023-09-05 07:30 PT to 2023-09-05 09:45 PT our web servers responsible for the Studio editor experienced issues resulting in intermittent timeouts.
Impact: 1-2% of requests failed resulting in a small number of users being unable to access various parts of the Studio Editor intermittently.
Recovery: Our Infrastructure team resolved the issue by resetting the connections to the database.
Preventative Measures taken: Our Infrastructure team is testing a fix to address this issue and prevent it from reoccurring.
Minor
Studio Editor Outage
Started
Wed, Jun 7, 2023, 07:00:00 PM
Updated
Tue, Aug 29, 2023, 10:49:16 PM
Resolved
Wed, Jun 7, 2023, 07:00:00 PM
Duration
0m
resolved
Hygraph, our 3rd party Avatar data source, went down for ~7 minutes between 11:54-12:01 PT. This impacted studio users providing a degraded editor experience for customers interacting with the avatar loader
Minor
Partial Studio Editor Outage
Started
Thu, Jun 8, 2023, 05:00:00 PM
Updated
Tue, Aug 29, 2023, 10:49:16 PM
Resolved
Thu, Jun 8, 2023, 05:00:00 PM
Duration
0m
resolved
Due to an issue with autoscaling logic, some users may have experienced a degraded Studio Editor experience while we addressed the issue