Intermittent DNS Resolution Errors for Upstash Vector in US East (us-east-1)
Started
Wed, Sep 9, 2026, 12:00:00 PM
Updated
Wed, Sep 9, 2026, 01:01:03 PM
Resolved
Wed, Sep 9, 2026, 12:00:00 PM
Duration
0m
resolved
Upstash Vector experienced a brief DNS resolution issue in the US East (us-east-1) region lasting approximately 20 minutes. During this period, new DNS resolution attempts may have failed.
Existing connections and clients with cached DNS records were not affected.
The issue has been resolved, and DNS resolution is operating normally.
None
Upstash Console login issue
Started
Thu, Sep 3, 2026, 08:00:00 AM
Updated
Thu, Sep 3, 2026, 08:29:36 AM
Resolved
Thu, Sep 3, 2026, 08:00:00 AM
Duration
0m
resolved
We experienced a brief issue affecting access to the Upstash Console. Our team quickly identified the cause and resolved the issue within minutes.
The console is now operating normally.
Major
Message Persistence Issue — QStash us-east-1
Started
Fri, Aug 28, 2026, 10:20:00 PM
Updated
Fri, Aug 28, 2026, 10:57:24 PM
Resolved
Fri, Aug 28, 2026, 10:20:00 PM
Duration
0m
resolved
Between 22:20 and 22:33 UTC, QStash clients in us-east-1 experienced disruptions related to message persistence. The team applied a fix, and service has been restored.
None
Fly.io infrastructure disruption affecting some Upstash Redis databases on Fly.io DFW Region
Started
Wed, Jul 22, 2026, 08:50:12 AM
Updated
Wed, Jul 22, 2026, 10:05:17 AM
Resolved
Wed, Jul 22, 2026, 10:05:17 AM
Duration
1h 15m
resolved
This incident has been resolved.
monitoring
The affected Fly.io-hosted machines recovered at 08:48 UTC, and service availability has been restored. We are monitoring the systems to confirm continued stability.
identified
We are investigating connectivity issues affecting a subset of Redis databases hosted on Fly.io. The issue is related to an ongoing Fly.io infrastructure disruption in the DFW region (https://status.flyio.net/incidents/n4z2my6qw4sb). We are working to restore availability and will provide updates soon.
None
QStash us-east-1 - URL Publish Errors
Started
Thu, Jul 16, 2026, 06:30:00 AM
Updated
Thu, Jul 16, 2026, 08:35:07 AM
Resolved
Thu, Jul 16, 2026, 06:30:00 AM
Duration
0m
resolved
Partial outage due to http parsing errors. Problem was resolved couple minutes later. If problem continues, refreshing the DNS cache is recommended.
Major
Upstash Redis Partial Service Disruption
Started
Thu, Jun 25, 2026, 03:23:44 PM
Updated
Thu, Jun 25, 2026, 04:36:00 PM
Resolved
Thu, Jun 25, 2026, 04:36:00 PM
Duration
1h 12m
resolved
This incident has been resolved, we will publish RCA soon.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently experiencing a partial service disruption affecting Upstash Redis in the following regions:
* eu-central-1
* us-east-1
* us-west-1
Our team is investigating the issue and working to restore full service as soon as possible. We will share updates as more information becomes available.
Minor
QStash EU Region — Degraded Performance
Started
Tue, Jun 23, 2026, 04:22:25 PM
Updated
Wed, Jun 24, 2026, 12:30:46 PM
Resolved
Tue, Jun 23, 2026, 05:14:46 PM
Duration
52m
postmortem
### **Summary**
On June 23, 2026, QStash users in the EU-CENTRAL-1 region experienced elevated latency and degraded performance. The issue was caused by a regression introduced in a recent enhancement to Flow Control scheduling logic. The change increased resource consumption under load, leading to reduced performance on affected shards. The incident was resolved by rolling back to the previous stable version.
### **Impact**
* Service: QStash \(EU-CENTRAL-1\)
* Start: 2026-06-23 16:22 UTC
* Resolved: 2026-06-23 17:14 UTC
* Duration: ~52 minutes
* Impact: Increased latency and degraded performance for workloads routed to the affected instance in the EU. Multiple customers in the region experienced slower request processing during the incident window.
### **Root Cause**
We recently deployed an enhancement to our Flow Control feature, designed to improve fairness between independent flow controls when unused global parallelism capacity was available.
Under production load, the new behavior introduced a regression that significantly increased resource utilization. The elevated resource consumption caused performance degradation on the affected QStash instamce, impacting all workloads sharing that instance.
### **Resolution**
After identifying the regression as the source of the slowdown, we rolled back the deployment to the last known stable version.
* Rollback completed: 2026-06-23 19:18:46.84 UTC
Following the rollback, system performance returned to normal levels and service stability was restored.
### **Preventive Actions**
To reduce the likelihood of similar incidents in the future, we are taking the following actions:
* Expand performance and load testing coverage for Flow Control changes.
* Add resource utilization regression checks to the deployment pipeline.
* Improve monitoring and alerting for abnormal CPU and memory consumption patterns.
* Introduce additional canary validation before wider production rollout of scheduler-related changes.
We apologize for the disruption and appreciate our customers’ patience while we resolved the issue.
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating this issue.
None
Box Access and Management Operations Unavailable
Started
Tue, Jun 23, 2026, 06:05:57 AM
Updated
Tue, Jun 23, 2026, 06:05:57 AM
Resolved
Tue, Jun 23, 2026, 06:05:57 AM
Duration
0m
resolved
Between 01:18 and 04:03 UTC, customers were unable to access running boxes or perform create, read, update, or delete operations. Running containers were not affected and continued operating throughout. The issue has been fully resolved and all box operations are functioning normally.
We are taking all necessary measures to prevent this from happening again. We apologize for the disruption.
Major
QStash Service Disruption in the EU Region
Started
Fri, Jun 19, 2026, 09:11:00 AM
Updated
Fri, Jun 19, 2026, 01:59:02 PM
Resolved
Fri, Jun 19, 2026, 09:11:00 AM
Duration
0m
postmortem
## What happened
A configuration change applied to QStash's EU networking layer introduced an outbound connectivity fault. Between **09:11 and 09:45 UTC**, a subset of QStash instances had degraded egress connectivity, and some dispatched messages may have failed on their initial attempt. A related side effect caused some servers to egress from IP addresses outside QStash's advertised outbound range until **11:24 UTC**; during that window, endpoints enforcing QStash IP allowlists may have rejected requests from those servers.
QStash's automatic retries delivered many affected messages on subsequent attempts once connectivity was restored, and QStash is now operating normally.
## Customer impact
Impact was limited to the EU region and to two effects within the window above:
* **Delivery delays / failures** _\(09:11 – 09:45 UTC\)_ — messages dispatched through affected instances may have failed on their first attempt. Thanks to automatic retries with backoff, most were re-delivered once connectivity recovered; messages that exhausted their retry policy during the window followed their configured failure path \(e.g., DLQ / failure callback\).
* **Allowlist rejections** _\(09:11 – 11:24 UTC\)_ — customers who restrict inbound traffic to QStash's outbound IP ranges may have seen requests from the affected servers rejected. Customers who do not enforce source-IP allowlisting were unaffected by this.
## Root cause
A routine networking configuration change contained an error that:
1. disrupted outbound connectivity on the affected nodes, and
2. caused affected nodes to acquire outbound IPs outside the advertised range.
The underlying gap was systemic: the change procedure had no automated validation gate to catch the faulty state before it reached production.
## What we've done
* **Fail-fast pre-checks** — automated validation now halts the change procedure _before it runs_ if a configuration would degrade connectivity or violate expected state.
* **Egress IP-range enforcement** — outbound IP assignments are now checked against the advertised QStash range, so a node can no longer come online with an unadvertised IP.
* **Tighter post-change verification** — completion now confirms outbound reachability and egress-IP conformance across affected nodes.
## Customer action
None required. All QStash traffic again originates from the advertised outbound IP ranges.
We apologize for the disruption.
resolved
We identified an outbound networking issue affecting a subset of QStash server instances in the EU region between 2026-06-19 09:11 UTC and 2026-06-19 09:45 UTC.
During this period, some requests dispatched by QStash may have failed.
The issue has been resolved, and QStash is currently operating normally. We will publish a postmortem with additional details as soon as it is available.
Update: We identified that one QStash server had been assigned an outbound IP address outside of the configured QStash IP range. This remained the case until 2026-06-19 11:24 UTC.
Endpoints that restrict access based on QStash outbound IP allowlists may have rejected requests from this server during this period.
None
Fly.io - Upstash Vector service disruption on IAD region
Started
Thu, Jun 4, 2026, 10:10:36 AM
Updated
Thu, Jun 4, 2026, 02:09:13 PM
Resolved
Thu, Jun 4, 2026, 02:09:13 PM
Duration
3h 58m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently experiencing a service disruption affecting Upstash Vector in the IAD region.
Initial investigation indicates that the issue is related to a Fly.io networking/routing problem impacting connectivity in the region. The Fly.io team is actively investigating the underlying cause.
We will continue to monitor the situation and provide updates as more information becomes available.
Minor
Intermittent slowness in Vector US-EAST-1 region
Started
Fri, May 29, 2026, 08:18:02 PM
Updated
Sat, May 30, 2026, 07:51:00 AM
Resolved
Sat, May 30, 2026, 07:51:00 AM
Duration
11h 32m
resolved
The issue affecting some Upstash Vector indexes in the US-EAST-1 region has been resolved.
Our team investigated the incident and identified the conditions that were contributing to elevated memory pressure on the affected servers. We mitigated those conditions by reducing memory utilization on the impacted nodes, rebalancing affected workloads where needed, and increasing available headroom capacity across the region.
All affected indexes should now be operating normally, and based on the mitigations applied, we do not expect this issue to recur. We apologize for any inconvenience this may have caused.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
Some Upstash Vector indexes in the US-EAST-1 region may be experiencing intermittent slowness. Initial investigation indicates this may be related to memory pressure on the affected servers.
Our team is actively investigating the root cause and working to restore normal performance. We will provide updates as soon as more information is available.
None
New Database Creation Failing Due to Upstream Provider Issue
Started
Fri, May 22, 2026, 11:27:48 PM
Updated
Sat, May 23, 2026, 12:00:03 AM
Resolved
Sat, May 23, 2026, 12:00:03 AM
Duration
32m
resolved
This incident has been resolved.
monitoring
Upstream issue resolved — new database creation is working again. We're monitoring to confirm full recovery.
investigating
New database provisioning is currently failing due to an ongoing upstream provider issue affecting DNS. Existing databases and connections are unaffected — only the creation flow is impacted. We're monitoring the situation and will update once the upstream issue is resolved.
Major
Fly.io Upstash Redis service distruption (FRA region)
Started
Mon, May 11, 2026, 03:05:43 PM
Updated
Fri, May 15, 2026, 12:49:12 PM
Resolved
Mon, May 11, 2026, 06:22:29 PM
Duration
3h 16m
postmortem
On May 12th and 13th at various times, a subset of Upstash Redis instances on [Fly.io](http://Fly.io) experienced intermittent hangs and elevated error rates. The Redis process would stall inside a logging syscall — alive but not making progress — which made the issue hard to spot from our usual telemetry. After investigating with Fly's team, we identified the root cause as a bad interaction between a recent guest kernel update on Fly's newer machines and an upstream Cloud Hypervisor bug \([cloud-hypervisor#7672](https://github.com/cloud-hypervisor/cloud-hypervisor/issues/7672)\) affecting log writes from inside the VM. We mitigated by disabling the affected logging paths, and Fly has since rolled out a hypervisor-side patch, fully resolving the issue. No data was lost. Sorry for the disruption.
resolved
The incident has been resolved. We are working with Fly team on RCA.
investigating
We are working with Fly team to investigate the root cause.
investigating
We are continuing to investigate the issue.
investigating
Some databases may experience increased latency or timeouts in Fly.io’s FRA region.
Critical
Fly.io Upstash Redis Service Distruption
Started
Tue, May 12, 2026, 09:49:32 AM
Updated
Fri, May 15, 2026, 12:48:59 PM
Resolved
Tue, May 12, 2026, 01:48:53 PM
Duration
3h 59m
postmortem
On May 12th and 13th at various times, a subset of Upstash Redis instances on [Fly.io](http://Fly.io) experienced intermittent hangs and elevated error rates. The Redis process would stall inside a logging syscall — alive but not making progress — which made the issue hard to spot from our usual telemetry. After investigating with Fly's team, we identified the root cause as a bad interaction between a recent guest kernel update on Fly's newer machines and an upstream Cloud Hypervisor bug \([cloud-hypervisor#7672](https://github.com/cloud-hypervisor/cloud-hypervisor/issues/7672)\) affecting log writes from inside the VM. We mitigated by disabling the affected logging paths, and Fly has since rolled out a hypervisor-side patch, fully resolving the issue. No data was lost. Sorry for the disruption.
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and the fix is being implemented.
investigating
Some regions are experiencing connectivity issues due to an ongoing network problem. We are currently investigating
Major
Upstash Redis – intermittent connection issues in some regions
Started
Thu, May 14, 2026, 03:47:23 PM
Updated
Thu, May 14, 2026, 05:31:50 PM
Resolved
Thu, May 14, 2026, 05:31:50 PM
Duration
1h 44m
resolved
Earlier today, unexpected load on our proxies caused intermittent connection issues for Upstash Redis in the following regions:
us-east-1, us-west-1, ap-southeast-2, and ap-south-1.
During this period, some clients may have seen connection timeouts or elevated error rates when reaching their databases.
Our team identified the issue quickly and applied workarounds to relieve pressure on the affected proxies. Connection health has since been restored and we've been monitoring the regions to confirm everything is stable. All systems are now operating normally.
We appreciate your patience and apologize for any disruption this may have caused.
monitoring
A fix has been implemented and we are monitoring.
identified
The issue has been identified and a fix is being implemented.
monitoring
We identified intermittent connection issues affecting some regions. Fix is being deployed. We'll update the list of affected regions shortly.
Critical
QStash US Region Service Disruption
Started
Fri, May 8, 2026, 09:46:14 AM
Updated
Tue, May 12, 2026, 08:25:40 AM
Resolved
Fri, May 8, 2026, 10:08:41 AM
Duration
22m
postmortem
**Root Cause Analysis**
On **April 24**, we deployed a more optimized scheduler implementation in the **US East \(N. Virginia\)** region.
On **May 8**, a user who had active schedules deleted their account.
Under normal behavior, scheduled tasks associated with a deleted account should wake up, detect that the account no longer exists, and exit after performing cleanup. Due to a bug introduced in the new scheduler implementation, this code path did not return early as intended. Execution continued and resulted in a nil pointer dereference.
A second issue then amplified the impact. When a panic occurs in the scheduler, it is designed to be recovered, logged, and isolated so that the process remains healthy. Because of another bug in the panic recovery path, the panic was not properly caught, which caused the worker process handling the scheduled job to terminate.
After that process exited, another worker picked up responsibility for delivering the same scheduled task. Since the same faulty execution path was still present, that worker also failed. This created a cascading failure pattern across workers attempting to process the affected schedules.
**Resolution**
We deployed two fixes:
* Added the missing early return in the deleted-account cleanup path, preventing the nil pointer dereference.
* Corrected the panic recovery logic so that future panics are safely recovered, logged, and reported without causing worker processes to terminate.
With these changes in place, the affected execution path is now safe. Even if a future bug triggers a panic in this area, it will be isolated and reported rather than causing process-level failure.
resolved
This incident has been resolved, we will publish RCA soon.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating the issue.
Minor
QStash US Region: Schedule Degradation
Started
Tue, May 5, 2026, 02:24:43 PM
Updated
Wed, May 6, 2026, 08:56:42 AM
Resolved
Wed, May 6, 2026, 08:05:27 AM
Duration
17h 40m
postmortem
# **Incident Postmortem: Scheduled Jobs Inconsistency in US Region**
On May 1, 2026, we experienced an incident affecting a subset of schedules in the US region following a recent infrastructure update.
The issue has been resolved, and all affected schedules have been restored.
## **Summary**
As part of an ongoing scalability improvement, we recently updated scheduling infrastructure in the US region to a new architecture. During this transition, a legacy execution path remained in the codebase as a fallback mechanism.
On May 1, a bug caused the system to revert to the legacy path. This resulted in inconsistent state between the old and new scheduling systems for some users.
## **Impact**
The incident affected a limited number of users in the US region.
**Most users were not affected**, and the vast majority of schedules continued operating normally throughout the incident.
Users who did not update schedules during the transition window continued operating normally throughout the incident.
**A subset of users who created, edited, paused, or deleted schedules between April 24 and May 1 may have experienced one or more of the following:**
* Schedule updates not being reflected
* Paused schedules becoming active again
* Deleted schedules reappearing
* Newly created schedules not executing as expected
Schedules created after the transition may have stopped executing briefly before recovery.
## **Root Cause**
During the transition, the new scheduling infrastructure became the source of truth for schedule state.
Due to a bug, the system unexpectedly reverted traffic to the legacy scheduling path, which began accepting updates independently from the new system.
This caused the two systems to diverge and resulted in inconsistent schedule state for affected users.
## **Resolution**
After identifying the issue, we:
1. Restored the new scheduling system as the active source of truth
2. Reconciled data between the legacy and new systems
3. Updated missing schedule changes back into the new infrastructure
4. Performed conflict resolution to preserve user data and schedule continuity
In some cases, schedules that had previously been paused or deleted were restored to avoid permanent data loss.
## **Preventive Measures**
We are implementing several changes to prevent similar incidents:
* Removing obsolete fallback execution paths after transitions complete
* Adding automated safeguards and alerts for unexpected system fallback behavior
* Improving consistency validation between systems
* Expanding rollback and reconciliation testing
We apologize for the disruption and appreciate everyone’s patience while we resolved the issue.
resolved
This incident has been resolved.
monitoring
Main schedule functionality is back to normal. We are currently checking if previously created schedules are delivered as expected before marking the incident as resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
We are continuing to work on a fix for this issue.
identified
We are currently experiencing issues in the US region.
- Duplicate Deliveries: During this period, some scheduled jobs may be executed twice.
- Schedule Disruption: Schedules created between April 24, 2026 and May 2, 2026 are currently not running.
Our team is actively working on a fix. Once the migration is complete, affected schedules will resume normal operation.
We will provide updates as progress continues.
Minor
Fly.io Upstash Redis – iad Region Elevated Latency and Temporary Read-Only State
Started
Thu, Apr 2, 2026, 06:45:17 PM
Updated
Thu, Apr 2, 2026, 07:53:23 PM
Resolved
Thu, Apr 2, 2026, 07:53:23 PM
Duration
1h 8m
resolved
Replication complete, incident resolved.
identified
Servers in the iad region experienced unexpected disk load, resulting in elevated latencies and a temporary read-only state. We are migrating replicas to new instances to mitigate the issue and expect to have it fully resolved shortly.
None
Upstash Redis: GCP Global Connectivity Problems
Started
Fri, Mar 27, 2026, 03:30:00 PM
Updated
Fri, Mar 27, 2026, 05:05:02 PM
Resolved
Fri, Mar 27, 2026, 03:30:00 PM
Duration
0m
resolved
Due to a race condition in a process that attaches static IPs to nodes, some of the IPs in the dns were detached from the nodes, causing timeouts.
Critical
Region ap-northeast-1 outage on Upstash Global
Started
Fri, Mar 6, 2026, 01:50:17 PM
Updated
Mon, Mar 9, 2026, 10:09:03 AM
Resolved
Fri, Mar 6, 2026, 02:49:57 PM
Duration
59m
postmortem
On March 6, between approximately 13:44–14:07 UTC, some databases experienced elevated latency and connection errors in the Tokyo \(ap-northeast-1\) region.
The issue was caused by a sudden spike in traffic that significantly increased network utilization and connection load on a subset of nodes.
Our team mitigated the incident by scaling up capacity in the region and redistributing load across additional nodes. Service recovered once the additional capacity was brought online.
Resolution
We have increased the number of machines in the Tokyo region to provide additional headroom and reduce the likelihood of similar incidents during traffic spikes.
Next Steps
We are continuing to review capacity safeguards and connection-handling limits to improve resilience against sudden traffic surges.
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating this issue.
Minor
Upstash Redis: Intermittent latency on us-east-1
Started
Fri, Jan 23, 2026, 03:30:00 PM
Updated
Fri, Jan 23, 2026, 05:05:42 PM
Resolved
Fri, Jan 23, 2026, 03:30:00 PM
Duration
0m
resolved
We identified the cause of elevated latency impacting some databases in us-east-1 region between 15:30–15:35 UTC as a sudden surge of connection attempts that hit OS-level connection limits on our proxy layer. This resulted in slower new connection establishment and increased latency for some requests. Databases were not impacted. We are implementing additional proxy-level metrics and safeguards to detect and manage similar edge cases earlier.
Minor
QStash - Message delays for Flow Control configurations
Started
Mon, Dec 1, 2025, 07:00:00 AM
Updated
Fri, Dec 12, 2025, 02:34:52 PM
Resolved
Mon, Dec 1, 2025, 07:00:00 AM
Duration
0m
resolved
We identified and fixed a bug that could cause messages with Flow Control enabled to be delayed longer than their configured delay, resulting in unexpectedly long pending times.
The fix is in place and the issue should not recur. If you’re still seeing unusually long-delayed messages, please contact support@upstash.com and we can help with remediation.
Major
Connectivity issue impacting Regional Databases in US-East-1
Started
Wed, Dec 10, 2025, 11:50:47 AM
Updated
Wed, Dec 10, 2025, 12:08:43 PM
Resolved
Wed, Dec 10, 2025, 12:08:43 PM
Duration
17m
resolved
Issue has been identified and replicas were successfully reconnected.
investigating
We identified an issue in Regional Databases in US-East-1 where database replicase have connectivity issues with each other.
Other regions are not impacted. Global databases are not impacted.
We are working on the issue.
Critical
Upstash Console and Context7 Console is Currently Experiencing issues
Started
Fri, Dec 5, 2025, 09:03:56 AM
Updated
Fri, Dec 5, 2025, 09:18:53 AM
Resolved
Fri, Dec 5, 2025, 09:18:53 AM
Duration
14m
resolved
This incident has been resolved.
investigating
Upstream provider confirmed an incident. We are investigating the impact and potential resolutions.
Major
Upstash Console issues
Started
Mon, Oct 20, 2025, 02:52:07 PM
Updated
Mon, Oct 20, 2025, 06:50:38 PM
Resolved
Mon, Oct 20, 2025, 06:50:38 PM
Duration
3h 58m
resolved
A fix has been deployed as a workaround so that our systems are not affected from the ongoing incident of the cloud provider
monitoring
Developer API (api.upstash.com) is having availability issues alongside Upstash Console. While we are monitoring the underlying cloud provider's status updates, we are also working on a remediation.
monitoring
Console is back to normal again. We are currently monitoring.
investigating
We are currently experiencing errors in the Upstash Console due to issues with one of our upstream providers. This may affect access to the dashboard and related operations.
Our team is actively monitoring the situation and working to mitigate the impact. We will provide updates as soon as more information becomes available.
Critical
Login Issues on Upstash Console
Started
Mon, Oct 20, 2025, 07:00:00 AM
Updated
Mon, Oct 20, 2025, 10:29:33 AM
Resolved
Mon, Oct 20, 2025, 07:00:00 AM
Duration
0m
resolved
As a side effect of an incident on the underlying cloud provider, Upstash Console has had availability issues between 07:00UTC and 09:23UTC.
Only Upstash Console is impacted, Upstash products remained operational.
Minor
Connectivity issue on us-east-1
Started
Fri, Oct 10, 2025, 03:46:00 PM
Updated
Fri, Oct 10, 2025, 04:49:10 PM
Resolved
Fri, Oct 10, 2025, 03:46:00 PM
Duration
0m
resolved
Between 15:46–15:55 UTC, some client connection attempts to databases in us-east-1 timed out due to unexpected high load on a server. The node was recovered at 15:50 UTC, and the updated DNS record propagated by 15:55 UTC. Services are operating normally.
Minor
Temporary Database Routing Issue
Started
Tue, Sep 9, 2025, 12:00:00 PM
Updated
Tue, Sep 9, 2025, 03:48:41 PM
Resolved
Tue, Sep 9, 2025, 12:00:00 PM
Duration
0m
resolved
Impact:
A subset of clients connecting through the eu-central-1 region experienced increased error rates and timeouts when accessing certain databases. Clients in us-west-2 were also briefly affected. The issue was limited in scope and did not impact other regions.
Root Cause:
During an ongoing migration to improve database routing reliability, a configuration step was applied inconsistently across regions.
Resolution:
Our monitoring alerted us within minutes, and the migration was promptly rolled back for the affected regions. Service definitions were restored, and normal database connectivity resumed by 15:08 UTC.
Next Steps:
We are reviewing our migration process to ensure consistency across all regions and adding additional safeguards to prevent similar issues in the future.
Minor
Connectivity Issues in us-east-1
Started
Mon, Aug 18, 2025, 12:37:02 PM
Updated
Mon, Aug 18, 2025, 01:30:54 PM
Resolved
Mon, Aug 18, 2025, 12:37:02 PM
Duration
0m
postmortem
Between 12:37 UTC and 12:50 UTC, an overload in the connection proxying system resulted in instability for a subset of databases.
The root cause was identified as a misconfiguration in the routing rules, which caused certain requests to experience timeouts during the initial phase of a gradual deployment. Upon detection, the deployment was immediately rolled back, restoring normal service.
We are reviewing our deployment and configuration validation processes to prevent similar issues in the future.
resolved
Between 12:37 UTC and 12:50 UTC, an overload in the connection proxying system resulted in instability for a subset of databases.
The root cause was identified as a misconfiguration in the routing rules, which caused certain requests to experience timeouts during the initial phase of a gradual deployment. Upon detection, the deployment was immediately rolled back, restoring normal service.
We are reviewing our deployment and configuration validation processes to prevent similar issues in the future.
Major
Login problems on Upstash Console
Started
Thu, Jun 26, 2025, 06:53:12 AM
Updated
Thu, Jun 26, 2025, 07:34:49 AM
Resolved
Thu, Jun 26, 2025, 07:34:49 AM
Duration
41m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
Some users may experience login issues due to a disruption in our authentication provider. We’re actively monitoring the situation.
Minor
QStash: Degraded performance
Started
Thu, Jun 12, 2025, 02:28:12 PM
Updated
Thu, Jun 12, 2025, 04:07:21 PM
Resolved
Thu, Jun 12, 2025, 04:07:21 PM
Duration
1h 39m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating this issue.
Critical
Degraded Performance
Started
Wed, Jun 11, 2025, 06:51:33 AM
Updated
Wed, Jun 11, 2025, 11:49:53 AM
Resolved
Wed, Jun 11, 2025, 10:02:14 AM
Duration
3h 10m
postmortem
A routine system maintenance operation at the OS level led to the application of system updates across multiple EC2 instances in our clusters in several AWS regions. These updates included changes to networking components, which inadvertently triggered restarts.
As a result, several EC2 nodes failed health checks and temporarily dropped out of the cluster, disrupting high availability and causing partial connectivity issues for some clients and operations.
We have since reproduced the issue in a controlled environment and verified the root cause. To prevent a recurrence, we are updating our node maintenance strategy to ensure greater control over the timing and impact of system-level changes and excluding networking components from automated upgrades.
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating this issue.
Minor
Performance Degradation
Started
Tue, Jun 10, 2025, 07:23:41 AM
Updated
Wed, Jun 11, 2025, 11:49:32 AM
Resolved
Tue, Jun 10, 2025, 12:08:51 PM
Duration
4h 45m
postmortem
A routine system maintenance operation at the OS level led to the application of system updates across multiple EC2 instances in our clusters in several AWS regions. These updates included changes to networking components, which inadvertently triggered restarts.
As a result, several EC2 nodes failed health checks and temporarily dropped out of the cluster, disrupting high availability and causing partial connectivity issues for some clients and operations.
We have since reproduced the issue in a controlled environment and verified the root cause. To prevent a recurrence, we are updating our node maintenance strategy to ensure greater control over the timing and impact of system-level changes and excluding networking components from automated upgrades.
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating this issue.
Minor
Global Ireland (eu-west-1) Degraded Performance
Started
Thu, May 22, 2025, 04:44:23 PM
Updated
Thu, May 22, 2025, 05:41:21 PM
Resolved
Thu, May 22, 2025, 05:41:21 PM
Duration
56m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
Minor
QStash: Degraded performance
Started
Fri, Apr 18, 2025, 12:36:14 PM
Updated
Fri, Apr 18, 2025, 01:23:31 PM
Resolved
Fri, Apr 18, 2025, 01:18:21 PM
Duration
42m
postmortem
An internal cleanup task for QStash events, coinciding with disc layer compaction task has caused performance degradation to some users.
Mitigation: Team has mitigated the event by pausing some of these tasks and monitored the status for a while.
Fixes: Improvements are being applied to these tasks to use resources more gracefully. Disk resources are increased to be able to handle a bigger burst of load.
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating this issue.
Minor
QStash: Degraded performance in request processing and event logs
Started
Wed, Apr 2, 2025, 07:15:34 AM
Updated
Thu, Apr 3, 2025, 03:58:24 PM
Resolved
Wed, Apr 2, 2025, 09:12:51 PM
Duration
13h 57m
postmortem
###
Product: QStash
### Incident Summary
Due to high load, the volume of QStash event logs reached to a point which caused latency in the underlying data store operations.
Event log creation was slowed down and lead to performance degradation in QStash request processing.
In order to resolve the performance degradation in QStash requests, event logging module was turned off temporarily.
After deploying a hot fix and configuration changes, we eventually turned on event logging and system went back to stable state again.
### Root Cause
At 07:15 UTC we received alerts on the performance degradation and started the investigation.
We discovered long running queries for synching event logs from main QStash servers to QStash event server.
In order to resolve performance degradation, we turned off event logging functionality as an immediate action.
This action turned the performance back to normal levels for QStash requests but left event log processing disabled.
We deployed a hotfix during the day to remove some redundant calls and alleviate the impact.
Around 16:20 UTC, we observed another performance degradation on QStash requests due to a load increase, and disabled event log processing again.
In the following hours, we deployed a configuration change to relax the job interval durations for event log tasks and turned on event logging again.
This configuration change helped to resolve the performance issues without any further issues.
### Impact
During the problematic timeframes, when the slow event log processing was observed, QStash requests experienced high latency and caused timeouts for customers.
No events were lost. Duplicate event deliveries were observed due to a number of restarts during the incident.
### Resolution
Improvements are applied to the event logging module to prevent the same issue from happening again.
Also, we have planned to upgrade underlying disks to stronger models.
resolved
QStash service and Event logs are fully functional without any remaning issues.
monitoring
Monitoring:
Main QStash service is back to normal.
Event logs service is back online but events will be lagging a few mins.
monitoring
Main QStash service is back to normal.
Event logs are still temporarily unavailable.
investigating
We are continuing to investigate this issue.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating this issue.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating this issue.
Minor
Performance degradation on QStash
Started
Wed, Mar 12, 2025, 09:43:07 AM
Updated
Wed, Mar 12, 2025, 12:36:09 PM
Resolved
Wed, Mar 12, 2025, 10:16:46 AM
Duration
33m
postmortem
On 09:43 UTC, QStash experienced degraded service when a high number of requests to a specific domain were throttled, resulting in timeouts during an unexpected phase of the TCP connection establishment. These requests and resulting retries triggered excessive consumption of network resources and negatively impacted all users. We have added more resources to QStash as a quick remediation and as for the resolution, we have improved the timeout mechanism to detect and fail-faster in such cases.
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
### Incident Summary
During a maintenance update to the regional Upstash Redis databases in AWS eu-west-1, several databases hosted in that region has unnecessarily triggered a full synchronisation between their primary and backup replicas.
### Root Cause
A full synchronisation is the invalidation of the whole data in the target replica and starting a fresh re-population from the source replica. Under normal circumstances, full synchronisation is required only in cases where the data integrity is lost in one of the replicas, which was not the case here.
### Impact
This incident impacted the performance of regional databases on AWS eu-west-1 only. Full synchronisation has caused a very high CPU load and caused a performance degradation on some of the databases that has a replica in this region. Moreover, our system throttles the databases that are going through this operation to allocate more CPU to the synchronisation to finish it sooner.
No data or consistency has been lost.
### Resolution
As a quick remediation, we have unthrottled affected databases on 15:06UTC and enabled more throughput, however high latency has still been observed until the full synchronisation is completed on 21:23UTC.
A fix has been prepared to avoid this unnecessary full synchronisation on regional databases, and will be deployed shortly.
This issue is not present on Upstash Global databases, which is our new generation infrastructure and is now our default offering. We will reach out to our regional users on how to migrate to Upstash Global going forward.
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
Regional AWS eu-west-1 cluster is experiencing performance degradation, and we are adding more resources to the cluster.
None
We experienced a very short period of API downtime for the incoming requests to QStash due to urgent maintenance to ensure the stability and performance of our services. Our team acted quickly to address the issue, and everything is now fully operational.
Started
Thu, Feb 20, 2025, 08:00:00 AM
Updated
Thu, Feb 20, 2025, 08:58:54 AM
Resolved
Thu, Feb 20, 2025, 08:00:00 AM
Duration
0m
resolved
We experienced a very short period of API downtime for the incoming requests to QStash due to urgent maintenance to ensure the stability and performance of our services. Our team acted quickly to address the issue, and everything is now fully operational.
None
QStash Workflow Run Failure
Started
Thu, Feb 6, 2025, 07:00:00 AM
Updated
Thu, Feb 6, 2025, 12:01:06 PM
Resolved
Thu, Feb 6, 2025, 10:10:00 AM
Duration
3h 10m
resolved
Latest QStash release caused some Workflow runs to fail due to a bug in Workflow URL detection mechanism.
Affected workflows did not start at all and moved directly to DLQ. These workflow runs, which show "detected non-workflow destination" message in the response body, can be retried from DLQ.
This was a partial failure, lasted from 10:10 to 11:00 UTC, not all users' workflows were affected.
Major
Disk failure in some Regional Databases
Started
Wed, Jan 29, 2025, 06:03:32 PM
Updated
Wed, Jan 29, 2025, 06:06:39 PM
Resolved
Wed, Jan 29, 2025, 06:03:32 PM
Duration
0m
resolved
We observed a disk failure on some instances on 16:52 UTC and restarted affected components to bring them back online. Issue is resolved on 17:13 UTC.
During the time of the issue databases reachability are affected, no data is lost.
Major
Performance degradation on QStash
Started
Wed, Nov 13, 2024, 02:14:40 PM
Updated
Wed, Nov 13, 2024, 08:48:41 PM
Resolved
Wed, Nov 13, 2024, 05:27:29 PM
Duration
3h 12m
postmortem
**Product:** QStash
**Impact:** Degraded performance, delayed processing of events, and duplicate event deliveries for some customers
## Incident Summary
QStash experienced an incident marked by a sudden and extreme load on our servers. This caused a degradation in performance, with extremely high latency for event processing for all users. We also noticed some of the events being delivered multiple times to some of the users. To mitigate the high load, we have increased the capacity as our initial response while investigation proceeds. Eventually, fixes for the issues are confirmed with an issue reproducer and deployed to production.
## Root Cause Analysis
In a certain type of usage, failure handling of [failureFunction](https://upstash.com/docs/workflow/basics/serve#failurefunction) can cause recursive calls which causes a leak in the queue of the tasks, causing a severe load on the QStash servers. This also triggered an edge case which caused some of the events to be delivered multiple times.
## Resolution
Two hotfixes to the QStash processes are deployed
1. Prevent recursive calls within the failure function.
2. Eliminate duplicate deliveries while keeping "at least once delivery" guarantee.
These are verified to successfully resolve the root cause, normalizing server load and restoring standard event processing operations.
## Impact on Customers
High latency of event processing is observed for all users. Some users received duplicate event deliveries. No events were lost, and all were delivered as part of our "at least once delivery" guarantee. Customers do not need to take any corrective action, as workflows have returned to normal and preventive fixes are deployed.
resolved
We will be sharing a postmortem about the incident soon.
identified
The issue has been identified and a fix is being implemented.
investigating
We are currently investigating this issue.
Minor
Performance degradation on QStash
Started
Wed, Nov 13, 2024, 09:31:00 AM
Updated
Wed, Nov 13, 2024, 12:49:01 PM
Resolved
Wed, Nov 13, 2024, 12:49:01 PM
Duration
3h 18m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
We are continuing to work on a fix for this issue.
identified
The issue has been identified and a fix is being implemented.
investigating
We are currently investigating this issue.
Minor
Performance degradation on QStash
Started
Mon, Sep 2, 2024, 12:18:32 PM
Updated
Mon, Sep 2, 2024, 02:53:42 PM
Resolved
Mon, Sep 2, 2024, 02:53:42 PM
Duration
2h 35m
resolved
This incident has been resolved.
monitoring
We are continuing to monitor for any further issues.
monitoring
Our data processing infrastructure is running behind. No data has been lost and the system should be caught up shortly.
Critical
QStash API Not Reachable
Started
Fri, Aug 23, 2024, 05:32:40 AM
Updated
Fri, Aug 23, 2024, 06:01:38 AM
Resolved
Fri, Aug 23, 2024, 06:01:38 AM
Duration
28m
resolved
This incident has been resolved.
identified
The issue has been identified and a fix is being implemented.
None
Partial Degraded Performance
Started
Wed, Jun 5, 2024, 02:02:00 PM
Updated
Wed, Jun 5, 2024, 03:56:37 PM
Resolved
Wed, Jun 5, 2024, 02:02:00 PM
Duration
0m
resolved
Some of the databases in us-east-1 region might have experienced increased latencies.
None
Degraded availability on Vector us-east-1
Started
Mon, Jun 3, 2024, 08:52:52 AM
Updated
Mon, Jun 3, 2024, 09:20:42 AM
Resolved
Mon, Jun 3, 2024, 09:20:42 AM
Duration
27m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
None
Degraded performance on us-east-1 Kafka
Started
Thu, May 9, 2024, 02:39:02 PM
Updated
Mon, May 13, 2024, 08:30:33 AM
Resolved
Mon, May 13, 2024, 08:30:33 AM
Duration
3d 17h
resolved
This incident has been resolved.
monitoring
We are working on a fix and monitoring the cluster performance.
investigating
We are currently investigating this issue.
None
Maintenance - Global Ap-Southeast-1
Started
Wed, Mar 13, 2024, 08:30:00 AM
Updated
Wed, Mar 13, 2024, 09:30:56 AM
Resolved
Wed, Mar 13, 2024, 08:30:00 AM
Duration
0m
resolved
We have taken actions to increase the capacity of the region. During the operation, clients might have experienced higher than usual latencies for about 15 minutes.