Platforms & Infra

status.inngest.com

Inngest

Event-driven AI workflows

All Systems Operational

Operational
Latency
378ms
Checked
just now
Active incidents
0
Components
5
Source
Statuspage.io API

Overview

Observed uptime · 1 day100%
2026-09-202026-09-20
Component health
5 up0 degraded0 down

Components5

5 operational0 degraded0 outage5 total
  • Event APIAPI for processing inbound events
    Operational
  • Inngest Dashboard
    Operational
  • Function executionRemote execution of of Inngest functions via HTTP
    Operational
  • ObservabilityMetrics, runs, and trace data
    Operational
  • API (REST and GraphQL)The GraphQL API powers the Inngest Cloud webapp and the developer REST API
    Operational

Incidents25

History 25

Major

Degraded Function Execution

Started
Fri, Sep 18, 2026, 04:31:21 PM
Updated
Fri, Sep 18, 2026, 09:17:13 PM
Resolved
Fri, Sep 18, 2026, 09:17:13 PM
Duration
4h 45m
  1. resolved

    The incident is now resolved and the system is full operational.

  2. monitoring

    A fix has been deployed and we're monitoring the system to ensure the system is fully operational.

  3. identified

    Monitoring revealed additional issues, and we're actively working on further fixes to restore normal system operation.

  4. monitoring

    A fix has been deployed and we're monitoring the system to ensure the system is fully operational.

  5. investigating

    We are actively investigating an issue with function execution. We will provider further updates as we identify the cause and resolve the issue.

Major

Degraded Function Execution

Started
Fri, Sep 4, 2026, 06:26:54 PM
Updated
Sat, Sep 5, 2026, 12:22:20 AM
Resolved
Fri, Sep 4, 2026, 07:10:00 PM
Duration
43m
  1. resolved

    The incident is now resolved and the system is full operational.

  2. monitoring

    We have identified the cause of the issue and we're monitoring the system as it is currently recovering. If users experienced any failed runs, you are able to replay failed runs in the UI. More info can be found here: <https://www.inngest.com/docs/platform/replay#how-to-create-a-new-replay>

  3. investigating

    We are actively investigating an issue with execution. Users may experience failed runs or delayed scheduling of runs from incoming events. We will provider further updates as we identify the cause and resolve the issue.

  4. investigating

    We are actively investigating an issue with function execution. We will provider further updates as we identify the cause and resolve the issue.

None

Connect Message Routing Delays - for Connect users only

Started
Thu, Aug 27, 2026, 04:03:14 PM
Updated
Thu, Aug 27, 2026, 04:03:14 PM
Resolved
Wed, Aug 26, 2026, 12:31:00 AM
Duration
  1. resolved

    Between August 26 00:31 to 01:13 UTC, a subset of customers using Inngest Connect experienced delayed and failed function runs due to an internal networking issue that disrupted connectivity between Connect services. We reverted the change and restored service. Although Connect buffers worker responses, stalled internal API calls delayed acknowledgements and subsequent replies. In some cases, request forwarding failed even after a buffered response had been received. We fixed the root cause, improved Connect’s failure handling, and enhanced our internal observability so we can detect and respond more quickly to similar problems in the future. Affected customers should review failed Connect runs during the incident window and replay them after confirming that doing so is safe and idempotent.

None

Connect Message Routing Delays - for Connect users only

Started
Tue, Aug 25, 2026, 12:25:19 AM
Updated
Tue, Aug 25, 2026, 12:25:19 AM
Resolved
Mon, Aug 24, 2026, 08:38:00 PM
Duration
  1. resolved

    Between August 24 at 23:38 UTC and August 25 at 00:08 UTC, a subset of customers using Inngest Connect experienced delayed and failed function runs due to an internal networking issue that disrupted connectivity between Connect services. We reverted the change and restored service. Although Connect buffers worker responses, stalled internal API calls delayed acknowledgements and subsequent replies. In some cases, request forwarding failed even after a buffered response had been received. We fixed the root cause, improved Connect’s failure handling, and enhanced our internal observability so we can detect and respond more quickly to similar problems in the future. Affected customers should review failed Connect runs during the incident window and replay them after confirming that doing so is safe and idempotent.

Minor

We are currently experiencing a delay with function runs being visible in the dashboard and have scaled up processing to work through the backlog.

Started
Thu, Aug 20, 2026, 09:44:40 PM
Updated
Thu, Aug 20, 2026, 10:35:00 PM
Resolved
Thu, Aug 20, 2026, 10:35:00 PM
Duration
50m
  1. resolved

    Run history delay has caught up and the incident is now resolved.

  2. monitoring

    A fix has been deployed and we're monitoring the system to ensure the system is fully operational.

Minor

Delays in Function Execution

Started
Thu, Aug 20, 2026, 01:44:50 PM
Updated
Thu, Aug 20, 2026, 03:04:12 PM
Resolved
Thu, Aug 20, 2026, 03:04:12 PM
Duration
1h 19m
  1. resolved

    The incident is now resolved and the system is full operational.

  2. monitoring

    A fix has been deployed and we're monitoring the system to ensure the system is fully operational.

  3. identified

    We are actively investigating an issue with delays with function execution. We have a fix that has been released and are monitoring the system.

Minor

Delayed function executions for a subset of customers

Started
Wed, Aug 19, 2026, 04:54:00 AM
Updated
Wed, Aug 19, 2026, 04:54:00 AM
Resolved
Wed, Aug 19, 2026, 04:41:00 AM
Duration
  1. resolved

    Execution latencies are back to normal across all customers, and the incident is now resolved.

  2. identified

    We’re investigating execution slowness affecting a subset of customers, where steps may be delayed by a few seconds to a couple of minutes in the worst cases, caused by elevated memory usage on one of our shards.

Minor

Delayed function execution for 15 minutes

Started
Mon, Aug 10, 2026, 11:06:44 PM
Updated
Mon, Aug 10, 2026, 11:06:44 PM
Resolved
Mon, Aug 10, 2026, 10:30:00 PM
Duration
  1. resolved

    We observed temporary function execution delays for a period of 15 minutes as we performed emergency maintenance to manage disk capacity to several nodes.

Major

Function processing delays

Started
Thu, Jul 23, 2026, 02:17:45 PM
Updated
Thu, Jul 23, 2026, 03:24:44 PM
Resolved
Thu, Jul 23, 2026, 03:24:44 PM
Duration
1h 6m
  1. resolved

    The incident is now resolved and the system is full operational. Delays began around 14:03 UTC caused by an issue publishing to a new Kafka topic. Recovery began around 14:40 UTC as the team was able to isolate this topic. The consuming services which schedules new function runs was scaled out and began to consume the backlog. No events were dropped during this incident, all functions were scheduled, but with delays during this time window. No action is needed for manual recovery.

  2. monitoring

    New function execution has resumed and nearly caught up from the incurred backlog. Wait for event and cancel on handlers are still in a backlog and are catching up with processing. No events were dropped during this issue, all functions will be executed, but with delays. The root cause has been confirmd and changes are underway to prevent reoccurrence.

  3. identified

    We have identified the cause of the issue - the team has deployed a fix and is scaling to catch up on function delays.

  4. investigating

    We are actively investigating an issue with function execution that is leading to delays starting function runs after events are received. We will provider further updates as we identify the cause and resolve the issue.

Major

Metrics Data Delays

Started
Thu, Jul 16, 2026, 05:52:05 PM
Updated
Thu, Jul 16, 2026, 07:29:07 PM
Resolved
Thu, Jul 16, 2026, 07:29:07 PM
Duration
1h 37m
  1. resolved

    Metrics in the UI are now restored. Users may notice metrics data continue to be unavailable during the time of the incident, from the timeframe 7/16/26 15:45 UTC until approximately 7/16/26 21:07 UTC

  2. monitoring

    A fix has been deployed and we're monitoring the system to ensure the system is fully operational. We confirmed all executions were functioning as expected, so this only impacts the metrics on the dashboard.

  3. investigating

    We are actively investigating an issue with our metrics service. Metrics displayed in the Inngest UI may not be up to date, but core platform functionality is unaffected and function execution is operating normally. We will provide further updates as we identify the cause and resolve the issue.

Minor

Latency and Execution Errors

Started
Thu, Jul 16, 2026, 11:45:59 AM
Updated
Thu, Jul 16, 2026, 12:57:46 PM
Resolved
Thu, Jul 16, 2026, 12:57:46 PM
Duration
1h 11m
  1. resolved

    Throughput and latency have recovered and the system should be fully operational again. The core issue is related to the part of the system that persists run state. Some operations experienced and increase in errors and retries, causing a backlog in some parts of the system and some failed operations like checkpoints or signals. We are continuing a thorough post-mortem investigation to prevent reoccurrance.

  2. monitoring

    We identified issues in the system and have deployed changes to mitigate the issues. Throughput has increased across the system. We continue to monitor latency as well as we expect it to reduce with the throughput increases. We are continuing to investigate for any additional issues that may have occurred.

  3. investigating

    We are actively investigating an issue with latency and execution errors. We will provider further updates as we identify the cause and resolve the issue.

Major

Issues Affecting Function Triggering from Events

Started
Fri, Jul 10, 2026, 05:44:39 PM
Updated
Mon, Jul 13, 2026, 07:44:57 PM
Resolved
Mon, Jul 13, 2026, 07:44:57 PM
Duration
3d 2h
  1. resolved

    The incident is now resolved and the system is full operational.

  2. monitoring

    We are continuing monitoring the system. Users affected by this incident may need to replay impacted events to ensure their associated functions are executed.

  3. monitoring

    We've made updates to the system, and function throughput is continuing to improve for users affected. The backlog is actively being processed, and we're closely monitoring recovery to ensure processing returns to normal.

  4. investigating

    We are investigating a partial outage affecting function triggering from incoming events. Some events may experience delays before triggering functions, or functions may not trigger as expected. Our team is actively investigating and will provide updates as more information becomes available.

Major

Degraded function execution performance

Started
Tue, Jul 7, 2026, 11:03:28 PM
Updated
Wed, Jul 8, 2026, 03:38:52 PM
Resolved
Wed, Jul 8, 2026, 03:38:52 PM
Duration
16h 35m
  1. resolved

    The incident is now resolved and the system is full operational.

  2. monitoring

    A fix has been deployed and we're monitoring the system to ensure the system is fully operational.

  3. investigating

    We are actively investigating degraded function execution performance. We will provider further updates as we identify the cause and resolve the issue.

Major

Function execution down for a small subset of customers

Started
Tue, Jul 7, 2026, 07:57:30 AM
Updated
Tue, Jul 7, 2026, 07:57:30 AM
Resolved
Tue, Jul 7, 2026, 06:45:00 AM
Duration
  1. resolved

    The incident is now resolved and the system is full operational.

  2. monitoring

    We’ve recovered the queue shard and it is now operational. We’re continuing to monitor it closely.

  3. identified

    We’ve identified that one of the queue shards is down and are currently working to recover it. Only a subset of customers is affected.

Major

Support Portal Unavailable

Started
Fri, Jun 19, 2026, 11:34:13 AM
Updated
Fri, Jun 19, 2026, 01:07:49 PM
Resolved
Fri, Jun 19, 2026, 01:07:49 PM
Duration
1h 33m
  1. resolved

    The Support Portal is now available.

  2. investigating

    We are currently investigating an issue causing the Inngest support portal to be unavailable. Our team is working on a fix, and in the meantime, any support questions can be emailed directly to [hello@inngest.com](mailto:hello@inngest.com).

Major

Downstream provider causing dashboard loading errors

Started
Tue, May 26, 2026, 09:31:17 PM
Updated
Tue, May 26, 2026, 11:46:14 PM
Resolved
Tue, May 26, 2026, 11:46:14 PM
Duration
2h 14m
  1. resolved

    The incident is now resolved.

  2. monitoring

    A fix has been deployed via Clerk and we're monitoring the system to ensure the system is fully operational.

  3. identified

    We have identified the cause of the issue and are waiting for the downstream provider to implement their fix

Major

Issues connecting to data store for events/runs/metrics pages

Started
Tue, May 26, 2026, 09:26:45 AM
Updated
Tue, May 26, 2026, 09:32:23 PM
Resolved
Tue, May 26, 2026, 09:32:23 PM
Duration
12h 5m
  1. resolved

    The previous incident with the observability database has beens resolved, but there is a separate incident with a downstream provider.

  2. monitoring

    A fix has been deployed and we're monitoring the system to ensure the system is fully operational.

  3. investigating

    We are actively investigating an issue with our backing database for several of our dashboard pages. REST API and Function execution are not currently impacted.

Minor

Elevated latency in runs and event details lookups

Started
Thu, May 21, 2026, 01:42:31 PM
Updated
Thu, May 21, 2026, 08:20:52 PM
Resolved
Thu, May 21, 2026, 08:20:52 PM
Duration
6h 38m
  1. resolved

    The incident is now resolved and the system is fully operational.

  2. monitoring

    We're monitoring the system to ensure the system is fully operational.

  3. monitoring

    A failover event in our data store provider changed the query plan behavior for some queries negatively and we had to adjust those plans to restore performance. Runs and event detail lookups are again performing normally. We're monitoring the system to ensure the system is fully operational.

  4. investigating

    We are actively investigating elevated latency with the backing data store powering runs and event lookups for the dashboard at [https://app.inngest.com/](https://app.inngest.com/. "https://app.inngest.com/.") as well as our REST APIs. We will provider further updates as we identify the cause and resolve the issue.

  5. investigating

    We are actively investigating elevated latency with the backing data store powering runs and event lookups for the dashboard at https://app.inngest.com/. We will provider further updates as we identify the cause and resolve the issue.

Minor

Increased function execution latency

Started
Tue, Apr 28, 2026, 07:11:25 PM
Updated
Tue, Apr 28, 2026, 10:48:42 PM
Resolved
Tue, Apr 28, 2026, 10:48:42 PM
Duration
3h 37m
  1. resolved

    Performance on all shards is back to normal levels. The degradation was caused by a change that aimed to improve concurrency metrics for Inngest accounts. The change, while out for a full 24 hours and fairly benign, began to compound and produced slowness for some queue shards within our system earlier today. We have reverted that change and are working to understand the performance impact of this change.

  2. monitoring

    Function execution latency has returned to normal for affected customers. Some customer shards were affected and we've pinpointed the cause of a slow degradation that compounded over time. We are working on adding new monitoring to catch performance regressions in this part of the system more quickly.

  3. investigating

    We are actively investigating increased function execution latency on a subset of customer shards starting around 5pm UTC. We will provider further updates as we identify the cause and resolve the issue.

Minor

Degraded performance for REST API (runs/events data)

Started
Thu, Apr 23, 2026, 08:06:08 PM
Updated
Sat, Apr 25, 2026, 06:04:36 PM
Resolved
Sat, Apr 25, 2026, 06:04:36 PM
Duration
1d 21h
  1. resolved

    The incident is now resolved and the system is full operational.

  2. monitoring

    We have fixed the issue affecting a subset of users utilizing our REST API for retrieving runs and events data for recent updates and are monitoring for issues with availability of historical data over the REST API.

  3. identified

    We are currently experiencing degraded performance affecting a subset of users utilizing our REST API for retrieving runs and events data. Responses may be delayed or temporarily return incomplete data, with recent updates not appearing immediately. We have identified the cause of the issue. We're actively working on implementing a fix to resume normal operation of the REST API. Function execution is not affected and is working as expected.

Minor

Delays in function run scheduling

Started
Wed, Apr 15, 2026, 12:15:58 PM
Updated
Wed, Apr 15, 2026, 11:21:38 PM
Resolved
Wed, Apr 15, 2026, 11:21:38 PM
Duration
11h 5m
  1. resolved

    The incident is now resolved and the system is full operational.

  2. monitoring

    Async "pause" operations (step.waitForEvent, step.invoke, cancelOn) should be running again at typical throughput. Event batching still has an backlog that we are actively working through. Status: • Function scheduling - Running as expected • Function execution - Running as expected, no queue backlogs • Async "pause" opts (waitForEvent, invoke, cancelOn) - Running as expected • Event batching - Significant backlog

  3. monitoring

    The function run backlog for functions has been resolved as of 11:10AM PT. Batched functions and `step.waitForEvent` may face delays as the backlog continues to process.

  4. identified

    Function execution scheduling and processing throughput is at normal levels. Async "pause" operations (step.waitForEvent, step.invoke, cancelOn) are severely backlogged which may cause delays in any of these operations from completing. This may cause issues with your function execution if you rely on them. The team is working on fixes and clear this backlog and fix the key issues. We do not yet have an ETA on resolving this specific issue.

  5. identified

    Function scheduled delays should be caught up. With events processed and new runs scheduled, your system my still see backlogs based on your function's flow control (e.g. concurrency) config and your account's concurrency. We still see backlogs in processing step.waitForEvent, step.invoke and cancelOn event expressions. We are continuing to work on this. We also are continuing our rollout of isolated batch processing as previously mentioned to further isolate parts of our system. EDIT - This was edited to include step.invoke as well for completeness.

  6. identified

    The system is consuming the event backlog as fast as possible, with an ETA of ~10-15 minutes until function scheduling is caught up. After function scheduling is caught up, function execution in your account may still be limited by your account concurrency or a given function's own flow control settings (concurrency, rate limit, etc.). We will continue to share more updates as soon as we can.

  7. identified

    There is an increase in throughput since 15:37 UTC (~15 min ago). We are continuing to apply changes and prepare a larger change to decouple parts of the system. **Event observability**: Events may be delayed when appearing in the dashboard as the database ingestion for these events is also related to this part of the system that handles function scheduling. Events continue to be ingested and the Event API remains unaffected.

  8. identified

    Changes have increase throughput, but not yet to typical levels. We are actively testing the new system change to decouple batch processing before enabling it for all accounts.

  9. identified

    We're deploying an in-memory optimization within the the part of the system the schedules new function runs. This optimization will alleviate pressure on underlying systems and increase throughput. The change will be rolled out momentarily. We're also working in parallel on a system change to create a dedicated service for processing for event batching which is the cause of the overall backlog on the system.

  10. identified

    We have scaled up several resources across the system and to handle a large increase in scale within the system. Services are scaled up and we have also added new function state shards, but rollout of those new shards can take up to ~30m. We are also working on networking improvements to improve efficiency of the system with this significantly higher load.

  11. investigating

    We are actively investigating delays with function run scheduling for a subset of customers. We will provide further updates as we identify the cause and resolve the issue.

Critical

Dashboard down

Started
Thu, Apr 2, 2026, 05:07:08 PM
Updated
Thu, Apr 2, 2026, 05:49:59 PM
Resolved
Thu, Apr 2, 2026, 05:49:59 PM
Duration
42m
  1. resolved

    The incident is now resolved and the system is full operational. Related to the Vercel incident, we updated our dashboard to Node 22.x to solve the issue. We continue to monitor Vercel's incident and react accordingly. https://www.vercel-status.com/incidents/5r9bp5y8rql2

  2. monitoring

    A fix has been deployed and the dashboard is back online. We will continue to monitor the Vercel status page for any changes to the incident itself.

  3. identified

    We area pushing a hotfix to the dashboard as recommended by Vercel's incident report. The rest of the Inngest system, function execution, API, etc. all remain functional.

  4. identified

    The Inngest dashboard is down due to an issue with our downstream provider. Vercel. We are working quickly to bring this back up

Major

Increased failures with step.fetch, step.ai.infer

Started
Tue, Mar 31, 2026, 11:14:43 AM
Updated
Tue, Mar 31, 2026, 11:59:31 AM
Resolved
Tue, Mar 31, 2026, 11:59:31 AM
Duration
44m
  1. resolved

    The incident is now resolved and the system is full operational. During this incident step.fetch and step.ai.infer were failing due to a bug causing empty request bodies to be returned. The root cause was determined, the system was rolled back and a fix will be rolled out today.

  2. monitoring

    A fix has been deployed for step.fetch and step.ai.infer and we're monitoring the system to ensure the system is fully operational. We continue to investigate the root cause.

  3. investigating

    We are actively investigating an issue with proxied requests via step.fetch or step.ai.infer. We will provider further updates as we identify the cause and resolve the issue. If you are not using these features your system should remain unaffected.

Minor

Function run scheduling delays

Started
Thu, Mar 26, 2026, 10:42:16 PM
Updated
Thu, Mar 26, 2026, 11:45:37 PM
Resolved
Thu, Mar 26, 2026, 11:45:37 PM
Duration
1h 3m
  1. resolved

    The incident is now resolved and the system is full operational. This was related to an issue caused by the part of the system powering the debounce feature. The internal event backlog is fully caught up and the two mitigations deployed have addressed the issue. The team is preparing a post-mortem to ensure this issue does not reoccur.

  2. monitoring

    The mitigations have fixed the issue and the event backlog is now caught up. We continue to monitor and evaluate other short and long term mitigations to add.

  3. identified

    We have deployed an additional hot fix. The earlier change rolled out have addressed the core issue and the system is now processing the backlog. The backlog is decreasing. We will provide another update as we have an estimate on time to recovery.

  4. identified

    We have identified the cause of the issue affecting a core system queue. We are rolling our a mitigation now and preparing follow up changes.

  5. investigating

    We are actively investigating delays with function run scheduling. We will provider further updates as we identify the cause and resolve the issue.

Major

Degraded function execution performance

Started
Thu, Mar 26, 2026, 11:12:59 AM
Updated
Thu, Mar 26, 2026, 02:55:42 PM
Resolved
Thu, Mar 26, 2026, 02:55:42 PM
Duration
3h 42m
  1. resolved

    After an extended monitoring period, we are resolving this incident. The system is full operational.

  2. monitoring

    Function execution has returned to normal levels as of 11:12 UTC. We are actively looking into the root cause and taking further measures to stabilize the system.

  3. investigating

    We are actively investigating an issue with internal networking. We will provider further updates as we identify the cause and resolve the issue. Function execution is picking up again and was degraded between 11:02 and 11:10 AM UTC.

  4. investigating

    We are actively investigating an issue with function execution and other core system health. We will provider further updates as we identify the cause and resolve the issue.

Watch Inngest
Email alerts on every status change — outages, degradations, new incidents, and resolutions.

Watching all 1 providers. Customize on the alerts page.

Details

Aliases
inngest ai
Indicator
none
Path
/inngest

Related in Platforms & Infra