Fast Inference

status.baseten.co

Baseten

Model deployment & Truss

All Systems Operational

Operational
Latency
88ms
Checked
just now
Active incidents
0
Components
6
Source
Statuspage.io API

Overview

Observed uptime · 1 day100%
2026-09-202026-09-20
Component health
6 up0 degraded0 down

Components6

6 operational0 degraded0 outage6 total
  • Dedicated Inference
    Operational
  • Model APIs
    Operational
  • Training
    Operational
  • Model Management API
    Operational
  • Web Application
    Operational
  • Homepage and DocsBaseten documentation website
    Operational

Incidents50

History 50

Major

Kimi K3 Model API outage

Started
Sun, Sep 20, 2026, 07:19:01 AM
Updated
Sun, Sep 20, 2026, 08:13:44 AM
Resolved
Sun, Sep 20, 2026, 08:13:44 AM
Duration
54m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    Calls to the US-hosted endpoint are resulting in errors. A fix is in progress.

  4. identified

    ~20% of calls are resulting in errors. A fix is in progress.

  5. identified

    The issue has been identified and a fix is being implemented.

Minor

Elevated errors for models in a single India-based cluster

Started
Tue, Sep 15, 2026, 01:22:38 AM
Updated
Tue, Sep 15, 2026, 02:07:38 AM
Resolved
Tue, Sep 15, 2026, 02:07:38 AM
Duration
44m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are currently investigating this issue.

Minor

Partial outage on Kimi K3 model API inference (non-US only)

Started
Mon, Sep 14, 2026, 04:21:34 PM
Updated
Mon, Sep 14, 2026, 05:03:20 PM
Resolved
Mon, Sep 14, 2026, 05:03:20 PM
Duration
41m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are currently investigating this issue.

Minor

Partial outage on Nemotron Ultra Model API. **Dedicated Inference is not affected**

Started
Wed, Sep 9, 2026, 05:42:43 PM
Updated
Wed, Sep 9, 2026, 05:52:20 PM
Resolved
Wed, Sep 9, 2026, 05:49:06 PM
Duration
6m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are currently investigating this issue.

Major

Elevated inference 500s for models in a single US-based cluster

Started
Mon, Aug 31, 2026, 06:38:34 AM
Updated
Mon, Aug 31, 2026, 07:02:44 AM
Resolved
Mon, Aug 31, 2026, 07:02:44 AM
Duration
24m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Fix is in place and traffic is back.

  3. investigating

    We are currently investigating this issue.

Minor

Elevated errors for models in a single US cluster due to provider outage

Started
Thu, Aug 27, 2026, 09:17:41 PM
Updated
Thu, Aug 27, 2026, 09:31:55 PM
Resolved
Thu, Aug 27, 2026, 09:31:55 PM
Duration
14m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Auto-recovery is in progress.

  3. investigating

    We are currently investigating this issue.

Minor

Elevated errors for deepseek-ai/DeepSeek-V4-Flash-0731 on Model APIs. **Dedicated Inference is not affected**

Started
Tue, Aug 25, 2026, 04:17:49 PM
Updated
Tue, Aug 25, 2026, 04:59:02 PM
Resolved
Tue, Aug 25, 2026, 04:59:02 PM
Duration
41m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix is in place and we are monitoring

  3. investigating

    We are currently investigating this issue.

Minor

Elevated inference 500s for models in an EU cluster due to provider outage

Started
Tue, Aug 25, 2026, 12:34:49 PM
Updated
Tue, Aug 25, 2026, 01:15:33 PM
Resolved
Tue, Aug 25, 2026, 01:15:33 PM
Duration
40m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    Failover in progress.

Minor

Elevated errors for moonshotai/Kimi-K3 for US-pinned traffic on Model APIs. **Dedicated Inference is not affected**

Started
Wed, Aug 19, 2026, 01:23:41 AM
Updated
Wed, Aug 19, 2026, 02:00:12 AM
Resolved
Wed, Aug 19, 2026, 02:00:12 AM
Duration
36m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are currently investigating this issue.

Minor

Elevated errors for moonshotai/Kimi-K3 for US-pinned traffic on Model APIs. **Dedicated Inference is not affected**

Started
Tue, Aug 18, 2026, 07:04:18 PM
Updated
Tue, Aug 18, 2026, 08:21:42 PM
Resolved
Tue, Aug 18, 2026, 08:21:42 PM
Duration
1h 17m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are currently investigating this issue.

Minor

Elevated errors for moonshotai/Kimi-K3 for US-pinned traffic on Model APIs. **Dedicated Inference is not affected**

Started
Tue, Aug 18, 2026, 05:24:01 PM
Updated
Tue, Aug 18, 2026, 05:43:58 PM
Resolved
Tue, Aug 18, 2026, 05:35:50 PM
Duration
11m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are currently investigating this issue.

Minor

Elevated errors for moonshotai/Kimi-K3 for US-pinned traffic on Model APIs. **Dedicated Inference is not affected**

Started
Tue, Aug 18, 2026, 04:27:04 PM
Updated
Tue, Aug 18, 2026, 05:14:44 PM
Resolved
Tue, Aug 18, 2026, 05:14:44 PM
Duration
47m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are currently investigating this issue.

Minor

Elevated errors for moonshotai/Kimi-K3 for US-pinned traffic

Started
Mon, Aug 17, 2026, 04:56:06 PM
Updated
Mon, Aug 17, 2026, 05:41:45 PM
Resolved
Mon, Aug 17, 2026, 05:41:45 PM
Duration
45m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are currently investigating this issue.

Minor

Metrics degradation in UI and API. **Inference not affected**

Started
Mon, Aug 17, 2026, 01:11:10 PM
Updated
Mon, Aug 17, 2026, 01:35:10 PM
Resolved
Mon, Aug 17, 2026, 01:35:10 PM
Duration
24m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are currently investigating this issue.

Minor

Elevated inference 502s for a subset of models in a single US cluster

Started
Sat, Aug 1, 2026, 11:57:54 PM
Updated
Sun, Aug 2, 2026, 12:56:47 AM
Resolved
Sun, Aug 2, 2026, 12:41:11 AM
Duration
43m
  1. resolved

    This incident has been resolved.

  2. identified

    The issue has been identified and a fix is being implemented.

  3. investigating

    This is due to a networking failure in a cloud provider.

Minor

Elevated inference 502s for a subset of models in a single US cluster

Started
Fri, Jul 31, 2026, 11:52:09 PM
Updated
Sat, Aug 1, 2026, 12:20:55 AM
Resolved
Sat, Aug 1, 2026, 12:20:55 AM
Duration
28m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are currently investigating this issue.

Minor

Slower than usual scale-ups for B200s

Started
Fri, Jul 31, 2026, 04:46:29 PM
Updated
Fri, Jul 31, 2026, 04:58:01 PM
Resolved
Fri, Jul 31, 2026, 04:58:01 PM
Duration
11m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are currently investigating this issue.

Minor

Elevated inference failures for models in an EU cluster

Started
Wed, Jul 22, 2026, 05:49:51 PM
Updated
Wed, Jul 22, 2026, 06:22:45 PM
Resolved
Wed, Jul 22, 2026, 06:22:45 PM
Duration
32m
  1. resolved

    This incident has been resolved.

  2. identified

    We are continuing to work on a fix for this issue.

  3. identified

    Failure in GPU provider's cooling system. Models are automatically migrating over to other EU clusters.

Minor

Slower than usual scale-ups in 1 us-central cluster

Started
Wed, Jul 22, 2026, 05:45:12 PM
Updated
Wed, Jul 22, 2026, 06:04:30 PM
Resolved
Wed, Jul 22, 2026, 06:04:30 PM
Duration
19m
  1. resolved

    This incident has been resolved.

  2. identified

    The issue has been identified and a fix is being implemented.

Major

GLM 5.2 Model API partial outage due to provider network failure

Started
Sat, Jul 18, 2026, 10:40:12 PM
Updated
Sat, Jul 18, 2026, 11:12:32 PM
Resolved
Sat, Jul 18, 2026, 11:12:32 PM
Duration
32m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We are continuing to monitor for any further issues.

  3. monitoring

    A fix has been implemented and we are monitoring the results.

  4. identified

    Escalated Dedicated inference is *not* affected.

Minor

Elevated inference errors in 1 Canada cluster due to provider network outage

Started
Sat, Jul 18, 2026, 12:22:55 AM
Updated
Sat, Jul 18, 2026, 12:28:34 AM
Resolved
Sat, Jul 18, 2026, 12:28:34 AM
Duration
5m
  1. resolved

    This incident has been resolved.

  2. identified

    The issue has been identified and a fix is being implemented.

Minor

Elevated errors in accessing the Baseten UI. **Inference is not affected**

Started
Thu, Jul 16, 2026, 09:06:26 AM
Updated
Thu, Jul 16, 2026, 10:25:13 AM
Resolved
Thu, Jul 16, 2026, 10:25:13 AM
Duration
1h 18m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We are starting to see recovery and are continuing to monitor.

  3. investigating

    We are continuing to investigate this issue.

  4. investigating

    We are currently investigating this issue.

Major

Partial Outage on Model APIs

Started
Mon, Jul 13, 2026, 12:04:01 AM
Updated
Mon, Jul 13, 2026, 02:10:34 AM
Resolved
Mon, Jul 13, 2026, 02:10:34 AM
Duration
2h 6m
  1. resolved

    Kimi-K2.6 recovered.

  2. monitoring

    Identified an infiniband outage in one of our providers. GLM-5.2 Recovered. Kimi-K2.6 recovering.

  3. identified

    Affecting Kimi-K2.6 and GLM-5.2.

Minor

Elevated inference 500s for models deployed in 1 US-east cluster

Started
Thu, Jul 9, 2026, 02:07:51 PM
Updated
Thu, Jul 9, 2026, 02:38:26 PM
Resolved
Thu, Jul 9, 2026, 02:28:16 PM
Duration
20m
  1. resolved

    This incident has been resolved.

  2. monitoring

    The provider had a network outage. Affected models are automatically moving to other clusters.

  3. identified

    The issue has been identified and a fix is being implemented.

Minor

Metrics degradation in UI and API. **Inference not affected**

Started
Wed, Jul 8, 2026, 04:50:15 PM
Updated
Wed, Jul 8, 2026, 11:03:09 PM
Resolved
Wed, Jul 8, 2026, 06:45:29 PM
Duration
1h 55m
  1. resolved

    This incident has been resolved.

  2. investigating

    Incorrect values recorded for some metrics

Minor

Elevated inference 500s — network outage in our cloud provider in Delaware

Started
Mon, Jul 6, 2026, 09:08:27 PM
Updated
Mon, Jul 6, 2026, 09:22:15 PM
Resolved
Mon, Jul 6, 2026, 09:22:15 PM
Duration
13m
  1. resolved

    This incident has been resolved.

  2. identified

    The issue has been identified and a fix is being implemented.

Major

Elevated errors for GLM 5.2 on Model APIs. **Dedicated Inference is not affected**

Started
Fri, Jul 3, 2026, 02:59:34 AM
Updated
Fri, Jul 3, 2026, 04:25:22 AM
Resolved
Fri, Jul 3, 2026, 04:25:22 AM
Duration
1h 25m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are continuing to investigate this issue.

  4. investigating

    We are currently investigating this issue.

Minor

Elevated inference 500s for a subset of models deployed in 1 cluster

Started
Fri, Jul 3, 2026, 12:21:01 AM
Updated
Fri, Jul 3, 2026, 12:36:48 AM
Resolved
Fri, Jul 3, 2026, 12:36:48 AM
Duration
15m
  1. resolved

    This incident has been resolved.

  2. identified

    We are continuing to work on a fix for this issue.

  3. identified

    The issue has been identified and a fix is being implemented.

Minor

Degraded metrics performance in dashboard and API. **Inference is not affected**

Started
Thu, Jul 2, 2026, 10:00:12 PM
Updated
Thu, Jul 2, 2026, 11:13:51 PM
Resolved
Thu, Jul 2, 2026, 11:13:51 PM
Duration
1h 13m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are currently investigating this issue.

Major

Elevated errors in a single EU cluster

Started
Wed, Jul 1, 2026, 10:30:46 AM
Updated
Wed, Jul 1, 2026, 08:53:51 PM
Resolved
Wed, Jul 1, 2026, 11:30:11 AM
Duration
59m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are continuing to investigate this issue.

  3. investigating

    Elevated errors due to degradation in a single EU cluster

Minor

Elevated error rates on GLM 5.2 Model APIs. **Dedicated Inference is not affected**

Started
Mon, Jun 29, 2026, 06:26:41 PM
Updated
Mon, Jun 29, 2026, 08:11:15 PM
Resolved
Mon, Jun 29, 2026, 07:01:52 PM
Duration
35m
  1. resolved

    This incident has been resolved.

  2. investigating

    Elevated error rates on GLM 5.2 Model APIs. **Dedicated Inference is not affected**

Minor

Elevated inference 500s for models in 1 cluster due to cloud provider outage

Started
Sat, Jun 27, 2026, 06:35:31 AM
Updated
Sat, Jun 27, 2026, 06:50:31 AM
Resolved
Sat, Jun 27, 2026, 06:50:31 AM
Duration
15m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Models are auto-moving to other clusters

  3. identified

    The issue has been identified and a fix is being implemented.

Minor

Elevated model deployment failures. **Inference is not affected**

Started
Fri, Jun 26, 2026, 11:41:14 PM
Updated
Sat, Jun 27, 2026, 12:55:40 AM
Resolved
Sat, Jun 27, 2026, 12:55:40 AM
Duration
1h 14m
  1. resolved

    The incident has resolved. New model deployments are successfully coming up.

  2. investigating

    New model deployments failing to become ready

Major

Elevated errors for dedicated inference in a subset of clusters

Started
Fri, Jun 26, 2026, 10:00:38 PM
Updated
Fri, Jun 26, 2026, 10:15:23 PM
Resolved
Fri, Jun 26, 2026, 10:15:04 PM
Duration
14m
  1. resolved

    This incident has been resolved.

  2. identified

    We are continuing to work on a fix for this issue.

  3. identified

    Nearly all clusters have recovered.

  4. identified

    The issue has been identified and a fix is being implemented.

  5. investigating

    We are currently investigating this issue.

Minor

Slower than usual scale-ups. High failure rate for new model deploys. Existing replicas aren't impacted.

Started
Tue, Jun 23, 2026, 12:10:14 AM
Updated
Tue, Jun 23, 2026, 12:54:43 AM
Resolved
Tue, Jun 23, 2026, 12:54:43 AM
Duration
44m
  1. resolved

    This incident has been resolved.

  2. identified

    We are continuing to work on a fix for this issue.

  3. identified

    The issue has been identified and a fix is being implemented.

Minor

Elevated errors for models in 1 Austin-based cluster affecting some models on B200s

Started
Thu, Jun 18, 2026, 12:46:34 PM
Updated
Thu, Jun 18, 2026, 02:30:17 PM
Resolved
Thu, Jun 18, 2026, 02:30:17 PM
Duration
1h 43m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    GLM 5.1 and Kimi k2.6 Model APIs are also affected.

Minor

Elevated errors for models in 1 Sweden-based cluster

Started
Wed, Jun 17, 2026, 04:40:04 AM
Updated
Wed, Jun 17, 2026, 06:03:13 AM
Resolved
Wed, Jun 17, 2026, 06:03:13 AM
Duration
1h 23m
  1. resolved

    This incident has been resolved.

  2. identified

    The issue has been identified and a fix is being implemented.

  3. investigating

    We are currently investigating this issue.

Minor

Elevated errors for models in a single EU cloud provider

Started
Mon, Jun 15, 2026, 04:14:00 PM
Updated
Mon, Jun 15, 2026, 04:58:16 PM
Resolved
Mon, Jun 15, 2026, 04:58:16 PM
Duration
44m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are currently investigating this issue.

Minor

Elevated errors for models in a single cloud provider

Started
Sat, Jun 13, 2026, 06:25:24 PM
Updated
Sat, Jun 13, 2026, 07:09:52 PM
Resolved
Sat, Jun 13, 2026, 07:09:52 PM
Duration
44m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    Network outage in single cloud provider (https://status.vultr.com/), elevated errors for some models

Major

Elevated errors in a single US-East cluster

Started
Fri, Jun 5, 2026, 02:10:16 AM
Updated
Fri, Jun 5, 2026, 03:36:17 AM
Resolved
Fri, Jun 5, 2026, 02:29:32 AM
Duration
19m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are currently investigating this issue.

Minor

Elevated latencies for a subset of models in a US-East cluster

Started
Thu, Jun 4, 2026, 08:53:22 PM
Updated
Thu, Jun 4, 2026, 09:59:44 PM
Resolved
Thu, Jun 4, 2026, 09:59:44 PM
Duration
1h 6m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    The issue has been identified and a fix is being implemented.

  4. investigating

    We are currently investigating this issue.

Minor

Slower than usual scale-ups for some models using H100s

Started
Wed, Jun 3, 2026, 02:56:57 PM
Updated
Wed, Jun 3, 2026, 06:49:26 PM
Resolved
Wed, Jun 3, 2026, 05:11:42 PM
Duration
2h 14m
  1. resolved

    This incident has been resolved.

  2. identified

    The issue has been identified and a fix is being implemented.

Minor

Elevated async inference errors in a single cluster

Started
Fri, May 29, 2026, 07:31:24 PM
Updated
Fri, May 29, 2026, 08:04:13 PM
Resolved
Fri, May 29, 2026, 08:04:13 PM
Duration
32m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are currently investigating this issue.

Minor

Elevated errors in a single cluster

Started
Fri, May 15, 2026, 01:00:00 AM
Updated
Tue, May 26, 2026, 06:55:13 PM
Resolved
Fri, May 15, 2026, 03:00:00 AM
Duration
2h
  1. resolved

    This incident has been resolved.

  2. investigating

    Elevated errors in our us-east1 cluster

Minor

Slower than usual scale-ups in 1 India cluster

Started
Tue, May 26, 2026, 04:21:51 PM
Updated
Tue, May 26, 2026, 04:35:27 PM
Resolved
Tue, May 26, 2026, 04:35:27 PM
Duration
13m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    The issue has been identified and a fix is being implemented.

Minor

Elevated errors for models in single region

Started
Wed, May 20, 2026, 01:40:54 AM
Updated
Wed, May 20, 2026, 01:54:13 AM
Resolved
Wed, May 20, 2026, 01:54:13 AM
Duration
13m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are investigating elevated errors and latency for models in a single region

None

Slower than usual scale-ups for L4s in US-East

Started
Mon, May 18, 2026, 03:20:36 PM
Updated
Mon, May 18, 2026, 03:59:17 PM
Resolved
Mon, May 18, 2026, 03:59:07 PM
Duration
38m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are currently investigating this issue.

None

Slower than usual scale-ups in EU due to a provider failure

Started
Sat, May 16, 2026, 03:32:02 PM
Updated
Sat, May 16, 2026, 06:07:08 PM
Resolved
Sat, May 16, 2026, 05:05:02 PM
Duration
1h 32m
  1. resolved

    This incident has been resolved.

  2. identified

    Root-cause has been identified and a fix is in progress.

None

Higher latencies for a subset of models in one cluster. Root caused and a fix is in progress

Started
Sat, May 16, 2026, 01:02:31 AM
Updated
Sat, May 16, 2026, 01:39:09 AM
Resolved
Sat, May 16, 2026, 01:39:06 AM
Duration
36m
  1. resolved

    The incident is resolved.

  2. monitoring

    Fix is in place. Monitoring.

  3. identified

    The issue has been identified and a fix is being implemented.

Major

Training cluster is down due to a provider failure

Started
Sun, May 10, 2026, 12:30:56 AM
Updated
Sun, May 10, 2026, 03:26:54 AM
Resolved
Sun, May 10, 2026, 03:24:36 AM
Duration
2h 53m
  1. resolved

    The issue has been resolved.

  2. monitoring

    Monitoring the fix. Most nodes are back online, but not all.

  3. identified

    Root-caused and fix is underway. Inference is not affected.

Watch Baseten
Email alerts on every status change — outages, degradations, new incidents, and resolutions.

Watching all 1 providers. Customize on the alerts page.

Details

Aliases
truss, model apis
Indicator
none
Path
/baseten

Related in Fast Inference