Platforms & Infra

status.braintrust.dev

Braintrust

LLM evals & monitoring

All Systems Operational

Operational
Latency
164ms
Checked
just now
Active incidents
0
Components
4
Source
Statuspage.io API

Overview

Observed uptime · 1 day100%
2026-09-202026-09-20
Component health
4 up0 degraded0 down

Components4

4 operational0 degraded0 outage4 total
  • Web UI / Control PlaneThe braintrust.dev UI and Control Plane API
    Operational
  • Centrally-Hosted Data Plane (EU)The EU backend which handles reads and writes to Braintrust. If you are self-hosting Braintrust, this service should not apply to you
    Operational
  • AI GatewayCentrally hosted AI Gateway for accessing LLM providers
    Operational
  • Centrally-Hosted Data Plane (US)The backend which handles reads and writes to Braintrust. If you are self-hosting Braintrust, this service should not apply to you
    Operational

Incidents25

History 25

Major

Slow or failed data loading in US region

Started
Wed, Sep 9, 2026, 09:30:35 PM
Updated
Wed, Sep 9, 2026, 10:10:23 PM
Resolved
Wed, Sep 9, 2026, 10:10:23 PM
Duration
39m
  1. resolved

    The issue causing slow or failed data loading in our US-hosted service has been resolved. We're sorry for the disruption.

  2. monitoring

    We've applied a mitigation for the data-loading issue affecting our US-hosted service and are monitoring performance. We'll share another update once we've confirmed stability.

  3. investigating

    We're investigating slow or failed data loading affecting some customers on our US-hosted service. Affected pages may load slowly, hang, or show errors.

Major

Elevated 5XXs on Control Plane API

Started
Fri, Sep 4, 2026, 03:19:49 PM
Updated
Fri, Sep 4, 2026, 03:52:23 PM
Resolved
Fri, Sep 4, 2026, 03:52:23 PM
Duration
32m
  1. resolved

    The issue causing elevated errors in the Braintrust control plane APIs has been resolved.

  2. monitoring

    We have mitigated the issue causing elevated errors in the Braintrust control plane APIs, we are continuing to monitor closely.

  3. investigating

    We are investigating elevated errors affecting the Braintrust control plane APIs. Some users may experience failures when accessing projects, prompts, settings, or creating new resources.

Major

Gateway 5xxs

Started
Fri, Aug 28, 2026, 08:38:26 PM
Updated
Fri, Aug 28, 2026, 08:38:26 PM
Resolved
Fri, Aug 28, 2026, 06:54:00 PM
Duration
  1. resolved

    The load spike has subsided and 5xxs are mitigated

  2. investigating

    The gateway encountered a load spike leading to 5xx errors.

Major

Gateway 5xxs

Started
Fri, Aug 28, 2026, 06:24:47 PM
Updated
Fri, Aug 28, 2026, 06:24:47 PM
Resolved
Fri, Aug 28, 2026, 03:55:00 PM
Duration
  1. resolved

    The deploy was rolled back and gateway 5xxs were mitigated.

  2. investigating

    A fresh gateway deploy was attempted and led to 5xx errors, and then was immediately rolled back.

Major

Gateway 5xxs

Started
Fri, Aug 28, 2026, 06:21:24 PM
Updated
Fri, Aug 28, 2026, 06:21:24 PM
Resolved
Thu, Aug 27, 2026, 01:15:00 PM
Duration
  1. resolved

    The faulty deploy was rolled back and 5xxs were mitigated.

  2. investigating

    A deployment caused the gateway to begin returning 5xxs for a portion of traffic.

None

Upstream package mirror is degraded, Brainstore is not currently affected

Started
Wed, Aug 19, 2026, 06:01:58 PM
Updated
Wed, Aug 19, 2026, 10:35:25 PM
Resolved
Wed, Aug 19, 2026, 10:35:25 PM
Duration
4h 33m
  1. resolved

    The upstream APT package mirror in AWS us-east-1 has recovered and we're no longer seeing impact when starting new brainstore instances.

  2. monitoring

    The upstream APT package mirror on AWS us-east-1 is degraded. **There is currently no downtime** and we are looking into alternative package mirrors in the meantime. Please note that for **both the Braintrust-hosted data plane and self-hosted data planes:** EC2 instance replacement due to healthcheck failures or deploys may result in failed instance initialization and downtime until the upstream issue is resolved.

Major

Gateway 5XXs

Started
Mon, Aug 3, 2026, 07:52:06 PM
Updated
Mon, Aug 3, 2026, 08:19:15 PM
Resolved
Mon, Aug 3, 2026, 08:19:15 PM
Duration
27m
  1. resolved

    The issue has been resolved and service has remained stable following the rollback. We will continue to monitor internally, but this incident is now considered resolved.

  2. monitoring

    The issue has been resolved. We are monitoring the service to confirm normal operation.

  3. identified

    We identified an issue causing 5XXs introduced by a recent Gateway deployment and are currently rolling it back.

Critical

High 504 error rate in EU data plane

Started
Thu, Jul 16, 2026, 08:32:44 AM
Updated
Thu, Jul 16, 2026, 09:01:09 AM
Resolved
Thu, Jul 16, 2026, 09:01:09 AM
Duration
28m
  1. resolved

    We have reverted to an older configuration to avoid an AWS outage with Cloudfront.

  2. investigating

    The hosted EU data plane is currently experiencing a very high rate of 504 errors.

Critical

US Data plane outage

Started
Fri, Jun 12, 2026, 06:38:22 PM
Updated
Fri, Jun 12, 2026, 06:44:50 PM
Resolved
Fri, Jun 12, 2026, 06:26:00 PM
Duration
  1. resolved

    A configuration change was identified as the root cause and reverted.

  2. investigating

    Our US data plane is encountering issues.

Major

Issues loading logs page

Started
Thu, Jun 4, 2026, 11:06:33 PM
Updated
Thu, Jun 4, 2026, 11:09:57 PM
Resolved
Thu, Jun 4, 2026, 11:07:00 PM
Duration
0m
  1. resolved

    We have pushed a fix and the loading issues have been remediated.

  2. investigating

    We're experiencing issues loading logs and some other pages in Braintrust. Our team is actively investigating and working to resolve.

Minor

Automation degradation

Started
Thu, Jun 4, 2026, 01:23:21 AM
Updated
Thu, Jun 4, 2026, 01:23:21 AM
Resolved
Sat, May 16, 2026, 06:50:00 PM
Duration
  1. resolved

    We determined the issue was caused by a role used by the automation service losing required permissions following a configuration change. The issue was reported on May 15 at approximately 1:00 PM PT and resolved on May 16 at 11:50 AM PT. To help prevent similar incidents, we have improved our alerting and added audit logging for role and permission changes.

  2. investigating

    On May 15, some automations stopped running which impacted scorer executions.

Major

Errors logging into Braintrust

Started
Tue, May 26, 2026, 09:33:09 PM
Updated
Tue, Jun 2, 2026, 06:08:17 AM
Resolved
Tue, May 26, 2026, 10:11:00 PM
Duration
37m
  1. resolved

    Our auth provider has recovered.

  2. investigating

    Our upstream auth provider is having issues. Fresh logins to Braintrust will fail with an error.

Major

High rate of errors on control plane

Started
Tue, May 5, 2026, 11:35:29 PM
Updated
Wed, May 6, 2026, 12:41:06 AM
Resolved
Wed, May 6, 2026, 12:23:00 AM
Duration
47m
  1. resolved

    The issue was caused by our firewall provider unexpectedly blocking traffic despite safeguards we had in place to prevent this behavior. We are actively working with the vendor on additional mitigation measures to reduce the likelihood of this occurring again.

  2. resolved

    We are no longer seeing issues.

  3. monitoring

    We are actively monitoring the fix.

  4. identified

    We have identified the cause and are working with an upstream provider on a fix.

  5. investigating

    We're observing a high rate of errors on the production control plane which is causing errors on braintrust.dev and data plane operations.

Critical

Control plane outage

Started
Sun, Apr 26, 2026, 07:17:19 AM
Updated
Sun, Apr 26, 2026, 07:23:10 AM
Resolved
Sun, Apr 26, 2026, 07:23:09 AM
Duration
5m
  1. resolved

    Provider has recovered. Error levels back to zero.

  2. investigating

    Our upstream auth provider is having an outage. Users unable to login to Braintrust.

Critical

Control Plane Database Connectivity Issues

Started
Mon, Apr 20, 2026, 03:33:53 AM
Updated
Mon, Apr 20, 2026, 03:49:52 AM
Resolved
Mon, Apr 20, 2026, 03:49:52 AM
Duration
15m
  1. resolved

    We're no longer seeing impact from the connectivity issues and we are still investigating the root cause.

  2. monitoring

    We've updated our database connection method and are observing recovery. We're monitoring to look for remaining impact.

  3. identified

    We're currently experiencing an issue with database connectivity. We're currently attempting to mitigate by connecting through an alternative method.

Critical

Control plane database degradation

Started
Mon, Apr 13, 2026, 10:21:24 PM
Updated
Mon, Apr 13, 2026, 10:21:24 PM
Resolved
Mon, Apr 13, 2026, 09:51:00 PM
Duration
  1. resolved

    Performance quickly returned to normal when the problematic query was identified and remediated, and we are implementing safeguards to prevent similar future issues.

  2. investigating

    We experienced a brief period of degraded performance on the Braintrust hosted control plane due to elevated database load. The issue was caused by an internal traffic spike to a specific query pattern.

Major

Elevated control plane errors

Started
Tue, Mar 31, 2026, 05:52:50 PM
Updated
Tue, Mar 31, 2026, 06:05:04 PM
Resolved
Tue, Mar 31, 2026, 06:05:04 PM
Duration
12m
  1. resolved

    Error rates have dropped and functionality is restored.

  2. investigating

    We're seeing a high rate of 5xx errors

Critical

Database load issues with control plane

Started
Fri, Mar 20, 2026, 11:07:40 PM
Updated
Sat, Mar 21, 2026, 12:23:37 AM
Resolved
Sat, Mar 21, 2026, 12:23:37 AM
Duration
1h 15m
  1. resolved

    Load issues have subsided

  2. monitoring

    DB upgrade complete. CPU usage appears to be at normal levels. Monitoring for now.

  3. monitoring

    Database size is being increased. This will cause a short outage.

  4. monitoring

    Error rates have recovered by database usage is still high.

  5. investigating

    The control plane is experiencing very high load.

Major

Authorization service in degraded state

Started
Fri, Mar 20, 2026, 10:01:03 AM
Updated
Fri, Mar 20, 2026, 10:32:23 AM
Resolved
Fri, Mar 20, 2026, 10:32:23 AM
Duration
31m
  1. resolved

    The authorization provider has completed mitigation work and is reporting a full recovery in their systems. We are no longer observing impact to users.

  2. monitoring

    The authorization provider has applied a mitigation and we're seeing signs of recovery. We're continuing to monitor until all errors have subsided.

  3. investigating

    Degraded performance of our authorization provider is causing issues with accessing the control plane.

Critical

High rate of errors on control plane

Started
Sun, Feb 15, 2026, 11:16:48 PM
Updated
Thu, Mar 12, 2026, 06:22:07 PM
Resolved
Mon, Feb 16, 2026, 12:59:23 AM
Duration
1h 42m
  1. resolved

    The root cause appeared to be an issue with an upstream dependency (database connection pooler), which we have replaced. We will be writing a full post-mortem with more details.

  2. resolved

    We are no longer seeing issues

  3. monitoring

    DB upgrades have completed, and we have seen error rates go down. Continuing to monitor.

  4. investigating

    We are upgrade DB resources, which may result in brief downtime

  5. investigating

    We are repairing the DB which may result in brief downtime.

  6. investigating

    We're observing a high rate of errors on the production control plane which is causing errors on braintrust.dev and data plane operations.

Minor

Control Plane UI Hanging

Started
Wed, Feb 18, 2026, 09:32:16 PM
Updated
Thu, Mar 12, 2026, 06:22:07 PM
Resolved
Wed, Feb 18, 2026, 09:35:30 PM
Duration
3m
  1. resolved

    The issue has been resolved and you should be seeing no slowness in the Braintrust Control Plane UI. Please let us know if you have any issues.

  2. investigating

    At the current moment in time we have reports of hanging in the UI particularly in dataset or trace views. We are currently narrowing down the code change causing the issue. Please stand by we are aware and currently working to resolve.

Minor

Authorization service in degraded state

Started
Thu, Feb 19, 2026, 04:34:52 PM
Updated
Thu, Mar 12, 2026, 06:22:07 PM
Resolved
Thu, Feb 19, 2026, 06:35:55 PM
Duration
2h 1m
  1. resolved

    We have not observed any recurrence of issues and our authorization provider has reported full recovery and stable performance across their entire system.

  2. monitoring

    Impact for Braintrust users appears to have been resolved for the past 30 minutes. We'll continue to monitor until our authorization provider reports full resolution.

  3. investigating

    We've observed improvement in authorization errors but our authorization provider is still reporting a degraded state. We're continuing to monitor.

  4. investigating

    Degraded performance of our authorization vendor is causing issues with user authorization accessing the control plane. The vendor has identified the issue and is working on a fix.

Major

S3 Export Automations failing

Started
Tue, Mar 3, 2026, 03:50:59 PM
Updated
Thu, Mar 12, 2026, 06:22:07 PM
Resolved
Tue, Mar 3, 2026, 04:24:51 PM
Duration
33m
  1. resolved

    Affected a small subset of users.

  2. investigating

    S3 automations have stopped running on the expected schedule they were set to.

Major

Authorization service in degraded state

Started
Tue, Mar 10, 2026, 04:18:23 PM
Updated
Thu, Mar 12, 2026, 06:22:07 PM
Resolved
Tue, Mar 10, 2026, 04:38:10 PM
Duration
19m
  1. resolved

    The authorization provider has completed mitigation work and is reporting a full recovery in their systems. We are no longer observing impact to users.

  2. monitoring

    The authorization provider has applied a mitigation and we're seeing signs of recovery. We're continuing to monitor until all errors have subsided.

  3. investigating

    Degraded performance of our authorization provider is causing issues with accessing the control plane. The service provider is investigating the root issue.

Critical

Control Plane Errors

Started
Tue, Mar 10, 2026, 07:25:01 PM
Updated
Thu, Mar 12, 2026, 06:22:07 PM
Resolved
Tue, Mar 10, 2026, 07:11:00 PM
Duration
  1. resolved

    Provider recovered

  2. investigating

    An upstream provider had a brief outage that caused high errors in our control plane

Watch Braintrust
Email alerts on every status change — outages, degradations, new incidents, and resolutions.

Watching all 1 providers. Customize on the alerts page.

Details

Aliases
brain trust
Indicator
none
Path
/braintrust

Related in Platforms & Infra