Compute services are partially unavailable for customers
Started
Thu, Sep 17, 2026, 04:26:17 PM
Updated
Thu, Sep 17, 2026, 05:09:26 PM
Resolved
Thu, Sep 17, 2026, 05:09:26 PM
Duration
43m
resolved
The root cause was found and this incident has been resolved.
investigating
We are currently experiencing a partial unavailability of our Compute services, which may affect some of your workloads. Our team is actively investigating the issue and working towards a resolution. We will provide further updates as soon as more information becomes available.
None
Resuming IAM accounts is unavailable
Started
Thu, Sep 17, 2026, 10:46:52 AM
Updated
Thu, Sep 17, 2026, 11:45:10 AM
Resolved
Thu, Sep 17, 2026, 11:45:10 AM
Duration
58m
resolved
This incident has been resolved.
monitoring
Root cause is located and fixed. We keep monitoring the resolution.
investigating
We are investigating an issue with IAM Tenants Status change.
Customers may not be able to resume their accounts from the suspended state.
Major
[eu-west1] Metrics for resources in this region are partially unavailable
Started
Wed, Sep 16, 2026, 04:52:49 PM
Updated
Wed, Sep 16, 2026, 06:21:30 PM
Resolved
Wed, Sep 16, 2026, 06:21:30 PM
Duration
1h 28m
resolved
This incident has been resolved.
investigating
We are currently investigating this issue.
Major
[us-central1] Object Storage is partially unavailable
Started
Wed, Sep 16, 2026, 10:45:23 AM
Updated
Wed, Sep 16, 2026, 03:42:08 PM
Resolved
Wed, Sep 16, 2026, 03:42:08 PM
Duration
4h 56m
resolved
This incident has been resolved.
identified
The team has identified the root cause of the issue and is currently applying a fix. We expect the incident to be fully mitigated within two hours.
investigating
Users might be experiencing a high rate of 5xx errors on upload requests and elevated latency on read requests.
Major
Managed Service for SkyPilot: cluster creation and update requests failing in eu-north1
Started
Mon, Sep 14, 2026, 03:00:00 PM
Updated
Tue, Sep 15, 2026, 03:33:28 PM
Resolved
Tue, Sep 15, 2026, 03:33:28 PM
Duration
1d
resolved
This incident has been resolved.
monitoring
The error rate has dropped to zero. Requests to create, update, start, and stop SkyPilot clusters in eu-north1 are completing successfully again. We are continuing to investigate the root cause and are monitoring the service closely. Instances that were already running have not been affected at any point. We will provide a further update as the investigation progresses.
investigating
We are investigating an issue affecting Managed Service for SkyPilot in eu-north1. Requests to create or update SkyPilot clusters, as well as start and stop requests, are failing with HTTP 503 errors. Instances that are already running are not affected and continue to operate normally.
Minor
Unstable compute VM create
Started
Wed, Sep 9, 2026, 12:17:14 PM
Updated
Wed, Sep 9, 2026, 12:41:15 PM
Resolved
Wed, Sep 9, 2026, 12:41:15 PM
Duration
24m
resolved
This incident has been resolved.
monitoring
The fix has been deployed and the team is looking into the instrumentation for any signs of any remaining abnormal behaviour.
investigating
In some cases VMs may fail to create or metadata inside may not be available. The team has identified the issue and is working on deploying the fix.
Major
Connectivity issues for newly created worker nodes
Started
Tue, Sep 1, 2026, 04:11:31 PM
Updated
Tue, Sep 1, 2026, 04:26:32 PM
Resolved
Tue, Sep 1, 2026, 04:26:32 PM
Duration
15m
resolved
This incident has been resolved.
monitoring
We are continuing to monitor for any further issues.
monitoring
The issue may have caused consequential impact to VPC functionality of VMs that are not Managed Service for Kubernetes worker nodes.
identified
Worker nodes of Managed Service for Kubernetes clusters created after 15:13 CET Sep 1 2026 could experience problems connecting to/from other nodes in the same cluster. The root cause is understood and mitigation is underway.
Most critical issues have been resolved and the region is operational, though a subset of nodes remains unavailable.
identified
The Managed Service for Kubernetes is nearly fully restored, though a subset of nodes remains unavailable.
identified
Most nodes are operating normally. The Managed Service for Kubernetes® is still degraded, but the situation is improving.
identified
Most nodes are operating normally. The Managed Service for Kubernetes® is experiencing a partial outage, and the team is actively working to restore it.
identified
Services are still gradually continuing to recover. Many of the virtual machines are already running, but not all of them.
identified
Services are still gradually continuing to recover. Some virtual machines are already running, but not all of them.
identified
Services are gradually continuing to recover.
identified
The root cause was fixed, and the situation is stabilizing. The services are recovering.
identified
We are still experiencing temperature issues in the us-central1. Team is working on it
identified
The root cause has been identified as overheating hardware. GPU performance is impacted, and some nodes may be unavailable.
investigating
We are currently investigating this issue.
Minor
Issues logging to the Management Console
Started
Mon, Aug 17, 2026, 11:00:00 AM
Updated
Tue, Aug 18, 2026, 04:01:51 PM
Resolved
Mon, Aug 17, 2026, 01:00:00 PM
Duration
2h
resolved
Some newly registered users may have experienced issues logging in to the management console between 13:00 UTC on 17 August 2026 and 14:00 UTC on 18 August 2026.
Major
Partial InfiniBand unavailability
Started
Sat, Aug 15, 2026, 06:25:42 PM
Updated
Sat, Aug 15, 2026, 10:27:30 PM
Resolved
Sat, Aug 15, 2026, 09:43:06 PM
Duration
3h 17m
resolved
This incident has been resolved.
investigating
We are currently investigating this issue.
Major
External VPC connectivity issues
Started
Wed, Aug 12, 2026, 01:30:52 PM
Updated
Wed, Aug 12, 2026, 03:23:49 PM
Resolved
Wed, Aug 12, 2026, 01:56:33 PM
Duration
25m
resolved
This incident has been resolved.
monitoring
According to our metrics, the situation with VPC has stabilised.
identified
We are investigating an issue with external VPC connectivity
None
Creation of new VMs is broken
Started
Wed, Aug 5, 2026, 02:27:20 PM
Updated
Wed, Aug 5, 2026, 03:45:34 PM
Resolved
Wed, Aug 5, 2026, 03:45:34 PM
Duration
1h 18m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are continuing to investigate this issue.
investigating
Due to issue with internal services the creation of new VM and other operations may be affected. We are investigating the issue.
Major
Network issues
Started
Mon, Jul 27, 2026, 06:30:09 PM
Updated
Mon, Jul 27, 2026, 08:03:14 PM
Resolved
Mon, Jul 27, 2026, 08:03:14 PM
Duration
1h 33m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
The eu-north1 region experienced network issues that may have affected the operation of several services. The issue has been resolved, and all services are operating normally.
Major
InfiniBand switch failure affecting GPU workloads in the eu-north1-c fabric in the eu-north1 region
Started
Tue, Jul 14, 2026, 10:10:18 PM
Updated
Tue, Jul 14, 2026, 10:32:34 PM
Resolved
Tue, Jul 14, 2026, 10:32:34 PM
Duration
22m
resolved
This incident has been resolved.
investigating
We are currently investigating this issue.
Minor
Managed SkyPilot OIDC Login is unavailable
Started
Fri, Jul 3, 2026, 02:50:15 PM
Updated
Fri, Jul 3, 2026, 02:56:34 PM
Resolved
Fri, Jul 3, 2026, 02:50:15 PM
Duration
0m
resolved
Managed SkyPilot OIDC Login was unavailable since 2 of July.
The problem has been identified and resolved.
Major
Monitoring, Object Storage and Token Factory services are partially unavailable
Started
Wed, Jul 1, 2026, 10:33:53 PM
Updated
Thu, Jul 2, 2026, 12:17:26 AM
Resolved
Thu, Jul 2, 2026, 12:17:26 AM
Duration
1h 43m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring results.
investigating
We are continuing to investigate this issue.
investigating
We are continuing to investigate this issue.
investigating
We are currently experiencing partial unavailability of our Monitoring, Object Storage, Token Factory services. Our team is actively investigating the issue and working towards a resolution. We will provide further updates as soon as more information becomes available.
Minor
New virtual machines are not being created
Started
Tue, Jun 30, 2026, 01:24:15 PM
Updated
Tue, Jun 30, 2026, 01:28:29 PM
Resolved
Tue, Jun 30, 2026, 01:28:29 PM
Duration
4m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
Major
Object Storage service partially unavailable in US-CENTRAL1 region
Started
Wed, Jun 24, 2026, 02:08:04 PM
Updated
Wed, Jun 24, 2026, 03:44:50 PM
Resolved
Wed, Jun 24, 2026, 03:44:50 PM
Duration
1h 36m
resolved
The Object Storage service in the US-CENTRAL1 region has been fully restored and is operating normally. The root cause has been identified, and corrective measures are in place.
monitoring
The Object Storage service in the US-CENTRAL1 region has recovered and is operating normally. We are monitoring to confirm stability. Root cause analysis is ongoing
investigating
We are continuing to investigate this issue.
investigating
We are investigating an outage affecting the Object Storage service in the US-CENTRAL1 region.
Minor
One InfiniBand switch failure
Started
Tue, Jun 16, 2026, 08:02:39 PM
Updated
Tue, Jun 16, 2026, 08:34:45 PM
Resolved
Tue, Jun 16, 2026, 08:34:45 PM
Duration
32m
resolved
This incident has been resolved.
identified
One InfiniBand switch in the fabric-3 cluster is down.
Minor
Network Connectivity Issues
Started
Mon, Jun 15, 2026, 06:27:57 PM
Updated
Mon, Jun 15, 2026, 06:49:11 PM
Resolved
Mon, Jun 15, 2026, 06:49:11 PM
Duration
21m
resolved
Connectivity issues are gone now. Intermittent failure is being investigated in the background with the uplink provider, while the Nebius Netowrking team continues to monitor the link closely.
investigating
We have noticed that there have been problems with the connection latency to external addresses in eu-north1.
None
Compute instances startup issues
Started
Wed, Jun 10, 2026, 08:56:13 PM
Updated
Wed, Jun 10, 2026, 11:29:09 PM
Resolved
Wed, Jun 10, 2026, 11:28:59 PM
Duration
2h 32m
resolved
This incident has been resolved.
investigating
We're currently investigating an issue affecting Compute instances startup. Some requests to create or start instances may fail. Already-running instances are not affected.
Our engineering team is actively investigating.
Minor
Some public IPs unreachable in eu-north1
Started
Tue, Jun 9, 2026, 10:30:00 AM
Updated
Tue, Jun 9, 2026, 01:25:28 PM
Resolved
Tue, Jun 9, 2026, 10:30:00 AM
Duration
0m
resolved
Some of public IPs became unreachable.
Major
Network Connectivity Issues
Started
Thu, Jun 4, 2026, 09:46:26 AM
Updated
Thu, Jun 4, 2026, 11:14:49 AM
Resolved
Thu, Jun 4, 2026, 11:14:49 AM
Duration
1h 28m
resolved
This incident has been resolved.
monitoring
The issue has been resolved, and affected services have recovered.
The situation continues to be monitored to ensure service stability.
investigating
An issue with Internet connectivity in the data center has been detected.
Impact:
- Logs and metrics from virtual machines are not being delivered.
- Object Storage is currently unavailable from internet.
- Container registry is currently unavailable from internet.
- Additional services relying on external connectivity may be affected.
The issue is under investigation.
Minor
Unable to allocate new public IPv4 addresses in eu-north1
Started
Tue, Jun 2, 2026, 10:53:19 AM
Updated
Wed, Jun 3, 2026, 11:06:18 AM
Resolved
Wed, Jun 3, 2026, 11:06:18 AM
Duration
1d
resolved
This incident has been resolved.
monitoring
Public IPv4 address capacity has been restored in eu-north1, and new public IPv4 addresses can be allocated to workloads again. We are monitoring the service to confirm full recovery before resolving this incident.
identified
We have identified the cause: the pool of available public IPv4 addresses in eu-north1 is currently exhausted, which is why new public IPv4 addresses cannot be allocated to workloads.
We are reclaiming and reallocating addresses to restore allocation capacity. Workloads that already have a public IPv4 address assigned remain unaffected. We will provide a further update once new allocations are succeeding.
investigating
We are investigating an issue in eu-north1 affecting Virtual Private Cloud (Networking). New public IPv4 addresses cannot currently be allocated to workloads. Customers may be unable to assign a public IP when creating or modifying resources — for example, attaching a public IP to a virtual machine or load balancer in this region.
Resources that already have a public IPv4 address assigned are not affected and continue to operate normally. We will provide an update as we learn more.
None
Power issues
Started
Fri, May 29, 2026, 09:22:00 PM
Updated
Fri, May 29, 2026, 10:20:19 PM
Resolved
Fri, May 29, 2026, 09:22:00 PM
Duration
0m
resolved
Several servers were rebooted.
Minor
Unable to perform operations on a small subset of virtual machines
Started
Tue, May 19, 2026, 01:45:19 PM
Updated
Tue, May 19, 2026, 03:40:15 PM
Resolved
Tue, May 19, 2026, 03:40:15 PM
Duration
1h 54m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
Minor
InfiniBand performance degradation
Started
Wed, May 13, 2026, 06:45:28 PM
Updated
Wed, May 13, 2026, 09:05:05 PM
Resolved
Wed, May 13, 2026, 09:05:05 PM
Duration
2h 19m
resolved
This incident has been resolved.
investigating
We are investigating reduced InfiniBand fabric performance affecting GPU workloads in the uk-south1 region. Customers may observe lower throughput and slower distributed training jobs.
None
Nebius console is unavailable
Started
Thu, May 7, 2026, 09:03:44 PM
Updated
Thu, May 7, 2026, 09:28:03 PM
Resolved
Thu, May 7, 2026, 09:28:03 PM
Duration
24m
resolved
Issue resolved 21:10.
investigating
We are continuing to investigate this issue.
investigating
Nebius console is unavailable
Major
Managed Kubernetes operations can take longer than expected
Started
Wed, May 6, 2026, 12:39:30 PM
Updated
Thu, May 7, 2026, 07:45:47 AM
Resolved
Thu, May 7, 2026, 07:45:47 AM
Duration
19h 6m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and the fix is being implemented
investigating
Managed Kubernetes operations can take longer than expected, leading to delays or timeouts. Deletion operations are particularly impacted.
The team is currently investigating the issue.
Major
Partial unavailability of external models in Token Factory
Started
Wed, Apr 29, 2026, 03:36:38 PM
Updated
Thu, Apr 30, 2026, 08:05:26 AM
Resolved
Thu, Apr 30, 2026, 08:05:26 AM
Duration
16h 28m
resolved
The external provider has restored operations. All models are now available.
investigating
The following serverless models are unavailable due to external provider downtime.
Qwen/Qwen3-235B-A22B-Thinking-2507-fast
Qwen/Qwen3-Next-80B-A3B-Thinking-fast
Qwen/Qwen3.5-397B-A17B-fast
deepseek-ai/DeepSeek-V3.2-fast
openai/gpt-oss-120b-fast
Minor
Partial degradation of Dedicated Endpoints in Token Factory
Started
Fri, Apr 24, 2026, 09:46:41 AM
Updated
Fri, Apr 24, 2026, 12:22:40 PM
Resolved
Fri, Apr 24, 2026, 10:30:09 AM
Duration
43m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating this issue.
Minor
TokenFactory DataLab unavailability
Started
Mon, Apr 20, 2026, 07:26:31 AM
Updated
Tue, Apr 21, 2026, 09:15:15 AM
Resolved
Tue, Apr 21, 2026, 09:15:15 AM
Duration
1d 1h
resolved
This incident has been resolved.
monitoring
Most of Data Lab functionality is restored. Data filtering functionality is disabled until future updates.
investigating
We are continuing to investigate this issue.
investigating
We are currently investigating this issue.
Minor
Metrics for VMs are unavailable
Started
Mon, Apr 20, 2026, 08:05:12 AM
Updated
Mon, Apr 20, 2026, 08:53:33 AM
Resolved
Mon, Apr 20, 2026, 08:30:13 AM
Duration
25m
resolved
Since 8:30 back to normal.
identified
The issue has been identified and a fix is being implemented.
Minor
TokenFactory — Some Serverless Models Partially Unavailable in us-central1 Region
Started
Wed, Apr 15, 2026, 08:50:16 AM
Updated
Wed, Apr 15, 2026, 10:26:17 AM
Resolved
Wed, Apr 15, 2026, 10:26:17 AM
Duration
1h 36m
resolved
This incident has been resolved.
monitoring
Impact is mitigated and we are monitoring the situation
investigating
We are continuing to investigate this issue.
investigating
Users may experience issues when working with the following serverless models in the us-central1 region:
- Qwen/Qwen3.5-397B-A17B
- MiniMaxAI/MiniMax-M2.5
We are currently investigating the issue and will provide updates as more information becomes available.
Minor
Degraded logs, metrics and audit trails availability
Started
Fri, Apr 10, 2026, 02:00:00 AM
Updated
Fri, Apr 10, 2026, 08:28:49 AM
Resolved
Fri, Apr 10, 2026, 02:00:00 AM
Duration
0m
resolved
During the night we experienced issues with storage service. While writing for data was not affected (just delayed), but reading logs, accessing metrics and audit trails in console was not possible while the incident was in active stage.
Minor
Observability data partially unavailable for Dedicated Endpoints in TokenFactory
Started
Tue, Apr 7, 2026, 11:50:48 AM
Updated
Tue, Apr 7, 2026, 01:05:08 PM
Resolved
Tue, Apr 7, 2026, 01:05:08 PM
Duration
1h 14m
resolved
This incident has been resolved.
monitoring
The problem was fixed, and we are monitoring the results.
investigating
We are noticing issues with some customers viewing metrics for Dedicated Endpoints in Token Factory. We are currently investigating this issue.
Minor
Slower project and tenant creation
Started
Wed, Apr 1, 2026, 01:58:13 PM
Updated
Wed, Apr 1, 2026, 03:11:54 PM
Resolved
Wed, Apr 1, 2026, 02:57:43 PM
Duration
59m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
investigating
We are currently investigating this issue.
Major
Managed k8s is partially unavailable for customers
Started
Thu, Mar 26, 2026, 10:20:56 AM
Updated
Thu, Mar 26, 2026, 10:57:23 AM
Resolved
Thu, Mar 26, 2026, 10:57:23 AM
Duration
36m
resolved
A fix has been implemented and we are monitoring the results
investigating
We are continuing to investigate this issue.
investigating
We are currently experiencing a partial unavailability of our mk8s services, which may affect some of your workloads. Our team is actively investigating the issue and working towards a resolution. We will provide further updates as soon as more information becomes available.
Major
Storage Metrics Ingestion Failure
Started
Tue, Mar 24, 2026, 11:30:00 PM
Updated
Wed, Mar 25, 2026, 06:26:29 PM
Resolved
Tue, Mar 24, 2026, 11:30:00 PM
Duration
0m
resolved
We had a problem with metrics ingestion which resulted in storage metrics not being ingested during the following intervals:
- eu-west: from 14:50 CET to 19:05 CET
- me-west1: from 17:00 CET to 19:05 CET
- eu-west1: from 15:50 CET till 19:05 CET
Our team has resolved this issue, new metrics are being ingested, but, unfortunately, data for that period is lost
Major
Observability data partially unavailable in TokenFactory
Started
Mon, Mar 23, 2026, 08:15:24 PM
Updated
Mon, Mar 23, 2026, 09:42:20 PM
Resolved
Mon, Mar 23, 2026, 09:42:20 PM
Duration
1h 26m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
We are currently investigating this issue.
Major
Power issues during planned maintenance
Started
Tue, Mar 10, 2026, 03:47:12 PM
Updated
Wed, Mar 11, 2026, 12:49:29 AM
Resolved
Wed, Mar 11, 2026, 12:49:29 AM
Duration
9h 2m
resolved
This incident has been resolved and all services and clusters should be operational now.
monitoring
Maintenance has been completed, and we do not expect any further power-related issues. We are now recovering services and closely monitoring system status.
investigating
We are continuing to investigate this issue.
investigating
We are continuing to investigate this issue.
investigating
We are currently experiencing another service disruption due to ongoing on-site maintenance. Our engineers are actively working to stabilize the affected clusters and maintain service throughout the maintenance period.
monitoring
We have identified failed switches and nodes. All major services were recovered, we are now collecting faulty instances and recovering them.
identified
We are continuing to work on a fix for this issue.
identified
We are continuing to work on a fix for this issue.
identified
We are continuing to work on a fix for this issue.
identified
During planned maintenance we experienced unexpected issues with power supply. We have identified the faulty circuits and already bringing nodes that are not part of planned maintenance back to work.
Minor
Some VMs scheduling can take more than 45 minutes
Started
Thu, Mar 5, 2026, 05:49:02 PM
Updated
Thu, Mar 5, 2026, 06:30:57 PM
Resolved
Thu, Mar 5, 2026, 06:30:57 PM
Duration
41m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
Major
Power issues in us-central1
Started
Mon, Mar 2, 2026, 06:45:43 PM
Updated
Mon, Mar 2, 2026, 08:02:54 PM
Resolved
Mon, Mar 2, 2026, 08:02:54 PM
Duration
1h 17m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and a fix is being implemented.
monitoring
Several racks failed.
Major
Power issues in eu-north1
Started
Thu, Feb 26, 2026, 02:40:46 PM
Updated
Fri, Feb 27, 2026, 12:56:18 AM
Resolved
Fri, Feb 27, 2026, 12:56:18 AM
Duration
10h 15m
resolved
This incident has been resolved.
monitoring
The root cause was identified, failed workloads restored.
investigating
We are continuing to investigate this issue.
investigating
We’re experiencing power issues in the eu-north1 region. Some resources (virtual machines, network disks and filestores) may be unavailable.
Major
Network issues in us-central1
Started
Fri, Feb 20, 2026, 10:35:03 PM
Updated
Fri, Feb 20, 2026, 11:30:20 PM
Resolved
Fri, Feb 20, 2026, 11:30:20 PM
Duration
55m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
identified
The issue has been identified and the team is implementing a fix.
investigating
We are currently investigating the issue
Major
Batch inference unavailability in Token Factory
Started
Wed, Feb 18, 2026, 06:26:32 PM
Updated
Thu, Feb 19, 2026, 04:00:38 PM
Resolved
Thu, Feb 19, 2026, 04:00:38 PM
Duration
21h 34m
resolved
This incident has been resolved.
investigating
Batch inference is temporarily unavailable.
Major
Partial unavailability of Managed Kubernetes LoadBalancers
Started
Wed, Feb 18, 2026, 01:26:27 PM
Updated
Wed, Feb 18, 2026, 04:31:47 PM
Resolved
Wed, Feb 18, 2026, 04:31:47 PM
Duration
3h 5m
resolved
This incident has been resolved.
monitoring
A fix has been implemented and we are monitoring the results.
investigating
Some Managed Kubernetes k8s LoadBalancers transitioned to the Pending state and could be temporarily unavailable. We are actively working on mitigation and clarifying the scope of the impact.
Minor
One Infiniband switch failure
Started
Mon, Feb 16, 2026, 11:44:44 PM
Updated
Tue, Feb 17, 2026, 12:02:31 AM
Resolved
Tue, Feb 17, 2026, 12:02:31 AM
Duration
17m
resolved
This incident has been resolved.
monitoring
A faulty IB switch was rebooted and sucesfully running. The team monitors the situation
identified
The issue has been identified and a fix is being implemented.
Minor
Partial loss of Infiniband connectivity
Started
Tue, Feb 10, 2026, 07:00:25 PM
Updated
Tue, Feb 10, 2026, 07:25:32 PM
Resolved
Tue, Feb 10, 2026, 07:25:32 PM
Duration
25m
resolved
The connectivity is fully restored.
monitoring
Links are back up, we are monitoring the equipment and bandwidth saturation.
identified
The switch in question has been cycled.
investigating
One of IB switches has failed. DC ops are working on restoring it.