Platforms & Infra

status.fly.io

Fly.io

Edge compute for AI apps

All Systems Operational

Operational
Latency
712ms
Checked
just now
Active incidents
0
Components
58
Source
Statuspage.io API

Overview

Observed uptime · 1 day100%
2026-09-202026-09-20
Component health
58 up0 degraded0 down

Components58

58 operational0 degraded0 outage58 total
  • AMS - Amsterdam, Netherlands
    Operational
  • Fly Machine .internal DNSDNS service discovery for application VMs and 6PN.
    Operational
  • Usage Metrics API
    Operational
  • Upstash for Redis
    Operational
  • Customer Applications
    Operational
  • Management Plane - ORD
    Operational
  • Stripe API Connection
    Operational
  • Dashboard
    Operational
  • ARN - Stockholm, Sweden
    Operational
  • Fly Machine External DNS
    Operational
  • Management Plane - IAD
    Operational
  • Machines API
    Operational
  • Management Plane - FRA
    Operational
  • Regional Availability
    Operational
  • *.flyio.net Nameservers
    Operational
  • Management Plane - GRU
    Operational
  • Persistent Storage (Volumes)Fly volumes and apps that use them
    Operational
  • BOM - Mumbai, India
    Operational
  • Management Plane - LAX
    Operational
  • flydns.netDNS servers for fly.dev and other Fly.io domains
    Operational
  • Deploymentsflyctl and GitHub Action app deployments
    Operational
  • CDG - Paris, France
    Operational
  • Management Plane - SYD
    Operational
  • Remote Builds
    Operational
  • Management Plane - AMS
    Operational
  • DFW - Dallas, Texas (US)
    Operational
  • Logs
    Operational
  • Management Plane - LHR
    Operational
  • MetricsCustomer metrics
    Operational
  • EWR - Secaucus, NJ (US)
    Operational
  • Management Plane - NRT
    Operational
  • SSL/TLS Certificate Provisioning
    Operational
  • Management Plane - SIN
    Operational
  • FRA - Frankfurt, Germany
    Operational
  • Anycast EdgeFly App HTTP, TCP, UDP services through our global load balancer.
    Operational
  • Management Plane - SJC
    Operational
  • Fly Machine Image Registry 1
    Operational
  • Management Plane - YYZ
    Operational
  • Fly Machine Image Registry 2
    Operational
  • GRU - Sao Paulo, Brazil
    Operational
  • Extensions
    Operational
  • DNS
    Operational
  • IAD - Ashburn, Virginia (US)
    Operational
  • Billing
    Operational
  • JNB - Johannesburg, South Africa
    Operational
  • CorrosionDisseminates state of creation and changes of resources on our platform (e.g apps, Machines, IP addresses)
    Operational
  • LAX - Los Angeles, California (US)
    Operational
  • LHR - London, United Kingdom
    Operational
  • Managed Postgres
    Operational
  • Phoenix.newPhoenix.new IDEs and preview sites
    Operational
  • Support Portal
    Operational
  • Sprites
    Operational
  • NRT - Tokyo, Japan
    Operational
  • ORD - Chicago, Illinois (US)
    Operational
  • SIN - Singapore
    Operational
  • SJC - San Jose, California (US)
    Operational
  • SYD - Sydney, Australia
    Operational
  • YYZ - Toronto, Canada
    Operational

Incidents50

History 50

Minor

Depot builder failures

Started
Tue, Sep 15, 2026, 04:08:58 AM
Updated
Tue, Sep 15, 2026, 05:34:09 AM
Resolved
Tue, Sep 15, 2026, 05:34:09 AM
Duration
1h 25m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring builds. Standard flyctl builds should be working again for customers in all regions.

  3. identified

    We have identified an issue with deploying via Depot for users connecting through our SYD and JNB regions. Affected customers in these regions can deploy successfully using the --depot=false or --buildkit arguments to flyctl.

  4. investigating

    We are investigating reports of Depot builds failing for some customers. Affected customers can deploy successfully using the --depot=false or --buildkit arguments to flyctl.

Minor

Network issues in US West Coast

Started
Sat, Sep 12, 2026, 09:22:52 PM
Updated
Sat, Sep 12, 2026, 10:12:29 PM
Resolved
Sat, Sep 12, 2026, 10:12:29 PM
Duration
49m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Private networking between Fly Machines is resolved, and most outbound connections are healthy. We're continuing to monitor the network, and some issues will still be expected from clients physically located in US West until upstream transit issues are resolved.

  3. investigating

    We are investigating upstream network issues from US West Coast (SJC, LAX). Apps hosted in US West regions may experience higher latency or packet loss, and requests from clients physically located in US West may experience higher latency.

Minor

Sprites API Partial Outage

Started
Wed, Sep 2, 2026, 11:03:57 PM
Updated
Thu, Sep 3, 2026, 12:41:41 AM
Resolved
Thu, Sep 3, 2026, 12:41:41 AM
Duration
1h 37m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Error rates have decreased. We are continuing to monitor the API health.

  3. investigating

    We're aware of a problem affecting a subset of Sprites users. We are investigating the source of the issue.

Minor

Upstream network issues

Started
Wed, Sep 2, 2026, 02:18:44 PM
Updated
Wed, Sep 2, 2026, 03:57:45 PM
Resolved
Wed, Sep 2, 2026, 03:57:45 PM
Duration
1h 39m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    We're updating the affected region list to also include SJC since this seems to be a wider upstream issue in US West Coast.

  4. identified

    We have put in some temporary mitigations along with our providers. However since the root cause of this issue lies within a bigger upstream transit provider, you may continue to see some elevated latency and connection issues in/around affected regions. We're still working closely with them to resolve the root cause.

  5. identified

    We have observed an upstream network issue in LAX. Connections to some destinations may see elevated latency and packet loss. We're working with our upstream to resolve this issue.

Major

API background job queue failure

Started
Wed, Sep 2, 2026, 07:21:44 AM
Updated
Wed, Sep 2, 2026, 07:44:39 AM
Resolved
Wed, Sep 2, 2026, 07:44:39 AM
Duration
22m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are investigating an issue with the background job runner for our API. Actions that require a background job, such as creating apps, assigning IP addresses, or creating/renewing certificates, may fail at this time.

None

HTTP/2 traffic disruptions

Started
Mon, Aug 31, 2026, 03:30:00 PM
Updated
Mon, Aug 31, 2026, 03:42:23 PM
Resolved
Mon, Aug 31, 2026, 03:30:00 PM
Duration
0m
  1. resolved

    A configuration update caused temporary failures for incoming HTTP/2 traffic for Fly Machines located on a subset of hosts for a few minutes. This incident has since been resolved. Managed Postgres depends on HTTP/2 and some control plane ops may have been affected as well. However, downstream Postgres connections were unlikely to have been affected by this since they do not use the HTTP2 handler.

Minor

Packet loss in ORD

Started
Mon, Aug 31, 2026, 06:26:25 AM
Updated
Mon, Aug 31, 2026, 08:10:37 AM
Resolved
Mon, Aug 31, 2026, 08:10:37 AM
Duration
1h 44m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Packet loss in ORD is improving and impacted services are recovering; we’re continuing to monitor for intermittent issues

  3. investigating

    Due to an upstream provider, we are seeing ~50% packet loss on a subset of hosts in ORD. Some MPG clusters in ORD are slow to replicate as a result.

None

Sprite deletion jobs failing

Started
Sun, Aug 30, 2026, 10:45:04 PM
Updated
Sun, Aug 30, 2026, 10:45:04 PM
Resolved
Sun, Aug 30, 2026, 10:45:04 PM
Duration
0m
  1. resolved

    We saw Sprite deletion jobs failing between 21:18 and 22:05 UTC. This issue has been resolved.

Minor

Networking Issues in GRU

Started
Fri, Aug 28, 2026, 09:56:35 PM
Updated
Fri, Aug 28, 2026, 10:36:45 PM
Resolved
Fri, Aug 28, 2026, 10:36:45 PM
Duration
40m
  1. resolved

    This incident has been resolved.

  2. identified

    Our upstream provider has implemented a fix. Network performance in GRU has normalized.

  3. identified

    We are seeing a recurrance in networking issues in GRU. Some apps in the region may experience increased latency or packet loss. We are working with our upstream networking provider to resolve.

  4. monitoring

    Networking performance in GRU has normalized and we are no longer seeing issues. We are continuing to monitor to ensure a full recovery.

  5. investigating

    We are investigating networking issues impacting some hosts in GRU (São Paulo, Brazil) region. Some apps in GRU may experience increased latency or packet loss.

Minor

Increased packet loss

Started
Fri, Aug 28, 2026, 08:09:03 AM
Updated
Fri, Aug 28, 2026, 10:22:35 AM
Resolved
Fri, Aug 28, 2026, 10:22:35 AM
Duration
2h 13m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are currently investigating this issue.

Minor

WireGuard gateway issues

Started
Wed, Aug 26, 2026, 06:14:23 PM
Updated
Wed, Aug 26, 2026, 06:43:07 PM
Resolved
Wed, Aug 26, 2026, 06:43:07 PM
Duration
28m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Our testing and monitoring indicates gateways should be back to normal; if you are still having problem using `flyctl ssh console`, try restarting the `flyctl` agent by `flyctl agent restart`.

  3. monitoring

    A fix has been implemented and we are monitoring the results.

  4. investigating

    We are investigating issues with our WireGuard gateways. Some CLI commands like `flyctl ssh console` or `flyctl proxy` may not work at this time. Apps continue to run.

Minor

Metrics in some regions are lagging behind

Started
Mon, Aug 24, 2026, 10:23:31 AM
Updated
Mon, Aug 24, 2026, 01:24:48 PM
Resolved
Mon, Aug 24, 2026, 01:24:48 PM
Duration
3h 1m
  1. resolved

    This is now resolved

  2. monitoring

    All hosts have caught up with metrics and we're monitoring the situation

  3. investigating

    We are currently experiencing some metrics lag on servers in some regions. We are provisioning more metric processing instances to accommodate the backlog and catch up.

Major

Network Issues in LAX Region

Started
Sun, Aug 23, 2026, 01:28:34 AM
Updated
Sun, Aug 23, 2026, 02:10:13 AM
Resolved
Sun, Aug 23, 2026, 02:10:13 AM
Duration
41m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Upstream networking issues have resolved.

  3. investigating

    We are investigating network issues in the Los Angeles region. Apps may experience higher latency or be unreachable at this time.

Minor

Temporary DNS resolution failure

Started
Thu, Aug 20, 2026, 07:30:00 PM
Updated
Thu, Aug 20, 2026, 08:08:00 PM
Resolved
Thu, Aug 20, 2026, 07:30:00 PM
Duration
0m
  1. resolved

    A BGP configuration error caused our Anycast DNS to route to some nodes without the proper DNS infrastructure. The issue was temporary and was resolved as soon as we removed that node from BGP.

None

Oauth/Macaroon Errors from flyctl

Started
Thu, Aug 20, 2026, 01:54:08 PM
Updated
Thu, Aug 20, 2026, 02:16:43 PM
Resolved
Thu, Aug 20, 2026, 02:16:43 PM
Duration
22m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been deployed and this error should no longer be occurring. We're monitoring to ensure full recovery.

  3. identified

    We have identified an issue causing authentication errors for some operations from `flyctl`. These operations are failing with an error like: `This endpoint no longer accepts legacy OAuth tokens (starting with `fo1_`). Please use a macaroon token (starting with `fm2_`) instead. We have identified the issue and are rolling out a fix

Major

MPG (v1) partially down in ORD

Started
Thu, Aug 20, 2026, 07:25:38 AM
Updated
Thu, Aug 20, 2026, 07:57:54 AM
Resolved
Thu, Aug 20, 2026, 07:57:54 AM
Duration
32m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are continuing to investigate this issue.

  4. investigating

    We had an issue with the ord-0 Fly Kubernetes cluster, and many MPG clusters are failing to restart. Our MPG team is actively working on it.

None

6PN Networking issue in YYZ

Started
Tue, Aug 18, 2026, 08:00:00 PM
Updated
Wed, Aug 19, 2026, 01:17:30 AM
Resolved
Wed, Aug 19, 2026, 12:00:00 AM
Duration
4h
  1. resolved

    6PN networking issues between some machines in YYZ during a rollout which was rolled back once we noticed errors. During this time some machines were unable to talk to internal resources like other DBs, other apps or MPG clusters.

None

No capacity in ARN

Started
Mon, Aug 17, 2026, 01:19:35 PM
Updated
Mon, Aug 17, 2026, 03:41:29 PM
Resolved
Mon, Aug 17, 2026, 03:41:29 PM
Duration
2h 21m
  1. resolved

    The capacity issue in the ARN region has been resolved.

  2. investigating

    New machines may fail to create in ARN because we lack capacity.

Major

Secrets service outage

Started
Fri, Aug 14, 2026, 07:30:44 PM
Updated
Fri, Aug 14, 2026, 08:33:45 PM
Resolved
Fri, Aug 14, 2026, 08:33:45 PM
Duration
1h 3m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We have failed over the secrets database to a replica, and the Machines API appears healthy now. We are monitoring for any further issues.

  3. identified

    We are working to recover our secrets service after a failed deployment. Apps continue to run, but it is not possible to create new apps or update secrets at this time.

Minor

IPv6 Networking Issues

Started
Thu, Aug 13, 2026, 05:45:09 PM
Updated
Thu, Aug 13, 2026, 07:41:10 PM
Resolved
Thu, Aug 13, 2026, 07:41:10 PM
Duration
1h 56m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We are continuing to monitor for any further issues.

  3. monitoring

    A fix has been implemented and we are monitoring the results.

  4. identified

    The issue has been identified and a fix is being implemented.

  5. investigating

    We are continuing to investigate this issue.

  6. investigating

    We are currently investigating degraded ipv6 networking on a subset of hosts

Major

Increased app-not-found errors

Started
Sun, Aug 9, 2026, 02:50:05 AM
Updated
Sun, Aug 9, 2026, 07:10:08 AM
Resolved
Sun, Aug 9, 2026, 07:10:08 AM
Duration
4h 20m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results

  3. identified

    We’ve deployed an additional mitigation to further reduce Corrosion retry pressure and are seeing improvement; we’re continuing to monitor while remaining affected nodes catch up.

  4. identified

    We’ve applied a mitigation to reduce the impact from Corrosion batch insertion retries and are continuing to monitor while affected nodes catch up.

  5. identified

    We have identified the issue as failed insertions in a subset of Corrosion batches. These failures trigger retries, which can cause timeouts for other batches.

  6. investigating

    We are currently investigating app-not-found errors returned by Machines API calls made shortly after creating new applications.

Minor

MPG IAD data plane degraded for new clusters

Started
Wed, Aug 5, 2026, 08:26:01 PM
Updated
Wed, Aug 5, 2026, 08:41:05 PM
Resolved
Wed, Aug 5, 2026, 08:41:05 PM
Duration
15m
  1. resolved

    Etcd is stable. Services are back to normal.

  2. identified

    High CPU pressure on a shared etcd instance is causing lags on MPG creation in the IAD region

Major

MPG creation is failing in GRU due to lack of capacity

Started
Tue, Aug 4, 2026, 07:03:41 PM
Updated
Wed, Aug 5, 2026, 03:22:51 AM
Resolved
Wed, Aug 5, 2026, 03:22:51 AM
Duration
8h 19m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We tweaked hosts to allow for more machine allocation. We'll be monitoring the region over the next hours.

  3. identified

    New MPG clusters may fail to create in GRU because we lack capacity.

Minor

Certificate issuance delays

Started
Tue, Aug 4, 2026, 02:53:40 PM
Updated
Tue, Aug 4, 2026, 11:58:38 PM
Resolved
Tue, Aug 4, 2026, 11:58:38 PM
Duration
9h 4m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    We have the size of TLS certificate issuance backlog under control, but are still seeing some remaining issues and are currently working to clean up the edge cases.

  4. identified

    We believe we have identified the issue and are releasing a fix.

  5. investigating

    We're currently investigating issues related to certificate issuance. Certificates may be delayed for new custom domains.

Critical

Inbound connection failure to Fly Apps

Started
Mon, Aug 3, 2026, 03:16:18 PM
Updated
Mon, Aug 3, 2026, 03:17:14 PM
Resolved
Mon, Aug 3, 2026, 06:50:00 PM
Duration
3h 33m
  1. resolved

    A BGP misconfiguration while provisioning new edge capacity caused most traffic from Europe endpoints to be dropped, between 14:50 UTC and 15:02 UTC. The misconfiguration has been fixed and we are implementing safeguards against this kind of issue in the future.

Minor

Managed Postgres v2 control plane issues in iad

Started
Sat, Aug 1, 2026, 05:01:50 PM
Updated
Sat, Aug 1, 2026, 06:30:10 PM
Resolved
Sat, Aug 1, 2026, 06:30:10 PM
Duration
1h 28m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are investigating an issue with the control plane for Managed Postgres v2 in the IAD region. Creating new v2 clusters in the IAD region may fail at this time. Existing clusters continue to run, but may experience lagging backups.

None

Capacity issues in CDG

Started
Fri, Jul 31, 2026, 04:44:41 PM
Updated
Fri, Jul 31, 2026, 08:11:51 PM
Resolved
Fri, Jul 31, 2026, 08:11:51 PM
Duration
3h 27m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    Creating machines in CDG region may fail at this time with an "no capacity available in cdg" message. Existing apps continue to run.

Minor

Outbound email issues

Started
Fri, Jul 31, 2026, 01:34:44 PM
Updated
Fri, Jul 31, 2026, 02:51:42 PM
Resolved
Fri, Jul 31, 2026, 02:51:42 PM
Duration
1h 16m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    Our dashboard is failing to send outbound email. Emails such as new account verification or password reset may fail to send at this time.

Minor

Increased API latency

Started
Wed, Jul 29, 2026, 12:41:43 PM
Updated
Wed, Jul 29, 2026, 01:26:08 PM
Resolved
Wed, Jul 29, 2026, 01:26:08 PM
Duration
44m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are investigating some database issues causing high latency on some API endpoints and dashboard operations. You may experience intermittent "503 service unavailable" errors at this time. Currently deployed apps continue to run.

Critical

High number of 5XX on the Machines API and dashboard

Started
Mon, Jul 20, 2026, 07:10:35 AM
Updated
Mon, Jul 20, 2026, 05:09:52 PM
Resolved
Mon, Jul 20, 2026, 05:09:52 PM
Duration
9h 59m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We are still working on fixing degraded Managed Postgres clusters.

  3. monitoring

    Some Managed Postgres v1 clusters are degraded. We are working on fixing them. Managed Postgres v2 is unaffected.

  4. monitoring

    A fix has been implemented and we are monitoring the results.

  5. identified

    We've identified an internal service providing authentication to our Machines API has failed, our team is currently looking at our options for restoring this service. Existing Machines/Apps will continue to run as normal. Thank you for your patience.

  6. identified

    We've identified an internal service providing authentication to our Machines API has failed, our team is currently looking at our options for restoring this service. Existing Machines/Apps will continue to run as normal. Thank you for your patience.

  7. investigating

    We are continuing to investigate this issue.

  8. investigating

    Existing machines are unaffected. We are investigating the issue.

Minor

Egress IPv6 issues in BOM

Started
Sun, Jul 19, 2026, 03:48:45 AM
Updated
Sun, Jul 19, 2026, 04:52:50 AM
Resolved
Sun, Jul 19, 2026, 04:52:50 AM
Duration
1h 4m
  1. resolved

    This incident has been resolved.

  2. identified

    We have identified an upstream issue that is preventing egress IPv6 addresses in BOM from reaching parts of the internet, and we're currently working with an upstream provider to resolve this issue. Normal IPv6 addresses remain unaffected.

None

Brief Flycast / MPG interruption in YYZ

Started
Fri, Jul 17, 2026, 06:30:00 PM
Updated
Fri, Jul 17, 2026, 07:13:35 PM
Resolved
Fri, Jul 17, 2026, 06:30:00 PM
Duration
0m
  1. resolved

    A bad deployment momentarily caused issues with Flycast connectivity, and, by extension, MPG, in our YYZ region. The deployment was immediately rolled back and connectivity was restored shortly after.

Minor

Edge proxy issues

Started
Thu, Jul 16, 2026, 11:34:24 AM
Updated
Thu, Jul 16, 2026, 01:44:28 PM
Resolved
Thu, Jul 16, 2026, 01:44:28 PM
Duration
2h 10m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are investigating increased connection latency and "connection reset" errors from our edge proxy. Apps continue to run, but requests may experience increased connection latency or fail at this time.

Minor

App creation failing

Started
Thu, Jul 16, 2026, 11:54:59 AM
Updated
Thu, Jul 16, 2026, 12:17:03 PM
Resolved
Thu, Jul 16, 2026, 12:17:03 PM
Duration
22m
  1. resolved

    This incident has been resolved.

  2. identified

    An issue with our Machines API is causing app creations to fail in some cases. We are working on a fix.

Major

Partial outage in SJC

Started
Tue, Jul 14, 2026, 02:51:39 PM
Updated
Tue, Jul 14, 2026, 04:59:18 PM
Resolved
Tue, Jul 14, 2026, 04:59:18 PM
Duration
2h 7m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results. Apps should be reachable at this point in time.

  3. identified

    A subset of hosts in SJC are currently offline. Some apps may be unreachable at this time.

Minor

Some DFW hosts offline

Started
Mon, Jul 13, 2026, 10:25:41 PM
Updated
Mon, Jul 13, 2026, 10:58:51 PM
Resolved
Mon, Jul 13, 2026, 10:58:51 PM
Duration
33m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    A subset of hosts in DFW are currently offline, and we're investigating the issue.

Minor

Delays starting Depot Builders in IAD

Started
Thu, Jul 9, 2026, 03:21:15 PM
Updated
Thu, Jul 9, 2026, 04:18:03 PM
Resolved
Thu, Jul 9, 2026, 04:18:03 PM
Duration
56m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are seeing improvements in builder performance, latency, and error rates. We are continuing to monitor for a full recovery. Customers still seeing issues can trigger a fly-hosted builder based deploy with `fly deploy --depot=false`.

  3. identified

    We have identified an issue causing delays or failures when starting depoting builders located in the IAD region. Customers with builders in IAD may see delays or timeouts starting builds during `fly deploy`. We are working on a fix. In the meantime customers can trigger a fly-hosted builder based deploy with `fly deploy --depot=false`. You can also change your builder region away from IAD via the `settings` tab of your Fly dashboard. We recommend DFW or ORD as alternate builder regions at this time. Please note, your builder region can differ from the region your machines are in.

Minor

Registry performance issues

Started
Thu, Jul 9, 2026, 03:33:31 PM
Updated
Thu, Jul 9, 2026, 04:04:49 PM
Resolved
Thu, Jul 9, 2026, 04:04:49 PM
Duration
31m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    The fly.io registry is currently experiencing capacity constraints that reduced performance and may lead to temporary high latency or failed pushes. We are currently working to add capacity and restore service.

Major

Partial Outage in ORD

Started
Fri, Jul 3, 2026, 05:00:06 PM
Updated
Fri, Jul 3, 2026, 06:11:33 PM
Resolved
Fri, Jul 3, 2026, 06:11:33 PM
Duration
1h 11m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    We've identified the issue as a networking hardware failure impacting a subset of hosts at one of our Upstream providers in ORD. We are working with our provider to restore connectivity.

  4. investigating

    We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly. Deploys with machines on these hosts may fail at this time. Some Managed Postgres clusters in ORD region may be unavailable or see connectivity issues at this time.

Major

Partial outage in ORD

Started
Fri, Jul 3, 2026, 12:11:11 AM
Updated
Fri, Jul 3, 2026, 05:59:18 AM
Resolved
Fri, Jul 3, 2026, 05:59:18 AM
Duration
5h 48m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Customer workloads are now starting and we're monitoring the affected hosts. Affected Managed Postgres instances will be investigated.

  3. identified

    Power restoration is ongoing and we're making sure the hosts are healthy before starting customer workloads to avoid issues. Customer impact remains and updates to come.

  4. identified

    Power restoration work in a subset of ORD is still in progress and impact remains ongoing for a subset of hosts and some Managed Postgres clusters.

  5. identified

    Our provider has advised us their facilities team is working on restoring the power, we'll provide another update as soon as we learn more.

  6. identified

    We've identified and reported power issues with one of our upstream providers in ORD. We're waiting for an update from our upstream for a resolution. Some Managed Postgres clusters in ORD will be unavailable due to placement.

  7. investigating

    We are investigating an issue with one of our upstream providers in ORD. Machines across a subset of hosts may be unreachable or not running correctly. Deploys with machines on these hosts may fail at this time. Some Managed Postgres clusters in ORD region may be unavailable at this time.

Critical

Errors issuing new SSL certificates

Started
Thu, Jul 2, 2026, 09:38:33 PM
Updated
Thu, Jul 2, 2026, 11:18:06 PM
Resolved
Thu, Jul 2, 2026, 11:18:06 PM
Duration
1h 39m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented upstream and certificates are being issued successfully. We will continue to monitor.

  3. identified

    The issue has been identified and we are awaiting a fix.

  4. investigating

    We are currently investigating errors when issuing new SSL certificates for hostnames.

Minor

Static Egress IPv6 issues in NRT

Started
Wed, Jul 1, 2026, 01:06:58 PM
Updated
Wed, Jul 1, 2026, 01:57:17 PM
Resolved
Wed, Jul 1, 2026, 01:57:17 PM
Duration
50m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    We are investigating issues with static egress IPv6 addresses in NRT region. Apps using static egress IPs may experience connectivity failures to some destinations.

Major

Elevated API Errors

Started
Wed, Jul 1, 2026, 06:14:49 AM
Updated
Wed, Jul 1, 2026, 07:50:42 AM
Resolved
Wed, Jul 1, 2026, 07:50:42 AM
Duration
1h 35m
  1. resolved

    This incident has been resolved.

  2. monitoring

    Background jobs have caught up and the API is fully operational. We are continuing to monitor service health.

  3. identified

    A fix has been put in place, and we are no longer seeing elevated API errors. Some dashboard actions will be delayed while background processing catches up.

  4. identified

    The cause of the errors has been identified and we are working on a fix.

  5. investigating

    We are investigating elevated errors with our GraphQL API and background job processing

Major

Delayed Metrics

Started
Mon, Jun 29, 2026, 01:03:08 AM
Updated
Tue, Jun 30, 2026, 10:19:56 PM
Resolved
Tue, Jun 30, 2026, 10:19:56 PM
Duration
1d 21h
  1. resolved

    This incident has been resolved.

  2. identified

    Almost all metrics have caught up aside from a small handful in `sin` and `syd`. We're continuing to monitor and expect these to complete in the next few hours.

  3. identified

    Backlogged metrics are still being processed. We're bringing extra processing capacity online to speed up the process.

  4. identified

    Backlogged metrics are still being processed and ingestion delays persist for some customers

  5. identified

    Backlogged metrics are still being processed and ingestion delays persist for some customers, but we’re continuing to see gradual improvement as the backlog continues to drain.

  6. identified

    Backlogged metrics data is still being processed, continue to monitor the situation while the system catches up.

  7. identified

    Metric ingestion is still delayed for some customers. We’re seeing gradual improvement and continue to monitor the situation while the system catches up.

  8. identified

    We are still in the process of increasing the metrics cluster throughput to catch up with metric backlog.

  9. identified

    We are continuing to see delayed metrics due to resource contention on a subset of metrics ingestion hosts, and we are working to rebalance ingestion traffic and reduce the backlog.

  10. identified

    We are continuing to see delayed metric exports from a number of hosts. Users will see delayed or missing metrics for machines on impacted hosts at this time.

  11. identified

    We've identified an issue on multiple hosts causing delayed metric ingestion into our hosted fly-metrics.net dashboards. We are working on a fix.

  12. investigating

    We are investigating issues with customer facing metrics in the fly-metrics.net dashboard. Users may see delayed or missing metrics at this time.

Minor

Egress IP issues in SIN and NRT

Started
Tue, Jun 30, 2026, 01:45:42 PM
Updated
Tue, Jun 30, 2026, 02:03:50 PM
Resolved
Tue, Jun 30, 2026, 02:03:50 PM
Duration
18m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    We are aware of egress IP issues in SIN and NRT and are working on a fix. Some machines in SIN and NRT using egress IPs may temporarily lose connectivity or otherwise see degraded performance.

Major

Metrics currently experiencing issues

Started
Sun, Jun 28, 2026, 08:10:51 AM
Updated
Sun, Jun 28, 2026, 09:20:51 PM
Resolved
Sun, Jun 28, 2026, 09:20:51 PM
Duration
13h 10m
  1. resolved

    This incident has been resolved.

  2. investigating

    We are currently investigating an issue with our metrics cluster.

Major

IPv6 Connectivity Issues in EWR

Started
Fri, Jun 26, 2026, 10:17:11 PM
Updated
Fri, Jun 26, 2026, 11:48:08 PM
Resolved
Fri, Jun 26, 2026, 11:48:08 PM
Duration
1h 30m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. investigating

    One of our upstream providers is experiencing IPv6 network connectivity problems in EWR. Apps with machines on affected hosts may have impacted connectivity to certain IPv6 destinations while they investigate and resolve this issue.

Minor

Deploys defaulting to Fly-hosted Builders

Started
Thu, Jun 25, 2026, 01:41:59 PM
Updated
Thu, Jun 25, 2026, 03:38:02 PM
Resolved
Thu, Jun 25, 2026, 03:38:02 PM
Duration
1h 56m
  1. resolved

    This incident has been resolved.

  2. monitoring

    We are seeing improvements in Depot builder provision times and are switching the default deploy strategy back to them. We will continue to monitor builder performance closely. Users with a preference can trigger a Fly builder based deploy with `fly deploy --depot=false` or use `fly deploy --depot=true` to force a depot-based deployment.

  3. investigating

    We are investigating delays provisioning Depot backed builders for deploys. We have switched the default `fly deploy` strategy to use fly hosted builders at this time. Users can still trigger a depot based deploy with `fly deploy --depot=true`

Minor

Elevated control plane latency

Started
Thu, Jun 25, 2026, 12:11:01 PM
Updated
Thu, Jun 25, 2026, 03:18:10 PM
Resolved
Thu, Jun 25, 2026, 03:18:10 PM
Duration
3h 7m
  1. resolved

    This incident has been resolved.

  2. monitoring

    A fix has been implemented and we are monitoring the results.

  3. identified

    The issue has been identified and a fix is being implemented.

  4. investigating

    We're addressing elevated control plane latency and saturation affecting the BOM and NRT regions. Apps with machines in this region might experience longer response times and possible timeouts (502 errors).

Minor

Degraded networking in North America

Started
Wed, Jun 24, 2026, 01:42:04 AM
Updated
Wed, Jun 24, 2026, 06:57:37 AM
Resolved
Wed, Jun 24, 2026, 06:57:37 AM
Duration
5h 15m
  1. resolved

    This incident has been resolved.

  2. identified

    Some 6PN Private Networking traffic remains impacted into and out of our LAX region, pending upstream resolution.

  3. identified

    Most networking is largely healthy between primary North American regions. Some Machines may see ongoing packet loss and higher latency communicating with other Machines on certain routes. We're continuing to monitor the backbone health upstream.

  4. investigating

    We are currently investigating degraded network performance between sites in NA due to an upstream incident

Watch Fly.io
Email alerts on every status change — outages, degradations, new incidents, and resolutions.

Watching all 1 providers. Customize on the alerts page.

Details

Aliases
flyio, fly machines
Indicator
none
Path
/fly

Related in Platforms & Infra