US region increased latencies
- resolved
This incident has been resolved.
- identified
The issue has been identified and a fix is being implemented.
status.bfl.ai
FLUX image generation API
All Systems Operational
This incident has been resolved.
The issue has been identified and a fix is being implemented.
After the maintenance on our US3 cluster, the regional endpoint api.us3.bfl.ai is no longer valid, however we are still receiving traffic from certain users. Please ensure to only ever use the global api.bfl.ai or regional api.<eu/us>.bfl.ai endpoints as specified in the docs, not cluster specific ones like api.us3.bfl.ai, to avoid using internal domains that may be removed at any time. https://docs.bfl.ai/api_integration/integration_guidelines#api-endpoints-overview
Full capacity restored to api.eu.bfl.ai
EU cluster outage leading to slow or failed generations - traffic has been shifted to other clusters on api.bfl.ai - customers using our European endpoint, api.eu.blf.ai, will still see impact
FLUX.2 [flex] restored - vendor mitigation in place with full recover observed
One of our inferencing clusters is down, which is impacting FLUX.2 [flex] traffic. The rest of our models should be unimpacted.
Observing full recovery.
Fix is in place and we are seeing full recovery
Cluster storage outage identified, vendor implementing fix
We are currently investigating this issue.
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
This is a minor issue isolated to a portion of our US infrastructure. We have isolated the problem and are working to restore api to full serving capacity
We have completed initial scaling and optimizations to enable more traffic through on the flux-3-video api. The BFL team are excited to see what everyone creates!
A fix has been implemented and we are monitoring the results.
We are experiencing significant traffic from our Flux 3 launch and are working to serve all customers. We appreciate your understanding and are working to lower generation times and give everyone the chance to experience our new video models!
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
The issue has been identified and a fix is being implemented.
A fix has been implemented and we are monitoring the results.
US3 has issues due to the 3d party experiencing service degradation.
Large increase in traffic to our services is causing delays in image generation. We are working to balance traffic and bring down the time it takes to receive you images!
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
We are continuing to investigate this issue.
We're experiencing higher latency on our EU2 cluster. We're investigating the root cause.
This incident has been resolved.
The issue has been identified and a fix is being implemented.
This incident has been resolved.
Third party provider restoring systems after power outage.
The issue has been identified and a fix is being implemented.
This incident has been resolved.
We are currently investigating this issue.
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
This incident has been resolved.
The error rate is small, we implemented several fixes to mitigate and work on the major update to resolve the issue.
The issue has been identified and a fix is being implemented.
This incident has been resolved.
The issue has been identified and a fix is being implemented.
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
Image generation is experience internal errors, issue has been identified and fix is underway
This incident has been resolved.
We are experiencing above average load and working to mitigate this.
This incident has been resolved.
We are currently investigating this issue.
This incident has been resolved.
We are currently investigating this issue.
Infrastructure maintenance operation in EU2 region caused temporary disruptions to request processing, resulting in elevated error rates for approximately two hours. The issue resolved automatically and all services are operating normally.
This incident has been resolved.
The issue is now partially affecting traffic on the primary global/regional endpoints. We're focusing all capacity on resolving this.
Noting that this should not affect customers using the primary global or regional endpoints noted in the Docs here: https://docs.bfl.ai/quick_start/generating_images#legacy-regional-endpoints.
We are currently investigating this issue.
This incident has been resolved.
We are currently investigating this issue.
Temporary disruption with a downstream provider in the moderation chain. The issue is resolved as the downstream service recovered.
A critical component in one regional cluster was failing which resulted in increased latency and error rates. A fix has been deployed to prevent this issue in the future.
This incident has been resolved.
We are currently investigating this issue.
This incident has been resolved.
We are currently investigating this issue.
This incident has been resolved.
We've identified the issue and implemented a fix, latency is returning to normal and we'll keep monitoring.
We're investigating an increase in latency in a single EU cluster, EU2, for a few of our models.
This incident has been resolved.
We are investigating an issue with some generations taking a longer amount of time than usual
This incident has been resolved.
Customers experience error 500 on some endpoints. We are looking into it.
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
We are continuing to investigate this issue.
Issue is identified and we are working on a fix. This affects only Flux.2.
This incident has been resolved.
A fix has been implemented and we are monitoring the results.
The issue has been identified and fix is being worked on.
We are currently investigating the issue.
This incident has been resolved.
We are also noted an increase in 403s for certain older endpoints, we've identified the issue and are deploying a fix.
We are continuing to investigate this issue.
Auth issues are resolved, still investigating credits purchase.
We noted that purchasing new credits is also encountering some issues, currently investigating.
We are encountering some authentication issues relating to Playground and Google login.
The Incident has been resolved on the cloud provider side, all systems operational.
We are currently experiencing elevated error rates on BFL Playground, Auth and Dashboard endpoints coming from one of our service providers that has a significant outage. We are investigating our options for mitigation. Thank you for your patience!
Incident has been resolved.
A fix has been implemented and we are monitoring the situation
We have resolved one issue, and are working on resolving another one only applying to some customers.
Investigating the same issue again
A fix has been implemented and we are monitoring now. You should be able to use the different endpoints again.
We are investigating an issue triggering 429 errors in some clusters.
The incident has now been resolved.
We have identified the issue and are working on a fix.
We are continuing to investigate this issue.
We are investigating an issue related to the `us2` cluster returning errors in some cases.
The issue has been resolved.
A fix has been put in place, we are monitoring.
We are continuing to investigate this issue.
We are investigating an issue with elevated latency.
The incident has been resolved
We are currently seeing elevate latency on multiple endpoints.
The incident has been resolved now.
A fix has been implemented and we are monitoring
We are currently investigating an issue on FLUX.1 Kontext about a latency increase.
The incident is now resolved.
We are continuing to investigate this issue.
We are investigating an issue with FLUX 1.1 [pro]
The incident has been resolved
A fix has been implemented and we're monitoring the situation
We are continuing to investigate this issue.
We are currently investigating this issue.
The incident has been resolved
We are currently investigating an latency increase issue on our clusters.
Our third party had resolved their issue be we kept monitoring closely, we trust the stability now.
Apologies, false status update
A fix has been implemented. Please use the global endpoint: https://api.bfl.ai. Regional endpoints are also now operational: `api.eu.bfl.ai`, and `api.us.bfl.ai` As part of our cloud-outage mitigation, we’ve temporarily moved delivery to the following Azure Blob endpoints. Please whitelist them: - *.blob.core.windows.net We intend to switch back to our standard domains (delivery.*.bfl.ai and delivery-*.bfl.ai) as soon as it’s viable. We recommend keeping those domains whitelisted as well, if they aren’t already. Note: there’s a difference between delivery.*.bfl.ai and delivery-*.bfl.ai. Please ensure both patterns are allowed.
A fix has been implemented. Please use the global endpoint: https://api.bfl.ai. Regional endpoints are also now operational: `api.eu.bfl.ai`, and `api.us.bfl.ai`
A fix has been implemented. Please use the global endpoint: https://api.bfl.ai. DNS cache needs to be flushed. Regional endpoints: api.us.bfl.ai and api.eu.bfl.ai are not functional at the moment.
We have implemented a mitigation to the problem. Please use regional endpoints for now: `https://api.eu2.bfl.ai` and `https://api.eu4.bfl.ai/`
Azure is currently experience issues resulting in a full outage in our image generation services. We are awaiting a status update from them.
We identified and resolved the issue.
We have seen an increase in latency on the US1 cluster and are investigating.
Scaling is stable again.
Issue has been brought under control, we're monitoring the fix.
Our US1 cluster is experiencing some scaling issue, resulting in some requests taking longer to process. We're working on a fix.
The third party has declared their incident resolved.
Due to a critical third party dependency we rely on for load balancing experiencing issues, customers within certain regions are experiencing an increase in latency and timeouts.
The critical external dependency has become stable again.
We are continuing to monitor for any further issues.
A critical external dependency outage has been identified, and triaged
A fix has been implemented and we are monitoring the results.
This incident has been resolved now.
We have identified the issue, implemented a fix and are now monitoring the results
We’ve narrowed the problem down to a single issue. Our team is working on a fix, and performance may still be degraded in the meantime. We’ll provide another update once the fix is deployed.
We are continuing to investigate the issue and are checking several potential root causes. Higher latency and occasional errors are still to be expected.
We are still investigating the issue.
We are currently experience degraded performance, you may face higher latency or timeouts. We are working on it.
This incident has been resolved
A fix has been implemented, we are monitoring the API now.
We are currently investigating a higher latency on Kontext Image to Image