OTEL ingestion queue delay over 10 minutes in US
- resolved
Issue is resolved and OTEL ingestion pipeline is back to regular ingestion speed.
- investigating
We're investigating degraded performance in the ingestion flow
status.langfuse.com
Open-source LLM observability
All Systems Operational
Issue is resolved and OTEL ingestion pipeline is back to regular ingestion speed.
We're investigating degraded performance in the ingestion flow
The incident has been resolved.
We're making consistent progress in our backlog and are currently seeing delays of about 60 minutes.
We considerably increased throughput of the batch job to write data into v4 data model. We are currently catching up processing data.
We currently experience an ingestion delay of old SDKs by 3 hours. This is due to performance issues in our batch job which moves data into the v4 data model (<https://langfuse.com/changelog/2026-08-17-langfuse-v4>). We currently work on improving the batch job. All v4 compatible SDKs are not impacted by this. Upgrading SDKs is recommended.
OTEL ingestion latency spike is resolved
We're investigating degraded performance in the ingestion flow
Completely back to normal
Backlog is almost completely gone. Ingestion delay is < 1 min.
A deployment rollout coincided with a traffic spike, resulting in an unusually large backlog. Adjusting our autoscaling to absorb the backlog faster.
We are observing OTEL ingestion backlog in US region
Issue resolved. Latency is back to normal levels.
Ingestion delay of up to multiple minutes on the ingestion path. We have scaled up and are investigating the root cause.
Issue resolved since 12:30. Legacy ingestion pipeline is fully operational.
The legacy ingestion queue is experiencing latency degradation up to 5 minutes. We have scaled up our infrastructure in response and are monitoring the situation. We recommend to switch to the new OTEL ingestion queue.
All is back to normal
After concurrency adjustments execution delay is back to normal. Monitoring.
Evaluations are taking longer to execute in EU
Back to normal propagation delay
Latest major SDK versions are not impacted. Find upgrade docs [here](https://langfuse.com/docs/observability/sdk/upgrade-path).
We have scaled our database infrastructure to match the throughput demand
Steady throughput for the last 30 minutes. Delay < 5min.
our autoscaling responded slower than expected, but responded nonetheless. the backlog is draining
evaluation throughput was overwhelmed by a slow tenant/model workload, with lock-expiry retries multiplying work; the concurrent worker rollout prolonged the incident.
Eval execution in EU region is taking longer than usual
We incurred a temporary delay on the legacy ingestion path for v4. This delay has been resolved. OTEL ingestion has not been impacted.
Currently the ingestion for the events table is delayed in the US environment. This happens only for data from non-latest major SDK versions. We are investigating this and have started a mitigation. The data should be up to date soon.
The issue is resolved now
Previous latency increase of the eval execution is resolved. Slight delays may still occur.
We are seeing an increase in the eval execution. We have scaled up our infrastructure and are monitoring.
This issue is resolved and we're back to regular processing times.
We see an improvement and the current delay is below 5min. We continue to investigate and scale.
We're currently observing about 10min ingestion delays. Our team is investigating.
We resolved the performance bottleneck. All APIs are responding as expected as of now.
We experience occasional API errors with 502 Bad Gateway. The team is investigating the root cause currently.
The issue has been resolved.
We're observing elevated latencies and error rates across read APIs and some frontend routes. Ingestion is not affected.
The incident has been resolved.
We found a patch and currently work through the backlog. Things should recover within the next 10min.
We're currently experiencing elevated ingestion times in the EU environment. Some events are processed with a delay of about 15 minutes. Our team is investigating.
All latencies are back to normal
A hiccup during a planned database maintenance. Identified and resolved. Presently monitoring.
We scaled our ingestion and are processing data in time.
We currently have a bottleneck in our ingestion pipeline which causes a delay of ingested data of 12 minutes. We currently work to scale the pipeline.
We scaled our infrastructure and are processing media uploads without errors.
We scaled our underlying infrastructure and are processing media uploads without errors.
We observe that our api/public/media endpoint currently returns 500 status errors as our underlying database has performance issues.
Upgrade successfully resolved.
We just executed a database upgrade which caused a downtime of 5 minutes.
This issue has been resolved.
We've changed infrastructure capacity and observe that this mitigates the problems. We continue to observe the situation.
We're investigating elevated loading times within the Langfuse UI and our APIs.
The underlying incident with our infrastructure (https://statuspage.incident.io/clickhousecloud/incidents/01KT1G25S9PBKM7VJEB146680G) has been resolved.
We're investigating an issue with elevated latencies and error rates in the EU environment.
We've fully caught up and process LLM as a judge in realtime again.
We're currently investigating delays in our LLM as a judge execution times.
We have scaled our infra to drain the queue backlog. Evals are executed without delay again.
We continue to monitor the situation
We believe only a single customer is affected and it is caused by LLM's provider rate limits. Since we monitor worst delay across our system this initially looked as a bigger issue.
We are investigating the root cause
The queue backlog has been fully processed for legacy trace and dataset targeted evals. There is no longer a delay for eval executions.
We have identified the root issue and are scaling our infrastructure to work through the evaluation backlog.
We currently see a delay of LLM as a judge execution for trace-level evaluations by 30 minutes. We currently look into the issue to make sure we execute evals in time. Observation level evaluations are not impacted as they are on a new code path and more scalable by an order of magnitude. Please migrate to observation level evaluations.
The backlog was processed and ingestion delays are below 30 seconds again.
We're currently observing ingestion delays of approximately 6min for our US environment. We're investigating.