Visual TL;DR. 30,000+ Flink Jobs drove need In-house Autoscaler (2019). In-house Autoscaler (2019) initially led to Reduced Resource Usage. Reduced Resource Usage but faced Limitations of External View. Limitations of External View prompted Consolidate Autoscalers. Consolidate Autoscalers into Open-Source Solution. Open-Source Solution resulting in Cost Savings. Open-Source Solution and Improved Stability.
- 30,000+ Flink Jobs: Netflix operates a massive scale of Flink jobs requiring efficient resource management
- In-house Autoscaler (2019): built an external observer system monitoring Flink jobs via Atlas telemetry
- Reduced Resource Usage: achieved 25-45% resource reduction for simpler, single-operator pipelines
- Limitations of External View: coarse container metrics and single TaskManager knob insufficient for complex jobs
- Consolidate Autoscalers: strategic shift to merge two distinct Flink autoscaling systems into one
- Open-Source Solution: adopting and contributing to a single open-source Flink autoscaler
- Cost Savings: achieving significant cost reductions through optimized resource allocation
- Improved Stability: enhanced system reliability and performance for complex Flink workloads
Visual TL;DR
