❌

Normal view

Netflix Reworks Conductor for 420 Million Monthly Workflow Executions and 10X Larger Workflows

11 September 2026 at 22:17

Netflix has reworked its Conductor workflow orchestration engine to handle larger workloads, increasing supported workflow size from about 2,500 to 30,000 tasks and reducing p99 workflow evaluation latency by about 40%. Conductor 4.0 separates workflow metadata from task data, moves evaluation to asynchronous processing, and introduces dynamic worker allocation and concurrency controls.

By Leela Kumili
  • βœ‡InfoQ
  • Netflix Moves toward Open Source Flink Autoscaler for 30,000+ Streaming Jobs Leela Kumili
    Netflix is moving toward the open-source Apache Flink Autoscaler for more than 30,000 streaming jobs across multiple AWS regions. The operator-level approach addresses limitations of Netflix’s cluster level autoscaler for complex, stateful pipelines. Netflix reports a 58% reduction in annualized Flink compute expenditure for one team, saving approximately $1.1 million annually. By Leela Kumili
     

Netflix Moves toward Open Source Flink Autoscaler for 30,000+ Streaming Jobs

7 September 2026 at 22:06

Netflix is moving toward the open-source Apache Flink Autoscaler for more than 30,000 streaming jobs across multiple AWS regions. The operator-level approach addresses limitations of Netflix’s cluster level autoscaler for complex, stateful pipelines. Netflix reports a 58% reduction in annualized Flink compute expenditure for one team, saving approximately $1.1 million annually.

By Leela Kumili
  • βœ‡InfoQ
  • Presentation: Latency: the Race to Zero...Are We There Yet? Amir Langer
    Amir Langer discusses the evolution of latency reduction, from the Pony Express to modern hardware. He explains how separation of concerns - decoupling business logic from I/O - and tools like Aeron and the Disruptor achieve single-digit microsecond speeds. He shares insights into replicated state machines, consensus protocols like Raft, and the future of low-latency sequencer architectures. By Amir Langer
     

Presentation: Latency: the Race to Zero...Are We There Yet?

10 April 2026 at 22:28

Amir Langer discusses the evolution of latency reduction, from the Pony Express to modern hardware. He explains how separation of concerns - decoupling business logic from I/O - and tools like Aeron and the Disruptor achieve single-digit microsecond speeds. He shares insights into replicated state machines, consensus protocols like Raft, and the future of low-latency sequencer architectures.

By Amir Langer
  • βœ‡InfoQ
  • Pinterest Reduces Spark OOM Failures by 96% Through Auto Memory Retries Leela Kumili
    Pinterest Engineering cut Apache Spark out-of-memory failures by 96% using improved observability, configuration tuning, and automatic memory retries. Staged rollout, dashboards, and proactive memory adjustments stabilized data pipelines, reduced manual intervention, and lowered operational overhead across tens of thousands of daily jobs. By Leela Kumili
     

Pinterest Reduces Spark OOM Failures by 96% Through Auto Memory Retries

6 April 2026 at 22:32

Pinterest Engineering cut Apache Spark out-of-memory failures by 96% using improved observability, configuration tuning, and automatic memory retries. Staged rollout, dashboards, and proactive memory adjustments stabilized data pipelines, reduced manual intervention, and lowered operational overhead across tens of thousands of daily jobs.

By Leela Kumili

Article: Replacing Database Sequences at Scale Without Breaking 100+ Services

3 April 2026 at 17:00

The article discusses the challenges faced during a migration from a relational database to NoSQL, focusing on the importance of database sequences for unique identifiers. It outlines the development of a new sequence service using DynamoDB and a two-tier caching architecture.

By Saumya Tyagi
❌