In-House LLM Serving at Netflix
By AI Platform’s Model Runtime team and Inference team
Jul 17, 2026 /Read More
Building Service Topology at Scale: Architecture, Challenges, and Lessons Learned
By Parth Jain, Rakesh Sukumar, Yingwu Zhao, Renzo Sanchez-Silva & Nathan FisherA deep dive into the engineering challenges of building a ...
Jul 13, 2026 /Read More
Measuring the Impact of Personalized Recommendations
By Kevin Zielnicki, Guy Aridor, Aurélien Bibaut, Allen Tran, Winston Chou, and Nathan Kallus
Jul 10, 2026 /Read More
GenPage: Towards End-to-End Generative Homepage Construction at Netflix
Authors: Lequn Wang, Jiangwei Pan, and Linas Baltrunas
Jun 29, 2026 /Read More
Toward More Controllable AI Video Editing: An Early Research Exploration at Netflix
By Zhuoning Yuan, Ta-Ying Cheng, Benjamin Klein, Bahareh Azarnoush
Jun 23, 2026 /Read More
How Netflix Simplified Batch Compute with Kueue
By Alvin Bao, Alex Petrov, Jennifer Lai, Aidan Sherr, and Samartha Chandrashekar
Jun 22, 2026 /Read More
VMAF v1: Good Is Not Good Enough
By Christos G. Bampis, Zhi Li, Kyle Swanson, Nil Fons Miret and Pavan Madhusudanarao
Jun 19, 2026 /Read More
A Human-Augmenting Agentic Workflow for Causal Inference
By Winston Chou, Adrien Alexandre, Lars Olds, Yi Zhang, Garrett Hagemann, and Nathan Kallus
Jun 08, 2026 /Read More
Thinking Fast & Slow for a Personalized Notification System
by Matthew Wood, Ishan Gupta, Kevin Mercurio, Devon Bryant, and Claire Dorman
Jun 05, 2026 /Read More
Dynamic Repartitioning for Time Series Workloads
By Rajiv Shringi, Kaidan Fullerton, Oleksii Tkachuk and Kartik Sathyanarayanan
Jun 03, 2026 /Read More
