Staff Software Engineer (L6) - Developer Productivity — Platform Systems, AIMS Engineering
Netflix
- Location
- USA - Remote
- Work model
- Remote
- Level
- Staff
- H-1B history
- 80 approvals (FY2023)
Skills
About this role
At Netflix, our mission is to entertain the world. Together, we are writing the next episode - pushing the boundaries of storytelling, global fandom and making the unimaginable a reality. We are a dream team obsessed with the uncomfortable excitement of discovering what happens when you merge creativity, intuition and cutting-edge technology. Come be a part of what’s next.
AI for Member Systems (AIMS) runs the AI systems behind every recommendation, search result, and personalized experience for 300M+ members. Hundreds of researchers and engineers across AIMS depend on a shared inner loop — build and test infrastructure, CI/CD, local and remote dev environments, and the tooling that takes an idea from a notebook to a production experiment — to get their work done every day. As the AIMS AI/ML stack scales and modernizes, that inner loop has to scale with it, or it becomes the bottleneck on everything else.
Platform Systems is the engineering foundation of AIMS, owning reliability, scalability, cost efficiency, and developer experience across the org. We're looking for a Senior/Staff AI Software Engineer to own developer productivity for AIMS: the build, test, and iteration loop that our ML researchers and engineers rely on daily. This is a cross-cutting, high-leverage role — improvements here compound across every team in the org, not just one.
Responsibilities
* Own the end-to-end developer experience for AIMS ML practitioners: local and remote dev environments, build and test infrastructure, CI/CD pipelines, and the tools researchers use to move from idea to production experiment.
* Identify friction in day-to-day engineering and research workflows firsthand, and design tooling and abstractions that remove it at the root rather than patching around it.
* Design, build, and operate large-scale build and CI/CD systems that keep build, test, and iteration times fast as the codebase, model count, and headcount grow.
* Partner directly with ML researchers and engineers embedded across AIMS teams to understand real workflows, prioritize the highest-leverage productivity investments, and ship tools people actually adopt.
* Build and maintain internal developer platforms and self-service tooling that reduce the operational burden on individual teams, so they can focus on ML work instead of infrastructure upkeep.
* Instrument developer workflows to measure productivity — build times, iteration speed, time-to-first-experiment — and use that data to prioritize where to invest next.
* Drive adoption of new tooling through documentation, migration support, and hands-on partnership with teams; treat launch as the start of the work, not the end.
* Set technical standards for developer tooling across AIMS and raise the engineering bar through design reviews and architectural guidance.
* Evaluate, integrate, and productionize GenAI-powered developer tooling — intelligent build/test selection, automated code review, triage automation — where it measurably improves velocity, and build the guardrails that make it safe to rely on.
What We're Looking For
* Significant experience building and operating developer productivity infrastructure — build systems, CI/CD, developer environments, or internal platforms — at scale.
* Strong software engineering fundamentals, with deep proficiency in Python and working proficiency in at least one JVM language (Scala, Java, or similar).
* Hands-on experience with distributed build systems (e.g., Bazel, Buck, Pants) and large-scale distributed data/compute frameworks (e.g., Spark, Beam).
* Working understanding of GenAI-powered developer tooling — AI coding assistants, automated code review, agentic coding workflows — and hands-on experience using these tools effectively in your own engineering practice, including judgment about where they help, where they don't, and how to validate their output.
* Comfort with parallel and distributed computing, and experience