yoinka

Software Development Manager, EC2 UltraServer Availability

Amazon

US, WA, SeattleFull TimeMid
Sign in to applyVerified 1h ago
Location
US, WA, Seattle
Employment
Full Time
Work model
On-Site
Level
Mid
Posted
48d ago

Skills

AWS

About this role

A NVIDIA UltraServer isn't a server: it's 18 compute nodes, 72 Blackwell GPUs, and an NVLink fabric that behaves as one giant, memory-coherent computer. When any part of it fails, you have to diagnose across racks, orchestrate hardware swaps with data center technicians, rebuild the NVLink fabric, re-run burn-in, and hand healthy capacity back to customers training frontier AI models. Every hour an UltraServer sits idle is some of the most sought-after compute on Earth going to waste. That's the problem our team owns, and we're looking for a Software Development Manager to lead it. The EC2 UltraServer Availability team builds the automated repair and recovery systems for Amazon's GB200 and GB300 fleet. One of the fastest-growing and most visible infrastructure domains at AWS. This is early-days territory: many repair flows that are fully automated for traditional EC2 hosts are still being invented for UltraServers. You'll set the technical direction that turns manual, expert-driven recovery into deterministic, self-healing automation, and you'll see your team's impact directly in fleet availability numbers that leadership watches weekly. Key job responsibilities You'll lead a two-pizza team of 8–12 engineers automating end-to-end repair of entire UltraServers: detection, control-plane teardown, chassis swap orchestration with hardware engineering and data center operations, fabric rebuild, burn-in testing, and return to customer. Your north-star metrics are fleet availability and repair dwell time: how fast a broken UltraServer gets back to serving customers. Examples of projects your team is working on 1. Automating deterministic repair: building the system that maps a failure signature directly to the right repair action, removing humans from the loop for well-understood failures 2. Orchestrating multi-team repair procedures that today require coordination across hardware engineering, data center technicians, and EC2 control-plane services 3. Cutting repair dwell time by parallelizing teardown, physical repair, and validation steps that currently run serially 4. Extending recovery workflows into new regions, including air-gapped (ADC) environments A day in the life You'll spend most of your time where an SDM should: growing engineers (1:1s, career development, hiring), making roadmap and architecture calls with your senior engineers, and unblocking the team. Because we sit at the intersection of software, hardware, and data center operations, you'll regularly partner with EC2 ML Supercomputing, hardware engineering, and DCO leadership. You'll also drive operational reviews - this is a fleet-facing team with an on-call rotation, and reducing that operational load through automation is itself a core part of the charter. You'll communicate progress and strategy to senior leadership through narratives and business reviews.

About the team

The EC2 UltraServer Availability team (part of EC2 Nitro) maintains the availability of NVIDIA-based ML infrastructure at scale. We own end-to-end recovery and repair for GB200 and GB300 UltraServers: from detecting an availability event through repair, testing, and return to service. We work closely with hardware engineering, data center operations, EC2 ML Supercomputing, and EC2 capacity teams. We value engineers who are curious about the full stack, from NVLink cabling to control-plane APIs, and we invest in reducing our own operational toil through automation.