Technical Program Manager, Availability Serviceability Team
Amazon
- Location
- US, VA, Herndon
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- 15h ago
Skills
About this role
The AWS Data Center Availability Serviceability team works with hundreds of AWS data centers globally to deliver the highest quality and lowest cost availability, capacity, and scaling results for our customers. We standardize operations globally by delivering tools, policy, processes, and procedures to our internal teams. Serviceability is the ability to maintain, repair, and manage infrastructure assets efficiently throughout their lifecycle -- encompassing maintenance standards, spares management, training and certifications, operational readiness, and the systematic use of data and automation to prevent disruptions and drive continuous improvement. We are seeking a Technical Program Manager to own the MBOM program as part of the AWS Data Center Availability Serviceability Critical Spares team. MBOMs define which parts make up each piece of critical infrastructure equipment and are foundational to programmatic spares planning and fulfillment -- enabling proactive stocking decisions, criticality assessments, and repair readiness across the global fleet.
What you will do
You will own the development and delivery of a multi-year roadmap to establish comprehensive MBOM coverage across the global AWS fleet. Working with stakeholders across Field Engineering, DCEO, vendors, and lifecycle data teams, you will scale AI-powered pipelines that extract MBOM data from vendor documentation, define per-equipment MBOM standards, and design the processes for establishing MBOMs where source documentation does not exist. Your work sets the foundation for programmatic spares identification, criticality assessment, and fulfillment -- ensuring the right parts are available at the right locations before failures occur. You will build the mechanisms that keep MBOM data current as the fleet grows, diversifies into new cooling technologies, and introduces new vendors. Why it matters: Spares readiness starts with knowing what's inside the equipment. MBOM is the data layer that connects equipment design to spares strategy to repair outcomes. With comprehensive MBOM coverage, we can programmatically identify critical parts, make proactive stocking decisions, and reduce time to repair -- directly improving data center availability for customers. Why you will love it: You will own a program that is genuinely novel -- building something that didn't exist two years ago, using AI tooling that is still maturing, against a problem (fleet-wide MBOM completeness) that no one has solved at this scale. You will operate in high ambiguity with real ownership to define what good looks like. You will work across PLM systems, enterprise asset management, spares strategy, and physical equipment -- giving you breadth that most TPM roles don't offer. And your output has direct, measurable impact: every MBOM you complete makes the fleet more repairable. Key job responsibilities - Own MBOM program strategy, execution, and completeness metrics across the global fleet - Operate and scale AI-assisted pipelines that convert vendor documentation into governed MBOM records - Define and execute the process for establishing MBOMs where vendor documentation is limited or unavailable - Ensure MBOM data flows into criticality assessments and stocking decisions within the Critical Spares program - Define per-equipment MBOM standards in partnership with Field Engineering and equipment owners - Build mechanisms to maintain MBOM accuracy as new equipment deploys and the fleet evolves (including liquid cooling) - Drive cross-functional coordination across Availability, Field Engineering, DCEO, and lifecycle data teams - Use data and metrics to measure program health, prioritize effort, and communicate progress - Write narratives (one to six pages) and present to Director-level leadership - Travel up to 10% A day in the life You will split your time between program execution and cross-functional coordination. On any given day you might be reviewing AI pipeline output quality with the