Principal AI Network Hardware Systems Engineer
Microsoft
- Location
- United States, Washington, Redmond; United States, California, Mountain View; United States, Oregon, Hillsboro
- Work model
- On-Site
- Level
- Principal
- H-1B history
- 2,066 approvals (FY2023)
- Posted
- 4h ago
Skills
About this role
Overview
Microsoft Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) powers the infrastructure behind Microsoft's Intelligent Cloud, delivering the foundational technologies that support services including Azure, Microsoft 365, Teams, Bing, Xbox Live, and more. As Microsoft continues to advance AI innovation, SCHIE is developing AI-native silicon and system-level solutions that enable next-generation AI training and inference at hyperscale. The Platform Systems Engineering (PSE) team is seeking a Principal AI Network Hardware Systems Engineer to lead the architecture, bring-up, validation, optimization, and deployment of networking infrastructure for Microsoft's MAIA AI platform. This role combines networking hardware, systems architecture, AI infrastructure, and large-scale deployment to deliver industry-leading AI performance and reliability. You will work across the networking stack, spanning high-speed SerDes, optics, cables, NICs, PHYs, switch silicon, AI communication frameworks, and distributed training systems. As a Principal engineer, you will influence architectural direction, guide technical strategy, and collaborate across silicon, firmware, hardware, software, validation, manufacturing, and Azure engineering teams. This is a unique opportunity to shape the future of AI networking infrastructure and drive technologies that power Microsoft's next generation of hyperscale AI systems. Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.
Responsibilities
AI Network Architecture & System Integration Define and develop networking requirements for large-scale AI training and inference clusters. Collaborate with silicon, system software, firmware, hardware, and Azure infrastructure teams to deliver scalable networking solutions from concept through datacenter deployment. Participate in architecture reviews and influence next-generation AI networking roadmaps. Define network concepts of operation, serviceability requirements, telemetry requirements, and operational models for AI infrastructure. Layer 3 / Layer 4 Networking Lead design and validation of IP-based AI networking solutions spanning TCP/IP, UDP, routing, congestion management, flow control, QoS, and traffic engineering. Analyze transport-layer behavior and performance characteristics across large-scale distributed AI workloads. Evaluate network protocol implementations and debug issues impacting latency, throughput, scalability, and reliability. Drive optimization of network communication paths supporting distributed AI training and inference. RDMA & AI Fabric Technologies Design, validate, and optimize RDMA-based networking solutions for AI clusters. Analyze RDMA performance, congestion behavior, packet loss, retransmissions, and collective communication efficiency. Work closely with networking vendors and software teams to optimize AI fabric performance and workload scalability. Develop validation methodologies for AI traffic patterns and collective communication workloads. Performance Characterization & Validation Develop and execute networking validation strategies covering functionality, performance, scale, interoperability, resiliency, and reliability. Characterize network behavior under AI training and inference workloads. Evaluate latency, bandwidth utilization, congestion events, flow distribution, and workload communication patterns. Create and automate network stress, scale, and performance qualification methodologies. Debugging & Root Cause Analysis Lead end-to-end troubleshooting of networking issues across physical, data link, network, and transport