Sr. SDM, AI Inference, Neuron SDK
Amazon
- Location
- US, CA, Cupertino
- Employment
- Full Time
- Work model
- On-Site
- Level
- Senior
- Posted
- 18h ago
Skills
About this role
AWS Utility Computing (UC) provides product innovations — from foundational services such as Amazon Elastic Compute Cloud (EC2), to new product innovations that continue to set AWS’s services and features apart in the industry. Come optimize and deploy inference models on Trainium, Amazon's custom cloud-scale machine learning accelerators that power the latest AI models As a Sr. SDM for the Inference Team, you will lead a strong team of managers and engineers to optimize models for best performance and robustness when deployed at scale on Trainium and Inferentia devices. You will be responsible for the full development life cycle of model onboarding, low-level performance optimization, feature support and efficient serving, including reliability and scalability. Preferred qualifications include an established background in optimizing and serving AI models under demanding, fast-changing priorities, and a strong technical ability to understand and manage a vertically integrated system stack that consisting of hardware, frameworks, serving stack, model and workflows. A day in the life You will work with the executive leadership and other senior management and technical leaders to define product directions and deliver them to customers. We build massive-scale distributed training and inference solutions, developing the full stack of software, servers and chips together with teams across the Annapurna organization to run the largest machine learning workloads.