Software Development Engineer II, AWS SageMaker AI
Amazon
- Location
- US, WA, Bellevue
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- 2h ago
Skills
About this role
At AWS SageMaker AI, we're making it easy to build state-of-the-art foundation models on the cloud. Model Factory is our platform for building, training, customizing, and evaluating foundation models at scale. Instead of hand-chaining data prep, distributed training, evaluation, and deployment across thousands of GPU and AWS Trainium devices, Model Factory lets teams express the whole lifecycle as a single, contract-validated workflow — orchestrated, reproducible, and fully managed. As LLMs and Generative AI scale, Model Factory is the platform that turns frontier training research into a reliable, repeatable pipeline for our internal teams and customers. We're looking for a Software Development Engineer to help design, build, and operate the distributed systems at the core of this platform — the orchestration engine, compute integrations, SDK and contract layer, and the infrastructure that runs large-scale training and customization jobs. You'll own components end-to-end, from design through delivery and on-call operations, and work closely with the ML scientists and platform teams who depend on Model Factory every day. You'll turn requirements into robust, scalable, supportable services that fit cleanly into the overall architecture, uphold a high engineering bar, and grow your scope and technical leadership as you go. A successful candidate has a strong software-engineering foundation, writes high-quality distributed-systems and services code, communicates clearly, and is motivated to deliver results in a fast-paced, ambiguous environment. Key job responsibilities As a Software Development Engineer on the SageMaker AI team, you will: - Design, build, test, and operate services that orchestrate foundation-model data preparation, training, evaluation, and deployment as reliable, contract-validated workflows. - Own delivery of individual components end-to-end — from design and implementation through deployment, monitoring, and on-call operations. - Build and extend compute-backend integrations and job launchers — submitting, monitoring, and recovering large-scale training jobs across SageMaker (Training/Processing/HyperPod), EMR, AWS Batch, and Kubernetes/EKS. - Improve the platform's resiliency and operability for long-running distributed jobs — checkpoint/resume, fault detection and recovery, retries, and observability (metrics, logging, experiment tracking). - Contribute to the SDK, workflow orchestration, and schema/contract layer that teams use to declare and run jobs, and to the CDK infrastructure that deploys the platform. - Integrate containerized training and evaluation frameworks (e.g., PyTorch/FSDP, verl, NeMo/Megatron) into the platform's task and recipe model. - Contribute to design and architecture discussions, write clear technical designs, and uphold engineering best practices (code review, testing, operational readiness). - Collaborate with ML scientists and internal customers to translate training requirements into reliable, self-service platform capabilities, and help onboard and mentor interns and new engineers as you grow.