SDM, ML Data Infra, Amazon Traffic Engineering
Amazon
- Location
- CA, BC, Vancouver
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- 22h ago
Skills
About this role
We are seeking an experienced Software Development Manager to lead a team of Software Development Engineers (SDEs) and Data Engineers (DEs) building the Core Data Infrastructure that underpins our ML and Science initiatives for Bot Management. You will own the end-to-end data platform—from source ingestion through transformation, feature engineering, and serving—ensuring Science and ML Platform teams have reliable, scalable, and timely access to the data they need for training, evaluation, and inference. This is a high-impact leadership role. Our Science teams are building increasingly sophisticated models—and each requires different data formats, latencies, and serving patterns. Your team will be the backbone that makes this possible: ingesting billions of events from diverse source systems, building production-grade pipelines that transform raw signals into ML-ready feature groups, and operating the Feature Store that serves these features consistently across all model types. You will partner closely with Science leadership to translate model requirements into data infrastructure investments, and with ML Platform leadership to ensure seamless integration between your data layer and their training/inference systems. This role demands a leader who can navigate ambiguity across organizational boundaries, drive technical alignment between data engineering, software engineering and science teams, and build systems that scale with the rapid pace of model innovation. Key job responsibilities Data Infrastructure & Feature Engineering — Own the Feature Store, feature pipelines, and data serving layer. Build versioned feature groups across multiple storage backends (S3 for tabular, OpenSearch for embeddings) and production pipelines that transform disparate datasets into ML-ready features for Science teams. Streaming & Real-Time Systems — Design and operate Apache Flink applications and live stream data processing for near real-time feature computation. Build event-driven architectures leveraging Kinesis and Kafka to support low-latency bot detection signals. Data Pipelines & Ingestion — Own batch and near real-time pipelines spanning Trails (raw + aggregated) and Non-Trails sources (AIT, Clickstream, Customer Segmentations, OPS). Evolve pipelines from Cradle/POC to production-grade using AWS Glue. Implement data drift detection and governance frameworks. Science & ML Platform Partnership — Serve as the primary data infrastructure partner to Applied Scientists and ML Platform. Define data contracts and SLAs, participate in model design reviews, and ensure the Feature Store integrates seamlessly with training and inference systems. People Leadership — Recruit, develop, and retain a high-performing team of SDEs and DEs. Set goals, manage roadmaps, and foster a culture of operational excellence.
About the team
Traffic Engineering's Bot Management organization protects Amazon's ecosystem by detecting and mitigating automated threats at scale. Our Core ML Data Infrastructure team is responsible for building and operating the foundational data infrastructure that powers bot detection, AI agent identification, and content exfiltration defense. We are building a unified, model-agnostic, production-grade ML platform that brings together training, evaluation, and inference pipelines into a cohesive system serving multiple model types across the organization.