Senior Technical Program Manager, RL Gyms, Frontier AI RL Gym Assets
Amazon
- Location
- US, WA, Bellevue
- Employment
- Full Time
- Work model
- On-Site
- Level
- Senior
- Posted
- 17h ago
About this role
We are seeking an experienced Senior Technical Program Manager to design and manage the program that drives execution of our Reinforcement Learning (RL) Gym development initiative. RL Gyms are standardized evaluation environments that measure AI agent performance across enterprise domains. In this role, you will define program structure, build mechanisms, and drive accountability to deliver RL Gyms at scale — partnering with scientists, engineers, and domain experts across multiple organizations. This is a high-ownership role where you will shape how the program operates end-to-end — from partner engagement and requirements definition through certification and delivery — rather than building gyms yourself. Key job responsibilities Program Execution: Own end-to-end delivery of RL Gym development programs — from requirements gathering and design through implementation, certification, and launch readiness Cross-functional Coordination: Drive alignment across engineering, science, and domain teams to define gym specifications, collection standards, and certification criteria Technical Depth: Understand gym harness architecture, evaluation frameworks, and agent benchmarking methodologies (e.g., SWE-bench style evaluations) to make informed tradeoffs and unblock teams Stakeholder Management: Partner with domain stakeholders across the company to identify high-value workflows suitable for gym development Certification & Quality: Define and enforce certification standards for RL Gyms, ensuring they meet bar for agent evaluation fidelity and reproducibility Capacity Planning: Track gym development targets, manage pipeline throughput, and forecast delivery timelines across multiple concurrent gym workstreams Tooling & Process: Drive improvements to gym tracking tools, development workflows, and operational processes (e.g., DPM tracking, standup cadences) Risk Management: Identify blockers early, escalate appropriately, and develop mitigation plans to keep programs on track About the team The RL Gym Program develops standardized evaluation environments ("gyms") that test AI agents on realistic enterprise tasks. Our gyms span multiple business domains and are used to certify agent readiness for production deployment. We operate at the frontier of AI evaluation, working closely with model teams, infrastructure teams, and business stakeholders to define what "good" looks like for AI agents.