Staff Engineer High Performance Computing
Pfizer
- Location
- United States New York New York City
- Work model
- On-Site
- Level
- Staff
- H-1B history
- 9 approvals (FY2023)
- Posted
- 22h ago
Skills
About this role
ROLE
SUMMARY Pfizer is committed to the application of computational science in the areas of drug discovery and development and has recently initiated a large-scale migration of computational infrastructure to cloud. This role provides technical vision and will drive the execution of high-performance computing (HPC) solutions that support computational workloads across the organization. We are seeking an experienced individual to lead the technical architecture of the cloud HPC platform. Key responsibilities include establishing go-forward cloud HPC platform computing technologies, implementing robust engineering practices, championing infrastructure as code (IaC), and configuring core services that support HPC at scale in the cloud environment. You will work with HPC engineers and scientific computing specialists to develop robust, scalable, high-performance cloud native infrastructure that underpins modernization of the scientific computing platform.
ROLE
RESPONSIBILITIES Strategic Leadership This role will lead development and operationalize cloud-based HPC infrastructure required for research, modeling, and large-scale data processing across multiple cloud environments. Serve as a primary technical expert; evaluate, advocate for, and drive consensus among senior managers and engineers for the go-forward technology platforms and toolkits used for HPC service delivery. Collaborate with stakeholders, users, and leaders to develop a long-term technical roadmap for cloud-based HPC services. Lead deep-dive discussions with technical partners at major cloud providers, defining HPC-related requirements and deliverables for Statements of Work. Drive a culture of shared ownership, transparency, and engineering excellence through mentoring, coaching, and example setting. Perform troubleshooting, system analysis, and benchmarking to manage escalated, difficult to resolve issues and maintain a high-performance environment. HPC Platform Architecture and Engineering Design and own robust and dependable high-throughput, parallel, low-latency infrastructure for HPC and ML/AI workloads in multiple cloud environments (AWS/GCP). Establish technical standards, best practices, architectural frameworks, and implementation guidelines for reproducible HPC platform and application deployments. Recommend cutting-edge HPC technologies including specialized accelerators, novel storage solutions, managed services, and open-source toolkits that will be integrated into the platform. Own OS image development, job scheduler configuration, high performance storage systems Ensure high performance, reliability, scalability, cost efficiency, and security. Automation and DevOps Drive adoption of infrastructure automation using IaC tools like Terraform and CloudFormation. Establish, promote, and enforce internal standards (naming, tagging, documentation, version control, and change procedures) to ensure repeatable environment provisioning and scaling. Establish infrastructure lifecycle management procedures, from provisioning to operations, support, updating, and teardown of production computing platforms. Monitoring and Reliability Determine KPIs to guide monitoring, logging, and alerting strategies for the infrastructure. Collaborate with stakeholders, users, and senior managers to develop meaningful user-facing dashboards, drive resource management, cost efficiency, and workload optimization. Design workflows, alerting systems and utilities to improve observability, user, or administrator experiences.
BASIC QUALIFICATIONS
B.S. in computer science, life science, data science or similar fields with 6+ years of experience in cloud infrastructure engineering. A proven track record of developing and supporting robust HPC frameworks in a cloud environment. Expert level experience with at least one of AWS and GCP, including knowledge of core compute and storage services relevant to HPC. Deep understanding of modern CI/CD practices,