Lead Software Engineer - AI Platform Reliability
JPMorgan Chase
- Location
- Seattle, WA, United States
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 1,524 approvals (FY2023)
- Posted
- 15h ago
Skills
About this role
Are you passionate about building resilient, scalable systems that power the future of AI? At JPMorganChase, we're pushing the boundaries of what's possible with artificial intelligence and machine learning — and we need engineers like you to help us do it reliably, securely, and at scale. As a Lead Software Engineer at JPMorganChase within the AI/ML Data Platforms organization, you will be a key member of the Reliability Engineering team, driving the design and delivery of trusted, market-leading technology products. You will apply your deep technical expertise and problem-solving skills to enhance the reliability and scalability of AI/ML platforms, build reusable services and tooling, and partner across teams to unblock high-impact AI use cases. This is an opportunity to shape how the firm delivers AI capabilities — with operational excellence at the core.
Job responsibilities
Design and implement solutions to enhance the reliability and scalability of AI/ML platforms and applications to accommodate fast-growing demands Develop secure, stable, and high-quality production code, and participate in code reviews, debugging, testing, and remediation of defects across AI Foundation Services components Build and enhance reusable platform services, APIs, SDKs, and libraries that standardize how application teams consume model hosting, inference, and AI/ML managed services Partner with Lines of Business application teams to implement AI Foundation Services capabilities that unblock generative AI and AI use cases, supporting delivery from technical design through build, launch, and early operational support Own and evolve non-functional requirements and build/enhance tooling for observability, resilience, security controls, infrastructure management, and cost optimization Establish and enforce standards and reference architectures for reliability, observability, automation, and operational readiness across services Partner with product and platform engineering teams to define and meet service reliability targets, including performance, availability, and recoverability Participate in on-call rotations, debug and resolve complex production issues; identify systemic gaps and drive durable remediation Mentor and guide engineers; raise the bar on engineering quality, documentation, and operational rigor Required qualifications, capabilities, and skills Formal training or certification on software engineering concepts and 5+ years applied experience Strong hands-on coding experience in Python with experience delivering production-grade services Demonstrated experience leading effective use of approved AI-assisted software development tools (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching engineers on safe, compliant adoption within delivery practices Hands-on practical experience with system design, automated testing, debugging, and operational stability for production software Experience implementing observability, logging, metrics, alerts, Service Level Objectives, incident response practices, and root-cause analysis for services in production Working knowledge of software application development and technical processes, with depth in one or more areas such as cloud platforms, artificial intelligence, machine learning platforms, distributed systems, or infrastructure engineering Ability to break down technical requirements into executable engineering tasks, manage dependencies, and deliver against milestones in partnership with product and application teams Strong written and verbal communication skills, with the ability to explain technical decisions, trade-offs, issues, and risks to