Senior Lead Infrastructure Engineer - Enterprise Technology Data Protection & Recovery
JPMorgan Chase
- Location
- OH, United States
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 1,524 approvals (FY2023)
- Posted
- 23h ago
Skills
About this role
Push the limits of what’s possible with us as an experienced senior member of our product team. As a Senior Lead Infrastructure Engineer at JPMorgan Chase within the Enterprise Technology Data Protection & Recovery product line, you will operate at the intersection of infrastructure engineering, product ownership, and risk. You will own the product vision, roadmap, and backlog for the Resiliency Evidence Service (RES), the platform of record for capturing, validating, and reporting resiliency evidence across the firm, while providing hands-on technical leadership to the platform, application, and infrastructure teams that produce and consume that evidence. You will translate control objectives, regulatory expectations, and recovery targets (RTO/RPO) into concrete engineering guidance, evidence contracts, and delivery plans, then partner with global development and infrastructure teams to make those outcomes real. Required qualifications, capabilities, and skills Own the product vision, roadmap, and quarterly plan for the Resiliency Evidence Service; translate firmwide resiliency, risk, and audit objectives into a prioritized backlog with clear outcomes, success metrics, and release milestones. Act as the single-threaded product owner: groom and prioritize the backlog, run sprint planning and reviews, manage cross-team dependencies, and communicate status, risks, and trade-offs to engineering, product, risk, and executive stakeholders. Provide hands-on infrastructure engineering leadership: define reference architectures, evidence contracts (schemas, APIs, event formats), and integration patterns that partner teams use to emit resiliency evidence into RES. Guide development and infrastructure teams on controls and control objectives from a risk perspective; interpret firmwide standards, recovery objectives, and audit findings into concrete engineering requirements, acceptance criteria, and evidence artifacts. Partner with platform, application, and infrastructure teams to help them design, instrument, and produce the telemetry, attestations, restore proofs, and control evidence consumed by RES; run working sessions, review designs, and unblock adoption. Steward the end-to-end architecture of RES ingestion, storage, validation, and reporting surfaces; define domain boundaries, service contracts, and cross-service standards, and maintain architectural decision records. Drive resilient service design for RES itself: manage availability SLOs, latency budgets, and error budgets; lead DR tests, restore validation exercises, and chaos/resilience drills for the platform and its dependencies. Champion the secure SDLC across the ecosystem: threat modeling, SAST/DAST integration, dependency and SBOM management, secrets hygiene, encryption in transit and at rest, and robust authentication/authorization patterns for evidence pipelines. Advance observability and SRE practices: instrument metrics, logs, and traces across ingestion and reporting paths; design actionable alerts; author runbooks; lead incident response and blameless postmortems for both RES and the evidence supply chain. Support CI/CD and test quality across the product line: maintain an end-to-end understanding of how RES services, evidence pipelines, and partner integrations fit together and are tested (unit, integration, contract, performance); set expectations for coverage, quality gates, and progressive delivery (canary/blue-green) with rollback plans; review results, triage failures, and guide engineers to the likely root cause when tests or pipelines break. Represent RES with Risk, Controls, Audit, and Compliance partners; produce data-backed status updates, escalation requests, and evidence packages that demonstrate control effectiveness and recovery posture. Support critical environments (Development, QA, Simulation, Production) and lead management of on-call obligations for owned services, ensuring operational stability and timely resolution of incidents. Contribute