Head of High Availability Systems Engineering
Citigroup
- Location
- Rutherford New Jersey United States
- Work model
- On-Site
- Level
- Staff
- Posted
- 1d ago
About this role
Citi, a leading global bank with approximately 200 million customer accounts in over 160 countries, provides a broad range of financial products and services to consumers, corporations, governments, and institutions. The bank's Enterprise Operations & Technology division underpins these offerings, delivering secure, reliable, and efficient technology solutions that are foundational to managing global resources, ensuring safety, and providing a first-class customer experience. This reflects Citi's mission to create economic value that is systemically responsible and in its clients’ best interests. Fostering a culture of diversity and inclusion, Citi is committed to a workforce that represents the clients it serves. The company values respect, promotes individuals based on merit, and ensures opportunities for personal development are widely available. Ideal candidates are passionate, innovative problem-solvers who contribute to a culture of delivering results with pride and are empowered to enable growth and progress together with the firm. This is a rare opportunity to lead a globally distributed engineering team at the forefront of mission-critical, High Availability infrastructure. As a Head of High Availability Systems Engineering , you will own the strategy, stability, and evolution of a large-scale z/TPF, Stratus, and I-Series estate that underpins enterprise-grade operations around the clock. If you thrive at the intersection of deep technical mastery and senior leadership, this role offers the scope, complexity, and impact to match your ambition.
Key Responsibilities
Team Leadership & Talent Development Lead, mentor, and grow a high-performing global team of z/TPF, Stratus, and I-Series Systems Programmers Foster a culture of engineering excellence, continuous learning, and accountability across geographically distributed teams Set clear performance expectations and develop career pathways for senior technical contributors Platform Engineering & System Management Oversee the installation, configuration, maintenance, and upgrades of z/TPF, Stratus, and I-Series software environments Drive end-to-end lifecycle management of complex, large-scale HA infrastructure — from hardware components to operating systems and related software Administer and optimize the IBM Security Portal, ensuring robust vulnerability management practices are embedded across the estate Performance Optimization & Capacity Planning Proactively monitor system performance, identify bottlenecks, and architect solutions that maximize throughput and reliability Lead capacity planning initiatives, translating business growth projections into actionable infrastructure roadmaps Implement performance tuning strategies across z/TPF, VOS, and IOS environments Security & Compliance Champion comprehensive security policies and procedures that protect sensitive data and preserve system integrity Maintain deep familiarity with mainframe security tooling, including Crypto, FICON, and platform-specific security utilities Ensure compliance with enterprise security standards and regulatory requirements Business Continuity & Disaster Recovery Architect and maintain enterprise-grade disaster recovery solutions, with expert-level command of storage replication technologies Lead continuity-of-business planning, testing, and execution to ensure zero-compromise availability for critical systems Cross-Functional Collaboration & Process Improvement Partner closely with application support, database administration, and network engineering teams to deliver integrated, resilient solutions Identify and implement process improvements that drive operational efficiency, reduce toil, and elevate engineering standards Maintain rigorous, up-to-date documentation of system configurations, runbooks, and troubleshooting procedures Skills Required: Technical Expertise Deep, hands-on expertise in z/TPF Administration , including system programming, tuning, and troubleshooting at