Lead Systems Operations Engineer - Technology Major Incident Management
Wells Fargo
- Location
- Bengaluru, India
- Work model
- On-Site
- Level
- Senior
- Posted
- 12h ago
Skills
About this role
About this role: Wells Fargo is seeking a Lead Systems Operations Engineer In this role, you will: Lead complex, broad impact initiatives including provision of high level systems consultation for the technology teams Work as key participant in large scale planning of computer systems and network infrastructure for Systems Operations functional area Review and analyze complex technical challenges, as well as escalated support issues related to core business solutions that require in depth evaluation of multiple factors, such as alternatives, enhancements, periodic systems reviews, or improvements to existing systems Make decisions on technical changes and enhancements Consult with engineering team on change design requiring solid understanding of technical process controls or standards that influence and drive new initiatives Collaborate and consult with technical peers, colleagues, and mid to more experienced level managers to resolve systems support issues and achieve goals Required Qualifications: 5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education Desired Qualifications: 5+ years of experience in Major Incident Management, Production Support, or IT Service Management. Strong experience with major incident triage, monitoring tools, escalation frameworks, and command center operations. Strong knowledge of ITIL Incident Management principles and service management frameworks. Proven experience leading customer-impacting and regulatory-sensitive incidents. Demonstrated hands-on experience leading complex incident bridge calls involving multiple technical towers, vendors, and stakeholders. Excellent verbal and written communication, executive communication, facilitation, and stakeholder management skills. Experience interacting with senior leadership and executive stakeholders during critical outages. Proven ability to make decisions under pressure and maintain control during high-impact situations. Ability to operate effectively in a 24x7, high-pressure environment. Experience in financial services or large regulated enterprises preferred. Experience as a Technical Writer or Technical Project Manager. Experience working in organizations with a large-scale IT footprint. Experience managing incidents across infrastructure, cloud, network, application, database, and end-user technology domains. Experience working with ServiceNow or equivalent ITSM platforms. Experience working with Microsoft Teams or equivalent collaboration tools. Experience in data analysis and reporting using Power BI, Excel, and similar tools. Experience building dashboards using Grafana, Splunk, or equivalent monitoring platforms. Working knowledge of Windows and Linux Server OS, networking, routing topology, and IT support fundamentals. Understanding of physical and virtual server technologies. Experience working with cloud technologies. Experience working in an Agile environment. Ability to analyze incident trends and drive service stability improvements. Experience mentoring or coaching junior team members is an advantage. Job Expectations Serve as the primary Incident Commander for Major Incidents (P1/P2) from identification through restoration and closure. Lead and facilitate major incident bridge calls, ensuring effective engagement and coordination across technical teams, vendors, and business stakeholders. Assess and validate incident severity, impact, urgency, and major incident criteria throughout the incident lifecycle. Drive timely resolution by coordinating cross-functional teams and maintaining focus on restoration activities. Establish clear ownership, accountability, and action plans during ongoing incidents. Identify and assign Recovery Leaders and ensure accountability for restoration activities. Manage escalations across technical, operational, and leadership teams to remove blockers