Lead Systems Operations Engineer - Windows / Unix
Wells Fargo
- Location
- CHARLOTTE, NC
- Work model
- On-Site
- Level
- Senior
- Posted
- 5h ago
Skills
About this role
About this role: Wells Fargo is seeking a Lead Systems Operations Engineer within the Branch Systems and Transformation Technology team. This role is aligned to modern Site Reliability Engineering (SRE) practices and is responsible for driving reliability, resiliency, observability, and operational excellence across critical platform and application services. The role is intended for senior engineers with deep expertise in one core platform domain, applying that expertise to proactively improve platform stability, scalability, and availability. In this role, you will: Lead complex, broad impact initiatives including provision of high level systems consultation for the technology teams Work as key participant in large scale planning of computer systems and network infrastructure for Systems Operations functional area Review and analyze complex technical challenges, as well as escalated support issues related to core business solutions that require in depth evaluation of multiple factors, such as alternatives, enhancements, periodic systems reviews, or improvements to existing systems Make decisions on technical changes and enhancements Consult with engineering team on change design requiring solid understanding of technical process controls or standards that influence and drive new initiatives Collaborate and consult with technical peers, colleagues, and mid to more experienced level managers to resolve systems support issues and achieve goals Required Qualifications: 5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education 5+ years of Site Reliability Engineering experience, or equivalent demonstrated through one or a combination of the following: work experience, training, education Desired Qualifications: Demonstrated experience leading Production Application Support at scale, including ITIL-aligned incident, problem, and change management Experience supporting AI/ML-enabled production systems, including reliability, scalability, and operational risk management for LLM- or agent-based workflows (e.g., model deployments, inference services, retrieval pipelines). Understanding of AI operational concerns such as model observability, prompt/configuration change management, data drift, failure modes, and integration of AI signals into incident response and SRE practices. Thorough understanding of application environment implementations, including on prem, client server, on prem cloud, hybrid cloud and public cloud Experience using and configuring Continuous Integration Continuous Deploy (CICD) tools including Jenkins, Aritifactory, Udeploy as well as Terraform Proven experience reducing operational toil through automation (Ansible or equivalent), including runbook automation and self‑healing patterns Experience implementing application Observability through tools such as AppDynamics, Splunk, Elastic, BigPanda (AIOPS), Grafana, Microsoft Application Insights Experience designing and operating resilient systems, including traffic management (F5, AVI), fault tolerance, and data replication strategies Experience supporting large‑scale Java and .NET applications; strong RDBMS expertise (Oracle, MSSQL); NoSQL experience (MongoDB) a plus Job Expectations: Periodic availability for evening and weekend On-Call Expectation of at least 3 days per week in office Serves as a subject matter expert for environment reliability and operational readiness. Provides technical guidance, promotes reliability engineering best practices, supports operational governance activities, and collaborates with engineering and testing teams to improve environment stability and delivery effectiveness. Pay Range Reflected is the base pay range offered for this position. Pay may vary depending on factors including but not limited to demonstrated examples of prior performance, skills, experience, or work