Staff Reliability Engineer
ServiceNow
- Location
- Santa Clara, CALIFORNIA, United States
- Employment
- Full Time
- Work model
- Remote
- Level
- Staff
- H-1B history
- 185 approvals (FY2023)
- Posted
- 1h ago
Skills
About this role
Company Description
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started. Join us to put AI to work for people.
Job Description
Join us to build the next generation of cloud-native reliability, release, and test platforms that enable engineering excellence, developer productivity, and high-confidence ServiceNow releases through automation, observability, and AI-driven operations. What you get to do in this role: Design, build, and operate cloud-native engineering platforms for software validation, release validation, and production readiness Design and maintain production-like release and test ServiceNow environments that improve release confidence and deployment readiness. Build and integrate automated test pipelines, observability, reliability signals, deployment intelligence, and quality gates into CI/CD workflows. Develop automation solutions that improve engineering productivity, streamline operations, and reduce manual toil through shift-left engineering practices. Build reusable frameworks, self-service engineering environments, test data management, mock services, and developer productivity tooling. Design and enhance Kubernetes-based platforms supporting scalable test infrastructure, release automation, cloud-native workloads, and developer self-service. Implement automated validation for failure detection, deployment verification, policy enforcement, security checks, resilience testing, and operational health assessments. Resolve complex platforms, infrastructure, and networking challenges through software engineering, systems design, and automation. Partner closely with engineering teams to improve platform reliability, release quality, cloud-native adoption, and engineering best practices. Participate in architecture reviews, technical design discussions, and implementation of scalable, automation-first engineering solutions. Influence technical decisions through strong engineering execution, collaboration, and delivery of high-quality platform capabilities. Mentor engineers through technical guidance, code reviews, knowledge sharing, and engineering best practices. Foster a culture of reliability, automation, operational excellence, continuous improvement, and customer-focused engineering.
Qualifications
To be successful in this role you have: Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry. 8+ years of experience in Site Reliability Engineering (SRE), DevOps, Platform Engineering, Software Engineering, or Infrastructure Engineering with a Bachelor's degree; or 6 years and a Master's degree; or a PhD with 3 years experience; or equivalent experience. Hands-on experience with Kubernetes across cluster operations, networking, storage, security, autoscaling, and multi-cluster environments. Experience building and operating cloud-native platforms supporting scalable, highly available services. Experience integrating Kubernetes with CI/CD, GitOps, automated test pipelines, deployment validation, and cloud-native deployment workflows. Experience designing and implementing automation to improve