Senior Site Reliability Engineer
Okta
- Location
- Bellevue, Washington; San Francisco, California
- Work model
- On-Site
- Level
- Senior
- Salary
- $165k/yr
- H-1B history
- 52 approvals (FY2023)
- Posted
- 12h ago
Skills
About this role
Secure Every Identity, from AI to Human
Identity is the key to unlocking the potential of AI. Okta secures AI by building the trusted, neutral infrastructure that enables organizations to safely embrace this new era. This work requires a relentless drive to solve complex challenges with real-world stakes. We are looking for builders and owners who operate with speed and urgency and execute with excellence.
This is an opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk.
The Technology, Data and Intelligence Team Message
Okta’s Technology, Data and Intelligence (TDI) team delivers the systems, tools, and services that power internal operations across the company. From core infrastructure to enterprise platforms, we partner across functions to drive scale, reliability, and innovation through technology.
The Senior Site Reliability Engineer Opportunity
Reporting to the Manager, Site Reliability Engineering, this role will help build, improve, and maintain our cloud platform services by designing and implementing complex cloud-based engineering enablement systems. With a strong focus on automation, testing, and operational excellence, you will deliver foundational infrastructure capabilities that enable corporate engineering teams to operate securely, reliably, and at scale.
What you'll be doing
• Secure Cloud Infrastructure & Pipelines: Design, build, and modernize scalable cloud environments and development tools while strictly enforcing security policies and standards for regulated environments.
• Cross-Functional Collaboration & Advocacy: Partner with software engineering teams to champion DevOps and SRE best practices, deliver excellent internal customer service, and actively contribute to Agile workflows (e.g., demos, architecture sessions).
• Technical Documentation & Operations: Create and maintain comprehensive technical documentation, including network diagrams, runbooks, and disaster recovery procedures to ensure system reliability and knowledge sharing.
What you'll bring to the role
• Professional Experience & Scale: 5+ years of experience in SRE, DevOps, or Systems Engineering roles with a proven track record of delivering complex, large-scale infrastructure projects.
• AWS Expertise & Centralized Governance: Expert in building and managing AWS multi-account environments (spanning hundreds of accounts), with deep proficiency in authentication, governance, and organization management (AWS Orgs, IAM, Identity Center, StackSets).
• Automation & CI/CD Pipelines: Highly skilled in infrastructure as code (Terraform), writing secure automation tools in Python, and building Git-based CI/CD workflows (GitLab, GitHub Actions).
• Containerization & Observability: Strong hands-on experience managing container orchestration environments (Kubernetes) and utilizing monitoring and logging tools (Splunk, CloudWatch, Grafana stack).
Extra credit if you have experience in the following
• Networking & AWS Infrastructure: Hands-on experience with general networking concepts (BGP and IPsec management) and leveraging core AWS networking services (VPCs, TGWs, and VPC endpoints).
• System Administration: Solid foundational knowledge and experience in Linux system administration.
• Security & Compliance: Proven experience operating within highly secure, regulated environments (e.g., FedRAMP), with a strong understanding of FIPS, STIGs, and data