Principal Site Reliability Engineer (Fastly, Cloudflare/Akamai)
Gartner
- Location
- Chennai
- Work model
- On-Site
- Level
- Principal
- H-1B history
- 30 approvals (FY2023)
- Posted
- 21h ago
Skills
About this role
About the role : The Cloud Center of Excellence (COE) is responsible for developing Gartner’s capabilities in automating and streamlining IT infrastructure processes and DevOps tasks while improving Gartner’s self-service capabilities using public cloud platforms and open-source technologies. Deliver world class availability, stability and performance of our digital products that guarantees a consistent and positive user experience for our customers What you will do: Deliver world-class availability, stability, and performance for our digital products, ensuring a consistent and positive user experience for customers. Establish multi-CDN architecture patterns, traffic-steering strategies, failover mechanisms, and operational standards to achieve five-nines (99.999%) availability targets. Lead CDN incident response, disaster recovery exercises, capacity planning, and performance optimization initiatives to ensure uninterrupted customer experiences during internet-scale events. Partner with Security, Infrastructure, and Application Development teams to implement and optimize edge security capabilities including DDoS protection, WAF policies, bot management, and edge compute services. Build and maintain Infrastructure as Code (IaC) solutions to manage CDN configurations, governance policies, observability, WAF rules, and deployment processes. Design, build, and maintain global traffic management policies, including DNS routing, certificate management, and load balancing, to meet disaster recovery (DR) and high-availability (HA) requirements. Design and implement edge-computing solutions using Cloudflare Workers, Lambda@Edge, or equivalent technologies to address key application use cases. Define and monitor CDN-specific SLOs, SLIs, and KPIs including cache hit ratio, origin offload, latency, availability, and traffic distribution across providers. Identify opportunities to improve the performance, scalability, and stability of applications across the full technology stack. Participate in operational support and on-call rotations to provide timely incident response for supported systems and applications. Lead production incident resolution efforts, triage application and system issues, identify root causes, and implement remediation measures to restore services quickly. Develop alerting strategies, dashboards, and reporting frameworks to communicate operational and business critical metrics. Identify opportunities to automate manual operational work (i.e., “toil”) using pipelines, new software, or any other appropriate mechanisms. Drive system design consulting, platform management, capacity planning, and launch reviews. • Partner with application teams to design and implement chaos engineering practices that improve resilience and fault tolerance. Establish strong relationships with development teams, product owners, and other stakeholders to define and manage Service Level Objectives (SLOs). Manage, collaborate and share lessons learned regarding CDN architecture and deployments with all stakeholders, including developers, other SREs, operations teams, and project management teams. Improve customer experience and increase product value by enhancing the reliability, availability, and performance of client-facing applications and service Internal relationships are critical to establish with application development teams and technology teams such as database, cloud/infrastructure, and enterprise architecture.
What you will need
Must Have Expert knowledge of one or two enterprise CDN platforms, including Fastly, Cloudflare, CloudFront, Akamai, or equivalent technologies, with hands-on experience designing, operating, and optimizing multi-CDN environments. Experience implementing multi-CDN routing and failover solutions using DNS, GSLB, traffic management platforms, or real-user-performance (RUM) driven traffic steering. Experience in managing large-scale CDN platforms using Terraform/OpenTofu, CI/CD pipelines,