Lead Site Reliability Engineer (SRE)
Gartner
- Location
- Gurgaon
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 30 approvals (FY2023)
- Posted
- 21h ago
Skills
About this role
About this role: We are seeking a Lead Site Reliability Engineer (SRE) to drive the reliability, performance, and operational excellence of Gartner's conference technology platforms and supporting infrastructure. You will lead the adoption of SRE best practices, champion observability and resilience initiatives, and partner with engineering teams to improve service reliability through automation, performance optimization, and chaos engineering. During live conferences, you will work closely with Network Operations, Application Development and Infrastructure teams to ensure the stability and performance of business-critical systems. Outside of event periods, you will focus on operational readiness, incident management, reliability improvements, capacity planning, and continuous optimization of our platforms and services.
What you will do
Define and evolve SRE strategy, standards, and best practices across multiple application teams, not just execute against existing playbooks Act as the senior technical escalation point during major incidents, driving root-cause analysis on the most complex, cross-system issues and coordinating swat teams across globally distributed groups Partner with senior stakeholders (engineering leadership, architects, product owners) to define SLOs/SLIs and error budgets, and hold teams accountable to them over time Identify systemic opportunities to improve performance, scalability, and stability across the application portfolio, and drive the roadmap to address them Set the standard for dashboards, alerting posture, and observability strategy, and drive consistent adoption across teams Participate in operational support and on-call rotation shifts for supported systems and products. Available to work flexible hours as required to prepare for and/or to provide operational support of select events like major app releases or Gartner hosted conferences, ensuring coordination among globally distributed teams Resolve incidents in production and help triage the application/system issues and identify root causes or remediations to help restore services quickly Conduct blameless postmortems of incidents to identify solutions, including automation, to reduce the probability and/or impact of problem recurrence. Drive advancements in alerting posture Create dashboards and reports to communicate key metrics. Oversee, design, implement, and manage DevOps capabilities using continuous integration/continuous delivery toolsets and automation Collaborate and share lessons learned regarding performance and reliability issues with all stakeholders including developers, other SREs, operations teams, and project management teams. Participate in continuous improvement in software quality and infrastructure, reliability, and resilience. Create templates, build, and maintain documentation for all assigned projects. Build and maintain performance testing frameworks, tools, and methodologies Identify opportunities to automate manual operational work (i.e., “toil”) using pipelines or by using new software or any other appropriate mechanisms Perform analytics on previous incidents to understand root causes and better predict and prevent future issues. Keep a proactive approach to spotting problems, areas for improvement, and performance of bottlenecks. Participate in system design consulting, platform management, capacity planning, and launch reviews. Design and implement chaos engineering principles with various application teams, so resilient patterns are designed and implemented.
What you will need
Must Have Exceptional analytical, problem-solving skills, oral and written communication skills Excels in adapting to changing circumstances. Interested and capable in continuously learning new skills and technologies. Proven track record of setting technical direction, mentoring engineers, and influencing outcomes across teams without formal management authority Excels on incident and response management. Proficiency in