yoinka

AVP, Site Reliability Engineer - CloudOps (L10)

Synchrony Financial

Hyderabad, INSenior
Sign in to applyVerified 2h ago
Location
Hyderabad, IN
Work model
On-Site
Level
Senior
Posted
1d ago

Skills

AWSArgoCDCI/CDCassandraDockerGitGitHub ActionsGrafanaJenkinsKubernetesMongoDBNew RelicPostgreSQLPrometheusSplunkTerraform

About this role

Job Description

Role Title : AVP, Site Reliability Engineer - CloudOps (L10) Company Overview: Synchrony (NYSE: SYF) is a leading consumer financing company that has been at the heart of American commerce and opportunity for nearly a century. Synchrony delivers credit and banking products that empower tens of millions of consumers to improve their financial lives and access what matters most. Leveraging innovative solutions that are shaping the future of retail commerce, Synchrony supports the growth and success of some of the nation’s most respected brands, alongside hundreds of thousands of small and midsize businesses, including health and wellness providers. Committed to excellence in service and culture, Synchrony is proud to be named as #3 as a Great Place to Work® in India and is honored to be ranked the #1 Best Company to Work For® in the U.S. by Fortune magazine and Great Place to Work®. For more information, visit www.synchrony.com. Organizational Overview: The CloudOps SRE team ensures reliability, stability, and performance of cloud and hybrid platforms through monitoring, incident response, and automation. Working closely with engineering and infrastructure teams, it enables secure, scalable operations and supports ongoing technology modernization.

Role

Summary/Purpose: We are seeking an experienced Site Reliability Engineer (SRE) – CloudOps at the AVP level to join our Technology organization. The incumbent will be responsible for ensuring the reliability, scalability, performance, and operational excellence of mission-critical fintech platforms hosted across AWS, Pivotal Cloud Foundry (PCF/Tanzu), and Hybrid On-Premise environments. The role demands a strong engineering mindset with a passion for automation, observability, and continuous improvement, while partnering closely with Development, Architecture, and Infrastructure teams to embed reliability into every layer of the product lifecycle. Key Responsibilities : Reliability Engineering : Define, measure, and govern Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets across business-critical applications; drive engineering practices that uphold them. Automation & Toil Reduction : Identify operational toil and engineer scalable automation solutions using Cloud Migration & Modernization : Lead and contribute to migration of workloads from on-premise / PCF estates to AWS-native and containerized (Kubernetes/Docker) platforms, ensuring resilience and zero-business-impact cutovers. Observability : Build and enhance the observability stack across Prometheus, New Relic, Grafana, Splunk, CloudWatch and create actionable dashboards, golden signals, and intelligent alerting that reduce noise and accelerate triage. Incident & Change Management : Own major incident response, root cause analysis, and post-incident reviews; uphold ITIL-aligned change, problem, and release management disciplines. Required Skills/Knowledge: Minimum 5+ years of hands-on experience in SRE, CloudOps, DevOps, or Production Engineering roles, preferably within the Fintech / Banking / Financial Services domain with overall 7+ years of industry experience. Minimum 5+ years of expertise across Cloud Platforms: AWS (EC2, EKS, S3, RDS, IAM, CloudWatch, VPC), PCF / Tanzu, and Hybrid On-Premise environments. Containers & Orchestration: Kubernetes and Docker at production scale. Infrastructure as Code: Terraform (modules, state management, governance). CI/CD Tooling: Jenkins, GitHub Actions, and ArgoCD (GitOps). Observability: Prometheus, Grafana, Splunk, Dynatrace, and the ELK stack. Databases: Oracle, PostgreSQL, MongoDB, Cassandra (operations and performance). Practical knowledge of ITIL-aligned Incident, Problem, and Change Management. Proven experience driving SLO/SLI frameworks, error budgets, and reliability KPIs in a regulated enterprise. Demonstrated success in cloud migration, modernization, and FinOps-led cost optimization programs. Strong

AVP, Site Reliability Engineer - CloudOps (L10) at Synchrony Financial — Hyderabad, IN | Yoinka