yoinka

Site Reliability Engineer

Procter & Gamble

Taguig CitySenior
Sign in to applyVerified 1h ago
Location
Taguig City
Work model
On-Site
Level
Senior
Posted
9h ago

Skills

AWSAzureDockerGCPGrafanaKubernetesLinuxPrometheusPythonSQLTerraform

About this role

Job Location Taguig City Job Description Information Technology (IT) at Procter & Gamble is where business, innovation and technology integrate to build a competitive advantage for P&G. Our mission is clear -- you deliver IT to help P&G win with consumers.     Do you love implementing continuous improvement in IT solutions to drive efficiency and agility in meeting constantly evolving business needs? Then this job might be for you!     As a Site Reliability Engineer, you will be instrumental in ensuring the high availability and reliability of our digital IT products in   P&G .    Your primary focus will be on enhancing system performance through faster detection, response, and resolution of issues, while also implementing strategies to prevent recurrence and reduce operational toil. You will use robust Observability and Monitoring tools, automate incident response systems, and optimize IT architecture to create a resilient and reliable infrastructure.   This is a Managerial position. Being a manager at P&G involves leading teams and / or end-to-end processes, managing P&G resources, and driving business results. Managers are responsible for overseeing various aspects of the business, including strategy, operations, and team performance. They play a crucial role in ensuring that P&G's brands continue to grow and succeed in the market. Managers at P&G are expected to have strong leadership skills, a growth mindset, and the ability to make data-driven decisions roles lead and initiatives, significantly impacting business results through independent judgment and minimal guidance.

Responsibilities

Implement and lead comprehensive monitoring solutions and tools to provide real-time insights into system performance, enabling proactive incident detection and ensuring accurate, actionable alerts for prompt responses.   Continuously refine monitoring strategies and develop automation scripts to address recurring issues, enhancing system visibility, resource optimization, and overall efficiency.   Establish and maintain Service Level Indicators (SLIs) and Service Level Objectives (SLOs) to improve service quality and reliability,   Collect and share data and insights from observability tools to drive continuous improvement initiatives.   Work closely with Software Engineers, Product Teams, and Infrastructure Teams to develop and implement initiatives that enhance IT reliability.   Engage with customers to understand their needs and difficulties regarding Observability and Monitoring tools, providing exceptional support in all interactions, including communications, updates, and feedback.   Stay updated on industry trends and effective strategies in Site Reliability Engineering while continuously enhancing technical skills in system architecture, automation, cloud technologies, and operational processes.   Job Qualifications Candidates must demonstrate strong leadership in the application of technical expertise to drive business results.   We are looking for candidates who possess the following core qualities:   A Bachelor's degree in related field such as Engineering, Information Technology and Computer Science discipline, and up to 5 years experience at most.   Experience or familiarity with monitoring and observability tools (e.g., Prometheus, preferably Grafana)   Knowledge and familiarity in system administration, including Linux/Unix environments, cloud platforms ( Azure or GCP preferred, but AWS is acceptable )   Experience with configuration management tools and infrastructure-as-code frameworks (e.g., Terraform)   Proficiency in at least one programming language (e.g., Python, C#) and a background in scripting for automation tasks   Understanding of networking protocols, network infrastructures, load balancing, and DNS management   Familiarity with containerization and Orchestration Technologies (e.g., Docker, Kubernetes)   Familiarity with databases and proficiency in writing SQL queries

Site Reliability Engineer at Procter & Gamble, Taguig City | Yoinka