Site Reliability Engineer Graduate (Global SRE- GMPT) - 2027 Start
TikTok
- Location
- Singapore, Singapore, Singapore
- Employment
- Full Time
- Work model
- On-Site
- Level
- New Grad
- H-1B history
- 148 approvals (FY2023)
About this role
Site Reliability Engineering (SRE) of Ads Infrastructure team works on building and running large-scale, globally distributed, fault-tolerant ads systems. SRE keeps ads systems at ByteDance up and running with the highest level of availability, ensuring our users have the best and fastest experience possible.
In the SRE team, you will build and improve multiple large scale ads services, while using your expertise in coding, performance optimization, troubleshooting, etc.
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth. Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume. Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.
Responsibilities - Own the reliability of ByteDance's global advertising systems, participate in the design of reliability architecture, and ensure high availability of large-scale advertising systems. - Design, implement, validate, and continuously optimize global disaster recovery (DR) solutions for the advertising systems, and lead disaster recovery preparedness and incident response. - Manage and plan infrastructure capacity for the advertising systems, ensuring resources scale efficiently with business growth while continuously improving infrastructure utilization through performance optimization - Develop reliability platforms and engineering tools for the advertising systems, including monitoring, alerting SLO management, risk governance, and resource management.