AI Infrastructure Engineer Graduate (TikTok Recommendation Architecture) - 2027 Start
TikTok
- Location
- Singapore, Singapore, Singapore
- Employment
- Full Time
- Work model
- On-Site
- Level
- New Grad
- H-1B history
- 148 approvals (FY2023)
Skills
About this role
Team Introduction Our team develops the core training and serving infrastructure that powers one of the world's largest recommendation systems, enabling billions of personalized recommendations every day. We are also advancing the next generation of AI infrastructure for foundation models and LLMs, driving innovation in large-scale model training, online inference, and GPU optimization. As part of the team, you will work on distributed training and inference systems, high-performance GPU computing, and scalable LLM infrastructure. You'll collaborate closely with experienced engineers and researchers to transform cutting-edge AI technologies into production systems that directly impact the experience of hundreds of millions of TikTok users. This role is ideal for candidates who are passionate about LLM systems, distributed computing, GPU programming, and building AI systems at massive scale. We are looking for passionate New Graduates to join our Model Infrastructure team, building the next generation of infrastructure for TikTok's For You recommendation system and Large Language Models (LLMs).
We are looking for talented individuals to join our team. As a graduate, you will get opportunities to pursue bold ideas, tackle complex challenges, and unlock limitless growth.
Successful candidates must be able to commit to an onboarding date by the end of the year. Please state your availability and graduation date clearly in your resume.
Candidates can apply to a maximum of two positions and will be considered for jobs in the order you apply. The application limit is applicable to our Company and its affiliates' jobs globally. Applications will be reviewed on a rolling basis - we encourage you to apply early.
Responsibilities - Build and optimize infrastructure for large-scale model training and online inference. - Develop distributed systems supporting large recommendation models and LLMs. - Improve training and inference performance through GPU optimization and efficient communication. - Collaborate with researchers to develop and deploy LLM training and serving solutions. - Analyze system bottlenecks and implement performance optimizations.