Machine Learning MLOps Engineer Intern (Global SRE) - 2027 Start
TikTok
- Location
- Singapore, Singapore, Singapore
- Employment
- Internship
- Work model
- On-Site
- Level
- Intern
- H-1B history
- 148 approvals (FY2023)
Skills
About this role
Team Introduction MLOps - Global SRE team is responsible for the reliability of the machine learning infrastructure powering TikTok’s global advertising systems. We ensure the stability, scalability, and operational efficiency of the entire machine learning lifecycle, including data pipelines, model development, training, deployment, inference, monitoring, and continuous optimization.
We are looking for talented individuals to join us for an internship. Our internship program offers students hands-on experience, industry exposure, and opportunities to apply their knowledge to real-world challenges while building a strong foundation for personal and professional growth. Interns will gain practical experience, explore potential career paths, and participate in social events, learning programs, and development workshops alongside industry professionals. Candidates may apply to a maximum of two positions across Our Company and its affiliates globally. Applications will be considered in the order they are submitted. Applications are reviewed on a rolling basis, so we encourage you to apply early. Please clearly state your availability in your resume, including your start and end dates.
Responsibilities - Define and drive Service Level Objectives (SLOs) for online machine learning inference systems, ensuring the reliability, availability, and performance of large-scale production inference services; - Ensure the reliability and operational excellence of offline machine learning training pipelines, continuously improving training job success rates; - Drive infrastructure capacity planning and resource management for machine learning workloads, ensuring compute resources meet evolving business demands while continuously improving GPU and CPU utilization through performance optimization; - Lead the enablement of new machine learning frameworks and GPU platforms, driving large-scale production deployment while maintaining model quality and business performance; - Design, build, and maintain MLOps platforms and automation tools, including quota management, job diagnostics, monitoring and observability, resource management, and operational tooling.