yoinka

Lead HPC Software Optimization Engineer - C++

AMD

Hyderabad, IndiaFull TimeSenior
Sign in to applyVerified 4h ago
Location
Hyderabad, India
Employment
Full Time
Work model
On-Site
Level
Senior

Skills

AgileC++Machine LearningPyTorchPython

About this role

WHAT YOU DO AT AMD CHANGES EVERYTHING   At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.   Together, we advance your career.

THE TEAM

Join AMD’s high-impact team at the heart of innovation in AI, ML, and high-performance computing (HPC). We’re a collaborative group of software architects and GPU engineers focused on pushing the boundaries of AI model performance across distributed, GPU-accelerated platforms. Our work drives the next generation of AMD’s AI software stack, enabling large-scale machine learning training and inference workloads in data centers and enterprise environment   THE ROLE: AMD is hiring a Senior Software Developer for its AI, ML, and high-performance computing team to lead GPU kernel optimization and distributed software for large-scale AI workloads. In this technical leadership role, you'll architect and implement optimized compute kernels, design multi-GPU/multi-node scaling strategies, profile systems to maximize hardware utilization, build benchmarking infrastructure, and guide agile teams across the product lifecycle. The ideal candidate is a deep systems thinker fluent in GPU architecture, parallel computing, and AI model execution, comfortable both writing performance-critical code and driving software architecture decisions. Required expertise includes GPU kernel optimization in C++ (17/20), hands-on CUDA and low-level GPU programming, distributed AI computing (multi-GPU, NCCL, MPI), and familiarity with frameworks such as PyTorch, vLLM, Cutlass, and Kokkos, plus strong performance tuning with profiling tools (Nsight, VTune, Perf) and Python automation. You should also bring proven software leadership experience—defining roadmaps with stakeholders, interfacing with executives, and shipping production software through open-source upstreaming or commercial rollouts.   THE PERSON: We’re looking for a highly skilled, deep systems thinker who thrives in complex problem domains involving parallel computing, GPU architecture, and AI model execution. You are confident leading software architecture decisions and know how to translate business goals into robust, optimized software solutions. You’re just as comfortable writing performance-critical code as you are guiding agile development teams across product lifecycles. Ideal candidates have a strong balance of low-level programming, distributed systems knowledge, and leadership experience—paired with a passion for AI performance at scale.

KEY RESPONSIBILITIES

GPU Kernel Optimization : Develop and optimize GPU kernels to accelerate inference and training of large machine learning models while ensuring numerical accuracy and runtime efficiency. Multi-GPU and Multi-Node Scaling: Architect and implement strategies for distributed training/inference across multi-GPU/multi-node environments using model/data parallelism techniques. Performance Profiling: Identify bottlenecks and performance limitations using profiling tools; propose and implement optimizations to improve hardware utilization. Parallel Computing : Design and implement multi-threaded and synchronized compute techniques for scalable execution on modern GPU architectures. Benchmarking & Testing: Build robust benchmarking and validation infrastructure to assess performance, reliability, and scalability of deployed software. Documentation & Best Practices: Produce

Lead HPC Software Optimization Engineer - C++ at AMD, Hyderabad, India | Yoinka