yoinka

Senior GPU Inference Performance Engineer

AMD

Santa Clara, CaliforniaFull TimeSenior
Sign in to applyVerified 1h ago
Location
Santa Clara, California
Employment
Full Time
Work model
On-Site
Level
Senior

Skills

DockerKubernetesLLMLinuxPython

About this role

WHAT YOU DO AT AMD CHANGES EVERYTHING   At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.   Together, we advance your career.

THE ROLE

We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose, and explain performance across the full stack from GPU silicon, communication libraries, networking fabrics, and operating systems through the software runtime and drive competitive positioning against other accelerator vendors. This role sits at the intersection of hardware, systems software, networking, and AI infrastructure, and requires someone who can go deep on a trace and present findings to product and executive stakeholders.   THE PERSON: A hands-on performance engineer who is equally comfortable reading a GPU trace, debugging distributed systems performance issues, and briefing executives. You are curious, evidence-driven, rigorous, and you don't stop at "X is faster" and you explain why, rooted in hardware and software evidence. You collaborate across hardware, systems software, networking, and AI infrastructure teams, communicate clearly in written reports and presentations, and thrive at the intersection of silicon, operating systems, communication libraries, networking, and AI. Experience with Linux systems, distributed GPU infrastructure, RDMA/RoCE networking, or communication libraries such as NCCL/RCCL is highly valued.

KEY RESPONSIBILITIES

Full-stack GPU profiling: Instrument and analyze inference workloads across AMD Instinct (ROCm, rocProfiler, ROCm Systems Profiler, RGP, rocprof-compute, rocprof-sys, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute, DCGM) GPUs. Identify bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, PCIe/Infinity Fabric data movement, and GPU runtime behavior. Systems and runtime performance analysis: Profile and diagnose performance interactions between GPU runtimes, Linux operating systems, device drivers, container runtimes, memory subsystems, CPU scheduling, NUMA topology, and I/O pathways. Identify system-level bottlenecks that impact throughput, latency, and GPU utilization. Competitive performance analysis: Design and execute head-to-head benchmarks (AMD vs. NVIDIA) on standardized AI and LLM workloads. Produce clear, data-backed explanations of why performance differs attributing gaps to hardware architecture, networking topology, communication libraries, software maturity, runtime behavior, or configuration differences. Multi-server inference networking: Profile and optimize distributed inference topologies including prefill-decode (PD) disaggregation, pipeline parallelism, and tensor parallelism across multi-node clusters. Analyze network-level bottlenecks using RDMA/RoCE traces, NCCL/RCCL collective profiling, GPUDirect RDMA, NIC-level counters (Pensando, ConnectX), and network performance tools. Quantify the impact of latency, bandwidth, congestion, and topology on end-to-end inference SLAs. GPU operator and Kubernetes stack: Profile the overhead introduced by GPU operators, device plugins, container runtimes (Docker, containerd), and Kubernetes scheduling on inference latency. Identify and resolve jitter, cold-start, resource contention, and infrastructure inefficiencies in production environments. Tooling

Senior GPU Inference Performance Engineer at AMD, Santa Clara, California | Yoinka