yoinka

Principal SoC Performance Architect-Microbenchmarks

AMD

Austin, TexasFull TimePrincipal
Sign in to applyVerified 1h ago
Location
Austin, Texas
Employment
Full Time
Work model
On-Site
Level
Principal

Skills

Python

About this role

WHAT YOU DO AT AMD CHANGES EVERYTHING   At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.   Together, we advance your career.

THE ROLE

AMD is looking for an outstanding technical contributor to drive performance analysis, characterization, and optimization of next-generation Data Center GPU (DCGPU) platforms. This role focuses on extracting maximum performance across the full system stack—including hardware, firmware, drivers, runtime, libraries, and workloads—through deep architectural understanding and data-driven methodologies. The engineer will develop and maintain microbenchmarks and system-level workloads spanning pre-silicon and post-silicon environments to enable performance validation, debug, and optimization. THE PERSON:     As a passionate and technically strong Principal SoC Performance Engineer, you will work on highly parallel SoC architectures, leveraging deep understanding of GPU compute, memory hierarchy, and interconnects to analyze and optimize performance across AI and HPC workloads. You will be responsible for building microbenchmark suites and workload-driven analysis frameworks that expose performance characteristics of GPU subsystems (compute, memory, IO, interconnect, collectives) and ensure continuity across pre-silicon models, emulation, and post-silicon systems. The ideal candidate combines strong hardware/software co-design expertise with hands-on experience in performance analysis, profiling, and system-level debugging. You are expected to identify bottlenecks across the entire stack—from kernels to runtime to hardware—and translate insights into actionable improvements for both current and future architectures. You thrive in a fast-paced environment, are highly data-driven, and have a deep curiosity for understanding “why” performance behaves the way it does.

KEY RESPONSIBILITIES

Performance Analysis & Optimization Analyze and optimize performance of DCGPU systems across AI training, inference, and HPC workloads Identify bottlenecks across hardware, firmware, drivers, runtime, libraries, and applications Perform deep kernel-level and system-level profiling to understand performance behavior Provide actionable insights to architecture, software, and design teams to improve performance Microbenchmark & Workload Development Design and develop targeted microbenchmarks to characterize GPU subsystems (compute, memory, interconnect, collectives) Build representative system-level workloads reflecting real-world AI/HPC use cases Ensure microbenchmarks correlate to application-level performance and architectural intent Maintain and evolve benchmark suites across multiple GPU generations Pre-Silicon & Post-Silicon Continuity Enable performance validation in pre-silicon environments (simulation/emulation/models) Correlate performance data across pre-silicon models and post-silicon measurements Develop methodologies to reuse workloads and microbenchmarks across the full lifecycle Support bring-up and early silicon performance characterization Full-Stack Performance Engineering Work across the entire software stack: compiler, runtime, libraries, drivers, and firmware Collaborate with ROCm / AI frameworks / kernel teams to improve performance Analyze interactions between workload characteristics and hardware execution Optimize key kernels (e.g., GEMMs, collectives,

Principal SoC Performance Architect-Microbenchmarks at AMD, Austin, Texas | Yoinka