yoinka

Senior Applied Research Engineer - Accelerator Programming Model and Compiler

NVIDIA

India, BengaluruSeniorH-1B sponsor company
Sign in to applyVerified 2h ago
Location
India, Bengaluru
Work model
On-Site
Level
Senior
H-1B history
394 approvals (FY2023)
Posted
23h ago

Skills

LinuxMachine Learning

About this role

As the world’s leading accelerated computing company, we are paving the way with innovations in self-driving cars, robotics, machine learning, super-computing, and visualization. We are building a team that will truly change the world and would love for you to join us. We are seeking a Sr Applied Research Engineer, Accelerator Programming Model and Compiler to work on the next generation programming model for NVIDIA’s Programmable Vision Accelerator (PVA) built for developers and AI coding agents. The PVA is a power-efficient, deterministic VLIW/SIMD processor used across NVIDIA DRIVE, Jetson, and IGX platforms. This role sits at the boundary between research and production systems software. You will explore and innovate new methods for expressing PVA workloads through higher-level declarative abstractions, domain-specific languages, and compiler/runtime interfaces. These interfaces will provide well-defined targets for both developer-written and AI-generated code. The successful candidate will be comfortable moving across levels of abstraction: shaping public APIs, debugging code generation deep in an LLVM-based backend, and researching higher-level programming models. They will also integrate the compiler with agentic coding harnesses, evaluate AI-generated code for correctness and performance, and improve the compiler-feedback loops to help agents produce optimized implementations.

What you will be doing

Work on the next-generation PVA programming model for developers and AI coding agents, making it easier to build optimized algorithms for PVA. Enable AI agent–driven PVA development by abstracting hardware-specific details into agent accessible, declarative interfaces, supported by improvements to compile time, emulator speed, and diagnostic tooling. Help define the architecture, and feature set of the PVA SDK, runtime APIs and programming model. Define benchmarks and evaluations for agent-generated PVA code and use them to improve the compiler optimizations and diagnostics. Develop and actively improve the LLVM-based VPU compiler backend targeting VLIW/SIMD architecture. Research and define efficient integration models for PVA workloads in CUDA-based heterogeneous pipelines, including execution, memory movement and synchronization. Drive deep technical integrations with internal and external customers to improve adoption and shape PVA runtime API for real-world workloads. What we need to see: BS, MS, or PhD in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience. 10+ years of experience building high-performance, low-level systems software, accelerator software, embedded software, or compiler/toolchain infrastructure. Experience building higher-level programming abstractions, DSLs or compiler IRs that abstract hardware complexity while preserving performance. Experience with compiler, debugger, linker, or toolchain development, particularly using LLVM. Experience developing code with agents — Claude Code, Cursor, Open AI and inference SDKs. Background integrating compilers and developer tools with AI coding agents or agentic development harnesses. Experience programming SIMD/VLIW processors. Excellent software development skills in C++, including low-level debugging and performance profiling. Experience with Linux or QNX development environments. Strong communication and social skills. Ways to stand out from the crowd: Experience with CUDA, especially integrating accelerators into a CUDA-based heterogeneous compute pipeline. Background with OpenCL, MLIR, Halide, TVM, Triton, graph compilers, image-processing DSLs, or other declarative/compiler-based programming systems. Familiarity with ISO 26262 and IEC 61508 or equivalent quality/safety standards. Experience building agent harnesses and orchestrating long-running, stateful, multi-agent workflows With competitive salaries and a generous benefits package, we are widely considered to be one of the

Senior Applied Research Engineer - Accelerator Programming Model and Compiler at NVIDIA — India, Bengaluru | Yoinka