Senior Storage Software Engineer - DGX Cloud
NVIDIA (Eightfold)
- Location
- US, CA, Santa Clara; Remote - US
- Work model
- Remote
- Level
- Senior
- Posted
- 23h ago
About this role
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. NVIDIA DGXC Storage team handles some of the fastest training and inference tasks. Every GPU cycle depends on a storage platform built to keep tens of thousands of accelerators continuously busy. It maintains exabytes of data securely and powers the largest AI workloads worldwide across cloud, neocloud, and on-prem setups. With the growth of accelerated computing, storage is essential. It can make the difference between effective GPU use and wasted potential, and between launching a frontier model on time or missing the deadline by months. We’re looking for a hands-on Storage Software Engineer to join the storage team as an individual contributor and technical lead. You will contribute to open-source parallel and distributed file systems and keep our largest GPU clusters fast, reliable, and durable. You will stay deeply hands-on: writing and reviewing production code, chasing root causes in the field, and setting the configuration and tuning standards our GPU fleets run on. This is a chance to do foundational storage engineering for the AI era at the company that introduced accelerated computing. What you’ll be doing: Contribute to open-source file systems. Contribute code to open-source parallel and distributed file systems, and distributed object storage. Upstream fixes and features, and engage directly with the upstream communities and maintainers. Serve as a hands-on storage software lead. Write and review production code yourself, and read kernel, NFS, NVMe-oF, or SPDK source when a bug requires it. Make the final technical calls on storage deliveries against measurable targets. Triage and troubleshoot at scale. Triage, troubleshoot, and root-cause large, complex storage issues across very large GPU clusters (tens of thousands of GPUs) — I/O and metadata performance, data corruption, and recovery. Validate architecture and capabilities. Validate storage architecture, capabilities, performance, and durability. Run scale tests, benchmarks, and recovery drills, and qualify new builds against measurable performance and durability targets. Recommend configuration, tuning, and guidelines. Define and recommend configuration, tuning, and operational best practices for high-performance file systems on GPU infrastructure, and help operators and internal customers apply them. Partner broadly. Work with training, inference, and accelerated-computing teams, site-reliability and operations, networking, and security, and collaborate with cloud providers, neocloud operators, and storage vendors on a common architecture. Work AI-first. Use modern AI coding and agentic tools day-to-day to accelerate building, debugging, validation, and operations. What we need to see: BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field — or equivalent experience. Over 12 years of direct experience in storage software engineering, including extensive involvement with a high-performance parallel or distributed file system handling multi-petabyte scale. Contributions to open-source projects involving a distributed or parallel file system. You are fully engaged in engineering tasks. You write and review production code, examine file system, kernel, NVMe-oF, or SPDK source to identify bugs, and personally conduct scale