Senior Network Software Validation Engineer – AI Clusters
NVIDIA (Eightfold)
- Location
- Remote - Switzerland; Switzerland, Zurich
- Work model
- Remote
- Level
- Senior
- Posted
- 19h ago
Skills
About this role
NVIDIA is the world leader in accelerated computing, powering the AI revolution across the world's largest data centers, clouds, and supercomputers. Within NVIDIA, the Networking Business Unit (NBU) builds the high-speed interconnect—including Ethernet, InfiniBand, NVLink, and BlueField DPUs—that connects thousands of GPUs into a single AI supercomputer, moving data at the scale and speed required by the world's most demanding AI workloads. NVIDIA is looking for a Senior Network Software Validation Engineer to join our Network Cluster Solutions (NCS) Validation group. You will work on validating advanced networking solutions across complex NVIDIA AI cluster environments. The group is a high-performance engineering force that treats validation as a first-class software problem. We build systems, frameworks, and benchmarks that prove the correctness, reliability, and performance of NVIDIA networking solutions at scale. This role combines the development of validation methodologies and automation tools with hands-on debugging, performance analysis, and investigation of cutting-edge AI networking technologies at scale.
What You'll be Doing
Design validation methodologies, review product requirements, and develop comprehensive test plans for networking technologies in large-scale AI cluster solutions. Develop and maintain automation tools and scripts for test execution, environment setup, log collection, and data analysis. Execute functional, regression, performance, scale, and reliability testing across networking and AI cluster solutions. Own end-to-end investigation of complex customer issues by reproducing real-world scenarios, analyzing logs, telemetry, packet captures, and system metrics to identify functional issues and performance bottlenecks, triaging problems across the hardware and software stack, and driving them to root cause and resolution. Read and understand source code (C/C++/Python) to investigate defects, validate fixes, and improve logging, instrumentation, and debugging capabilities. Collaborate closely with software and hardware development teams to debug networking technologies, including NCCL, RoCE, RDMA, and related software components using targeted experiments and code inspection. Profile and benchmark AI training and inference workloads, correlating application behavior with network and system telemetry to identify scalability and performance limitations. Document findings, communicate technical results, and continuously improve validation methodologies, automation, and engineering processes. What we need to see: B.Sc. or M.Sc. in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience. 8+ years of hands-on experience in networking, system validation, software testing, or system-level debugging on Linux. Strong Linux systems knowledge and networking debugging experience in scale. Proven experience debugging complex production systems by forming hypotheses, designing experiments, and driving issues to root cause. Ability to read and reason about source code and collaborate closely with development teams to validate fixes. Strong scripting and automation experience using Python, Bash, and/or Ansible. Excellent analytical, troubleshooting, communication, and collaboration skills with a strong sense of ownership. Fast learner with curiosity for AI infrastructure, distributed systems, and modern AI-assisted engineering tools. Ways to Stand Out from the Crowd: Strong knowledge of AI networking technologies including NCCL, RoCE, and RDMA, with experience debugging both correctness and performance issues at scale. Deep understanding of congestion control and lossless Ethernet technologies for AI workloads, including DCQCN, ECN, and PFC. Experience debugging issues spanning multiple layers, including networking (L2/L3), transport protocols, operating systems, AI frameworks, and distributed applications. NVIDIA is widely considered to be one of