GPU Cluster Architect
Nebius
- Location
- India; Singapore
- Work model
- On-Site
- Level
- Senior
- Posted
- 2h ago
Skills
About this role
About Nebius
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The role
We are seeking a GPU Cluster Architect to drive the design of our next-generation AI infrastructure. In this high-impact, hands-on role, you will make end-to-end architectural decisions across compute, networking, and storage — ensuring our platforms can meet the massive scale, performance, and reliability requirements of modern AI workloads.
This is a high-impact, hands-on architecture role where you’ll define how tens of thousands of GPUs are interconnected, cooled down, powered, and optimized across multiple data center sites.
Your responsibilities will include
• Cluster Design: Architect scalable GPU cluster topologies including compute nodes, interconnect (InfiniBand, Ethernet), storage, and control planes.
• Performance Modeling: Analyze AI/ML workloads (e.g. LLM training, inference) to inform design tradeoffs across latency, bandwidth, and GPU density.
• Network Architecture: Align with network architect relevant design and validate low-latency, high-throughput interconnects (e.g., InfiniBand HDR/NDR, RoCEv2) at POD and DC scale.
• Storage Integration: Work with storage teams to optimize performance for training datasets, checkpointing, and others.
• Reliability & Monitoring: Understand and analyze signal from monitoring systems to the detect flows in design
• Collaboration: Partner with site reliability, networking, storage, and DC engineering teams to operationalize and scale your architecture.
We expect you to have
• 5+ years of experience designing clusters.
• Deep understanding of modern GPU architecture (NVIDIA, AMD, etc.).
• Experience with HPC interconnects (InfiniBand & RoCE).
• Solid background in systems architecture, networking, and hardware reliability.
• Experience in scripting for