Senior System Integration Engineer, Memory
NVIDIA
- Location
- US CA Santa Clara
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 394 approvals (FY2023)
- Posted
- 19h ago
Skills
About this role
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology and amazing people. Today, we're harnessing the boundless possibilities of AI to build the next era of computing. An era in which our GPU acts as the brain of computers, robots, and self-driving cars that can understand the world. Accomplishing unprecedented goals calls for imagination, inventiveness, and exceptional talent from around the world. As a NVIDIAN, you'll be immersed in a diverse, encouraging environment where everyone is inspired to do their best work. Join our team and discover how you can build a lasting impact on the world. NVIDIA's Silicon Co-Design Group (SCG) leads the full product development lifecycle, from early architecture definition through silicon bringup to product release. The ArchDev team is the hub for silicon and system-level feature development, driving tradeoff analysis, system integration, and POR alignment across the entire organization. This is where ideas become chips, and chips become products that define the state of the art — and we're building that future with some of the most motivated engineers in the industry. We're looking for a Senior Memory Systems Engineer to own HBM and LPDDR integration in sophisticated SoCs. This role covers the full stack, including silicon, package, embedded software, testing, and product development. The engineer will resolve the toughest system-level memory challenges throughout the process, building solutions that hold up at scale.
What you'll be doing
HBM & LPDDR System Integration and Bringup: Drive HBM and LPDDR system integration, bringup, characterization, and debug for next-generation SoCs — taking memory subsystems through the full arc from first silicon to production-ready at scale. Full-Stack Memory Closure: Own memory performance, power management, thermal, and reliability closure across silicon, package, board, and firmware — translating characterization results into signed-off operating points and release criteria that the full program depends on. Post-Silicon Margining, VF Shmoo & Correlation: Lead post-silicon margining, VF shmoo, eye, and correlation work across voltage, temperature, and frequency — building the characterization foundation that underpins every product decision downstream. Memory Debug: Training, Calibration, SI/PI & Stability: Debug and resolve memory training, calibration, SI/PI, and system-level stability issues — the class of problems that sit at the intersection of electrical, physical, and software behavior and demand deep cross-domain expertise to resolve at scale. Validation & Characterization Planning: Define validation, characterization, and issue-tracking plans across chip programs — building the framework that ensures the right tests exist, the right data gets collected, and issues are tracked with enough fidelity to close. What we need to see: BS or MS in EE/CE — or equivalent experience. 12+ years in HBM, LPDDR, or high-speed memory systems, with hands-on depth in silicon bringup, characterization, and debug at scale. PHY controller development experience is a significant plus and will set you apart. Deep command of SI/PI, timing, margining, and memory training behavior — and the multi-functional fluency to work across design, package, firmware, validation, and product teams without losing the thread. HBM PHY and controller architecture familiarity is useful context, but what this role actually demands is rarer: the system-level instinct to see where memory behavior is about to become a product problem — and the inventiveness to resolve it before it does. Ways to stand out from the crowd: AI-Accelerated Debug & Root Cause Analysis: Use LLM-assisted tools and ML models to pattern-match failure signatures across VF shmoo, eye diagrams, and margining datasets — converging on root cause and outlier detection faster