Systems Software Engineer - Infrastructure
NVIDIA
- Location
- US, CA, Santa Clara
- Work model
- On-Site
- Level
- Mid
- H-1B history
- 394 approvals (FY2023)
- Posted
- 1d ago
Skills
About this role
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. Do you derive more satisfaction from eliminating a manual process than from completing one? Are you the resident tool-builder and creative solution specialist on your team? We're looking for the engineer who invents, crafts and overall automates everything they touch. AI, scripts, workflow redesigns, documentation - it's all at your disposal to empower teams making the world's best firmware for the world's best GPUs. Join us, the GPU Firmware Infrastructure team, whose mission is crafting extraordinary automation, tools, and practices that accelerate NVIDIA's GPU innovation. We're responsible for the substrate: dev tools, build compute, test farms, artifact and release pipelines, secure signing services, and the observability that proves it all works. When our platforms are fast and trustworthy, every product ships faster! This is your chance to make waves in the industry while working alongside some of the most top-valued diverse set of minds in the business, building the best AI Supercomputers. If you're up for the task, we'd like to hear from you!
What you'll be doing
Build, supervise, and improve the core infrastructure our firmware teams run on: build systems, regression farms, CI/CD pipelines, developer tooling, and the frameworks and web services that sit on top of them Debug and resolve issues across hardware, software, infrastructure, and team processes - solving not just the present problem, but for that entire class of issue in the future Invent and build AI-powered workflows that scale to multiply your own, your team's, or the whole company's output Learn, document, and automate processes, services, and tooling for teams via many internal-facing and external-facing projects Seek out, identify and eradicate toil company-wide by growing your ideas for improvements from a proof-of-concept to fully shipped and scalable solutions What we need to see: BS or MS degree in EE/CS/CE (or equivalent experience) 5+ years of software, infrastructure, DevOps, or SRE engineering Hands-on experience with modern CI/CD and test automation: architecting and debugging pipelines, working inside the frameworks that orchestrate the tests (pytest or equivalent) History in, or curiosity about device BIOS, firmware, or other low-level embedded software, and interest in building tooling for the engineers who write it Solid understanding of the services and data behind infrastructure: web services, REST APIs, relational databases and schema design (SQL, PostgreSQL) Scalability thinking Strong Python and comfort working in Linux shell environments Strong communication skills: ability to articulate the sharp questions, write clear requirements, review critically, and explain trade-offs Exceptional interpersonal and empathy skills, and the curiosity to match as you’ll work closely with both hardware and software engineers to design, develop, and debug process and tooling enhancements Ways to stand out from the crowd: Sense of humor heavily encouraged, but not required A side project, open source contribution, or internal tool you built because something annoyed you enough Track record of writing clear technical proposals, design docs, or architecture decisions that others have acted on independently Experience with containers, orchestration,