GenAI Software Development Architect
AMD
- Location
- Santa Clara, California
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
Skills
About this role
WHAT YOU DO AT AMD CHANGES EVERYTHING At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond. Together, we advance your career.
THE ROLE
We are building an AI-native hardware and firmware validation platform from the ground up — one where LLMs, RAG pipelines, autonomous agents, and knowledge graphs are the core of how the system works, not an add-on. As the Software Development Architect, you will own the end-to-end technical design of this platform: multi-agent orchestration, retrieval-augmented knowledge systems, MCP server infrastructure, and the engineering standards that make all of it reliable at scale. This role sits within the Global Cluster Engineering organization, where you will develop software that powers distributed infrastructure at global scale. You will work closely with validation engineers, hardware teams, and leadership to translate domain requirements into a production-grade AI-native system. This is a hands-on role — you will write code, drive technology decisions, and directly mentor engineers. THE PERSON: Experience: software development experience, with at least 4 years in architecture, staff, or principal engineer role AI-Native Systems: Deep, hands-on experience designing and shipping production AI-native systems — not just LLM API integration, but the full stack: RAG pipelines, agent orchestration, tool use, multi-agent coordination, and LLM evaluation LLM Fundamentals: Strong understanding of how LLMs work in practice — context windows, grounding, hallucination failure modes, prompt engineering, model selection, and how behavior changes across providers and versions Retrieval Systems: Proven experience with vector search, embedding models, hybrid retrieval, reranking pipelines, and knowledge graph-augmented RAG Core Skills: Strong proficiency in one or more modern programming languages such as Python, TypeScript/Node.js, Go, Java, C#, or Rust, with demonstrated ability to build and operate production-scale services. Python experience is preferred due to the AI/ML ecosystem Engineering excellence: Async programming, API design, distributed systems, clean code practices. Experience designing for reliability in automated/unattended environments — crash recovery, audit trails, state management, observability. Strong written communication — architecture docs, design specs, and engineering standards that outlast your tenure. Track record of setting engineering standards that teams follow Hardware Affinity: Experience working closely with hardware teams — servers, networking equipment, or compute infrastructure — with an understanding of how software interacts with physical systems Cloud Infrastructure: Experience with AWS, Azure, or GCP — infrastructure provisioning, managed services, networking, and deploying production workloads at scale AI Tooling: Demonstrated use of AI coding assistants and LLM-powered developer tools (Claude Code, GitHub Copilot, Cursor, etc.) to accelerate design, development, and documentation KEY RESPONSIBILITIES: Platform Architecture: Design and own the architecture of an AI-native validation platform where autonomous LLM agents plan, execute, and analyze hardware and firmware test campaigns end-to-end — without a human in the loop RAG System Design: Architect the full retrieval-augmented generation stack — document