yoinka

Site Reliability Senior Staff

Synopsys

Posted 18-May-2026StaffH-1B sponsor company
Sign in to applyVerified 2h ago
Location
Posted 18-May-2026
Work model
On-Site
Level
Staff
H-1B history
112 approvals (FY2023)

Skills

GenAILLMLinuxOAuth

About this role

Descriptions & Requirements

Job Description and Requirements

Site Reliability Engineer, EDA Engineering Compute Job Description Synopsys is expanding Synopsys.ai alongside traditional EDA engineering compute for semiconductor customers in Vietnam. We are looking for an experienced Unix/Linux infrastructure professional to keep these platforms reliable, secure, and operable in Synopsys sites and at major customer environments (including customer-hosted, air-gapped, and high-security on-prem deployments). You will join the regional EDA Engineering Compute / platform operations team working with global IT, AE engineering, and product teams. Your work directly enables R&D and customer engineering teams to run EDA and GenAI workloads—from HPC batch grids through containerized AI gateways, vector data services, and agent/MCP-based tool execution on the grid. The successful candidate will support Synopsys Vietnam engineering compute and data center services, and partner on in-country deployment and sustainment of the related platform components for key semiconductor accounts. How This Role Contributes Improves availability, operability, and utilization of business-critical EDA and Synopsys.ai infrastructure in Vietnam. Reduces operational risk for customer air-gapped and customer-hosted deployments through strong runbooks, monitoring, and incident practices. Connects customer HPC schedulers, shared storage, GPU compute, and platform gateways so Synopsys tools and GenAI services run predictably at scale. Extends APAC shared-support coverage with Japan, Korea, Taiwan, and China peers.

Responsibilities

Engineering compute and data center Maintain Synopsys Vietnam engineering compute and data center environments per corporate IT, data center, and security standards. Support server hardware lifecycle activities: installation, provisioning, maintenance, upgrade planning, retirement, and decommissioning. perform capacity planning, performance troubleshooting, and vendor/customer coordination at customer sites. Operate and troubleshoot Linux-based EDA compute: virtualization, engineering resource management, job schedulers, license connectivity, and remote access services. Run HPC operations: cluster health, scheduler integration (LSF, Slurm, or similar), InfiniBand where deployed, and performance-related incidents. Collaborate with corporate network and information security on data center connectivity, switching, firewalls, routing, and circuits. Synopsys.ai platform operations (customer and Synopsys environments) Support deployment and day-2 operations for Synopsys EDA and AI reference patterns in Vietnam: customer-hosted and air-gapped container environments, API/AI gateways, LLM gateway integration (cloud or on-prem inference endpoints per customer policy), and customer-controlled vector database/GPU compute where applicable. Partner on install, upgrade, and configuration of platform components that connect user/agent workflows to the engineering grid—e.g., job services, unified MCP server patterns, and tool execution on LSF/Slurm-backed clusters (in coordination with products and global platform teams). Implement and maintain observability, alerting, authentication/authorization integration (e.g., customer IdP, OIDC/SSO/OAuth patterns), and operational documentation for platform services. Support agentic AI platform rollouts where components run locally, centrally, or on the grid; escalate design questions to global architecture and engineering teams. Reliability, improvement, and customer engagement Own or co-own runbooks, monitoring/alerting policies, and automation (scripts, IaC, or equivalent) to reduce toil across the APAC team. Apply AI-assisted workflows to improve troubleshooting, knowledge management, and routine operations. On-call and incident response: shared rotation (weekday nights, weekends, and holidays); lead or support incident bridges; produce post-incident reports in English. Customer air-gapped and on-site

Site Reliability Senior Staff at Synopsys, Posted 18-May-2026 | Yoinka