yoinka

Manager - AI Inference Engineer

EY

Atlanta, GA, US, 30308 +80 more…Mid
Sign in to applyVerified 1h ago
Location
Atlanta, GA, US, 30308 +80 more…
Work model
On-Site
Level
Mid

Skills

Kubernetes

About this role

Location: Anywhere in Country   At EY, we’re all in to shape your future with confidence.    We’ll help you succeed in a globally connected powerhouse of diverse teams and take your career wherever you want it to go.  Join EY and help to build a better working world.

The opportunity We are seeking an experienced AI Inference Engineer to help design, deploy, and operate private AI inference infrastructure at rack scale . This role is focused on the engineering of a production-grade inference platform: architecture, serving stack design, benchmarking, integration, reliability, and operational readiness. The ideal candidate has deep hands-on experience building AI inference systems and the platform layers that support them, including multi-GPU inference, Kubernetes-based deployment, autoscaling, observability, routing, access control, and production support.   Your key responsibilities Private AI inference architecture

Design and deploy secure private inference architectures for enterprise AI inference. Translate workload requirements into technical architecture decisions across compute, caching, interconnect, storage & networking. Define reference architectures that support high-throughput, low-latency, enterprise inference at scale.

Inference platform engineering

Design the model-serving stack, including inference runtime, orchestration, model registry, artifact management, API layer, routing, observability, deployment and fallback. Evaluate and recommend inference technologies and platform components for production use. Configure and harden the target environment for secure, reliable inference workloads.

Benchmarking and performance evaluation

Build and run benchmark suites for enterprise workloads. Measure latency, throughput, concurrency, utilization, reliability, and cost efficiency. Tune serving configurations and system parameters to improve production performance. Produce evaluation results and recommendations based on objective testing.

Enterprise integration

Support integration of inference platforms with client IT systems. Work with infrastructure and security stakeholders to validate technical requirements and remediate issues.

Model onboarding and operationalization

Onboard models into the target inference platform. Configure serving endpoints, routing behavior, monitoring, and operational controls. Support deployment readiness, runbooks, and operational handoff for production use.

Skills and attributes for success

Ability to turn business and workload needs into concrete infrastructure and software design decisions. Clear technical communication and strong architecture documentation skills.

To qualify you must have

Bachelor’s Degree in a relevant field

Strong experience designing and operating AI inference systems in production. Deep familiarity with multi-GPU serving and high-performance deployment patterns. Hands-on experience with Kubernetes and cloud-native platform technologies. Strong understanding of scaling, routing, autoscaling, observability, and reliability for production AI workloads. Experience integrating AI platforms into enterprise security and access environments.

Ideally, you'll also have

Experience with private, sovereign, or enterprise AI infrastructure.

Experience benchmarking large-model inference on GPU clusters. Familiarity with service mesh, policy enforcement, secrets management, and RBAC. Experience working across infrastructure, security, and application teams in a regulated environment. Experience supporting pilot-to-production AI platform rollouts.

What we offer you At EY, we’ll develop you with future-focused skills and equip you with world-class experiences. We’ll empower you in a flexible environment, and fuel you and your extraordinary talents in a diverse and inclusive culture of globally connected teams. Learn  more .

We offer a comprehensive compensation and benefits package where you’ll

Manager - AI Inference Engineer at EY, Atlanta, GA, US, 30308 +80 more… | Yoinka