Senior Associate/Manager - Applied AI Engineer, Technology Consulting
EY
- Location
- SG, 048583
- Work model
- On-Site
- Level
- Mid
Skills
About this role
At EY, we develop you with future-focused skills and equip you with world-class experiences. We empower you in a flexible environment, and fuel you and your extraordinary talents in a diverse and inclusive culture of globally connected teams. We work together across our full spectrum of services and skills powered by technology and AI, so that business, people and the planet can thrive together. We’re all in, are you? Join EY and shape your future with confidence. We’re hiring an Applied AI Engineer to own the AI stack end-to-end — from the cloud infrastructure it runs on, through the model layer, to the systems and evaluations that keep it working in production. This is a builder role that flexes with the situation: some weeks you’re heads-down shipping as an individual contributor, others you’re setting architectural direction and leading a small team. We want someone who can do both well and read which the moment calls for. You’ll work across the full model landscape. That means getting the most out of frontier APIs (Claude, GPT, and their peers) and deploying, fine-tuning, and operating open-source models when that’s the right call — on cost, latency, control, or privacy. Knowing which to reach for, and why, is a core part of the job. We’re looking for someone who can apply first principles about why a transformer produces the output it does, drive real results out of frontier models, run open source models in production, and design the cloud infrastructure to serve them reliably and cost-effectively at scale. The opportunity:
Own the architecture of production AI systems — inference stacks, fine-tuning pipelines, retrieval and evaluation infrastructure, monitoring Build on frontier models (Claude, GPT, and peers) with real rigor: tool use, structured outputs, context and cost management, evals, and guardrails — not just prompt-and pray Deploy and operate open-source models (Llama, Qwen, Mistral, DeepSeek, and whatever comes next) on our cloud environment — including quantization, serving frameworks (vLLM, TGI, SGLang, TensorRT-LLM), and multi-GPU inference. Make the frontier-vs-open-source call deliberately, on cost, latency, control, and data sensitivity grounds — and be able to defend it Design the cloud infrastructure underneath it all: GPU orchestration, autoscaling, cost controls, VPC/networking, IAM, observability. This is not a “hand it to DevOps” role Fine-tune, distill, and evaluate models against real task metrics — not vibes, not leaderboards. Track the current research literature (arXiv, major labs, key conferences) and make judgment calls on what’s ready for production versus what’s still noise Depending on the project, contribute independently as a senior IC or lead and mentor a small team — and switch between the two as needed Partner with product and leadership to translate ambiguous problems into systems that actually ship
To qualify for the role, you must have: Cloud & infrastructure engineering
6+ years of software / infrastructure engineering, with deep production experience on at least one major cloud (AWS, GCP, or Azure). Strong command of GPU infrastructure: instance selection, driver / CUDA stack, containerization, Kubernetes or an equivalent orchestrator, autoscaling patterns for inference workloads. IaC discipline (Terraform / Pulumi / CDK), CI/CD, monitoring (Prometheus / Grafana / OpenTelemetry), and cost management as second nature. Fluent in Python; comfortable in at least one systems-adjacent language (Go, Rust, C++) for the parts that need it.
ML / AI depth
Genuine understanding of transformer internals: attention (including variants like MQA, GQA, sliding-window, and flash attention), positional encodings (RoPE, ALiBi), tokenization, KV cache mechanics, sampling, and where each part contributes to latency, memory, and quality Proven ability to get production-grade results from frontier models —