yoinka

Applied Scientist II, Core Shopping Data Science

Amazon

US, WA, SeattleFull TimeMid
Sign in to applyVerified 1h ago
Location
US, WA, Seattle
Employment
Full Time
Work model
On-Site
Level
Mid
Posted
20h ago

Skills

LLM

About this role

Some CX shopping defects might be straightforward to detect and track. The interesting ones are not because they depend on what a customer perceives. For example, a search page may return legitimately different results, yet a shopper has no way to tell apart. We turn these ambiguous perception questions into a measurable artifact using LLMs, and we build the frameworks to know exactly where model judgment can be trusted and where a human must decide and proving it, against ground truth, at Amazon scale. This is one example of an LLM based measurement pipeline you will own, but that’s not all. You will extend that measurement to other parts of the shopping experience like the homepage and the detail page, where the same customer problem looks nothing like it does in search results, and where you will design the measurement from scratch. The larger goal is what makes this role unusual. Teams across Amazon are each independently figuring out how to label quality with LLMs, hitting the same problems alone: prompts that break on the next model version; golden sets nobody audited, accuracy that collapses in other locales. Through the work above, you will set the standard and build the production tooling behind it. Reusable labeling pipelines, evaluation frameworks, and inference infrastructure that hold up against Amazon-sized data and get adopted by teams who did not have to use them. Key job responsibilities - Design and improve LLM-based labeling for perception-driven defects: prompt design, sampling strategy, and the split between model judgment and human annotation. - Validate labeling quality against human ground truth, and build and maintain the golden datasets that make that validation possible. - Extend perceived-duplicate measurement to other parts of the shopping experience, designing the methodology where none exists and evaluating approaches already in use where one does. - Build reusable, production-grade labeling and evaluation tooling, batch inference, quality sampling, prompt and model version control, that operates on Amazon-scale data. - Define and publish the standards other teams adopt for using LLMs to measure customer experience. - Partner with science, engineering, and product teams across Stores to make quality measurement usable in their decisions, and present metric results and methodology changes to stakeholders. A day in the life We are a small team, which means your work is visible and your scope grows as fast as you do. There are existing partnerships with other teams and a paved way to cross org influence.

Applied Scientist II, Core Shopping Data Science at Amazon, US, WA, Seattle | Yoinka