Sr Applied Scientist, Core Shopping Data Science
Amazon
- Location
- US, WA, Seattle
- Employment
- Full Time
- Work model
- On-Site
- Level
- Senior
- Posted
- 17h ago
Skills
About this role
A customer may encounter a recommendation that is irrelevant, two widgets offering similar products, and confusing labels. Customers notice all of it. The ranking systems that chose those impressions largely do not, because until recently there was no way to turn a judgment about how a page feels into a signal a model could learn from. LLMs changed that. We can now take an ambiguous statement about customer perception, make it a judgment that holds up consistently at scale, and validate it against what shoppers actually do next. What we cannot do is call a large model in a ranking request path at large scale. Every shopping session passes through those systems under latency and cost budgets that leave no room for one. So we can describe quality far better than we can optimize for it, and closing that gap is what this role exists to do. You will distill LLM quality judgments into models compact enough to serve online, model how customers respond to the defects they encounter so we know which ones are worth trading engagement to prevent, and work with Search, Homepage and Detail Page ranking teams to get those signals into online objectives. Scientists on this team own the measurement side of the problem. You own the half that turns their judgments into systems that act and test them in controlled experiments. It is an unusual combination of problems: research-grade modeling with an unambiguous production bar, on surfaces where the change you ship is visible to nearly every Amazon customer. Key job responsibilities - You will model how customers respond to the defects they encounter. Using newly instrumented logging data, you will build representations of customer sessions from page sequences, intent signals, and quality exposures, and quantify what changes downstream when a customer meets an irrelevant recommendation, a set of near-duplicate widgets, or a confusing label early in a journey. You will identify the contextual factors that mediate that impact, including session intent, category, device, and prior interactions, and turn the results into a ranking of which defects are worth trading engagement to prevent, on which surfaces, for which customers. - You will distill LLM quality judgments into models compact enough to serve online. That means training compact models against LLM-generated labels, characterizing where the student diverges from its teacher and on which segments, and holding accuracy under the latency and cost budgets of Search, Homepage, and Detail Page ranking. Where a distilled model cannot meet that bar, you will say so early and propose what would. - You will work with Search, Homepage, and Detail Page ranking teams to get those signals into online objectives. You will analyze which of those systems offers the most leverage, recommend where to invest first, and design quality-aware objective formulations that trade impression quality against engagement deliberately, replacing the current pattern of suspending a strategy after a problem surfaces. - You will prove all of it in controlled experiments. You will design the experiment, choose the right success metrics metrics and make the call on what ships.
About the team
Core Shopping Data Science owns the measurement of shopping quality across Amazon’s Homepage, Search, and Detail Page experiences, including the company-level defect metrics reviewed by Amazon’s most senior leadership. We build the metrics, tools, and datasets that teams across Stores and Advertising use to decide what to ship. We are a small team, which means your work is visible and your scope grows as fast as you do.