Data Scientist – AI Applications Engineer
Amgen
- Location
- India - Hyderabad
- Work model
- On-Site
- Level
- Mid
- H-1B history
- 137 approvals (FY2023)
- Posted
- 12h ago
Skills
About this role
Career Category Process Development Job Description Job Summary Transformative Digital Capabilities (TDC) drives Process Development (PD) and Operations digital transformation through pragmatic digital innovation that systematically delivers for the present while also building for the future. Our mission is to empower Process Development staff to harness their transformative potential through the rapid convergence of chemistry, biology, engineering, and computing, accelerated by innovative digital capabilities. TDC's core competencies span clear thought leadership and deep domain expertise in Data Products, Artificial Intelligence, Digital Twins, Product Management, and Organizational Change Management following Scaled Agile practices. As part of our team expansion at Amgen India (AIN), we are seeking a Level 4 AI Engineer to contribute to the development, evaluation, and deployment of AI-enabled workflow capabilities for Process Development. The successful candidate will work within defined workstreams to prepare scientific data, implement machine learning and generative AI components, test solution performance, and document results. This role will collaborate closely with senior data scientists and AI engineers, product owners, AIN colleagues, U.S. TDC partners, scientists, data strategy teams, and software/platform partners. The role requires strong hands-on execution, learning agility, clear communication, and effective collaboration across time zones.
What You Will Do
Contribute to AI-enabled use cases: work with senior data scientists, AI engineers, product owners, and scientific partners to turn defined requirements into practical data and AI workflow components for Process Development. Prepare data and build workflow components: clean, transform, and integrate structured and unstructured scientific data and metadata, and implement retrieval, automation, context, or human-in-the-loop components using established designs and standards. Develop and evaluate models and AI systems: implement machine learning, generative AI, retrieval-augmented generation, or decision-support approaches; prepare test data; run benchmarks; perform error analysis; and compare results against defined acceptance criteria. Contribute to reliable software delivery: develop maintainable Python code, notebooks, scripts, pipelines, or services using team practices for version control, testing, review, documentation, and reproducibility. Collaborate in agile product delivery: support requirements clarification, backlog refinement, implementation planning, demonstrations, user testing, feedback collection, and incremental improvement of AI-enabled capabilities. Document and communicate results: clearly describe data sources, methods, assumptions, model or agent behavior, limitations, test evidence, and appropriate use to technical and non-technical stakeholders.
Basic Qualifications
Master's or Bachelor's degree and minimum 4 years of related experience in engineering, computer science, data science, applied mathematics, statistics, life sciences, or a related technical field Required Skills/Experience Hands-on programming experience in Python for data analysis, modeling, automation, or application development. Experience preparing, cleaning, exploring, and integrating structured or unstructured data for analytical or machine learning use cases. Experience with one or more of machine learning, generative AI, retrieval-augmented generation, workflow automation, or decision-support systems. Experience applying standard evaluation methods, including train/test design, performance metrics, benchmark execution, error analysis, and quality checks. Familiarity with software engineering practices such as version control, code review, testing, documentation, and reproducible workflows. Experience with experimental data, scientific metadata, laboratory workflows, process development data, data pipelines, or data products. Understanding of biopharmaceutical