yoinka

Sr. Principal Data Engineer - Lakehouse Architecture

Eli Lilly

Indianapolis, Indiana, United States of AmericaFull TimePrincipal
Sign in to applyVerified 1h ago
Location
Indianapolis, Indiana, United States of America
Employment
Full Time
Work model
On-Site
Level
Principal
Posted
20h ago

Skills

AWSAirflowAzureCI/CDCloudFormationDatabricksDockerKafkaKubernetesPythonRedshiftSQLScalaSnowflakeSparkdbt

About this role

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.  Tech at Lilly is seeking a highly skilled Senior Data Engineer who can implement and optimize large-scale Lakehouse solutions and drive the evolution of our modern data platform while providing technical leadership to a growing team. The ideal candidate will have hands-on experience with modern data engineering technology stack and a proven track record of leading engineering talent in fast-paced environments .

What you will be doing

Design and implement comprehensive Lakehouse architecture solutions using technologies like Databricks and Snowflake platforms    Define and carry out medallion architecture standards (Bronze / Silver / Gold) across data domains, ensuring data quality, lineage, and discoverability   Lead Unity Catalog governance design: schemas, access control policies, and data contracts   Build and maintain real-time and batch data processing systems using Apache Spark (PySpark/Scala), Kafka, and Databricks Structured Streaming, and similar technologies   Architect scalable data pipelines that handle structured, semi-structured, and unstructured data to deliver AI ready data.   Develop data transformation workflows using tools like DBT, Airflow, or Databricks    Implement data governance frameworks, including data quality monitoring, lineage tracking, data time travel and security protocols.   Build data pipeline testing frameworks: unit tests, data quality assertions (Great Expectations / dbt tests), and schema validation   Define and publish data SLAs/SLOs in collaboration with data product owners; own incident response and root-cause analysis for pipeline failures   Drive adoption of modern data engineering standard processes including Infrastructure as Code, CI/CD, and automated testing   Collaborate with data scientists, analysts, and business collaborators to translate requirements into robust technical solutions   Mentor a team of 3-5 data engineers   Foster a collaborative team culture focused on continuous learning and innovation   How You Will Succeed   Proven ability to mentor junior engineers and facilitate knowledge sharing   Strong project management skills with experience leading multi-functional initiatives   Coordinate multi-functional projects and ensure effective communication between technical and business teams   Demonstrated ability to make architectural decisions and drive technical consensus   Embrace a growth mindset and actively seek opportunities to expand your leadership capabilities and technical mastery   What You Should Bring   Knowledge in the pharmaceutical o r life sciences domain   Experience with streaming data technologies ( Kafka ,Databricks Structured Streaming)   Familiarity with data cataloging tools   Familiarity with high performance data service framework ( Arrow Flight )   Expert-level proficiency in Python and SQL for data transformation and pipeline development   Strong experience with Apache Spark (PySpark or Scala) for big data processing and analytics   Hands-on experience with cloud platforms ( AWS or Azure ) and their data services, including S3, Glue, Redshift, IAM, and CloudWatch   Proficiency with Infrastructure as Code tools ( CloudFormation )   Experience with containerization ( Docker , Kubernetes ) and orchestration platforms   Knowledge of data modeling techniques for both

Sr. Principal Data Engineer - Lakehouse Architecture at Eli Lilly, Indianapolis, Indiana, United States of America | Yoinka