Senior Data Engineer
UnitedHealth Group
- Location
- Gurgaon, Haryana
- Work model
- On-Site
- Level
- Senior
Skills
About this role
Optum is a global organization that delivers care, aided by technology to help millions of people live healthier lives. The work you do with our team will directly improve health outcomes by connecting people with the care, pharmacy benefits, data and resources they need to feel their best. Here, you will find a culture guided by inclusion, talented peers, comprehensive benefits and career development opportunities. Come make an impact on the communities we serve as you help us advance health optimization on a global scale. Join us to start Caring. Connecting. Growing together.
Primary Responsibilities
Design, develop, test, and support scalable data pipelines and ETL/ELT solutions using Azure Databricks, PySpark, SQL, and cloud-native technologies Contribute to the NHI Cloud Modernization program by migrating and modernizing legacy Teradata, DataStage, Unix, and SAS-based data processing workloads to Azure Databricks and Snowflake Develop and maintain data ingestion, transformation, and data quality frameworks supporting Bronze, Silver, and Gold data layers within the enterprise lakehouse architecture Build reusable and optimized PySpark components, notebooks, and workflows to process large-scale healthcare claims, membership, provider, and ancillary datasets Implement monitoring, logging, alerting, exception handling, and operational automation to improve platform reliability and supportability Partner with product owners, business stakeholders, architects, and data consumers to understand requirements and deliver high-quality data products Support data governance, security, lineage, and access control processes in alignment with enterprise standards Perform performance tuning, workload optimization, and cost management of Databricks environments and data pipelines Participate in code reviews, technical design discussions, release activities, production support, and root-cause analysis of data platform issues Contribute to engineering best practices including CI/CD, automated testing, documentation, and operational excellence Leverage emerging AI-assisted engineering capabilities and automation opportunities to improve delivery speed, quality, and platform efficiency where appropriate Builder Responsibility: Design, develop, and deploy AI-powered solutions using no-code, low-code, and advanced platforms, translating business needs into scalable applications that enhance products, workflows, and decision-making Comply with the terms and conditions of the employment contract, company policies and procedures, and any and all directives (such as, but not limited to, transfer and/or re-assignment to different work locations, change in teams and/or work shifts, policies in regards to flexibility of work benefits and/or work environment, alternative work arrangements, and other decisions that may arise due to the changing business environment). The Company may adopt, vary or rescind these policies and directives in its absolute discretion and without any limitation (implied or otherwise) on its ability to do so Required Qualifications: Bachelor's degree in Computer Science, Information Technology, Engineering, or related field 3+ years of experience in Data Engineering, ETL development, or Big Data platforms 3+ years of hands-on experience with Azure Databricks and PySpark development Hands-on experience with cloud data platforms such as Azure Databricks, Azure Data Lake Storage (ADLS), and Snowflake Experience building large-scale batch and incremental data processing solutions Experience with orchestration and scheduling tools such as Airflow, Azure Data Factory (ADF), or equivalent platforms Experience implementing data quality validations, monitoring, logging, and operational support processes Knowledge of data warehousing concepts, dimensional modeling, and lakehouse architectures Familiarity with Git-based source control and CI/CD practices Demonstrated solid programming skills in Python, PySpark, and SQL Demonstrated solid