yoinka

Data Architect - Data Engineering 4D

Genpact

1401-G-India: 1-6 FL, Surya Park, STPI, BangaloreStaffH-1B sponsor company
Sign in to applyVerified 1h ago
Location
1401-G-India: 1-6 FL, Surya Park, STPI, Bangalore
Work model
On-Site
Level
Staff
H-1B history
63 approvals (FY2023)
Posted
21h ago

Skills

AWSAzureCI/CDDatabricksDockerFlinkGCPGitJavaKafkaKubernetesPythonSQLScalaSnowflake

About this role

Data Architect Ready to turn bold ideas into real-world impact?  At Genpact, we don’t just adapt to change, we lead it. AI and digital innovation are transforming the way businesses work, and we’re at the forefront of it. Genpact’s AI Gigafactory , our industry-first accelerator, exemplifies how we scale advanced technology solutions to help global enterprises work smarter, grow faster, and transform at scale. Whether tackling complex challenges through large-scale models or agentic AI , our breakthrough solutions tackle companies’ most complex challenges.   If you thrive in a fast-moving, innovation-driven environment, love building and deploying cutting-edge AI solutions, and want to push the boundaries of what’s possible, this is your moment.    Genpact (NYSE: G) is an agentic and advanced technology solutions company. We leverage process intelligence and artificial intelligence to deliver measurable outcomes. With a strong partner ecosystem and decades of client trust, we provide innovative solutions that transform how businesses run. Powered by a team with an active learning mindset and client centricity at its core, we deliver lasting value for the world’s leading enterprises.  Get to know us at genpact.com and on LinkedIn , YouTube , X , and Facebook .

Job Description

Key Responsibilities   Design, develop, and maintain high-performance real-time data processing applications using Apache Flink.   Develop in-flight transformation logic leveraging Apache Flink, Kafka Streams, KSQLDB, and Kafka Single Message Transforms (SMTs).   Build and maintain scalable mapping frameworks for transforming source system data into harmonized enterprise data models.   Implement complex event processing, enrichment, filtering, aggregation, and windowing logic for streaming workloads.   Optimize streaming pipeline performance, ensuring low latency, high throughput, and efficient resource utilization.   Design solutions to handle late-arriving and out-of-order events using event-time processing, watermarks, checkpointing, and state management.   Integrate streaming applications with Kafka, Schema Registry, APIs, and downstream analytical platforms.   Implement robust error handling, retry mechanisms, dead-letter queues (DLQ), and monitoring for streaming applications.   Collaborate with Solution Architects, Data Architects, API teams, and business stakeholders to define transformation rules and streaming integration patterns.   Develop automated testing frameworks and CI/CD pipelines for streaming applications.   Monitor production streaming jobs and troubleshoot performance, scalability, and reliability issues.   Required Skills & Experience   6–10+ years of experience in Data Engineering with strong expertise in real-time streaming solutions.   Hands-on experience with Apache Flink for enterprise-scale stream processing.   Strong knowledge of Apache Kafka, Kafka Streams, KSQLDB, Kafka Connect, and Single Message Transforms (SMTs).   Experience designing event-driven data architectures and real-time data integration pipelines.   Strong understanding of event-time processing, watermarks, windowing, checkpointing, stateful stream processing, and fault tolerance.   Experience developing data transformation and harmonization frameworks for enterprise data platforms.   Proficiency in Java and/or Scala; Python knowledge is desirable.   Experience with Avro, Protobuf, JSON, Schema Registry, and serialization techniques.   Strong SQL skills and experience working with large-scale structured and semi-structured datasets.   Experience with cloud platforms such as Azure, AWS, or GCP is preferred.   Familiarity with CI/CD, Git, Docker, Kubernetes, and monitoring tools is an advantage.

Preferred Qualifications

Experience with Confluent Platform and enterprise Kafka deployments.   Exposure to modern data lakehouse platforms such as Databricks, Snowflake, or Delta Lake.   Knowledge of CDC patterns,

Data Architect - Data Engineering 4D at Genpact — 1401-G-India: 1-6 FL, Surya Park, STPI, Bangalore | Yoinka