Senior Data Engineer - Streaming
Coca-Cola
- Location
- US - GA - Atlanta
- Work model
- On-Site
- Level
- Senior
- H-1B history
- 3 approvals (FY2023)
- Posted
- 14h ago
Skills
About this role
Job Description
Summary: We are seeking a highly specialized engineer in Real-time Streaming to evolve Coca- Cola’s enterprise data estate from batch-reliant processing to a continuous, event-driven ecosystem. In this role, your primary mission is to design, standardize, and codify streaming and Kappa Architecture patterns, minimizing redundant batch layers . Rather than just building isolated data pipelines, you will architect the foundational frameworks, reference implementations , and low-code configurations that enable engineering teams across the organization to instantly deploy robust, real-time data streams. The target destination for our real-time data ecosystem is Microsoft Fabric (Spark Streaming and /or Fabric Real-Time Intelligence ) . The ideal candidate is a champion of log-centric data design, possesses deep expertise in enterprise Apache Kafka environments.
Core Responsibilities
Build a reusable streaming architecture playbook for the enterprise, with a strong focus on Kappa-style patterns, event-driven processing, and near-real-time analytics on Fabric. Architect, document, and templatize standardized replay/backfill mechanisms (offset reset + checkpoint/state management + sink rebuild/dedupe strategy) . Contextualize Processing Latency by c reat ing distinct architectural blueprints and guardrails tailored to business SLAs : sub-second real-time and near-real-time Abstract Streaming Complexities by packaging advanced stream concepts ( stateful processing, watermarking, late-arriving data remediation, and complex windowing ) into reusable templates. Codify strict design patterns for effectively-once / end-to-end idempotency with deterministic replay . Build metadata-driven frameworks that gracefully handle upstream schema drift without interrupting or crashing 24/7 downstream streaming queries. Act as the internal authority on streaming design patterns across Kafka/Event Hubs, designing scalable ingestion into Fabric . Build standardized dashboards to track critical stream health metrics, including end-to-end data lag, processing throughput, backpressure, and poison-pill routing via Dead Letter Queues (DLQ). Required Qualifications & Experience 8+ years of experience in data platform architecture, with a dedicated minimum of 3+ years architecting and operating production-grade Streaming and Kappa Architectures. 4+ years of experience running Spark Streaming and Apache Kafka pipelines . Must understand topic partitioning strategies, consumer group mechanics, schema registries (Avro/JSON), and performance tuning for high-throughput streams. Deep familiarity with stream processing engines such as Spark Structured Streaming, Kafka Streams, or Apache Flink. Strong Scala and/or PySpark experience with stateful Structured Streaming. Advanced knowledge of the internal mechanics of Delta Lake storage, specifically regarding append-only streaming sinks and checkpointing behaviors. The "Framework" Mindset: A strong portfolio demonstrating the creation of developer-facing frameworks, custom connectors, or configuration-driven streaming engines that abstract infrastructure. Ability to articulate complex streaming topologies and real-time concepts to both application developers and executive stakeholders.
Preferred Qualifications
Confluent Certified Developer or Kafka Administrator certification. Proven experience modernizing Lambda architectures toward Kappa-first designs . Hands-on experience utilizing Fabric Real-Time Intelligence, including Fabric Eventstreams and KQL Databases. (Equivalent deep experience with Azure Event Hubs, Stream Analytics, or Databricks Delta Live). The Coca-Cola Company will not offer sponsorship for employment status