Lead Eventing AWS Engineer (Kafka/Kinesis)
State Farm
- Location
- Bloomington, Illinois; Dunwoody, Georgia; Richardson, Texas; Tempe, Arizona
- Employment
- Full Time
- Work model
- On-Site
- Level
- Senior
- Salary
- $81k – $142k/yr
Skills
About this role
Overview
Being good neighbors – helping people, investing in our communities, and making the world a better place – is who we are at State Farm. It is at the core of how we operate and the reason for our success. Come join a #1 team and do some good! HYBRID: Qualified candidates must live or relocate within a 180-mile radius of a hub location listed below and should plan to spend time working from home and some time working in the office as part of our hybrid work environment. HUB LOCATIONS: Bloomington, IL; Dunwoody, GA; Richardson, TX; or Tempe, AZ SPONSORSHIP: Applicants for this position are required to be eligible to lawfully work in the U.S. immediately; employer will not sponsor applicants for U.S. work authorization (e.g. H-1B visa) for this opportunity Grow Your Skills, Grow Your Potential Responsibilities In State Farm's contact center environments, the quality of our data determines the quality of every decision we make across tens of millions of contacts annually. Our data engineering team builds the canonical models, event streams, and pipelines that turn raw contact-center activity into trusted, governed, and analytics-ready data products used across the enterprise. We’re building an event-driven backbone to connect streaming data and intelligence to contact center leaders. We’re focused on routing signals in near real time to the right leaders with the right context to enable rapid decision making. We work at the intersection of cloud-native data infrastructure, event-driven architecture, modern lakehouse architecture, and AI-ready data design. Our stack is AWS-native today and evolving toward an enterprise lakehouse on Databricks. If you are an engineer who believes that great data and eventing is the foundation of everything that matters, come build the foundation with us. In This Role, You Will: Define Canonical Data and Event Schemas - Contribute to our shared canonical model — the versioned types that govern every domain event crossing an API or system boundary Design and Deliver Event-Driven Architectures - Architect and implement scalable, resilient event-based systems using AWS-native messaging and streaming services including Kinesis, SQS, SNS, and EventBridge. Help shape the platform’s evolution toward Confluent Kafka as it becomes available to our workloads. Design and Build Data Pipelines - Architect and implement scalable event-driven and batch pipelines using Kinesis and Kafka, Lambda, AWS Glue, Step Functions, and open-source tooling to ingest, transform, and deliver enterprise data for use in AI and analytics applications Build Change-Data-Capture Pipelines - Implement CDC pipelines from operational stores (DynamoDB Streams, Neptune Streams, RDS logical replication) into canonical domain events on the event backbone Build Materialized Read Models – Design and build operational read projections in DynamoDB, OpenSearch, and Aurora that turn streaming state into sub-second queryable derived state for leader-facing applications, APIs, and near real-time dashboards. Model Enterprise Data - Design and maintain dimensional models, entity-relationship models, and semantic layer definitions using dbt to power analytics, reporting, and AI/ML workloads Engineer the Lakehouse Transition - Help design and execute the path from our current AWS-native analytics stack toward an enterprise lakehouse on Databricks, Delta Lake, and Iceberg Develop Data Access APIs - Build data service APIs that expose governed, well-documented data products to downstream consumers including analytics tools and AI systems Engineer Derived-Fact Write-Back - Build the surface that lets ML models and AI agents write inferred facts back to the event backbone with confidence and provenance fields, consumable like any other state Enable AI/ML Data Readiness - Prepare and deliver feature-engineered datasets, vector embeddings, and data pipelines for AI and GenAI workloads Champion DevSecOps for Data - Integrate data