yoinka

Senior Data Engineer

Alpaca

RemoteRemote - North America - LATAMSenior
Sign in to applyVerified 2h ago
Location
Remote - North America - LATAM
Work model
Remote
Level
Senior
Posted
7h ago

Skills

AirflowAnsibleArgoCDDockerGCPKafkaKubernetesLookerPythonSQLSparkTerraformdbt

About this role

Who We Are

Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.

Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.

Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.

Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.

Our Team Members

We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!

We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.

Your Role: We are seeking a Senior Data Engineer to help design and build the next generation of our Data Platform as we continue to scale to larger customers and new jurisdictions. At Alpaca, Data Engineering encompasses financial transactions, customer data, API logs, system metrics, augmented data, and third-party systems that impact decision making for both internal and external stakeholders. We process hundreds of millions of events daily, and this number continues to grow as we onboard new customers and products.

We prioritize open-source technologies in our data stack while leveraging Google Cloud Platform (GCP) as the foundation for our data infrastructure. This spans batch and stream ingestion, transformation, and consumption layers for BI/Reporting, AI/agent interfaces (MCP), and external third-party sinks. We also oversee data experimentation, cataloging, and monitoring/alerting systems.

Our team is 100% distributed and remote.

Responsibilities

• Design, build, and evolve the core data platform infrastructure e.g., distributed query engines, orchestration, warehousing, cataloging, and more.

• Own our lakehouse infrastructure as code, managing deployments through Terraform and Ansible on Kubernetes.

• Build and maintain low-latency streaming and CDC ingestion pipelines, as well as batch ingestion paths landing in Iceberg.

• Develop and scale our BI landscape so downstream teams and agents get performant, self-serve access to lakehouse data.

• Enforce platform reliability best practices, including monitoring and alerting, on-call rotations, incident response, maintenance windows, runbooks, and SLAs.

• Partner with DevOps, Analytics Engineering, and other stakeholders to close infrastructure gaps and support new data requirements.

Must-Haves

• 5+ years of experience in Data Engineering, including 2+ years building and operating scalable, low-latency data platforms handling > 100M events/day.

• Strong hands-on experience running data infrastructure on Kubernetes,