Software Development Engineer, Sahale
Amazon
- Location
- US, WA, Seattle
- Employment
- Full Time
- Work model
- On-Site
- Level
- Mid
- Posted
- 10d ago
Skills
About this role
Amazon's Big Data Technologies (BDT) organization builds the data platform that connects millions of businesses of all sizes to hundreds of millions of customers across Amazon marketplaces worldwide. The PartiQL/HubSchema team sits at the heart of this platform: we build the query language, schema modeling, and data validation technologies that make exabytes of data discoverable, trustworthy, and queryable across Amazon. We own PartiQL, Amazon's open-source, SQL-compatible query language for semi-structured and nested data, and HubSchema, Amazon's unified schema modeling and operation solution. HubSchema implements a hub-and-spoke architecture with a canonical schema model that serves as the universal intermediary for schema operations - conversion, validation, and compatibility analysis - across diverse compute and storage systems including AWS Glue, Andes (Amazon's data catalog), Apache Iceberg, Apache Avro, Parquet, Redshift, DynamoDB, and more. When a BDT service needs to convert, validate, or reason about schemas, HubSchema provides a single consistent answer - detecting type compatibility issues and data precision loss before an actual data transformation even occurs. Our systems define how tens of thousands of datasets are modeled, validated, evolved, and queried. HubSchema is integrated across the BDT ecosystem - Cradle (the data loading engine), Maestro (the orchestration platform), DataCraft v3, External Tables, and Andes Views all rely on HubSchema converters to ensure consistent schema interpretation and unified conversion logic. We are actively driving adoption across remaining BDT services and eliminating legacy schema definition approaches in favor of a single unified standard (Andes Schema Spec v1.1+). We are looking for a passionate and innovative engineer with a solid technical background to join the team. You will design and build core language and schema infrastructure used by virtually every data producer and consumer at Amazon: - Extending HubSchema's canonical model and spoke converters to support new storage formats and compute engines - Evolving the PartiQL specification, its Kotlin/JVM and Rust implementations, and runtime performance for latency-sensitive use cases - Building schema validation, conversion, and compatibility-checking services that guard data quality at Amazon scale - Delivering HubSchema service APIs that power schema operations across the BDT platform - Contributing to PartiQL as an open-source project The successful candidate will have a background in building distributed systems or data infrastructure, strong computer science fundamentals, good communication skills, and the motivation to achieve results in a fast-paced environment. Experience with query engines, compilers, type systems, data serialization formats, or schema management is a strong plus - but curiosity and rigor matter more than prior exposure to any specific technology. Key job responsibilities - Design, implement, and operate core components of PartiQL and HubSchema used across Amazon's data platform - Build and extend HubSchema spoke converters (Iceberg, Avro, Parquet, Glue, Redshift, Ion) and the canonical schema model - Develop schema validation and compatibility APIs that detect type mismatches, precision loss, and breaking changes before they reach production - Enhance PartiQL runtime performance (lazy evaluation, async execution, zero-copy Ion integration) for latency-sensitive consumers - Drive HubSchema integration across BDT services, replacing legacy one-off conversion logic with unified library converters - Contribute to the open-source PartiQL specification and reference implementations - Collaborate with partner teams across BDT (Catalog, Compute, Cradle, Maestro, DataCraft) to deliver end-to-end customer experiences - Raise the bar on operational excellence, testing, and engineering quality for systems in the critical path of Amazon's data ecosystem About the team The Sahale team owns PartiQL and