Principal Software Development Engineer - Observability
Expedia Group
- Location
- USA California San Jose
- Work model
- On-Site
- Level
- Principal
- Posted
- 8h ago
Skills
About this role
At Expedia Group, we help travelers explore the world, one journey at a time. As a global travel company powered by passionate people, trusted partnerships, and leading technology, we connect travelers, partners, and advertisers through our consumer brands, B2B network, and travel advertising business. Here, you'll do meaningful work that helps millions of people discover, book, and experience travel with more ease, confidence, and joy. Our five Behaviors-Traveler First, Think Big, Operate with Excellence, Ownership Mindset, and Succeed Together-help foster a supportive environment where people can grow their careers and have the flexibility, benefits, and support to do their best work. Join us and build for travelers everywhere. Principal Software Engineer, Observability Introduction to the Team: Our Technology Team partners with teams across Expedia Group to create innovative products, services, and tools to deliver high-quality experiences for travelers, partners, and our employees. A singular technology platform powered by data and machine learning provides secure, differentiated, and personalized experiences that drive loyalty and traveler satisfaction. As a Principal Engineer, you will be part of an agile development team with deep expertise in cloud, distributed systems, and observability. You will play a pivotal role in crafting the strategic technical goals for our group. The main effort will involve leading the architecture, design, and implementation of a centralized, scalable, and cost-effective observability platform used by all engineering teams across Expedia. You will provide technical leadership for a dynamic engineering organization and work alongside talented product managers and other technical leaders to deliver best-in-class capabilities to our developer community. In this role, you will: Architect and Build Core Telemetry Pipelines : Lead the design and implementation of highly scalable and resilient telemetry pipelines for logs, metrics, and traces. Evolve our platform to handle a 10x increase in data volume while maintaining performance and cost-effectiveness Drive OpenTelemetry Adoption : Spearhead the strategy, rollout, and support for the OpenTelemetry collector across thousands of services. Develop best practices and automated configurations to ensure seamless and consistent data collection Implement Platform Governance and Optimization : Design and build capabilities for data governance, cost allocation, and resource management within the observability platform. Define and implement SLOs for the platform itself and create tools to help teams manage their observability costs Elevate the Practice of Observability : Act as a thought leader, driving the adoption of observability best practices across the engineering organization. Improve the developer experience by unifying tooling (e.g., Grafana, Datadog, Splunk), documentation, and service lifecycle management within our internal developer portal Automate Infrastructure Lifecycle : Author and maintain production-grade Infrastructure as Code ( IaC ) using tools like Terraform and/or Crossplane . Eliminate manual toil by automating cluster provisioning, dependency upgrades, and incident remediation workflows Technical Leadership and Mentorship : Act as a force multiplier. Mentor senior engineers on the team, lead architecture review sessions, and author RFCs to build consensus on significant technical decisions. Your influence will extend beyond the team to application developers and SREs Production Debugging : Serve as the final escalation point for complex, cross-cutting production incidents related to the observability platform, from telemetry agent bugs to data correlation failures in our distributed systems Collaborate and Innovate : Explore and utilize a wide variety of technologies and