Principal Observability Platform Engineer
Optiver
- Location
- Sydney, Australia
- Work model
- On-Site
- Level
- Principal
- Posted
- 1h ago
Skills
About this role
WHO WE ARE
Optiver is a tech-driven trading firm and leading global market maker. For over 35 years, Optiver has been improving financial markets worldwide, making them more transparent and efficient for all participants. With more than 1,400 employees in offices around the world, we’re united in our commitment to improving the market through competitive pricing, execution and thorough risk management. By providing liquidity on multiple exchanges across the world, we actively trade on 70+ exchanges, where we’re trusted to always provide accurate buy and sell pricing – no matter the market conditions.
WHAT YOU'LL DO
We are looking for a Principal Observability Platform Engineer to help evolve observability as a business-critical platform capability at Optiver. You will work on the shared platform behind metrics, logs, traces, events, alerts, dashboards, diagnostics, instrumentation and service health.
This is a platform engineering role for someone who enjoys building reliable systems used by other engineers. You will help turn a capable but heterogeneous observability foundation into a globally consistent, regionally federated platform that is reliable at scale, easy to adopt, and deeply embedded in how Optiver builds and operates production systems.
As a Senior Observability Platform Engineer, you will design, build, and operate components that help engineers, operators, trading teams, automated systems, and future agent-based workflows collect, query, understand, and act on production signals. You will work across platform and production domains: building high-scale telemetry pipelines, improving instrumentation quality, creating golden paths for adoption, and making observability more useful during real production investigations.
In this role, you will
• Design, build, and operate components of Optiver’s shared observability platform across telemetry collection, ingestion, storage, query, visualisation, alerting, diagnostics, and service health.
• Build software, services, APIs, integrations, libraries, dashboards, automation, and reusable patterns that make observability easier to adopt and more reliable to operate.
• Improve the scalability, reliability, performance, cost-effectiveness, and operational quality of high-volume telemetry systems.
• Improve developer and operator experience through self-service workflows, golden paths, documentation, investigation tooling, and practical platform abstractions.
• Work with engineering, infrastructure, trading systems, research, and regional operations teams to understand production debugging needs and improve observability adoption.
• Own the reliability and operational quality of the components you build, including service health, failure modes, monitoring, incident learnings, and continuous improvement.
• Raise the standard for telemetry quality, instrumentation, alerting, dashboards, diagnostic workflows, and service health across