Site Reliability Engineer in Network Infrastructure
Nebius
- Location
- Amsterdam, Netherlands; Remote - Europe
- Work model
- Remote
- Level
- Mid
- Posted
- 1h ago
Skills
About this role
About Nebius
Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.
Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.
The Role
We’re looking for a Site Reliability Engineer to help build and run the fundamental part of Nebius - the Network - the infrastructure everything else depends on. This is an engineering-first SRE role: you’ll set clear reliability targets, build the tooling and automation to meet them, and make the network safer to operate as we scale quickly.
Your responsibilities will include
• Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense)
• Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards
• Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting)
• Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents
• Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes
• Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast
We expect you to have
• Strong production Linux fundamentals and a structured approach to debugging complex systems
• Solid understanding of networking basics and how real networks fail (control plane vs data plane, latency/loss, failure domains, etc.)
• Hands-on experience operating high-availability systems and improving them over time (not just “keeping lights on”)
• Ability to write and maintain software/automation (Go is common for us; Python is also welcome)
• Experience with modern infrastructure tooling (e.g., IaC, CI/CD, container platforms) and comfort automating operational workflows
It will be an added bonus if you have:
• Experience with high-throughput traffic processing: load balancers, tunneling/decap, NAT64, or similar datapath-heavy