yoinka

Site Reliability Engineer in Network Infrastructure

Nebius

RemoteAmsterdam, Netherlands; Remote - EuropeMid
Sign in to applyVerified 1h ago
Location
Amsterdam, Netherlands; Remote - Europe
Work model
Remote
Level
Mid
Posted
1h ago

Skills

CI/CDLinuxPython

About this role

About Nebius

Nebius is leading a new era in cloud infrastructure for the global AI economy. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.

Built by engineers, for engineers. From large-scale GPU orchestration to inference optimization, we own the hard problems across compute, storage, networking and applied AI.

Listed on Nasdaq (NBIS) and headquartered in Amsterdam, we have a global footprint with R&D hubs across Europe, the UK, North America and Israel. Our team of 1,500+ includes hundreds of engineers with deep expertise across hardware, software and AI R&D.

The Role

We’re looking for a Site Reliability Engineer to help build and run the fundamental part of Nebius - the Network - the infrastructure everything else depends on. This is an engineering-first SRE role: you’ll set clear reliability targets, build the tooling and automation to meet them, and make the network safer to operate as we scale quickly.

Your responsibilities will include

• Define and own reliability goals for network services and critical paths (SLIs/SLOs, availability targets, error budgets where it makes sense)

• Drive reliability improvements across the whole network: not only services, but also site readiness, inter-site connectivity (DCI), and operational standards

• Own incident response for your areas, lead investigations/postmortems, and turn failures into durable fixes (not repeated firefighting)

• Build and evolve observability: actionable metrics/logs/traces, alerting, and faster debug loops during and after incidents

• Design safer change workflows: automation, CI/CD, test/staging environments, canarying, rollbacks, and auditability for network changes

• Work closely with network engineers and platform teams to embed operability into designs and keep operations practical and fast

We expect you to have

• Strong production Linux fundamentals and a structured approach to debugging complex systems

• Solid understanding of networking basics and how real networks fail (control plane vs data plane, latency/loss, failure domains, etc.)

• Hands-on experience operating high-availability systems and improving them over time (not just “keeping lights on”)

• Ability to write and maintain software/automation (Go is common for us; Python is also welcome)

• Experience with modern infrastructure tooling (e.g., IaC, CI/CD, container platforms) and comfort automating operational workflows

It will be an added bonus if you have:

• Experience with high-throughput traffic processing: load balancers, tunneling/decap, NAT64, or similar datapath-heavy

Site Reliability Engineer in Network Infrastructure at Nebius, Amsterdam, Netherlands; Remote - Europe | Yoinka