Rightmovecareers

Open

Engineering Manager (Site Reliability)

Location
London, UK
Posted
Jul 14, 2026
Last seen
Aug 7, 2026

About the role

Our vision is to give everyone the belief they can make their move. We aim to make moving simpler, by giving everyone the best place to turn to and return to for access to the tools, expertise, trust, and belief to make it happen.

We’re home to the UK’s largest choice of properties and are the go-to destination for millions of people planning their next move, reading the latest industry news, or just browsing what’s on the market.

The Role: Engineering Manager (Site Reliability)

Location: Reporting to: London / Hybrid (2 days per week in office)

Reporting to: Head of Technology Operations

The Role

The Platform and Reliability Engineering Teams are responsible for the services that underpin the Rightmove website and enable all of our product development teams to ship functionality rapidly and safely. We strive to deliver annual availability of at least 99.99% (less than 5 mins downtime a month).

The Site Reliability Engineering Manager’s role is to ensure operational excellence, drive observability and reliability at scale, and own the incident management processes and tools.

This position blends people leadership, full stack reliability engineering, service management, influencing without authority. The successful candidate brings strong technical experience in reliability engineering, monitoring, alerting and observability for product‑led technology companies combined with strong customer empathy and communication skills.

Monitoring, Alerting, and Observability

  • Product teams have high reliability confidence, incident detection and resolution are smooth due to proactive monitoring, well‑maintained alerts/logs and high levels of observability coverage.
  • Clear reliability expectations between platform, security, product & business. Prioritisation based on reliability risk and real data.

Incident Management

  • Consistency and standardisation of incident management resulting fast incident detection and resolution <li