Reliability Engineer SME (M&E)

Location
London
Posted
Aug 24, 2026
Last seen
Aug 29, 2026

About the role

**Regular site travel required across EMEA**

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About the Role (Job Purpose)

The Reliability Engineer (Mechanical & Electrical) provides cross-discipline M&E engineering expertise and hands-on technical support across Nscale's EMEA data centre estate — owned, operated, and third-party colocation. This role is the technical escalation point for mechanical and electrical systems on site, and plays a key part in setting and maintaining the engineering standards that keep Nscale's infrastructure reliable and predictable as the business scales.

A particular focus of this role is supporting high-density AI compute. As GPU rack densities and power draw increase, cooling performance and power resilience become two of the most critical reliability factors across the estate — and neither can sensibly be engineered in isolation from the other.

What You’ll be Doing (Responsibilities)

  • M&E Technical Support: Act as the technical escalation point for on-site Operations teams across all aspects of managing mechanical and electrical systems.
  • Set and maintain maintenance standards, defining the technical baseline that all sites operate to.
  • Provide technical input into procurement, supporting technical services contracts to ensure scope and requirements are fit for purpose.
  • Design Review & Engineering Assurance: Carry out M&E design reviews at concept, detailed design and pre-handover stages — challenging resilience, maintainability and single points of failure before they are built in.
  • Review vendor and consultant submissions, drawings, schematics and O&M documentation against Nscale engineering standards.
  • AI Infrastructure Support: Provide technical support to AI Infra teams, with particular focus on AI rack cooling and power distribution systems.
  • Change & Compliance: Own change approval — technical review and sign-off across internal and third-party sites.
  • Act as technical approver for colocation change requests: assess supplier-raised works on live systems for risk to Nscale infrastructure, and challenge or reject where method, timing or resilience impact is not acceptable.
  • Run technical audits as part of the engineering-standards audit programme across the estate.
  • Assist in technical site audits of colocation sites — assessing partner infrastructure, maintenance regimes and standards compliance against contracted and Nscale requirements, and tracking findings through to closure.
  • Incidents, Problem Solving & RCA: Provide technical support to site teams during live incidents, remote or on site, helping diagnose faults and guiding safe recovery of M&E systems.
  • Own Root Cause Analysis (RCA) creation and closure, and drive Corrective and Preventive Actions (CAPA) through to completion.
  • Lead structured problem solving on complex and recurring faults, using trend, alarm and maintenance data to resolve underlying causes rather than repeat symptoms.
  • Standards, Training & Planning: Maintain the engineering standards library — the single source of truth for M&E standards as the team scales.
  • Develop and deliver technical training and knowledge-sharing for site Operations teams, raising M&E capability, standardising how faults are handled, and closing gaps identified through audits and RCAs.
  • Provide technical due diligence for new sites — early input into Site Mobilisation, assessing M&E design and build before handover.
  • Support capacity and thermal engineering — modelling whether site power and cooling can support the next generation of racks.
  • Support asset lifecycle management — tracking end-of-life on critical M&E equipment so failures are planned for, not discovered.
  • Contribute to energy efficiency / PUE improvement — continuous improvement of site energy performance.

About You (Skills / Qualifications)

  • M&E Engineering Background: Strong technical background in building services engineering within critical environments, with deep hands-on expertise in either mechanical or electrical systems and credible working knowledge of the other.
  • Mechanical: data centre cooling systems — chillers, CRAC/CRAH units, air handling, piping and secondary cooling loops.
  • Electrical: data centre power systems — LV/HV distribution, UPS, generators, switchgear and protection systems.
  • Experience carrying out M&E design reviews and reviewing engineering standards compliance in a critical infrastructure environment.
  • Experience conducting technical audits, including of third-party or colocation facilities.
  • Comfortable providing technical sign-off on change requests, understanding the operational risk of works on live M&E systems.
  • Experience leading or contributing to Root Cause Analysis (RCA) and Corrective and Preventive Action (CAPA) following M&E system incidents, with a structured approach to problem solving.
  • Confident supporting site teams under pressure during live incidents, and comfortable coaching and training operational engineers.
  • Strong communication skills, able to explain technical electrical concepts clearly to non-technical stakeholders and hold vendors/contractors to account.

Desirable:</strong&g