Principal Network Engineer

Location
UK
Posted
Aug 26, 2026
Last seen
Aug 28, 2026

About the role

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud bolsters technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

At Nscale, our Engineering team plays a critical role in designing, deploying, and operating the infrastructure and software platforms that power our customers and enable AI workloads at scale.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you'll be contributing to building the technology that powers the future.

About the Role

We are hiring a Principal Network Engineer to act as a senior technical authority for Nscale's AI-optimised network infrastructure.

Our Network Engineering team is responsible for the design, validation, and ongoing operation of the networking services underpinning both our internal management platform and customer-facing cloud infrastructure. This includes high-performance Ethernet fabrics, InfiniBand, RoCE, WAN connectivity, and large-scale data centre networking.

In this role, you'll set technical direction across the low-latency, high-bandwidth networks supporting large-scale AI training and inference workloads. You'll own critical technical domains end-to-end and help raise the bar for architecture, automation, operational rigour, and engineering standards across Nscale.

This is a deeply technical Principal-level role combining hands-on engineering with broad architectural influence. You'll define reference architectures, drive consistency across sites, lead complex technical decisions and escalations, and mentor engineers while partnering closely with Deployment, Data Centre Operations, Platform Engineering, Systems, Storage, and technology vendors.

What you'll be doing

AI & High-Performance Network Architecture

• Define, design, validate, and evolve large-scale InfiniBand, RoCE, and Ethernet fabric architectures at rack, row, and data centre scale.

• Design networks that integrate closely with bare-metal provisioning and cluster management systems.

• Own technical direction for high-performance Ethernet fabrics, including BGP, EVPN, VXLAN, LACP, and QoS.

• Establish reference architectures and engineering standards that can be implemented consistently across Nscale's data centre estate.

• Identify systemic risks and architectural gaps and drive durable solutions that improve scalability, reliability, and operational simplicity.

Network Automation & Infrastructure as Code

• Lead Nscale's network automation strategy using a GitOps operating model.

• Build and guide Python and Ansible tooling for provisioning, configuration validation, compliance, and operational workflows.

• Drive version-controlled configuration and CI/CD-based network change across multi-vendor environments.

• Apply Infrastructure-as-Code and Network-as-Code principles to reduce manual intervention and improve operational consistency.

• Continuously identify opportunities to automate repetitive operational tasks and reduce reactive toil.

Network Security & Edge Infrastructure

• Design and engineer perimeter and network security infrastructure across WAN and data centre edge environments.

• Own architecture across firewalls, NAT, VPN, security policies, and multi-tenant segmentation.

• Design highly available and scalable security architectures appropriate for mission-critical AI infrastructure.

Reliability, Observability & Operations

• Lead complex technical escalations and root-cause analysis for network performance, reliability, and stability issues.

• Establish measurable SLOs and operational standards for network services.

• Set technical direction for network observability, telemetry, monitoring, and alerting.

• Ensure clear visibility into fabric health, traffic patterns, performance, and capacity.

• Develop runbooks, automation, and engineering improvements that systematically reduce operational toil.

• Act as a senior 3rd/4th line escalation point for complex networking issues.

Network Data & Configuration Management

• Ensure the accuracy and reliability of source-of-truth network inventory and configuration data.

• Establish structured engineering and change-management practices for network configuration.

• Ensure network changes are controlled, auditable, repeatable, and scalable across multiple sites.

Technical Leadership & Collaboration

• Partner with Deployment, Data Centre Operations, Platform Engineering, Systems, Storage, and vendors on new site delivery and platform evolution.

• Lead architecture and design reviews for significant network initiatives.

• Mentor engineers and raise technical capability across the wider networking organisation.

• Lead complex technical decisions and incidents spanning networking, systems, storage, and AI/HPC workloads.

• Influence engineering strategy and standards across teams without relying on formal authority.

About You

Required Experience

• 10+ years of network engineering experience, with significant depth in HPC, AI, hyperscale, or large-scale data centre environments.

• Extensive hands-on experience with RDMA-aware networking for AI/HPC workloads, including InfiniBand and/or RoCE.

• Experience with subnet managers and fabric orchestration technologies such as OpenSM or NVIDIA UFM.

• Expert-level understanding of modern data centre routing and control planes, including BGP, EVPN-VXLAN, and Clos/spine-leaf architectures.

• Production experience with network platforms such as Cumulus, Nokia, or Arista EOS.

• Strong network automation expertise using Python and Ansible.

• Experience with Git-based wo