Submer
OpenSenior Platform Support Engineer
- Location
- Malaysia
- Posted
- Aug 3, 2026
- Last seen
- Aug 7, 2026
About the role
About Radian Arc. Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks. Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.
What impact you will have
Mission: Provide advanced technical support for customers running workloads on the GPU cloud platform, ensuring reliable operation of edge- and large-scale GPU clusters and infrastructure services.
The Senior Cloud Support Engineer acts as a technical escalation point for complex incidents, helping diagnose and resolve issues across compute, networking, storage, and orchestration layers. This role works closely with engineering and operations teams to improve platform reliability, reduce incident frequency, and enhance the overall customer experience.
What you’ll do
Customer Support & Incident Management
- Provide advanced technical support for customers operating workloads on bare-metal and virtualized GPU infrastructure.
- Diagnose and resolve complex customer issues affecting GPU clusters, compute nodes, networking, and storage.
- Investigate incidents across multiple layers of the stack including firmware, drivers, operating systems, and platform services.
- Perform root cause analysis (RCA) for major incidents and contribute to long-term remediation efforts.
- Serve as a technical escalation point for complex or high-priority support cases.
Infrastructure Troubleshooting
- Troubleshoot issues affecting:
- GPU compute nodes
- Kubernetes clusters
- Networking infrastructure
- Local NVMe, hyperconverged and distributed storage systems
- Analyze logs, telemetry, and monitoring signals to identify underlying causes of platform instability.
- Investigate issues related to GPU drivers, firmware, networking, and system performance.
Security Monitoring & Incident Triage
- Monitor and investigate security alerts generated by the platform security stack.
- Analyze and triage alerts generated by
- Wazuh
- TheHive
- Cortex
- Validate alerts, determine impact, and escalate potential security incidents to the security engineering team.
- Assist in collecting system telemetry, logs, and forensic data required for incident investigations.
- Improve alert runbooks and operational procedures to reduce false positives and improve response time.
GPU HPC Workload Support
- Provide advanced support for large-scale GPU workloads running distributed training and inference jobs.
- Diagnose failures affecting multi-GPU and multi-node workloads.
- Investigate performance issues impacting distributed workloads, including:
- GPU utilization
- Communication latency
- Storage bottlenecks
- networking congestion
- Support scheduling systems used for GPU workloads, including troubleshooting:
- Job queue failures
- Scheduling constraints
- Cluster resource fragmentation
High-Performance Networking Troubleshooting
<li
