Logo for Radian Arc

Senior Platform Support Engineer (Remote)

Role overview

Qualifications

  • 5+ years of hands-on experience in infrastructure engineering/support, cloud/platform operations, systems engineering or production support
  • Advanced Linux administration and troubleshooting skills
  • Strong production Kubernetes experience
  • Hands-on experience operating and troubleshooting NVIDIA GPU infrastructure in production

Responsibilities

  • Own complex and high-priority technical incidents
  • Troubleshoot across Linux, Kubernetes, GPU compute, networking, storage and supported platform services
  • Drive recovery actions and workarounds where appropriate
  • Mentor Platform Support Engineers and help improve their troubleshooting capability

Key facts

Hard skills

Other skills

  • Problem Solving

About the company

Radian Arc logo

Radian Arc

Computer Software / SaaS

Radian Arc delivers a carrier-embedded GPU edge computing platform already deployed across 70+ telecom and edge customers worldwide, enabling low-latency, sovereign AI inference directly within telecom networks. Radian Arc is a core pillar of inferX, Submer’s AI cloud and AI-as-a-Service platform, which empowers AI workloads to be deployed, operated, and monetized across enterprise, hyperscale, and telecom environments. Built on over a decade of leadership in liquid cooling, Submer designs, builds, and manages AI-ready datacenter infrastructure, taking full accountability – end to end. Together, Radian Arc, inferX, and Submer deliver a unified core-to-edge AI platform, combining carrier-edge GPU compute, AI cloud services, and modular, high-density infrastructure. Learn more at radianarc.io and inferX.com

Company details

Company typeSME
IndustryComputer Software / SaaS
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Radian Arc provides an infrastructure-as-a-service (IaaS) platform for running cloud gaming, artificial intelligence and machine learning applications inside telecommunication carrier networks.

Our teams across the USA, Australia, Central Europe, Malaysia, Singapore and Japan offer telecom operators a GPU-based edge computing platform without the need for capital expenditure, facilitating low latency and improved economics for value-added services and the monetization of 5G investments.

What impact you will have?

As a Senior Platform Support Engineer, you’ll be one of the main technical escalation points for Radian Arc’s production AI and GPU infrastructure. You’ll take on the incidents that need deeper investigation across Linux, Kubernetes, NVIDIA GPU infrastructure, networking and storage. The expectation is that you can get into the detail, work across multiple layers and own the technical investigation through to recovery or a clear engineering/vendor escalation. This is not a coordination-only support role. We need someone who has already operated and troubleshot GPU and Kubernetes environments in production and can help raise the technical capability of the wider support team.

You’ll work closely with Engineering, Data Center Operations, Service Delivery and our technology partners.

What you’ll do

  • Own complex and high-priority technical incidents.
  • Troubleshoot across Linux, Kubernetes, GPU compute, networking, storage and supported platform services.
  • Use logs, monitoring data, hardware diagnostics and recent changes to work out where the problem sits.
  • Drive recovery actions and workarounds where appropriate.
  • Validate service recovery and contribute to root-cause analysis for major or recurring incidents.
  • Assess customer and SLA impact during major incidents and provide the SDM with clear technical status, recovery actions and realistic next steps. 
  • Diagnose GPU node failures, degraded GPU health, driver issues and infrastructure-related GPU performance problems.
  • Use NVIDIA diagnostic and telemetry tools to investigate GPU, driver, firmware and hardware issues.
  • Troubleshoot production Kubernetes issues across nodes, pods, scheduling, networking and the container runtime.
  • Diagnose problems that span Kubernetes, Linux and the underlying infrastructure rather than treating each layer separately.
  • Troubleshoot network issues across L2/L3, routing, VLAN/VXLAN and high-performance networking.
  • Investigate storage availability, connectivity and performance issues.
  • Escalate platform defects and systemic problems to Engineering with enough evidence to make the escalation useful.
  • Work with Data Center Operations where hardware needs to be replaced or checked onsite.
  • Engage OEMs and technology partners when specialist support is required.
  • Stay technically engaged through the escalation rather than simply handing the issue off.
  • Build or improve Bash/Python tooling for diagnostics and repetitive support tasks.
  • Work with Engineering and Observability teams to improve alerts, telemetry and supportability.
  • Turn repeatable L2 investigations into better runbooks, tooling or L1 procedures.
  • Mentor Platform Support Engineers and help improve their troubleshooting capability.
  • Contribute to readiness reviews for new platforms, infrastructure and customer environments.

What you’ll need 

  • Showcase 5+ years of hands-on experience in infrastructure engineering/support, cloud/platform operations, systems engineering or production support.
  • Advanced Linux administration and troubleshooting skills.
  • Strong production Kubernetes experience, including cluster, node, pod, scheduling, networking and container-runtime troubleshooting.
  • Hands-on experience operating and troubleshooting NVIDIA GPU infrastructure in production.
  • Experience diagnosing GPU hardware, drivers and system-level GPU issues using NVIDIA or equivalent tooling.
  • Strong networking knowledge, including TCP/IP, VLANs, routing and structured L2/L3 troubleshooting.
  • Experience troubleshooting storage availability, connectivity and performance.
  • Strong experience using observability platforms and working with logs, metrics and telemetry during complex incidents.
  • Strong Bash and/or Python scripting skills.
  • Proven experience owning significant production incidents through diagnosis, recovery and escalation.
  • Ability to troubleshoot across multiple infrastructure layers and identify the actual source of a problem.
  • Experience working directly with Engineering teams, infrastructure specialists and/or hardware vendors.
  • Clear written and spoken English.

Nice to Have:

  • Hands-on experience with NVIDIA HGX/DGX-class systems, including H100, H200, B200/B300 or comparable platforms.
  • NVIDIA DCGM and other GPU diagnostic or telemetry tooling.
  • RDMA/RoCE and high-performance Ethernet.
  • NVIDIA Spectrum-X or Cumulus networking.
  • NCCL and multi-GPU/multi-node communication troubleshooting.
  • Large-scale bare-metal GPU environments.
  • Distributed or high-performance storage such as DDN, WEKA, VAST or similar.
  • BMC, BIOS and firmware troubleshooting.
  • Infrastructure automation using Ansible, Terraform or similar.
  • HPC, distributed AI training or large-scale inference environments.
  • Experience with private cloud and virtualisation technologies such as CloudStack, KVM/KubeVirt, OVS/OVN, VyOS and Citrix NetScaler/WAF.

Location: Malaysia or a comparable APAC time zone preferred 

Employment type:

  • Contractor or Employee of Record
  • Average 40 hours per week
  • Paid as a monthly rate

What we offer
• Attractive compensation package reflecting your expertise and experience.
• A great work environment characterised by friendliness, international diversity, flexibility, and a hybrid-friendly approach.
• You'll be part of a fast-growing scale-up with a mission to make a positive impact, offering an exciting career evolution.

Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.

Our inclusive responsibility
Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Platform Engineer Related jobs

Other jobs at Radian Arc

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.