Logo for STN Incorporated

Senior Platform Engineer

Role overview

Qualifications

  • 6+ years in platform engineering, SRE, or cloud engineering at scale
  • Deep Kubernetes expertise including CRDs, operators, and multi-tenant patterns
  • Strong programming skills in Go, Python, or both
  • Experience operating GPU clusters or AI infrastructure at production scale

Responsibilities

  • Design and build the orchestration layer (Kubernetes, Slurm, Run:ai, or comparable)
  • Manage multi-tenant isolation including namespaces, networking, storage, and quotas
  • Build customer-facing platform APIs, CLIs, web portals, and SDKs
  • Implement and operate image management, GPU operator, and node provisioning automation

Key facts

Other skills

  • Problem Solving
  • Teamwork

About the company

STN Incorporated logo

STN Incorporated

Company details

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Senior Platform Engineer

Platform and software ยท shared across customers

Reports to: Director, Platform Engineering (or Chief Architect)

Location: Remote (US) or Pleasanton, CA (hybrid)

Department: Cloud Platform Engineering / GPU Platform Engineering

Position summary

The Senior Platform Engineer builds and operates the multi-tenant orchestration, scheduling, and customer-facing platform layer that turns raw GPU infrastructure into a usable cloud service. This role is the software backbone of GPU One (GPUaaS).

Key responsibilities

  • Design and build the orchestration layer (Kubernetes, Slurm, Run:ai, or comparable)

  • Manage multi-tenant isolation including namespaces, networking, storage, and quotas

  • Build customer-facing platform APIs, CLIs, web portals, and SDKs

  • Implement and operate image management, GPU operator, and node provisioning automation

  • Drive infrastructure-as-code and automation across the platform stack

  • Partner with SRE on platform reliability, SLO definition, and observability

  • Support TAM and Support engineers on customer-impacting platform issues

  • Maintain customer environment templates, configuration management, and rollout tooling

  • Participate in architecture review, design discussions, and technical roadmap

  • Drive continuous platform improvement and reduce operational toil

Required qualifications

  • 6+ years in platform engineering, SRE, or cloud engineering at scale

  • Deep Kubernetes expertise including CRDs, operators, and multi-tenant patterns

  • Strong programming skills in Go, Python, or both

  • Experience operating GPU clusters or AI infrastructure at production scale

  • Bachelor's degree in computer science or equivalent experience

Preferred qualifications

  • Experience with NVIDIA GPU Operator, MIG, MPS, and NCCL operator patterns

  • Familiarity with Slurm operator, Run:ai, KubeRay, or comparable AI orchestration

  • Service mesh experience (Istio, Linkerd) and multi-cluster networking

  • Open source contributions in the cloud-native or AI infrastructure ecosystem

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
ยท

Platform Engineer Related jobs

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.