Logo for Orion Placement

GPU Infrastructure NOC Engineer

Role overview

Qualifications

  • 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment.
  • Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments.
  • Hands-on Python or Bash scripting experience.
  • Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms.

Responsibilities

  • Monitor live GPU cluster health, power, cooling, networking, and infrastructure status across production deployments.
  • Triage, troubleshoot, and resolve incidents while maintaining SLA requirements.
  • Build and improve internal monitoring tools, scripts, and automation to reduce repetitive manual work.
  • Participate in incident reviews and turn recurring problems into permanent tooling, process, or automation improvements.

Key facts

Other skills

  • Troubleshooting (Problem Solving)
  • Communication
  • Problem Solving

About the company

Orion Placement logo

Orion Placement

Staffing & Recruiting

Unknown

Company details

IndustryStaffing & Recruiting
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Pay: $75,000.00 - $140,000.00 per year

Why This Is a Great Opportunity

  • Monitor and support cutting-edge GPU and AI infrastructure powering demanding enterprise workloads.
  • Go beyond traditional NOC monitoring by building automation, improving tooling, and making operations smarter.
  • Get hands-on exposure to GPU, HPC, networking, data center infrastructure, and AI-enabled operations.
  • Play a direct role in protecting customer uptime and meeting demanding SLA commitments.
  • Help create and continuously improve runbooks, SOPs, monitoring processes, and incident-response workflows.
  • Work alongside technical teams, OEMs, data center operators, and network providers to solve real infrastructure problems.
  • Join a fast-growing environment where your ideas can directly improve how the NOC operates.
  • Receive bonus and equity opportunities in addition to competitive base compensation.

Location: Remote nationwide. This is a fully remote role supporting a 24/7 global operations environment through rotating shifts, including nights and weekends.

Note: Must have 2+ years of NOC, network operations, or infrastructure monitoring experience, hands-on GPU or HPC infrastructure experience, and Python or Bash scripting experience. Candidates must be willing to work rotating 24/7 shifts, including nights and weekends. These are hard requirements.

About Us

We are building next-generation AI infrastructure designed to give enterprise customers fast, flexible access to high-performance GPU compute. Our operations team plays a critical role in keeping these environments reliable, responsive, and continuously improving. Confidential Employer.

Job Description

  • Monitor live GPU cluster health, power, cooling, networking, and infrastructure status across production deployments.
  • Triage, troubleshoot, and resolve incidents while maintaining SLA requirements.
  • Identify infrastructure issues early and take proactive action before they become customer-impacting incidents.
  • Escalate appropriate issues to OEMs, data center operators, network providers, or other Tier 3 partners.
  • Build and improve internal monitoring tools, scripts, and automation to reduce repetitive manual work.
  • Use Python, Bash, or similar scripting tools to automate monitoring, triage, reporting, and operational workflows.
  • Explore and implement AI-enabled workflows that improve NOC speed, accuracy, and efficiency.
  • Create, maintain, and continuously improve technical runbooks and standard operating procedures.
  • Participate in incident reviews and turn recurring problems into permanent tooling, process, or automation improvements.
  • Track SLA and incident metrics and identify opportunities to improve reliability and response times.
  • Communicate clearly and proactively with customers and internal stakeholders during incidents.
  • Coordinate with data center operators, OEMs, network providers, and other third parties to resolve customer-impacting issues.
  • Support a 24/7 rotating operations schedule, including nights, weekends, and other assigned shifts.

Qualifications

  • 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment.
  • Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments.
  • Hands-on Python or Bash scripting experience.
  • Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms.
  • Strong troubleshooting and incident-response skills.
  • Experience working with network, compute, storage, or data center infrastructure.
  • Ability to understand technical issues quickly and communicate effectively during incidents.
  • Demonstrated interest in automation, scripting, tool-building, and continuous operational improvement.
  • Must be comfortable working rotating 24/7 shifts, including nights and weekends.
  • Ability to work independently in a remote environment while collaborating effectively with global technical teams.
  • Additional languages beyond English are a plus.

Why You Will Love Working Here

  • Work directly with advanced GPU and HPC infrastructure rather than generic IT systems.
  • Build technical depth across AI compute, networking, data centers, monitoring, and automation.
  • Have a voice in how the NOC operates and help improve processes rather than simply following them.
  • Work with modern monitoring, automation, and AI-enabled operations tools.
  • Gain exposure to complex enterprise infrastructure and high-availability environments.
  • Remote nationwide flexibility with multiple shift options.
  • Medical, dental, and vision insurance.
  • 401(k).
  • Paid maternity and paternity leave.
  • Bonus and equity opportunities.

JPC-1916

Benefits:

  • Dental insurance
  • Paid time off
  • Retirement plan
  • Vision insurance

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Infrastructure Engineer Related jobs

Other jobs at Orion Placement

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.