Logo for KMC Solutions

XTN-DAD3686 | SITE RELIABILITY ENGINEER

Role overview

Qualifications

  • Strong hands-on experience administering and troubleshooting Linux systems
  • Confident use of CLI tools for diagnostics, including analysis of kernel logs, drivers, and system services
  • Excellent written and verbal English communication skills
  • High standards for system reliability, consistency, and documentation

Responsibilities

  • Validate GPU clusters of varying sizes to ensure hardware and system integrity prior to production release
  • Perform functional and reliability testing of GPUs, servers, and associated components
  • Maintain and extend the automated validation framework built using Python and Ansible
  • Produce clear, accurate documentation of test results, hardware states, and remediation actions

About the company

KMC Solutions logo

KMC Solutions

Outsourcing & Offshoring

The #1 flexible office space and fastest-growing EOR provider in the Philippines #DefyLimits 🚀

Company details

Company typeLarge
IndustryOutsourcing & Offshoring
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

You will be working on validating and testing GPU clusters prior to production release, ensuring hardware integrity, system reliability, and optimal performance. This role involves provisioning clusters, executing performance benchmarks, maintaining automated validation frameworks, and troubleshooting Linux-based systems in high-performance compute environments. You will collaborate closely with engineering and operations teams to ensure seamless handovers and production readiness. 

.•  Health Insurance/HMO 
•  Enjoy unlimited MadMax Coffee
•  Diverse learning & growth opportunities
•  Accessible Cloud HR platform (Sprout)
•  Above standard leaves

Cluster Validation & Testing

  • Validate GPU clusters of varying sizes to ensure hardware and system integrity prior to production release

  • Perform functional and reliability testing of GPUs, servers, and associated components

  • Verify network connectivity and performance, including InfiniBand where applicable

Orchestration & Benchmarking

  • Provision and configure GPU clusters using automated workflows

  • Execute and analyse performance and stability benchmarks orchestrated via Slurm

  • Validate results against expected performance and reliability thresholds

Test Framework & Automation

  • Maintain and extend the automated validation framework built using Python and Ansible

  • Integrate new test cases to support additional hardware platforms and GPU generations

  • Improve test reliability, coverage, and execution efficiency

Remediation & System Integrity

  • Diagnose and remediate unhealthy nodes through configuration changes or software fixes

  • Coordinate with on-site support and Smart Hands teams for hardware replacements when required

  • Ensure all issues are resolved and documented prior to handover to production operations

Documentation & Handover

  • Produce clear, accurate documentation of test results, hardware states, and remediation actions

  • Ensure smooth handovers to operations and engineering teams

  • Maintain up-to-date runbooks and validation procedures

Essential
• Strong hands-on experience administering and troubleshooting Linux systems (Prio)
• Confident use of CLI tools for diagnostics, including analysis of kernel logs, drivers, and system
services
• Excellent written and verbal English communication skills
• High standards for system reliability, consistency, and documentation
Preferred / Desirable
• Experience working with GPU-based or high-performance compute environments
• Familiarity with Slurm or other workload schedulers
• Understanding of datacenter hardware lifecycle and server validation processes
• Exposure to InfiniBand or high-speed networking technologies
• Experience working with distributed or remote infrastructure teams
• Proficiency in Python for automation, test execution, and parsing results (Preferred)
• Proven experience writing and maintaining Ansible playbooks (Preferred)

.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at KMC Solutions

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.