Match score not available

High Performance Computing Cluster Administrator

extra holidays - fully flexible

Remote:

Full Remote

Salary:

148 - 230K yearly

Experience:

Mid-level (2-5 years)

Work from:

United States

Offer summary

Qualifications:

BA, BS, or MS in CS, EE, CE or equivalent, 4+ years deploying and administrating HPC clusters, Familiarity with resource scheduling managers and scripting languages, Experience with containers, operating systems, computer networks, Ability to work well with developers and test engineers.

Key responsabilities:

Administer Linux systems from servers to embedded systems
Coordinate storage solutions and automate configuration management
Connect with management for problem resolution and plan new systems
Design and implement GPU compute cluster for deep learning and scientific computing
Support NVIDIA DL Software and provide strategic solutions

NVIDIA XLarge http://www.nvidia.com/

10001 Employees

See more NVIDIA offers

Job description

NVIDIA's Deep Learning Optimized Frameworks Group is looking for a deeply technical HPC cluster administrator to lead a diverse cluster of GPU-accelerated systems and provide architectural mentorship to product teams in the deep learning and scientific computing domains. As a member of the DLFW Infrastructure team, you will provide leadership in the design and implementation of groundbreaking GPU compute cluster that runs demanding deep learning, high performance computing, and computationally intensive workloads. We are looking for an expert to identify architectural changes and/or completely innovative approaches for our GPU Compute Cluster. In this role, you will help us with the strategic challenges we encounter, including compute, networking, and storage design for large-scale, high-performance workloads and effective resource utilization in a heterogeneous compute environment.

What you'll be doing:

Administer Linux systems, ranging from powerful DGX servers to embedded systems, bringup hardware to publicly available systems.
Coordinate Storage Solutions and plan for growth.
Automate configuration management, software updates, and maintenance and monitoring of system availability using modern DevOps tools (Ansible, Gitlab, etc.)
Actively connect with management regarding any problems with the equipment and propose resolution.
Plan, build and install/upgrade new systems that support NVIDIA DL Software

What we need to see:

You have a BA, BS, or MS in CS, EE, CE or equivalent experience
4+ years of previous experience deploying and administrating HPC clusters
Familiar with resource scheduling managers (Slurm (preferred), LSF, etc!
Proven track record to script in bash, Perl or python
Experience with containers (Docker, Singularity, LXC)
Deep understanding of operating systems, computer networks, and high-performance applications
Ability to work well with developers & test engineers
Hard-working dedication to provide quality in support for your users

Ways to stand out from the crowd:

Familiarity and prior work experience with technologies such as: Ansible, GIT, Slurm, Zabbix, Prometheus, Grafana and Docker
Familiarity with GPU usage in Compute Cluster and Cuda
Experience with mobile and embedded systems
Basic knowledge of Deep Learning.
Experience coding/scripting in Perl/Python/bash

The base salary range is 148,000 USD - 230,000 USD. Your base salary will be determined based on your location, experience, and the pay of employees in similar positions.

You will also be eligible for equity and benefits. NVIDIA accepts applications on an ongoing basis.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.