Logo for Uvation

Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)

Role overview

Qualifications

  • Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
  • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
  • Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
  • Advanced Linux storage administration

Responsibilities

  • Designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure
  • Managing infrastructure from the hardware layer up, including servers, networking, and storage
  • Supporting GPU clusters, AI training environments, and high-performance computing workloads
  • Creating operational documentation, runbooks, and infrastructure standards

About the company

Uvation logo

Uvation

Cloud Computing & Infrastructure (IaaS/PaaS)

We understand that the world is already overpopulated with IT consultancies, each offering the same types of services. This is why Uvation takes the role of ‘consultant’ one step further. We don’t just want to provide clients with advice about software and hardware, but rather serve as your comprehensive and strategic IT and security partner. Modern companies are requiring distinctive IT and security solutions within their niches. We know that meeting these unique needs, through trusted consultation and fully integrated solutions specifically tailored to your budget, is critical for businesses to scale effectively and navigate safely within an ever-transformative digital landscape. With a physical presence across North America, the EU and the Asia Pacific, we aim to be an end-to-end company that provides transformative services to address the most pressing technology challenges through each department of our business, which we call our focus areas. To meet today’s wide scope of IT and security needs, we’ve created devoted teams and resources to focus on specific sectors like aerospace and defense, the public sector, healthcare and the nonprofit field. Get Rewarded With Us By partnering with Uvation, you now get more out of every purchase with Uvation Rewards. Yes, our primary focus is on the success of our clients and their communities by delivering tangible results, not vague promises. In equal measure, we want our clients to get rewarded for investing in the future of their companies and the communities in which they thrive. Every time you invest in your company with Uvation, you now earn back cloud credits, gift cards and loyalty points that can even be converted into cash donations to support the missions of some of the world’s best charities. Contact us today for a free 30-minute consultation to see how we can provide your company with the full-stack services, innovative products, and pioneering technologies you deserve.

Company details

Company typeTPE
IndustryCloud Computing & Infrastructure (IaaS/PaaS)
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Job Overview

We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms. This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments.

This is not a DevOps-focused role. We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms.

The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical.

Key Responsibilities & Required Skills

Linux & Bare Metal Infrastructure

  • Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred)
  • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management
  • Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments
  • Strong understanding of server hardware, including:
    • BIOS/UEFI
    • RAID controllers
    • Firmware management
    • iLO/iDRAC/IPMI
    • NICs and SmartNICs
    • HBA cards
    • Hardware diagnostics and troubleshooting
  • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale

AI Factory & GPU Infrastructure

  • Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads
  • Understanding of NVIDIA GPU technologies including:
    • A100, H100, H200, B200, or equivalent GPU platforms
    • NVIDIA DGX and OEM GPU servers
    • GPU provisioning and lifecycle management
    • GPU monitoring and performance optimization
  • Knowledge of AI Factory architecture and infrastructure requirements
  • Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads
  • Understanding of:
    • GPU resource allocation and scheduling
    • Multi-GPU systems
    • GPU networking requirements
    • High-bandwidth, low-latency infrastructure design
  • Familiarity with NVIDIA ecosystem technologies such as:
    • CUDA
    • NCCL
    • GPUDirect Storage
    • NVIDIA Fabric Manager
    • NVIDIA Base Command (preferred)

Enterprise Storage & Data Platforms

  • Advanced Linux storage administration:
    • LVM
    • XFS, EXT4
    • NFS
    • iSCSI
    • Fibre Channel SAN
    • Multipath I/O
  • Strong hands-on experience with Ceph, including:
    • Cluster architecture
    • MON, OSD, MDS
    • RBD, CephFS, RGW
    • Capacity planning
    • Performance tuning
    • Failure recovery
  • Experience with high-performance AI storage platforms such as:
    • WEKA
    • VAST Data
    • Dell PowerScale
    • Pure Storage FlashBlade
    • NetApp
  • Understanding of:
    • NVMe-over-Fabrics (NVMe-oF)
    • RDMA
    • GPUDirect Storage
    • Parallel file systems
    • AI data pipelines

Networking & Infrastructure

  • Strong networking knowledge:
    • Bonding
    • VLANs
    • Routing
    • MTU optimization
    • DNS
    • DHCP
  • Experience with high-performance data center networking:
    • 100G/200G/400G Ethernet
    • RoCE
    • RDMA
    • Spine-Leaf architectures
  • Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies
  • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting

Operations & Reliability

  • Experience with high availability, clustering, and disaster recovery
  • Strong troubleshooting skills across:
    • Linux operating systems
    • Hardware platforms
    • GPU infrastructure
    • Networking
    • Enterprise storage
  • Experience supporting mission-critical production environments
  • Bash and Python scripting for automation and operational efficiency
  • Experience creating operational documentation, runbooks, and infrastructure standards

Nice to Have

  • Kubernetes infrastructure (especially AI/ML and GPU integration)
  • KVM, VMware, OpenShift Virtualization, or similar virtualization platforms
  • Ansible automation
  • NVIDIA Base Command Manager
  • Slurm or HPC workload schedulers
  • Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry)
  • Data Center Infrastructure Management (DCIM) tools
  • IPAM solutions
  • AWS, Azure, or hybrid cloud exposure

We Are Not Looking For

  • Candidates whose experience is primarily CI/CD pipeline engineering
  • Engineers focused mainly on Terraform, GitOps, or application delivery pipelines
  • Cloud-only administrators with limited bare metal, storage, or hardware experience
  • Professionals whose primary expertise is software development rather than infrastructure engineering

Ideal Candidate

Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Infrastructure Engineer Related jobs

Other jobs at Uvation

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.