Logo for Mirantis

AI Infrastructure Engineer

Role overview

Qualifications

  • Proven experience managing and operating large-scale production systems (bare-metal and/or cloud)
  • Solid working knowledge of Kubernetes with excellent, demonstrable troubleshooting skills
  • Experience configuring, customizing, and extending logging and monitoring tools (e.g., Prometheus, Grafana, ELK, or similar)
  • Experience with infrastructure automation technologies and Infrastructure-as-Code practices (e.g., Ansible, Terraform, or similar)

Responsibilities

  • Manage and operate production AI infrastructure environments
  • Lead incident response and troubleshooting efforts and deliver timely service restoration during outages or performance degradations
  • Troubleshoot infrastructure and networking issues across bare-metal and/or cloud environments with multiple vendors
  • Conduct root cause analysis and drive product and operational improvements

Key facts

Other skills

  • Troubleshooting (Problem Solving)
  • Analytical Skills
  • Communication
  • Problem Solving

About the company

Mirantis logo

Mirantis

Cloud Computing & Infrastructure (IaaS/PaaS)

Mirantis helps organizations ship code faster on public and private clouds. The company provides a public cloud experience on any infrastructure to the data center to the edge. With Lens and Docker Enterprise Container Cloud, Mirantis empowers a new breed of Kubernetes developers by removing infrastructure and operations complexity and providing one cohesive cloud experience for complete app and devops portability, a single pane of glass, and automated full-stack lifecycle management with continuous updates. Mirantis serves many of the world’s leading enterprises, including Adobe, DocuSign, Liberty Mutual, Nationwide Insurance, PayPal, Reliance Jio, Splunk, and STC. Learn more at www.mirantis.com.

Company details

Company typeSME
IndustryCloud Computing & Infrastructure (IaaS/PaaS)
Company size501 - 1000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Company Description

About Mirantis

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environmentβ€”on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.

Job Description

We are looking for an AI Infrastructure Engineer to join our global team responsible for managing and supporting large-scale AI infrastructure environments. You will help ensure the availability, performance, and operational stability of critical AI infrastructure platforms, working closely with a distributed team across regions to provide continuous coverage and support.

This is an opportunity to work hands-on with some of the most advanced Kubernetes-based AI infrastructure in production today, while contributing to the platforms and processes that keep it running reliably at scale.

Responsibilities:

  • Manage and operate production AI infrastructure environments.

  • Lead incident response and troubleshooting efforts and deliver timely service restoration during outages or performance degradations .

  • Troubleshoot infrastructure and networking issues across bare-metal and/or cloud environments with multiple vendors.

  • Conduct root cause analysis and drive product and operational improvements.

  • Contribute and improve operational documentation and knowledge base.

  • Collaborate with global team members across time zones to ensure continuous operational coverage, including occasional work during weekends and holidays.

Qualifications

  • Proven experience managing and operating large-scale production systems (bare-metal and/or cloud).

  • Solid working knowledge of Kubernetes with excellent, demonstrable troubleshooting skills.

  • Experience configuring, customizing, and extending logging and monitoring tools (e.g., Prometheus, Grafana, ELK, or similar).

  • Experience with infrastructure automation technologies and Infrastructure-as-Code practices (e.g., Ansible, Terraform, or similar).

  • Effective verbal and written communication skills in English.

  • Strong analytical and problem-solving skills, with the ability to work through complex, ambiguous technical issues.

  • Willingness to occasionally work weekends and holidays.

Nice to have:

  • Previous experience building, scaling, and running High-Performance Computing (HPC) environments.

  • Hands-on experiences with managing large scale Kubernetes platforms in production. 

  • A good understanding of NVIDIA GPU technologies and the associated software stack.

  • Proficiency in scripting languages (e.g., Python, Bash, Go).

Additional Information

What does Mirantis offer you?

- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.


It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to isamoylova@mirantis.com

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

We are a Leader for Container Management in G2 (#2 after AWS)!

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
Β·

Infrastructure Engineer Related jobs

Other jobs at Mirantis

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.