Logo for Mindbeam AI

Machine Learning Engineer - Pre Training

Role overview

Qualifications

  • Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or related field—or equivalent experience
  • 2+ years of experience with large-scale model training and distributed systems
  • Strong coding skills in Python and familiarity with ML frameworks (PyTorch, TensorFlow, JAX)
  • Experience with GPU scheduling, memory optimization, and parallelism strategies

Responsibilities

  • Build scalable pre-training pipelines for foundation models, optimizing throughput and efficiency
  • Implement distributed training strategies across GPUs/TPUs and high-performance clusters
  • Collaborate with researchers to translate experimental setups into production-ready workflows
  • Develop monitoring and fault-tolerance systems to ensure reliable large-scale training

About the company

Mindbeam AI logo

Mindbeam AI

Company details

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Mindbeam

We are building the next-generation AI infrastructure for open source and enterprise. Our work is deeply research-oriented and passionate about developing ground-breaking innovations to take state-of-the-art AI applications to the next level.

What drives us is not only advancing technology, but empowering the people behind it. We are a community of researchers, engineers, and visionaries who believe that collaboration, curiosity, and openness fuel progress. If you’re motivated by impact and inspired to build tools that others can build upon, you’ll be in the right place.

Mission

Design and optimize large-scale pre-training systems that power Mindbeam’s generative AI models.

Role Expectations

• Build scalable pre-training pipelines for foundation models, optimizing throughput and efficiency.

• Implement distributed training strategies across GPUs/TPUs and high-performance clusters.

• Collaborate with researchers to translate experimental setups into production-ready workflows.

• Develop monitoring and fault-tolerance systems to ensure reliable large-scale training.

• Continuously benchmark and tune performance across hardware and software stacks.

Background

• Bachelor’s, Master’s, or PhD in Computer Science, Engineering, or related field—or equivalent experience.

• 2+ years of experience with large-scale model training and distributed systems.

• Strong coding skills in Python and familiarity with ML frameworks (PyTorch, TensorFlow, JAX).

• Experience with GPU scheduling, memory optimization, and parallelism strategies.

• Comfort with containerized and orchestrated environments (Docker/Kubernetes).

• Understanding of high-performance computing and networking bottlenecks.

About You

You thrive on scale and complexity. You enjoy solving system-level bottlenecks, pushing hardware and software to their limits, and working closely with researchers to accelerate cutting-edge AI development.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Machine Learning Engineer Related jobs

Other jobs at Mindbeam AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.