Logo for Robots and Pencils

Staff ML Engineer – AWS Trainium & SageMaker

Role overview

Qualifications

  • Strong, hands-on PyTorch experience, ideally including distributed or multi-device training
  • Production experience with Amazon SageMaker for training and/or inference
  • Comfort working close to the hardware layer
  • Solid Python fundamentals and comfort operating in a client-facing, production engineering environment

Responsibilities

  • Train and operate models on Amazon SageMaker with AWS Trainium
  • Write and optimize PyTorch training code with understanding of how it compiles and executes on Trainium
  • Diagnose training run issues that show up specifically because of the hardware
  • Translate a request for 'a Trainium job' into an actual working, cost-aware training pipeline

Key facts

Hard skills

Other skills

  • Problem Solving

About the company

Robots and Pencils logo

Robots and Pencils

IT Services & IT Consulting

Robots & Pencils develops digital strategies and products that deliver exponential impact to our clients. We design and build solutions that unlock data and insights, infuse intelligent automation, and accelerate product innovation across the organization. Everything we do starts by blending the sciences with the humanities, the Robots with the Pencils. Our top-tier talent fuses creativity + technology to help brands transform their businesses, deliver delightful customer and employee experiences, and maintain a competitive edge amidst a constantly changing industry landscape.

Company details

Company typeSME
IndustryIT Services & IT Consulting
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Robots & Pencils is an AWS Partner building production AI systems for enterprise clients who need real engineering, not proofs of concept that never ship. We work forward-deployed, embedded directly with client teams, solving the problems that are too new or too specialized for a typical vendor relationship to handle.
The Role
We're looking for an engineer who can operate and train models on Amazon SageMaker running on AWS Trainium, AWS's custom silicon built specifically for large-scale model training. This isn't a role where you call an API and wait. You'll be walking up the stack: understanding what a training request actually looks like at the Trainium hardware and compiler level, then carrying that understanding all the way up through PyTorch training code and into a production SageMaker pipeline.
PyTorch is the backbone of this work. If you know the framework deeply and you're comfortable reasoning about how your code actually behaves on custom accelerator hardware rather than treating it as a black box, this role is built around that skill set specifically.   What You'll Do
  • Train and operate models on Amazon SageMaker with AWS Trainium as the underlying compute
  • Write and optimize PyTorch training code with a real understanding of how it compiles and executes on Trainium (NeuronCore architecture, compiler behavior, memory and throughput tradeoffs)
  • Diagnose training run issues that show up specifically because of the hardware, not just the model, distinguishing a data or code problem from a compiler or device-level one
  • Translate a request for "a Trainium job" into an actual working, cost-aware training pipeline, end to end
  • Tune distributed training runs for throughput and cost on SageMaker's training infrastructure
  • Work directly with client and internal engineering teams to scope and deliver real production training workloads, not experiments that stay in a notebook
What You'll Bring
  • Strong, hands-on PyTorch experience, ideally including distributed or multi-device training
  • Production experience with Amazon SageMaker for training and/or inference
  • Comfort working close to the hardware layer: you understand device-specific compilation and can debug issues that are actually about the accelerator, not just the model
  • AWS Trainium or Inferentia (Neuron SDK) experience is a strong plus; if you don't have it yet but have deep PyTorch and a track record of picking up new hardware targets fast, we want to talk to you
  • Solid Python fundamentals and comfort operating in a client-facing, production engineering environment

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Engineering Manager Related jobs

Other jobs at Robots and Pencils

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.