Logo for Anyone AI

AWS Trainium / NKI Kernel Expert

Role overview

Qualifications

  • 2+ years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface (NKI)
  • Experience working with AWS Trainium and/or Inferentia2
  • Strong understanding of tile-based computation, SBUF / PSUM / HBM memory hierarchy, and partition dimension constraints
  • Ability to evaluate CUDA → NKI migrations

Responsibilities

  • Review and evaluate technical tasks involving NKI kernel correctness and Trainium-specific development patterns
  • Manage CUDA → NKI kernel migrations and Trainium performance optimization
  • Analyze memory management across SBUF, PSUM, and HBM
  • Provide technical feedback and quality assessment of kernel implementations

Key facts

Hard skills

Other skills

  • Analytical Skills

About the company

Anyone AI logo

Anyone AI

E-Learning / EdTech

We invest in talent from Latam to bridge the talent gap in AI. Join our AI community: www.anyoneai.com We are AI / ML experts and second-time entrepreneurs, members of the founding team at Deep Vision AI (acquired in early 2020). We've worked with many Fortune 500 companies completing multiple projects in the early days of AI while leveraging remote talent from LatAm. We are VC-backed from day 1 by top global investors like GFC -Global Founders Capital- (investors in Facebook, LinkedIn, Slack, Canva, Trivago, etc), Canvas Ventures (early investor in Coursera), Latitud Fund (the largest community of angel investors for LatAm including investments in QuintoAndar, La Haus, Clara, Platzi, OnTop, Pomelo, etc), among other investors.

Company details

IndustryE-Learning / EdTech
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Anyone AI is recruiting experienced AWS Trainium / Neuron Kernel Interface (NKI) engineers for a specialized project focused on evaluating and improving kernel development tasks for AI workloads.

We’re looking for engineers with hands-on experience building or optimizing NKI kernels on AWS Trainium or Inferentia2 hardware who understand how Trainium’s architecture differs from traditional GPU programming.

What You’ll Work On

You’ll review and evaluate technical tasks involving:

  • NKI kernel correctness and Trainium-specific development patterns

  • CUDA → NKI kernel migrations

  • Trainium performance optimization and benchmarking

  • Memory management across SBUF, PSUM, and HBM

  • Tile-based computation and DMA scheduling

  • Cross-platform numerical correctness between CUDA/Triton and NKI

  • Trainium-specific performance bottlenecks and optimization opportunities

  • Technical feedback and quality assessment of kernel implementations

The work involves determining whether implementations are not only technically correct, but also idiomatic and optimized for Trainium hardware rather than simply translated from GPU-based approaches.

What We’re Looking For

  • 2+ years of hands-on experience developing or optimizing kernels with the Neuron Kernel Interface (NKI)

  • Experience working with AWS Trainium and/or Inferentia2

  • Strong understanding of:

    • Tile-based computation

    • SBUF / PSUM / HBM memory hierarchy

    • Partition dimension constraints

    • DMA orchestration

    • Trainium-specific optimization techniques

  • Ability to evaluate CUDA → NKI migrations

  • Experience profiling and optimizing workloads on Trainium

  • Understanding of numerical differences across GPU and Trainium backends

  • Strong ability to analyze complex technical implementations and provide clear written feedback

Nice to Have

  • Experience with the AWS Neuron SDK or Neuron Compiler

  • CUDA or Triton kernel development experience

  • Knowledge of NeuronCore-v2 architecture

  • Experience with FP32, BF16, FP8, and INT8 workloads

  • Experience benchmarking workloads on Trn1 or Trn2 instances

  • Familiarity with nki.language, @nki.jit, or XLA custom calls

  • Experience with technical evaluation, AI/ML data projects, RLHF, or rubric-based assessment

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: AWS Trainium / NKI kernel engineering and technical evaluation

This is a strong fit for engineers who have worked deeply with AWS Trainium infrastructure and low-level ML kernel optimization and are interested in applying that expertise to technically challenging AI projects.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Anyone AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.