Logo for Anyone AI

GPU Kernel Engineer – CUDA, Triton & Accelerator Performance

Role overview

Qualifications

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels
  • Strong experience with at least two of the following: CUDA, Triton, NKI/AWS Neuron, Pallas/JAX
  • Strong understanding of GPU performance optimization
  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers

Responsibilities

  • Reviewing GPU and accelerator kernel implementations for correctness
  • Comparing outputs against reference implementations
  • Evaluating numerical tolerance thresholds
  • Identifying performance bottlenecks and optimization opportunities

Key facts

Hard skills

Other skills

  • Problem Solving
  • Detail Oriented

About the company

Anyone AI logo

Anyone AI

E-Learning / EdTech

We invest in talent from Latam to bridge the talent gap in AI. Join our AI community: www.anyoneai.com We are AI / ML experts and second-time entrepreneurs, members of the founding team at Deep Vision AI (acquired in early 2020). We've worked with many Fortune 500 companies completing multiple projects in the early days of AI while leveraging remote talent from LatAm. We are VC-backed from day 1 by top global investors like GFC -Global Founders Capital- (investors in Facebook, LinkedIn, Slack, Canva, Trivago, etc), Canvas Ventures (early investor in Coursera), Latitud Fund (the largest community of angel investors for LatAm including investments in QuintoAndar, La Haus, Clara, Platzi, OnTop, Pomelo, etc), among other investors.

Company details

IndustryE-Learning / EdTech
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Anyone AI is recruiting experienced GPU Kernel Engineers for a specialized project focused on reviewing, debugging, and evaluating high-performance compute kernels used in AI workloads.

We’re looking for engineers with hands-on experience writing and optimizing kernels across frameworks such as CUDA, Triton, NKI, or Pallas, with a strong understanding of numerical correctness, GPU performance, memory optimization, and benchmarking.

What You’ll Work On

You’ll work with GPU and accelerator kernel tasks involving:

  • Kernel implementation and debugging

  • CUDA and Triton optimization

  • Translation between kernel frameworks

  • Hardware migration

  • Operator fusion

  • Performance profiling and benchmarking

  • Numerical correctness verification

  • Compilation and runtime debugging

  • Memory hierarchy optimization

  • Kernel-level AI workload performance

You’ll assess whether implementations are technically correct, efficiently designed, reproducible, and appropriately optimized for the target hardware.

What We’re Looking For

  • 3+ years of hands-on experience developing, optimizing, or debugging GPU or accelerator kernels

  • Strong experience with at least two of the following:

    • CUDA

    • Triton

    • NKI / AWS Neuron

    • Pallas / JAX

  • Strong understanding of GPU performance optimization

  • Experience with kernel profiling tools such as Nsight, NCU, roofline analysis, or framework-native profilers

  • Understanding of:

    • Memory bandwidth

    • Compute throughput

    • GPU occupancy

    • Shared memory

    • Register pressure

    • Memory coalescing

    • Bank conflicts

  • Strong understanding of floating-point numerical correctness and tolerance thresholds

  • Experience debugging kernel compilation and runtime issues

  • Ability to distinguish software defects, environment problems, and genuine optimization challenges

Relevant Experience

Candidates should have experience with several of the following types of work:

  • Writing kernels from technical specifications

  • Translating kernels between CUDA, Triton, or other frameworks

  • Migrating kernels across hardware platforms

  • Debugging incorrect kernel implementations

  • Optimizing kernel performance

  • Fusing multiple operations into optimized kernels

Nice to Have

  • Experience across both NVIDIA GPU and custom accelerator ecosystems

  • Experience with AWS Trainium, TPU, JAX, or other accelerators

  • Compiler engineering experience

  • Familiarity with MLIR, XLA, or intermediate representation lowering

  • Contributions to GPU or ML kernel libraries

  • Experience with cuBLAS, cuDNN, Triton community kernels, or JAX/XLA custom calls

  • Experience with AI model evaluation, RLHF, or technical benchmark development

What You’ll Be Responsible For

  • Reviewing GPU and accelerator kernel implementations for correctness

  • Comparing outputs against reference implementations

  • Evaluating numerical tolerance thresholds

  • Reviewing kernel benchmarks and determining whether comparisons are fair

  • Identifying performance bottlenecks and optimization opportunities

  • Assessing whether performance targets are realistic given hardware limits

  • Reviewing kernel translations and hardware migrations

  • Identifying compilation, driver, memory, shape, and runtime issues

  • Determining whether technical tasks are genuinely difficult or incorrectly configured

  • Providing clear, actionable technical feedback

Engagement

Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: GPU kernels, performance engineering, debugging, and technical evaluation

This role is ideal for engineers who enjoy working close to the hardware, optimizing GPU workloads, debugging low-level performance issues, and pushing AI compute systems toward their performance limits.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Anyone AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.