We are sharing a specialised part-time consulting opportunity for experienced kernel engineers with hands-on expertise in the Neuron Kernel Interface (NKI), AWS Trainium/Inferentia2 hardware, low-level performance optimisation, and CUDA-to-NKI migration.
This role focuses on evaluating NKI kernel-development tasks for technical correctness, hardware appropriateness, numerical fidelity, and performance quality. Selected experts will review Trainium-specific implementations, migration decisions, memory-management strategies, profiling results, and cross-platform numerical behaviour while providing clear, rubric-based technical feedback.
Key Responsibilities
NKI Kernel Development Review
-
Evaluate kernels developed using the Neuron Kernel Interface (NKI)
-
Assess whether implementations appropriately target AWS Trainium and Inferentia2 hardware
-
Review low-level computation patterns for technical correctness
-
Identify inefficient, incorrect, or hardware-inappropriate implementation choices
-
Apply practical judgement grounded in hands-on NKI development experience
CUDA-to-NKI Migration
-
Review migrations of existing CUDA kernels to NKI
-
Assess whether computational semantics are preserved across platforms
-
Identify translation errors, unsupported assumptions, or inefficient migration strategies
-
Evaluate whether NKI implementations appropriately account for Trainium architecture
-
Distinguish faithful migrations from implementations that merely reproduce surface-level CUDA structure
Tile-Based Computation
-
Assess tile decomposition and computation strategies
-
Review partitioning decisions against NKI execution constraints
-
Evaluate whether kernels make effective use of available compute resources
-
Identify inefficient tiling or data-movement patterns
-
Assess whether implementation choices align with NKI programming requirements
Memory Hierarchy Management
-
Review use of SBUF, PSUM, and HBM
-
Assess data placement and movement across Trainium memory hierarchies
-
Evaluate memory-bandwidth utilisation and locality
-
Identify unnecessary transfers or memory bottlenecks
-
Review implementation decisions affecting on-chip memory efficiency
DMA & Data Movement
-
Evaluate DMA orchestration within NKI kernels
-
Review sequencing of computation and data-transfer operations
-
Identify stalls, inefficient transfer patterns, or synchronisation issues
-
Assess whether data movement appropriately overlaps with computation
-
Evaluate implementation choices affecting pipeline utilisation
Trainium Performance Optimisation
-
Review Trainium-specific optimisation strategies
-
Assess NeuronCore pipeline utilisation, tensor-engine throughput, and memory-bandwidth behaviour
-
Identify performance bottlenecks within kernel implementations
-
Evaluate whether optimisation decisions are supported by profiling evidence
-
Review trade-offs affecting throughput, latency, and resource utilisation
Numerical Correctness
-
Evaluate numerical consistency between GPU and Trainium implementations
-
Review differences caused by accumulation order, rounding behaviour, and mixed-precision semantics
-
Assess appropriate tolerances for cross-platform comparisons
-
Identify numerical discrepancies that indicate implementation defects
-
Distinguish expected hardware-level variation from substantive correctness problems
Precision & Data Types
-
Review kernels using supported formats such as FP32, BF16, FP8, and INT8
-
Assess precision choices against computational requirements
-
Evaluate mixed-precision behaviour and numerical stability
-
Identify inappropriate casting or accumulation strategies
-
Review whether performance gains are achieved without compromising required correctness
AWS Neuron Ecosystem
-
Evaluate implementations using the AWS Neuron SDK
-
Review interactions between kernel code, compilation, and Trainium execution
-
Assess compiler-related behaviours where relevant
-
Apply familiarity with NKI kernel libraries and Neuron tooling
-
Identify implementation issues arising from platform-specific constraints
Benchmarking & Validation
-
Review benchmark results for Trainium workloads
-
Assess performance comparisons and experimental methodology
-
Evaluate workloads running on Trn1 or Trn2 instances where applicable
-
Determine whether claimed performance improvements are supported by evidence
-
Identify benchmarking methodologies that could produce misleading conclusions
Rubric-Based Technical Evaluation
-
Assess assigned kernel-development tasks against structured technical criteria
-
Provide clear written explanations supporting evaluation decisions
-
Reference specific implementation, profiling, or numerical evidence
-
Apply evaluation standards consistently across assignments
-
Distinguish valid optimisation alternatives from technically flawed approaches
Ideal Profile
-
2+ years of hands-on experience developing or optimising kernels using the Neuron Kernel Interface (NKI)
-
Professional experience targeting AWS Trainium or Inferentia2 hardware
-
Strong understanding of tile-based computation
-
Deep familiarity with SBUF, PSUM, and HBM memory management
-
Strong knowledge of partition-dimension constraints and DMA orchestration
-
Demonstrated experience evaluating or performing CUDA-to-NKI migrations
-
Familiarity with Trainium-specific performance profiling
-
Experience assessing NeuronCore pipeline utilisation, tensor-engine throughput, and memory-bandwidth bottlenecks
-
Strong understanding of cross-platform numerical correctness and mixed-precision behaviour
-
Direct experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel libraries is preferred
-
Prior CUDA or Triton kernel development experience is advantageous
-
Familiarity with NeuronCore-v2 architecture and supported numerical formats is preferred
-
Experience benchmarking ML workloads on Trn1 or Trn2 instances is advantageous
-
Strong written communication and ability to provide precise technical feedback
Engagement Details
-
Part-time independent contractor engagement
-
Fully remote within the United States
-
Flexible scheduling based on project requirements
-
Compensation: $60–$80/hour
-
Work focuses on NKI kernel development, Trainium performance optimisation, CUDA migration, numerical correctness, and technical quality evaluation
-
Projects may be extended, shortened, or concluded based on project needs and performance
-
Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
-
H1-B and STEM OPT support is unavailable for this engagement
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.