We are sharing a specialised consulting opportunity for experienced CUDA Engineering Experts with strong expertise in CUDA, C++, GPU kernel optimisation, GLSL, WebGPU, profiling, and high-performance computing to contribute to an advanced GPU-engineering and AI training project.
Selected professionals will analyse and optimise GPU kernels, improve CUDA and C++ codebases, work with shader and WebGPU workflows, and document performance improvements across modern GPU architectures. The work requires strong low-level performance engineering, rigorous profiling, and the ability to translate technical findings into clear, actionable recommendations. No prior experience in AI is required.
Key Responsibilities
CUDA Kernel Analysis & Optimisation
-
Analyse and optimise GPU kernels using CUDA
-
Identify performance bottlenecks affecting computational throughput
-
Improve kernel efficiency across modern GPU hardware
-
Apply targeted optimisation techniques based on profiling results
-
Validate performance gains using quantitative benchmarks
GPU Profiling & Performance Engineering
-
Profile GPU workloads using tools such as Nsight, Visual Profiler, or comparable platforms
-
Diagnose memory, compute, occupancy, and execution bottlenecks
-
Evaluate kernel performance across different hardware generations
-
Develop data-driven optimisation strategies
-
Document measurable changes in latency, throughput, and resource utilisation
C++ & CUDA Codebase Refactoring
-
Refactor CUDA and C++ code for improved maintainability and efficiency
-
Improve architecture and organisation within performance-critical codebases
-
Reduce unnecessary complexity while preserving functionality
-
Adapt implementations for portability across GPU architectures
-
Apply high-performance computing best practices throughout development
GLSL & WebGPU Development
-
Implement shader logic using GLSL
-
Develop graphics and compute workflows using WebGPU
-
Integrate shader-based processing into existing systems
-
Evaluate performance trade-offs across GPU execution environments
-
Maintain compatibility and consistency across pipeline components
Technical Evaluation & Design
-
Contribute expertise to GPU architecture and performance-design discussions
-
Evaluate new GPU-based implementation approaches
-
Assess technical trade-offs across performance, maintainability, and scalability
-
Support definition of meaningful performance metrics
-
Recommend practical approaches based on profiling and benchmarking evidence
Technical Documentation & Collaboration
-
Document optimisation strategies, benchmark results, and technical findings
-
Produce clear reports describing performance improvements
-
Communicate complex GPU behaviour to technical stakeholders
-
Collaborate with remote and cross-disciplinary project teams
-
Share relevant developments in GPU programming and performance engineering
Ideal Profile
-
Demonstrated expertise in CUDA programming
-
Strong track record of GPU kernel performance optimisation
-
Advanced C++ development experience
-
Experience working in high-performance computing environments
-
Hands-on experience with GLSL and WebGPU
-
Strong understanding of graphics or compute shader development
-
Proficiency with GPU profiling tools such as Nsight, Visual Profiler, or equivalent
-
Ability to reason about memory behaviour, kernel execution, and hardware utilisation
-
Strong analytical skills for evaluating performance across GPU architectures
-
Experience refactoring performance-critical codebases
-
Excellent written and verbal technical communication skills
-
Comfortable collaborating in remote, cross-disciplinary environments
-
No prior AI-training experience is required
Engagement Details
-
Independent contractor engagement
-
Fully remote
-
Compensation: $60–$100/hour
-
Work will involve CUDA kernel optimisation, GPU profiling, C++ development, GLSL, WebGPU, and performance analysis
-
Strong low-level GPU performance expertise is central to this engagement
-
Assignments may involve kernel benchmarking, bottleneck analysis, codebase refactoring, shader development, and architectural evaluation
-
Project scope, workload, hardware targets, and performance requirements may evolve depending on project needs
-
Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy