Logo for 24-MAG

Remote | Member of Technical Staff, Research Engineering — $400,000–$800,000/year

Role overview

Qualifications

  • Deep professional or research experience in reinforcement learning
  • Strong understanding of RL environment design, reward structures, training dynamics, and evaluation
  • Demonstrated experience building and scaling RL systems, training pipelines, or experimentation frameworks
  • Strong experience with automation and synthetic data-generation workflows

Responsibilities

  • Architect self-contained reinforcement-learning environments that capture complex real-world tasks
  • Design and scale episode pipelines and multi-component training processes
  • Build automated data-generation systems using synthetic data to accelerate training cycles
  • Fine-tune and optimise open-source reinforcement-learning and machine-learning models

Key facts

Hard skills

Other skills

  • Quality Assurance
  • Collaboration
  • Communication
  • Problem Solving

About the company

24-MAG logo

24-MAG

Business Consulting & Services

Company details

IndustryBusiness Consulting & Services
Company size2 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

We are sharing a specialised full-time opportunity for experienced Research Engineers with deep expertise in reinforcement learning, ML-oriented data systems, evaluation infrastructure, and scalable experimentation to contribute to advanced AI research and development.

Selected professionals will operate at the intersection of research and production, building reinforcement-learning environments, training pipelines, synthetic data systems, automated evaluation frameworks, and scalable experimentation workflows. The role focuses on translating experimental ideas into robust technical systems that improve model capability, reliability, and research velocity.

Key Responsibilities

Reinforcement Learning Environment Design

  • Architect self-contained reinforcement-learning environments that capture complex real-world tasks
  • Design reward functions, verifiers, evaluation logic, and supporting environment components
  • Structure environments to support reliable experimentation and measurable model improvement
  • Translate research objectives into technically rigorous RL workflows
  • Ensure environments remain reproducible, testable, and suitable for iterative model development

Training Pipelines & Experimentation Systems

  • Design and scale episode pipelines and multi-component training processes
  • Build reproducible experimentation workflows supporting reinforcement-learning research
  • Develop systems for running, tracking, and analysing large-scale training experiments
  • Improve reliability and efficiency across RL training infrastructure
  • Support rapid iteration between environment design, training, evaluation, and model refinement

Synthetic Data & Automated Evaluation

  • Build automated data-generation systems using synthetic data to accelerate training cycles
  • Develop AI-driven evaluation and quality-assurance systems for grading, validation, and feedback
  • Establish automated feedback loops that improve training-data and model quality
  • Design verification systems that distinguish strong model behaviour from superficially plausible outputs
  • Apply rigorous quality standards throughout data-generation and evaluation pipelines

Model Optimisation & Benchmarking

  • Fine-tune and optimise open-source reinforcement-learning and machine-learning models
  • Apply internally generated datasets and custom training strategies to improve model performance
  • Develop benchmarking frameworks measuring capability, robustness, and data quality
  • Analyse model behaviour across internal and external evaluation environments
  • Contribute to the development, release, and interpretation of research evaluations and benchmark results

Ideal Profile

  • Deep professional or research experience in reinforcement learning
  • Strong understanding of RL environment design, reward structures, training dynamics, and evaluation
  • Demonstrated experience building and scaling RL systems, training pipelines, or experimentation frameworks
  • Strong experience with automation and synthetic data-generation workflows
  • Familiarity with automated evaluation, model validation, and quality-assurance systems
  • Experience fine-tuning and evaluating open-source machine-learning models
  • Strong technical writing and communication skills
  • Ability to operate effectively in fast-paced, research-driven, and highly collaborative environments
  • Experience publishing benchmarks, evaluations, or research artifacts is advantageous
  • Familiarity with modern evaluation ecosystems and benchmarking frameworks is beneficial
  • Experience with scalable infrastructure supporting large-scale RL experimentation is strongly valued

Engagement Details

  • Full-time engagement
  • Fully remote
  • Compensation: $400,000–$800,000/year
  • Work will involve reinforcement-learning environment design, training pipelines, synthetic data generation, automated evaluation, model optimisation, and benchmarking
  • Responsibilities will span both research experimentation and production-oriented technical implementation
  • Research priorities, evaluation systems, and experimentation workflows may evolve as project requirements develop
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Technical Product Manager Related jobs

Other jobs at 24-MAG

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.