Logo for 24-MAG

Remote | Machine Learning Engineer — $60–$80/hour

Role overview

Qualifications

  • 3+ years of hands-on applied or experimental machine learning experience
  • Strong practical experience with experiment design, model selection, hyperparameter tuning, and evaluation methodology
  • Deep understanding of data leakage, metric gaming, train/test methodology, and cross-validation hygiene
  • Proficiency with standard ML frameworks such as PyTorch, TensorFlow, scikit-learn, or XGBoost

Responsibilities

  • Evaluate applied machine learning experiments for methodological soundness
  • Review experimental hypotheses, assumptions, and modelling choices
  • Identify weaknesses that could invalidate or distort conclusions
  • Reproduce or validate experimental results where required

Key facts

  • Remote from: New York (USA)
  • Freelance
  • Mid-level (2-5 years)
  • Machine Learning Engineer
  • English

Hard skills

Other skills

  • Communication

About the company

24-MAG logo

24-MAG

Business Consulting & Services

Company details

IndustryBusiness Consulting & Services
Company size2 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

We are sharing a specialised part-time consulting opportunity for experienced Machine Learning professionals with strong hands-on expertise in experiment design, model selection, hyperparameter tuning, data quality, and rigorous evaluation methodology.

This role focuses on reviewing applied machine learning tasks for technical correctness, methodological rigour, and reproducibility. Selected experts will assess experimental design, modelling decisions, validation practices, metrics, and supporting evidence while identifying issues such as data leakage, metric gaming, or unreliable conclusions.

Key Responsibilities

Machine Learning Experiment Review

  • Evaluate applied machine learning experiments for methodological soundness
  • Review experimental hypotheses, assumptions, and modelling choices
  • Assess whether experiments are appropriately designed to answer the stated question
  • Identify weaknesses that could invalidate or distort conclusions
  • Apply practical judgement grounded in hands-on experimental ML experience

Model Selection & Tuning

  • Review model-selection decisions and supporting rationale
  • Evaluate hyperparameter-tuning approaches and search strategies
  • Assess whether model comparisons are fair and methodologically appropriate
  • Identify overfitting, cherry-picking, or poorly justified modelling choices
  • Determine whether conclusions are supported by experimental results

Training & Validation Methodology

  • Evaluate train/test splits, cross-validation strategies, and validation procedures
  • Identify inappropriate data partitioning or evaluation practices
  • Review whether datasets and experimental protocols support reliable generalisation
  • Detect contamination between training, validation, and test data
  • Assess whether evaluation procedures match the underlying ML problem

Data Quality & Leakage Detection

  • Review datasets and preprocessing workflows for potential quality issues
  • Identify data leakage, target leakage, or unintended information exposure
  • Assess feature-engineering and preprocessing decisions
  • Detect methodological shortcuts that could inflate reported performance
  • Evaluate whether data handling supports reliable experimentation

Metrics & Performance Evaluation

  • Review metric selection against the task objective
  • Assess whether reported metrics appropriately capture model performance
  • Identify metric gaming, misleading optimisation targets, or incomplete evaluation
  • Review performance comparisons and supporting statistical evidence
  • Determine whether claimed improvements are meaningful and defensible

Reproducibility & Evidence Review

  • Reproduce or validate experimental results where required
  • Assess whether reported findings can be independently reproduced
  • Review code, configuration, experimental settings, and supporting evidence
  • Identify inconsistencies between claims and observed results
  • Determine whether conclusions follow logically from the available evidence

Machine Learning Frameworks

  • Work with experiments built using standard ML frameworks and libraries
  • Review implementations involving PyTorch, TensorFlow, scikit-learn, and XGBoost
  • Evaluate model-training and experimentation workflows
  • Identify implementation choices that may compromise experimental validity
  • Apply framework-specific knowledge when assessing technical quality

Benchmark & Challenge Evaluation

  • Review applied ML tasks and benchmark-style challenges
  • Assess whether tasks measure the intended modelling capability
  • Evaluate competition-style experimental approaches where relevant
  • Identify task-design issues that could reward shortcuts rather than genuine modelling quality
  • Apply experience from benchmarking or competitive ML environments where applicable

Rubric-Based Technical Evaluation

  • Assess assigned ML tasks against structured technical criteria
  • Provide clear written explanations supporting evaluation decisions
  • Identify specific methodological or implementation evidence behind each judgement
  • Apply grading standards consistently across assignments
  • Distinguish genuine methodological problems from reasonable alternative approaches

Ideal Profile

  • 3+ years of hands-on applied or experimental machine learning experience
  • Strong practical experience with experiment design, model selection, hyperparameter tuning, and evaluation methodology
  • Deep understanding of data leakage, metric gaming, train/test methodology, and cross-validation hygiene
  • Proficiency with standard ML frameworks such as PyTorch, TensorFlow, scikit-learn, or XGBoost
  • Strong ability to critique machine learning claims against experimental evidence
  • Comfortable reproducing results and diagnosing discrepancies
  • Strong quantitative and analytical judgement
  • Clear written communication and ability to provide precise technical feedback
  • Kaggle, ML competition, or benchmark experience is advantageous
  • Graduate research or publication experience in applied machine learning is preferred
  • Previous task-grading, technical peer-review, or ML evaluation experience is advantageous
  • This opportunity is focused on applied and experimental machine learning, rather than LLM application development or MLOps

Engagement Details

  • Part-time independent contractor engagement
  • Fully remote within the United States
  • Flexible scheduling based on project requirements
  • Compensation: $60–$80/hour
  • Work focuses on applied machine learning experimentation, model evaluation, methodological review, reproducibility, and technical quality assessment
  • Projects may be extended, shortened, or concluded based on project needs and performance
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
  • H1-B and STEM OPT support is unavailable for this engagement

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Machine Learning Engineer Related jobs

Other jobs at 24-MAG

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.