We are sharing a specialised full-time opportunity for experienced technical professionals to operate at the intersection of enterprise AI, applied research, machine-learning evaluation, and real-world AI system performance.
Selected professionals will work directly within enterprise AI workflows to identify real-world failure modes, design high-signal datasets and evaluation frameworks, and run rapid experimental cycles that improve system performance. The role combines forward-deployed research, ML-oriented data design, agentic workflow evaluation, technical analysis, and close collaboration across research, product, domain, and enterprise teams.
Key Responsibilities
Enterprise AI Research & Failure Analysis
-
Embed within enterprise AI workflows as a technical research collaborator
-
Work alongside domain experts and enterprise teams to understand real-world system behaviour
-
Identify, formalise, and prioritise failure modes emerging from deployed AI systems
-
Translate operational issues into structured research questions and measurable technical problems
-
Produce clear analyses of system behaviour, limitations, and opportunities for improvement
ML-Oriented Data & Evaluation Design
-
Design high-signal datasets targeting identified model and system weaknesses
-
Develop evaluation protocols, quality criteria, and structured assessment frameworks
-
Apply strong judgement to data selection, evaluation design, and research-signal quality
-
Identify gaps in existing datasets and evaluation coverage
-
Structure research workflows to support measurable improvements in model performance
Experimentation & Agentic Workflow Evaluation
-
Run rapid experimental cycles to test hypotheses and quantify system improvements
-
Develop and benchmark agentic workflows with a focus on robustness, reliability, and scalability
-
Evaluate AI systems operating across complex enterprise workflows
-
Analyse experimental results and determine whether observed improvements are meaningful and reproducible
-
Iterate on datasets, evaluations, and system configurations based on research findings
Research Tooling & Cross-Functional Collaboration
-
Build lightweight tooling to support evaluation, data curation, experimentation, and rapid iteration
-
Collaborate across research, engineering, product, domain, and enterprise-facing teams
-
Translate research findings into clear, decision-oriented recommendations
-
Contribute to research artifacts including reports, benchmarks, evaluation documentation, and technical analyses
-
Communicate complex findings clearly to both technical and non-technical stakeholders
Ideal Profile
-
Master's degree in Computer Science, Machine Learning, Artificial Intelligence, or a closely related technical discipline
-
Strong judgement regarding research-signal quality, data selection, and evaluation design
-
Experience designing datasets, evaluation frameworks, or QA processes for machine-learning systems
-
Ability to translate ambiguous operational issues into structured research and evaluation problems
-
Familiarity with reinforcement-learning environments, agentic systems, or AI-system evaluation
-
Strong analytical skills and ability to produce concise, actionable technical insights
-
Proven ability to execute effectively within rapid iteration cycles and high-ambiguity environments
-
Strong written and verbal communication skills
-
Collaborative experience across research, product, engineering, and domain teams
-
Client-facing experience within technical or research-focused environments is advantageous
-
Experience building internal research or evaluation tooling is beneficial
-
Contributions to benchmarks, research publications, or open research initiatives are advantageous
-
Exposure to enterprise AI deployments or forward-deployed research environments is strongly valued
Engagement Details
-
Full-time engagement
-
Fully remote
-
Compensation: $300,000–$700,000/year
-
Work will span enterprise AI research, evaluation design, ML-oriented data systems, experimentation, and agentic workflow analysis
-
Responsibilities will involve direct collaboration with research, product, technical, domain, and enterprise stakeholders
-
Research priorities, datasets, evaluation frameworks, and system requirements may evolve based on experimental findings and deployment needs
-
Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy