We are sharing a specialised part-time consulting opportunity for experienced software engineers with strong expertise in Python, Java, Rust, C++, Go, TypeScript, algorithms, debugging, feature implementation, refactoring, and performance optimisation to contribute to an advanced AI training and software-engineering evaluation project.
Selected professionals will create reinforcement-learning environments that test an AI model's ability to solve complex software-engineering problems using Model Context Protocol (MCP) tools. The work combines realistic codebase tasks, tool-based reasoning, deterministic verification, and the creation of high-quality reference solutions. No prior experience in AI is required.
Key Responsibilities
Reinforcement Learning Environment Development
-
Create reproducible environments for evaluating advanced software-engineering capability
-
Design tasks requiring agents to interact with and reason over real MCP servers
-
Build scenarios that test practical engineering ability rather than isolated code generation
-
Ensure environments accurately measure both tool use and software-engineering performance
-
Maintain consistency and reproducibility across evaluation runs
Software Engineering Task Design
-
Create realistic tasks involving bug fixing, feature implementation, refactoring, and performance optimisation
-
Develop scenarios based on existing codebases and practical engineering constraints
-
Design tasks that require meaningful reasoning across algorithms, data structures, and system behaviour
-
Define clear success criteria and expected outcomes
-
Balance technical complexity with reliable evaluability
Golden Solutions & Verification
-
Develop high-quality golden reference solutions
-
Create deterministic verification logic for task completion
-
Validate expected behaviours, edge cases, and failure conditions
-
Ensure evaluation systems distinguish correct from partially correct implementations
-
Maintain stable and reproducible grading criteria
Debugging, Refactoring & Performance
-
Diagnose and resolve complex software issues
-
Implement maintainable features in existing codebases
-
Refactor code while preserving intended functionality
-
Identify performance bottlenecks and scalability issues
-
Apply strong engineering judgement to maintainability, efficiency, and code quality
Technical Review & Engineering Standards
-
Review task quality, code correctness, and verification robustness
-
Contribute to software-engineering best practices
-
Participate in rigorous technical and code-review workflows
-
Document assumptions, design decisions, and implementation details clearly
-
Collaborate effectively with distributed technical teams
Ideal Profile
-
Strong proficiency in one or more of C++, Python, Java, Go, TypeScript, or Rust
-
Deep understanding of algorithms, data structures, and performance tuning
-
Demonstrated experience debugging complex software issues
-
Strong background in feature implementation and codebase refactoring
-
Proven ability to optimise software for performance and scalability
-
Experience working with large-scale or distributed codebases is highly valuable
-
Familiarity with rigorous code-review processes and software-engineering standards
-
Strong written and verbal communication skills
-
High attention to technical detail and reproducibility
-
Familiarity with modern AI or machine-learning systems is beneficial but not required
-
Prior experience in AI training or model evaluation is not required
Engagement Details
-
Part-time independent contractor engagement
-
Fully remote
-
Compensation: $60–$120/hour
-
Expected commitment: approximately 15 hours per week
-
Schedule is flexible, including evenings or weekends if preferred
-
Compensation is output-based, with payment made for tasks that meet project specifications
-
Minimum weekly submission requirements apply
-
Work will involve reinforcement-learning environment design, MCP tool use, software-engineering task creation, deterministic verification, and golden reference solutions
-
The selection process may include screening questions, an approximately 30-minute AI interview, a technical assessment, and hiring-manager review
-
Selected professionals should be prepared to begin their first tasks within approximately 24–48 hours of completing onboarding
-
Roles are typically filled within approximately 48 hours
-
Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy