We are sharing a specialised part-time consulting opportunity for experienced Software Engineers with hands-on open-source contribution or maintainer experience and strong expertise in repository-level code review, testing, debugging, and software quality evaluation.
This role focuses on reviewing software-engineering benchmark tasks for correctness, reproducibility, and grading integrity. Selected engineers will audit repository-level assignments, reference patches, test harnesses, containerised environments, and evaluation logic while identifying technical flaws, unintended shortcuts, and weaknesses in task design.
Key Responsibilities
Repository-Level Code Review
-
Review software engineering tasks built around real code repositories
-
Assess whether task requirements are technically clear, complete, and reproducible
-
Evaluate repository state, dependencies, configuration, and expected behaviour
-
Identify ambiguities or implementation issues that could affect task validity
-
Apply practical engineering judgement to realistic codebase-level problems
Reference Patch Auditing
-
Review reference patches for correctness and completeness
-
Determine whether proposed solutions appropriately address the underlying software issue
-
Identify unintended behavioural changes, incomplete fixes, or unsupported assumptions
-
Compare reference implementations against task requirements and expected outcomes
-
Assess whether alternative valid implementations are treated fairly
Test Harness & Grading Review
-
Audit test runners and automated evaluation logic
-
Assess whether tests accurately measure the intended behaviour
-
Identify missing coverage, brittle assertions, or grading inconsistencies
-
Verify that evaluation criteria appropriately distinguish correct from incorrect solutions
-
Review benchmark tasks for reliable and repeatable scoring
Reproducibility & Environment Validation
-
Evaluate whether tasks can be reproduced consistently across clean environments
-
Review dependency installation, build processes, configuration, and runtime requirements
-
Assess Docker-based isolation and containerised execution
-
Identify environmental dependencies or hidden assumptions affecting reproducibility
-
Verify that tasks execute reliably under their intended setup
Benchmark Integrity
-
Identify potential answer leakage, unintended shortcuts, or reward-hacking opportunities
-
Evaluate whether benchmark structure exposes information that makes tasks artificially easy
-
Review task and grading design for loopholes or exploitable behaviours
-
Assess whether successful completion genuinely demonstrates the intended engineering capability
-
Recommend improvements where benchmark integrity is compromised
Software Testing & Debugging
-
Investigate failing or inconsistent benchmark tasks
-
Review stack traces, logs, test failures, and repository behaviour
-
Identify root causes of technical issues
-
Distinguish task defects from legitimate implementation failures
-
Assess whether debugging and validation processes follow sound engineering practices
Open-Source Engineering
-
Apply experience from contributing to or maintaining open-source software
-
Evaluate repository conventions, contribution patterns, and realistic development workflows
-
Review patches with the perspective of an experienced contributor or maintainer
-
Assess whether proposed changes would meet reasonable code-review expectations
-
Apply practical judgement derived from real-world pull request and repository experience
Multi-Language Code Evaluation
-
Review software written in Python
-
Evaluate tasks involving at least one additional ecosystem such as Java, Go, TypeScript, or C++
-
Assess code structure, tests, implementation choices, and repository conventions across languages
-
Identify language-specific implementation or testing issues
-
Apply consistent engineering standards across different technology stacks
Rubric-Based Evaluation
-
Assess benchmark tasks against structured technical criteria
-
Provide clear written explanations supporting evaluation decisions
-
Reference specific code, tests, patches, or execution behaviour when identifying issues
-
Apply grading standards consistently across assignments
-
Distinguish substantive benchmark defects from minor implementation differences
Ideal Profile
-
3+ years of professional software engineering experience
-
Demonstrated open-source contribution or maintainer experience, such as merged pull requests, committer responsibilities, or maintainer roles
-
Strong ability to review repository-level software changes
-
Experience auditing reference patches, test runners, and automated test suites
-
Comfortable evaluating Docker-based isolation and reproducible development environments
-
Strong understanding of software testing, debugging, and code-review practices
-
Ability to identify answer leakage, reward hacking, or other benchmark-integrity issues
-
Strong proficiency in Python
-
Professional fluency in at least one additional language such as Java, Go, TypeScript, or C++
-
Familiarity with SWE-Bench Verified or similar repository-level software engineering benchmarks is preferred
-
Maintainer or contributor history with established Python open-source projects is highly valued
-
Previous code-review, software evaluation, or task-grading experience is advantageous
-
Strong written communication and ability to provide precise technical feedback
Engagement Details
-
Part-time independent contractor engagement
-
Fully remote within the United States
-
Flexible scheduling based on project requirements
-
Compensation: $60–$80/hour
-
Work focuses on repository-level software evaluation, reference-patch review, testing, reproducibility, benchmark integrity, and technical quality assessment
-
Projects may be extended, shortened, or concluded based on project needs and performance
-
Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
-
H1-B and STEM OPT support is unavailable for this engagement
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.