Logo for Gramian Consulting

Technical AI Evaluation Analyst (LATAM)

Role overview

Qualifications

  • Comfortable reading and interpreting Python, SQL, shell scripts, structured data, and execution logs.
  • Experience reviewing technical workflows, software behavior, data outputs, or evaluation logic.
  • Strong analytical skills, including the ability to verify calculations and reconcile conflicting evidence.
  • Demonstrated ability to assess the correctness and completeness of technical deliverables.

Responsibilities

  • Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.
  • Review AI agent execution traces, tool calls, and deliverables to determine whether outcomes are justified.
  • Audit grading logic to identify brittle checks, incorrect expected answers, unsupported rubric criteria, and unfair penalties for valid alternative solutions.
  • Investigate discrepancies between model performance, grader results, and expected outcomes.

Key facts

Hard skills

Other skills

  • Analytical Skills
  • Quality Assurance
  • Detail Oriented
  • Communication

About the company

Gramian Consulting logo

Gramian Consulting

IT Services & IT Consulting

Unknown

Company details

IndustryIT Services & IT Consulting
Company size2 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Gramian

Gramian Consultancy is a boutique consultancy specializing in IT professional services and engineering talent solutions. With a strong background in software engineering and leadership, we help companies build high-performing teams by matching them with professionals who truly fit their needs.

About the Role

We are seeking a detail-oriented AI Evaluation & Quality Assurance Specialist to review the quality, correctness, and fairness of tasks designed to evaluate AI agents. You will inspect reference solutions, grading logic, execution traces, and generated deliverables to identify task defects, evaluation errors, and unjustified model failures. This role requires strong technical fluency, independent analytical judgment, and the ability to produce clear, evidence-based feedback.

LOCATION: Remote – Latin America (LATAM)

CONTRACT: Hourly Contractor

COMMITMENT: 40 hours per week

TIME OVERLAP: 8 hours of mandatory PST overlap

DURATION: 10 weeks

START DATE: Immediately

Key Responsibilities

  • Validate task quality by checking instructions, source materials, reference solutions, and evaluation criteria for consistency and completeness.
  • Review AI agent execution traces, tool calls, and deliverables to determine whether outcomes are justified.
  • Audit grading logic to identify brittle checks, incorrect expected answers, unsupported rubric criteria, and unfair penalties for valid alternative solutions.
  • Investigate discrepancies between model performance, grader results, and expected outcomes.
  • Distinguish genuine model limitations from task defects, grader errors, and environment or tool failures.
  • Independently assess automated QC findings rather than accepting them without verification.
  • Document concise, evidence-backed findings and provide actionable, reproducible feedback.
  • Flag uncertainty and verify that implemented revisions resolve previously identified issues.

Requirements

  • Comfortable reading and interpreting Python, SQL, shell scripts, structured data, and execution logs.
  • Experience reviewing technical workflows, software behavior, data outputs, or evaluation logic.
  • Strong analytical skills, including the ability to verify calculations and reconcile conflicting evidence.
  • Demonstrated ability to assess the correctness and completeness of technical deliverables.
  • Strong written English with experience providing clear, specific, and reproducible feedback.
  • High attention to detail when identifying inconsistencies, missing information, and evaluation defects.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Gramian Consulting

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.