Logo for SASH

Research Engineer - Neo

Role overview

Qualifications

  • Strong Python and general software-engineering skills
  • Experience working with language models, agentic systems, or model evaluations
  • The ability to turn an underspecified research question into a reliable experimental setup
  • Strong engineering practices, particularly around reproducibility, testing, observability, and data integrity

Responsibilities

  • Design, implement, and run evaluations of misalignment and loss-of-control risks in frontier models
  • Build realistic agent environments, scaffolds, and tool integrations for studying behaviour under increasing autonomy
  • Develop infrastructure for long-horizon experiments involving many model calls, actions, tools, and environment states
  • Own experimental reliability and reproducibility across model access, sampling, configuration, environment management, logging, and analysis

Key facts

Hard skills

Other skills

  • Communication
  • Problem Solving
  • Adaptability

About the company

SASH logo

SASH

Computer Software / SaaS

Singapore AI Safety Hub is a co-working, events and community space for people working on or interested in AI safety. SASH’s mission is to strengthen and grow the AI safety community in Singapore through community, building awareness, upskilling talent and facilitating international cooperation. Although only launched in February 2025, SASH already houses established researchers from internationally renowned AI safety organisations such as FAR.AI,Truthful AI, Apart Research, The Future Society and Impact Academy working on technical and governance research. Find out more at www.aisafety.sg

Company details

Company typeTPE
IndustryComputer Software / SaaS
Company size1 - 1

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About the team

Neo Research (新衡) is an independent AI safety research organization based in Singapore. We study frontier risks in increasingly capable AI systems, with a particular focus on the open-weight model ecosystem and the rapidly growing frontier-model ecosystem in Asia.

Some of the world’s most capable open-weight models are now being developed in Asia and deployed globally. Yet, they remain poorly understood from a frontier-safety perspective. We want to understand how to ensure their safety, and what new risks become important as the frontier changes.

Our current work focuses on misalignment, loss of control, and harmful manipulation. We study questions such as whether models pursue unintended objectives, conceal problematic behaviour, recognize and adapt to evaluations, evade oversight, or become less safe as they are given greater autonomy. Our goal is to produce rigorous empirical evidence about risks that are important but difficult to measure.

We are looking for Research Engineers to build the experimental systems needed to evaluate increasingly capable models: realistic agent environments, model and tool integrations, long-horizon evaluation infrastructure, and the systems needed to make complex experiments reliable and reproducible.

Why join Neo Research

Evaluating increasingly capable models requires more than running benchmarks. It means building realistic environments, giving models tools and autonomy, running experiments that may unfold over hundreds of interactions, and capturing enough information to distinguish genuine model behaviour from quirks or failures of the evaluation itself.

You will work closely with researchers while owning substantial parts of the experimental system,from agent scaffolds and tool integrations to model sampling, observability, trajectory analysis, and reproducibility. In this role, you will go beyond implementing evaluations and actively help shape evaluation methodology.

Your work will contribute directly to published research shared with AI Safety Institutes, frontier labs, policymakers, and model developers. We have presented our work to most major Chinese model developers and run a joint evaluations project with an AI Safety Institute. Our mission statement describes our current research directions.

In this role, you would:

  • Design, implement, and run evaluations of misalignment and loss-of-control risks in frontier models.

  • Build realistic agent environments, scaffolds, and tool integrations for studying behaviour under increasing autonomy.

  • Develop infrastructure for long-horizon experiments involving many model calls, actions, tools, and environment states.

  • Own experimental reliability and reproducibility across model access, sampling, configuration, environment management, logging, and analysis.

  • Work with research scientists to turn open-ended questions into tractable experiments, and identify cases where an apparent model behaviour is actually an artefact of the evaluation.

  • Build tools for inspecting and analysing large collections of model trajectories and comparing behaviour across models and experimental conditions.

  • Contribute experimental methodology, infrastructure, and technical analysis directly to published research.

About you

Essentials

  • Strong Python and general software-engineering skills.

  • Experience working with language models, agentic systems, or model evaluations.

  • The ability to turn an underspecified research question into a reliable experimental setup.

  • Strong engineering practices, particularly around reproducibility, testing, observability, and data integrity.

  • The ability to investigate unexpected results across both the model and the surrounding infrastructure.

  • Clear technical writing and communication skills.

  • Comfort working in an early-stage environment where requirements and research directions evolve.

Helpful, but not required

  • Experience with AI safety or dangerous-capability evaluations.

  • Experience with evaluation frameworks such as Inspect.

  • Experience building agent scaffolds or tool-use environments.

  • Experience with distributed inference, large-scale API-based experimentation, Docker, Kubernetes, or related infrastructure.

  • Experience analysing large collections of model trajectories or transcripts.

  • Familiarity with frontier-model safety reports.

  • Mandarin reading or writing ability.

You don't need to have worked in AI safety specifically. We're also interested in strong ML and research engineers whose technical expertise and engineering judgment could transfer strongly to this work.

Role logistics & benefits

Location:

This role can be based in Singapore or remotely.

Our team is globally distributed, so remote team members should be comfortable maintaining some working-hour overlap with colleagues across regions.

Compensation:

Our compensation takes location, experience, and level into account, with indicative salary ranges of $120,000–$180,000+. We may offer above this range for exceptional candidates.

Benefits:

Competitive benefits and leave policies.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Research Engineer Related jobs

Other jobs at SASH

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.