Find your next role
Strengthen your profile
Intergral
Computer Software / SaaS
See how your profile stacks up against this role.
We compared the job requirements to your profile to show where you're strong and where you fall short.
AI Evaluation & Benchmarking Engineer
Location: Remote (United Kingdom)
Salary: £60,000–£75,000 DOE
Contract: Full-time, 40 hours per week
Employer: Intergral UK
Reports to: Director of Engineering
You must already have the right to work in the UK, as we're unable to sponsor visas for this role.
About the Role
We're looking for an AI Evaluation & Benchmarking Engineer to determine how we measure the quality of OpsPilot, and to build the systems that do it.
More important than any individual technology is how you approach measurement. We're looking for someone who questions whether a system is actually achieving the outcome it was designed for, works out how to measure that objectively, and builds what's needed to keep measuring it as the product changes.
You'll have real autonomy over the technical approach. We set the goals and check in regularly, but the expertise on how to get there is yours — we're hiring you because we need someone who can take an ambiguous problem and deliver a working system without being handed the steps. This is a hands-on engineering role with no reports and no QA silo. Our engineering structure is flat: you'll report to the Director of Engineering and work alongside the other engineers as a peer.
This isn't a traditional QA role. You won't be manually testing tickets, acting as a release gatekeeper or simply checking whether features technically work.
About OpsPilot
OpsPilot is an AI-led observability platform helping engineering and operations teams move from monitoring data to evidence-backed operational understanding — so they can investigate problems faster and act with greater confidence.
Intergral has more than 20 years of experience in application performance and observability, with an established customer base built around FusionReactor. OpsPilot is expanding beyond its historic Java and ColdFusion roots into the wider observability market, supporting OpenTelemetry-based metrics, logs and traces.
At the heart of OpsPilot is Coworker, an AI operations capability that continuously investigates telemetry, identifies situations that need attention and provides evidence-backed findings and recommended next steps.
What You'll Do
Evaluate the agent and the product it runs on
Coworker's findings are only as good as the system underneath them. An investigation can fail because the model reasoned badly, because retrieval surfaced the wrong evidence, because ingestion dropped a trace, or because the finding was presented in a way no engineer could act on. Evaluating the agent in isolation would tell us very little, so this role covers both.
On the agentic side, you'll:
Across the wider product, you'll:
When we change a model, prompt, tool or agent workflow, we want to know what became better, what became worse and why — including the impact on quality, reliability, latency and cost.
Turn what we learn into continuous improvement
Longer term, this evaluation system becomes the harness for controlled self-improvement: identifying weaknesses, testing potential changes and objectively determining whether they should be retained. That's the direction of travel rather than the first year's work, but it's why we're building this properly.
Work with the rest of engineering
Benchmarking should provide continuous feedback that helps engineering improve the product. This role is not a release gatekeeper.
What We're Looking For
Above everything else: the ability to take an ambiguous technical problem, develop an approach and deliver a working system independently.
We're more interested in demonstrated ability than an exact number of years, but we'd generally expect around 3+ years of relevant technical experience. Your background might be as an AI engineer, software engineer, SRE, platform engineer, performance engineer, SDET or similar.
Alongside that, we're looking for:
Desirable
Any of the following would be useful, but none are required:
What Success Looks Like
We have the beginnings of an evaluation harness, but the design and expertise are what we're hiring for.
We'd expect the first few months to go into the core evaluation harness and a starting corpus for Coworker's investigation quality, then extend outward across the rest of the product as that proves itself. How you sequence it is your call.
By six months, we should be able to objectively answer questions such as:
Our evaluation corpus should keep growing as we encounter new problems.
Success isn't measured by the number of tests written or percentage test coverage. It's measured by our ability to understand how well OpsPilot is doing its job, where it isn't, and whether the changes we're making are actually making it better.
What We Offer
A small company rather than a large one, with the trade-offs that implies. Under ten people in engineering, a flat structure, and decisions made in a conversation rather than across three meetings. You'll have genuine influence over how this is done, and very little bureaucracy to work through to get there.
Our Interview Process
Straightforward: usually two or three conversations, with no technical coding tests.
If this sounds like the kind of challenge you're looking for, we'd love to hear from you.
Compensation: £60,000–£75,000 DOE
After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.
Marcus Rivera
Chief Revenue Officer

Mercor

Mercor

Mercor

Mercor

Mercor

Intergral