Logo for Intergral

AI Evaluation & Benchmarking Engineer

Role overview

Qualifications

  • 3+ years of relevant technical experience
  • Practical experience working with LLMs, AI agents or AI evaluation
  • Strong software engineering skills, particularly Python or a similar language
  • Experience building automated evaluation, benchmarking, testing or experimentation infrastructure

Responsibilities

  • Evaluate the agent and the product it runs on
  • Build automated evaluations for our agentic AI capabilities
  • Turn failures and real-world problems into new evaluation scenarios
  • Work directly with AI, software, platform and SRE engineers to investigate findings and improve the product

Key facts

Hard skills

Other skills

  • Problem Solving
  • Analytical Thinking
  • Collaboration
  • Communication

About the company

Intergral logo

Intergral

Computer Software / SaaS

Intergral builds observability and AI-driven operations software for engineering and IT operations teams. For more than 20 years, we have helped organizations improve application reliability, accelerate troubleshooting and gain deeper operational visibility across complex software environments. Our platforms help organizations gain deep operational visibility across modern applications, cloud-native infrastructure, distributed services and enterprise workloads running both on-premises and in the cloud. Intergral is the company behind: • FusionReactor APM - application performance monitoring for Java and ColdFusion applications • OpsPilot - an AI-led observability platform and AI SRE teammate designed to help teams investigate incidents, reduce operational noise, and turn telemetry into action Built around OpenTelemetry, Grafana and modern observability standards, our solutions help teams move beyond dashboards and alerts toward proactive operational intelligence. Engineering and operations teams use Intergral products to: • Detect and resolve issues faster • Reduce downtime and improve reliability • Accelerate root cause analysis • Optimize observability and infrastructure costs • Improve operational efficiency across distributed systems Intergral serves customers worldwide across enterprise, SaaS, financial services, education and digital commerce sectors. Headquartered in Germany, Intergral operates internationally through offices and entities across Europe and North America.

Company details

Company typeStartup
IndustryComputer Software / SaaS
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

AI Evaluation & Benchmarking Engineer

Location: Remote (United Kingdom)
Salary: £60,000–£75,000 DOE
Contract: Full-time, 40 hours per week
Employer: Intergral UK
Reports to: Director of Engineering

You must already have the right to work in the UK, as we're unable to sponsor visas for this role.

About the Role

We're looking for an AI Evaluation & Benchmarking Engineer to determine how we measure the quality of OpsPilot, and to build the systems that do it.

More important than any individual technology is how you approach measurement. We're looking for someone who questions whether a system is actually achieving the outcome it was designed for, works out how to measure that objectively, and builds what's needed to keep measuring it as the product changes.

You'll have real autonomy over the technical approach. We set the goals and check in regularly, but the expertise on how to get there is yours — we're hiring you because we need someone who can take an ambiguous problem and deliver a working system without being handed the steps. This is a hands-on engineering role with no reports and no QA silo. Our engineering structure is flat: you'll report to the Director of Engineering and work alongside the other engineers as a peer.

This isn't a traditional QA role. You won't be manually testing tickets, acting as a release gatekeeper or simply checking whether features technically work.

About OpsPilot

OpsPilot is an AI-led observability platform helping engineering and operations teams move from monitoring data to evidence-backed operational understanding — so they can investigate problems faster and act with greater confidence.

Intergral has more than 20 years of experience in application performance and observability, with an established customer base built around FusionReactor. OpsPilot is expanding beyond its historic Java and ColdFusion roots into the wider observability market, supporting OpenTelemetry-based metrics, logs and traces.

At the heart of OpsPilot is Coworker, an AI operations capability that continuously investigates telemetry, identifies situations that need attention and provides evidence-backed findings and recommended next steps.

What You'll Do

Evaluate the agent and the product it runs on

Coworker's findings are only as good as the system underneath them. An investigation can fail because the model reasoned badly, because retrieval surfaced the wrong evidence, because ingestion dropped a trace, or because the finding was presented in a way no engineer could act on. Evaluating the agent in isolation would tell us very little, so this role covers both.

On the agentic side, you'll:

  • Build automated evaluations for our agentic AI capabilities.
  • Create realistic synthetic scenarios, datasets and workloads with meaningful ground truth and evaluation criteria.
  • Measure task success, diagnostic accuracy, evidence quality, reliability, consistency, latency and cost.
  • Account for the non-deterministic nature of AI systems through repeated runs, variance analysis and determining whether changes are meaningful rather than noise.
  • Benchmark models, prompts, tools, retrieval strategies and agent workflows against repeatable baselines.
  • Use relevant industry benchmarks and standards, including established SRE practices and emerging AI-agent, AIOps and incident-response benchmarks, and build our own where existing approaches don't represent real operational work.

Across the wider product, you'll:

  • Extend evaluation across important customer journeys, APIs, backend services and UI.
  • Use OpenTelemetry, including its Semantic Conventions, to make benchmark environments representative and portable.
  • Use product telemetry and correlate benchmark results with metrics, logs and traces to understand why failures occur.
  • Build end-to-end measurements focused on customer outcomes rather than isolated components.

When we change a model, prompt, tool or agent workflow, we want to know what became better, what became worse and why — including the impact on quality, reliability, latency and cost.

Turn what we learn into continuous improvement

  • Turn failures and real-world problems into new evaluation scenarios.
  • Identify recurring failure patterns and capability gaps.
  • Test potential improvements against established baselines, holdouts and unseen scenarios.
  • Detect regressions, benchmark overfitting and improvements that don't generalise.

Longer term, this evaluation system becomes the harness for controlled self-improvement: identifying weaknesses, testing potential changes and objectively determining whether they should be retained. That's the direction of travel rather than the first year's work, but it's why we're building this properly.

Work with the rest of engineering

  • Work directly with AI, software, platform and SRE engineers to investigate findings and improve the product.
  • Make evaluation failures clear, reproducible and actionable.
  • Make straightforward fixes yourself where that's the most efficient approach.
  • Build tooling that makes evaluations easy for other engineers to create, run and understand.
  • Use AI-assisted engineering where it improves the speed or quality of your work.

Benchmarking should provide continuous feedback that helps engineering improve the product. This role is not a release gatekeeper.

What We're Looking For

Above everything else: the ability to take an ambiguous technical problem, develop an approach and deliver a working system independently.

We're more interested in demonstrated ability than an exact number of years, but we'd generally expect around 3+ years of relevant technical experience. Your background might be as an AI engineer, software engineer, SRE, platform engineer, performance engineer, SDET or similar.

Alongside that, we're looking for:

  • Practical experience working with LLMs, AI agents or AI evaluation.
  • Strong software engineering skills, particularly Python or a similar language.
  • Experience building automated evaluation, benchmarking, testing or experimentation infrastructure.
  • Experience creating synthetic workloads, datasets or evaluation scenarios.
  • An understanding of non-deterministic evaluation, including repeated measurement, variance and distinguishing meaningful changes from noise.
  • The ability to turn complex system behaviour into measurable criteria.
  • Comfort working across APIs, distributed systems and multiple layers of a software product.

Desirable

Any of the following would be useful, but none are required:

  • Agentic AI evaluation, tool use and multi-step workflows.
  • LLM evaluation frameworks and model-based evaluation techniques.
  • Automated experimentation or self-improving systems.
  • Dataset, ground-truth and holdout evaluation design.
  • Statistical experimentation and performance benchmarking.
  • OpenTelemetry, metrics, logs and distributed tracing.
  • SRE, incident response or observability.
  • Production SaaS and distributed systems.

What Success Looks Like

We have the beginnings of an evaluation harness, but the design and expertise are what we're hiring for.

We'd expect the first few months to go into the core evaluation harness and a starting corpus for Coworker's investigation quality, then extend outward across the rest of the product as that proves itself. How you sequence it is your call.

By six months, we should be able to objectively answer questions such as:

  • Is Coworker getting better at investigating operational problems?
  • Where does it perform well or poorly, and why?
  • Are its conclusions supported by the right evidence?
  • How do different models, tools and agent configurations compare?
  • What quality, latency and cost trade-offs are we making?
  • How do we compare against relevant external benchmarks?
  • Have improvements introduced regressions elsewhere?
  • Where should we focus improvement next?

Our evaluation corpus should keep growing as we encounter new problems.

Success isn't measured by the number of tests written or percentage test coverage. It's measured by our ability to understand how well OpsPilot is doing its job, where it isn't, and whether the changes we're making are actually making it better.

What We Offer

A small company rather than a large one, with the trade-offs that implies. Under ten people in engineering, a flat structure, and decisions made in a conversation rather than across three meetings. You'll have genuine influence over how this is done, and very little bureaucracy to work through to get there.

  • Fully remote within the UK.
  • Flexible working hours.
  • 25 days holiday plus bank holidays.
  • Real autonomy over your technical approach and how you deliver the role.

Our Interview Process

Straightforward: usually two or three conversations, with no technical coding tests.

If this sounds like the kind of challenge you're looking for, we'd love to hear from you.


Compensation: £60,000–£75,000 DOE

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

AI Engineer Related jobs

Other jobs at Intergral

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.