Logo for Empirical

OB - Lead AI Engineer

Role overview

Qualifications

  • 10+ years in software engineering, including deep CI/CD and platform experience
  • Real security instincts without needing to be a security engineer
  • Testing judgment, not just testing habits
  • Excellent communication skills in English (written and spoken)

Responsibilities

  • Own the knowledge plane: project conventions, reusable agent procedures, and source indexing
  • Build gates that read artifacts: verification that inspects what actually happened rather than trusting an agent’s account of its own work
  • Design authority and identity: implementing least-privilege and short-lived agent credentials
  • Own intake and measurement: the request surface with real reproduction context attached, deduplication and routing

Key facts

  • Remote from: Latin America
  • Full time
  • Senior (5-10 years)
  • AI Engineer
  • English

Hard skills

Other skills

  • Communication
  • Problem Solving
  • Time Management

About the company

Empirical logo

Empirical

Software Development

We are product development company focused on helping entrepreneurs and their teams build successful products. We help them succeed by transforming product ideas into high quality working software solutions through innovation, embracing agile methodologies and delivering each project milestone on time and budget. We provide a professional and high quality option for companies in need of product development. Why us? We are product developers experts who understand the business side of your company as well as the technical aspects of your project. We become partners with our clients and we have full transparency throughout the process. We have high standards of the quality we deliver. We take every project as if it was our own.

Company details

IndustrySoftware Development
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Empirical

Empirical is a new kind of product development partner — people-first at our core, AI-driven by design. We help tech companies build products that people love across four areas: AI-enabled product & tech leadership (advisory, coaching, and fractional models), end-to-end product development (concept to launch), nearshore AI-enabled team augmentation (high-performing Latin American talent that integrates with your workflows), and AI strategy & implementation consulting (finding where AI drives the most value in your product, operations, or customer experience).

We’re a people-first company grounded in empathy, integrity, and meaningful collaboration. From strategic planning to hands-on execution, we work alongside founders and teams to bring great products to life—without the noise.

Our Core Values

  • Care about people first

  • Strive to add value always

  • Lead with integrity

  • Have fun every step of the way

The Opportunity

We are building a development pipeline for one of our clients where business users initiate work, AI agents implement it, and engineers review outcomes rather than steer execution. You will lead that build and own the machinery that makes it safe enough to run: the conventions, gates, identities, autonomy tiers, and measurements that determine whether AI-generated code is safe to merge.

Most engineering roles are measured on what you ship. This one is measured on what the pipeline ships. If a good week is one where you personally closed a lot of tickets, this role will frustrate you. If a good week is one where forty changes went through cleanly and you spent Thursday deleting a rule that had stopped earning its place, keep reading. The hard part of this job is not building the automation; it is refusing to trust it.

What You’ll Do

  • Own the knowledge plane: project conventions, reusable agent procedures, and source indexing so agents find every relevant call site, plus the enforcement that makes it all stick. This is the highest-leverage surface in the system.

  • Build gates that read artifacts: verification that inspects what actually happened (did the test file change, does coverage include the new branch) rather than trusting an agent’s account of its own work, including a protected path around tests and CI config.

  • Design authority and identity: least-privilege, short-lived agent credentials where an agent’s authority is the intersection of its own and its sponsor’s, never the union, with attribution from every commit back to a person, an agent, and a task.

  • Calibrate autonomy tiers: a classifier that routes changes by how hard they are to undo and fails toward caution. Copy changes merge themselves; schema migrations get a plan and a human. You will move the boundaries based on incident history, not optimism.

  • Engineer containment: circuit breakers that stop a repair loop from amplifying its own failures, per-workflow ceilings on tokens, wall-clock, retries and concurrency, progressive exposure, and rollback.

  • Own intake and measurement: the request surface with real reproduction context attached, deduplication and routing, plus a gate log that records every evaluation (passes included) and an evaluation harness so configuration changes are shown to be improvements, not assumed to be.

  • Run the improvement ritual: dated incident records, a threshold before a pattern becomes a rule, and a recurring pass that asks what each rule cost versus what it prevented. Removal is a first-class outcome.

Who You Are

  • Based in Latin America, with enough timezone overlap for daily collaboration with a U.S.-based engineering team.

  • 10+ years in software engineering, including deep CI/CD and platform experience: you have built pipelines other engineers depend on, and you have been on the hook when one broke at 6pm on a Friday.

  • A genuinely strong engineer first: you cannot govern code you can’t read, and the person setting the standard for agent output has to be better than the agent, not worse.

  • Real security instincts without needing to be a security engineer: least privilege, credential scope, blast radius, untrusted input. “Any authenticated user can submit free text that an agent acts on with commit rights” should immediately worry you.

  • Testing judgment, not just testing habits: you have opinions about what makes a suite worth trusting, you have measured a flake rate rather than estimated one, and you distrust green builds until the artifacts back them up.

  • Hands-on agentic tooling experience: you have driven coding agents on real work, not demos, and you have a specific, unsentimental view of where they are strong and where they quietly fail.

  • Comfortable operating cross-functionally: several of the things you will need are owned by SRE, security, legal, and support teams that have never been in a pipeline conversation.

  • A willing writer: much of this job is prose (conventions, procedures, incident records, the argument for a rule). Think of it as documentation with a compiler attached.

  • Excellent communication skills in English (written and spoken).

This is probably not the role for you if

  • Your best weeks are the ones where you personally shipped the most.

  • You want to “work on AI” in the sense of models and prompts. This is mostly plumbing, policy, and measurement; the model is somebody else’s product.

  • You would accept “the tests passed” as evidence that the tests ran.

  • You need the thing you own to be finished. It won’t be; it will be calibrated continuously against incidents that haven’t happened yet.

The Engagement

  • Full-time, dedicated client engagement with genuine ownership: this is a system, not a ticket queue, and the person who owns it makes the calls about how it works.

  • Executive sponsorship: the client has committed to this as an operating model change, not a tooling upgrade, which means protected time and no expectation that you also carry a feature roadmap.

  • An honest scope: low-risk changes end to end, higher-risk changes gated, and irreversible changes staying with humans for the foreseeable future.

  • Leverage that compounds: a rule you write today is enforced against every change for as long as it earns its place. Very little engineering work has that shape.

Why Join Empirical

  • Be part of a curated network of high-caliber, values-driven professionals

  • Do meaningful, strategic work on your terms

  • Stay in motion without burnout

  • Access Empirical’s community, engineering talent, and shared knowledge base

  • Work with clients who value clarity, velocity, and outcomes—not just hours

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

AI Engineer Related jobs

Other jobs at Empirical

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.