Logo for Firmable

Senior Data Engineer - Data Platform

Role overview

Qualifications

  • 5+ years building production data platforms in business-critical environments
  • Worked with billions of rows in data pipelines
  • Shipped LLMs inside data pipelines as production systems
  • Expert in Python and SQL

Responsibilities

  • Architect the transformation and warehouse layer for B2B dataset
  • Design patterns for extraction, enrichment, and validation of records
  • Implement observability for LLM calls and pipeline alerting
  • Make architecture calls on model design and materialization

Key facts

Hard skills

About the company

Firmable logo

Firmable

Computer Software / SaaS

Firmable is the AI-native sales intelligence platform for B2B sales teams, providing the most complete prospect data, buying signals, and agent-driven actions to help you win more deals. As the salesperson's favorite teammate, we don't just give you data - we give you direction, whatever tool you work in. Firmable's agentic technology, built from the ground up, sources, assembles, and continuously refreshes a proprietary map of the market. It delivers the richest company and contact details, including data you won't find anywhere else, to sales teams across the United States, Canada, and APAC. With Firmable, your team always knows who's ready to buy, when, and what to do next.

Company details

IndustryComputer Software / SaaS
Company size51-200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Firmable is the market-leading B2B sales intelligence platform in Asia Pacific — and we're scaling that success globally at pace. Backed by leading investors and 2,000+ customers strong, we exist to give sales teams an unfair advantage: the deepest company and people data of any platform, enriched with real-time signals, served at the right moment by intelligent agents.

Our data is the product. Building it now means building with LLMs in the pipeline, and setting the standard for when to trust them.

The Role

As Senior Data Engineer — Data Platform, you'll architect the transformation and warehouse layer that turns billions of raw records into the world's most accurate B2B dataset, and design how LLM-based extraction, enrichment, and validation run inside it as production infrastructure.

You make the architecture calls: model design and materialisation, warehouse cost and performance, orchestration patterns the team builds on, and the rule-vs-LLM standard for every quality and enrichment step. You know when a dbt test does the job, when an LLM check is needed, and when both belong together.

This is hands-on senior engineering with sharp architecture judgement. You write the reference implementation, then the team builds on it.

What You'll Own

  • LLM-in-the-pipeline architecture — the patterns for extraction, enrichment, entity resolution, and semantic validation over millions of records a day: structured outputs, retries, human-review fallbacks, and the abstractions others reuse

  • Eval harnesses — design the labelled sets, scorers, judge calibration, and regression suites that gate every prompt or model change; own precision/recall targets per check

  • The rule-vs-LLM standard — deterministic checks (dbt tests, data contracts, SQL) wherever structure allows; LLMs only where semantic judgement is needed; codified so the team can apply it without you

  • Cost and drift — token budgets per pipeline, model routing (cheap models for classification, stronger ones for hard cases), drift detection on vendor updates, and the token math behind the ceilings

  • Observability — every LLM call logged with prompt version, model, cost, latency, and decision, alongside pipeline alerting that surfaces problems before they cascade

  • Warehouse and transformation architecture — dbt layers that encode real business logic; Snowflake performance, clustering, materialisation, and cost over billions of rows

  • Orchestration and infrastructure — Airflow patterns for recovery and cost-aware scheduling; AWS infrastructure as code

  • Matching and deduplication — embeddings and retrieval patterns for company and people entity resolution across 13 markets

  • Cross-functional voice — technical voice for the platform in architecture decisions with sourcing, data quality, product, and analytics

What We're Looking For

Must Haves

  • 5+ years building production data platforms in business-critical environments, with architecture you've owned end to end

  • Worked with billions of rows — you know what breaks at that scale and how to design pipelines and checks that don't

  • Shipped LLMs inside data pipelines as production systems — extraction, enrichment, or validation with structured outputs, versioned prompts, and labelled eval sets. You can show the repo

  • Designed eval harnesses others depend on — scorers, labelled sets, judge calibration, regression suites; you can talk through what they caught and what they missed

  • Sharp judgement on rules vs. LLMs — you reach for a dbt test or SQL first, defend the call either way, and have codified the standard for a team

  • Expert Python and SQL — production-grade, performance-aware, comfortable at very large scale

  • Expert dbt and Snowflake — modular transformation layers with real business logic; warehouse design, query optimisation, clustering, RBAC, cost at scale

  • Extensive Airflow — orchestration, dependency management, recovery patterns, cost optimisation in production

  • Solid AWS — S3, Lambda, Glue, ECS, RDS

  • AI coding tools are how you work — Claude Code, Cursor, or equivalent, daily, with real shipped work to show for it

  • Architecture judgement and a product mindset — when to refactor, optimise, ship, or redesign; you care how data quality lands with customers, not just uptime

Highly Valued

  • Eval frameworks (Braintrust, Promptfoo, Inspect) and LLM tracing (Logfire, OpenTelemetry)

  • Embeddings, vector search, or fuzzy matching for entity resolution at scale

  • Fine-tuning or distilling small models to replace expensive LLM calls

  • dbt Cloud and CI/CD for data pipelines

  • Spark or PySpark; streaming (Kafka, Kinesis, Snowpipe Streaming)

  • B2B data: firmographics, people data, company registries

  • Data privacy and compliance (GDPR, SOC2, CCPA)

How We Build

Firmable is an AI-native organisation. AI coding tools, automated testing, and AI-assisted review are how we work by default. Every LLM check ships with a labelled eval set, measured precision/recall, and a prompt version you can roll back. Every LLM call in production is logged from day one; retrofitting later is not the plan.

We run lean and ship fast — small senior teams, no layers, minimal process, weekly releases moving toward daily. Teams own their stack end to end. There are no fixed hours and no handholding. If you're not already working this way, this role isn't right for you.

Why This Role

  • LLMs as production infrastructure — not a demo, not a notebook; models making millions of decisions a day on data customers pay for, and you set how they ship

  • Greenfield architecture — eval harnesses, drift detection, cost controls, and the rule-vs-LLM standard for the warehouse are largely unbuilt; you'll architect them

  • Scale that matters — billions of rows, 13 markets, and a dataset nobody else has

  • Small team, massive leverage — your architecture reaches every Firmable customer, every day

  • Competitive base + meaningful equity — a share in the upside we're building toward

Firmable is an equal opportunity employer. We believe diverse teams build better products.

Ready to architect the AI-native data platform behind the world's smartest B2B sales intelligence platform? Apply now — let's talk!

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Engineer Related jobs

Other jobs at Firmable

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.