Logo for Firmable

Lead Data Engineer - Sourcing

Role overview

Qualifications

  • 6+ years shipping production extraction, ETL, or data pipeline systems
  • Expert Python and advanced SQL skills
  • Extensive Airflow experience or equivalent for orchestration
  • Experience with agentic IDEs and LLMs in production systems

Responsibilities

  • Design and own the end-to-end extraction and ETL pipeline
  • Architect extractor frameworks and agentic workflows
  • Set standards for sourcing layer coverage and accuracy
  • Optimize cost and performance of extraction systems

Key facts

Hard skills

Other skills

  • Teamwork
  • Problem Solving
  • Communication

About the company

Firmable logo

Firmable

Computer Software / SaaS

Firmable is the AI-native sales intelligence platform for B2B sales teams, providing the most complete prospect data, buying signals, and agent-driven actions to help you win more deals. As the salesperson's favorite teammate, we don't just give you data - we give you direction, whatever tool you work in. Firmable's agentic technology, built from the ground up, sources, assembles, and continuously refreshes a proprietary map of the market. It delivers the richest company and contact details, including data you won't find anywhere else, to sales teams across the United States, Canada, and APAC. With Firmable, your team always knows who's ready to buy, when, and what to do next.

Company details

IndustryComputer Software / SaaS
Company size51-200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Firmable is the market-leading B2B sales intelligence platform in Asia-Pacific. Our competitive moat is our data: the deepest, most localised company and people dataset in every market we operate in. We've proven the model in ANZ. Now we're scaling it across Southeast Asia and the US.

Every record Firmable sells starts in a sourcing pipeline. This role architects that pipeline.

The Role

Lead Data Engineer: Sourcing designs and owns the end-to-end extraction and ETL pipeline that turns unstructured web data into the world's most accurate B2B dataset across 13 markets. You set the architectural standards: extractor patterns, proxy strategy, LLM infrastructure as production systems, agentic escalation workflows, and cost-aware orchestration.

This is not a pipeline maintenance or incremental optimisation role. You're architecting the framework other engineers ship into, building for the hardest extraction problems (anti-bot defences, JS-heavy rendering, schema drift), and owning the sourcing layer end to end.

~80% hands-on architecture and reference implementation: designing extractor frameworks, shipping the hardest extractors, building agentic pipelines, writing eval harnesses and labelled datasets. ~20% cross-functional voice: technical input on sourcing decisions with data platform, product, and analytics; setting standards for coverage, schema, and accuracy.

What You'll Own

Sourcing and ETL architecture

  • End-to-end pipeline design: extraction, normalisation, deduplication, validation, and load; plus cost, performance, and reliability of the whole layer

  • Sourcing layer standards: coverage, accuracy, schema design; the rule-vs-LLM standard codified so the team applies it without you

  • Cost and performance optimisation: cost ceilings with the token math behind them, incremental processing, recovery design, scheduling

Extraction systems and frameworks

  • Extractor framework: patterns, abstractions, and tooling other engineers ship into; new extractors are fast to build and reliable to run

  • Hard-source extractors: anti-bot defences, JS-heavy rendering, schema drift, low-quality structure; plus the proxy and IP rotation strategy behind them

  • Agentic extraction pipelines: rule-based triage, LLM escalation, structured-output validation, retries, human-review queues

LLM infrastructure and observability

  • LLMs as production systems: versioned prompts, labelled eval sets, measured precision and recall, prompt versioning you can roll back, judges debugged on real data

  • Eval and observability scaffolding: eval frameworks, prompt versioning, traces, drift detection when a vendor silently updates a model

  • Model-choice playbook: cheap models for classification, stronger models for nuanced extraction, frontier models for hard edge cases; revised as model economics shift

Skills library and orchestration

  • Skills library: versioned SKILL.md specs for recurring extraction patterns, invocable by engineers and agents alike

  • Orchestration: Airflow or equivalent patterns that scale with data volumes and source counts; dependency management, recovery, cost-aware scheduling

What We're Looking For

Must Haves

  • 6+ years shipping production extraction, ETL, or data pipeline systems in business-critical environments; deep web extraction at scale with experience designing around anti-bot defences, proxy architecture, JS rendering, schema drift, and recovery

  • Shipped LLMs inside extraction pipelines as production systems: structured outputs, versioned prompts, labelled eval sets, logged traces, judges debugged on precision/recall, drift detected on vendor model updates. You can show the repo.

  • Built agents and tool-calling pipelines: architected agent workflows, written SKILL.md specs others depend on, run tool-calling systems in production

  • Sharp judgement on rules vs. LLMs: reach for deterministic logic first, defend the call either way, and have codified the standard for a team

  • Expert Python: production-grade, performance-aware, comfortable with concurrency and scale; plus advanced SQL for complex transformations and performance optimisation

  • Extensive Airflow or equivalent: orchestration, dependency management, recovery patterns, cost optimisation at production scale

  • Shipped with agentic IDEs: Claude Code, Cursor, or equivalent as your default mode, with shipped extraction systems to show for it

  • Architecture judgement and a product mindset: when to refactor, optimise, ship, or start again; you care about coverage and accuracy landing with customers, not uptime metrics

  • You live and breathe AI tools. Structured outputs, evals, traces, and LLM tracing (Logfire, OpenTelemetry, similar) as your default way of working, not a productivity experiment. In 2026, this is how data engineers build extraction at scale and you need to already be doing it.

Highly Valued

  • Cloud data platforms: Snowflake, Redshift; AWS for pipeline deployment (Lambda, S3, ECS, Glue)

  • Vector databases, embeddings, or retrieval patterns for matching and deduplication

  • Eval frameworks like Braintrust, Promptfoo, or Inspect; LLM tracing tooling

  • Data quality frameworks with automated testing and anomaly detection at scale

  • B2B data: firmographics, people data, company registries across markets

  • Data privacy and compliance: GDPR, CCPA

  • Startup or scaleup experience where you shipped fast and owned outcomes end to end

The Environment

Firmable runs lean and ships fast: small senior teams, no layers, minimal process. Sourcing is a core competitive advantage; this role sets the pace for how fast and accurate the entire data pipeline moves. Weekly releases moving toward daily; no fixed hours, full autonomy on architecture choices.

We are an AI-native organisation. That means AI isn't a tool we reach for: it's the default operating mode. Agentic extraction pipelines, LLM-powered classification with eval sets, structured-output validation, labelled datasets versioned with prompt runs, traces logged from day one with cost and latency baked in. If you're not already working this way, this role isn't right for you.

Why This Role

  • Own the sourcing engine: every record Firmable sells starts in the pipeline you architect; your design shapes the accuracy and cost of every customer interaction

  • Greenfield AI-native architecture: the extractor framework, skills library, eval harnesses, agentic orchestration, and production LLM infrastructure are largely unbuilt; you'll architect them from first principles

  • Frontier technical problems: agentic extraction at production scale, LLM-as-extractor with full eval coverage, vendor model drift detection, rule-vs-LLM orchestration across 13 markets

  • Small team, massive leverage: your architecture reaches every Firmable customer, every day; you write the reference implementation daily and own outcomes end to end

  • Real career runway: this role is a bottleneck role; progress is measured by pipeline speed and accuracy, not tenure

Who This Role is NOT For

  • Anyone looking to optimise an existing pipeline rather than architect from first principles

  • Anyone who treats infrastructure as someone else's problem; sourcing architecture lives in your code every day

  • Anyone who sees LLMs as a parsing shortcut rather than production systems that need eval sets, versioning, and drift detection

  • Anyone uncomfortable shipping with minimal process or without an established playbook to follow

  • Anyone who treats AI as something they'll learn on the job rather than already use daily in infrastructure work

  • Anyone coming from a pure ops or data analytics background without shipped systems engineering at scale

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Engineer Related jobs

Other jobs at Firmable

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.