Logo for Oxford Data Plan

Senior Data Engineer (Web Scraping)

Role overview

Qualifications

  • Demonstrated professional experience building and operating production web-scraping systems at scale.
  • Proven ability to independently own substantial scraping projects from initial investigation through production operation.
  • Strong production Python engineering skills, including building maintainable applications rather than standalone scripts.
  • Strong practical experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright or Selenium.

Responsibilities

  • Own the development and operation of web-scraping and web-data ingestion pipelines.
  • Design and establish a scalable web-scraping framework, including common patterns for extraction, scheduling, storage, monitoring, validation and failure handling.
  • Design, build and maintain reliable production scrapers for new and existing data sources.
  • Evaluate and make build-versus-buy recommendations for scraping infrastructure and third-party services.

Key facts

Hard skills

Other skills

  • Problem Solving
  • Teamwork
  • Communication

About the company

Oxford Data Plan logo

Oxford Data Plan

Market Research

Oxford Data Plan delivers institutional-grade alternative data and daily KPI estimates for 250+ global equities, enabling hedge funds and asset managers to identify inflections ahead of consensus. We combine proprietary and exclusive datasets—including a global receipt panel, exclusive advertising agency partnerships, —with multi-signal modeling to produce point-in-time estimates across 500+ KPIs. Our coverage spans TMT, consumer, financials, and real economy sectors, with daily delivery designed for systematic and fundamental workflows. Core capabilities include: • Daily revenue, orders, and operational metrics with 12-hour delivery lag • Advertiser-level digital spend tracking across 15,000+ brands and major platforms • Sector insights covering digital advertising, food delivery, airlines, and classifieds • Historical backtesting and out-of-sample validation for every tracker Built by former buy-side analysts, our platform is optimized for alpha generation, not data exploration. Clients receive structured feeds with full point-in-time integrity. Oxford Data Plan is trusted by leading quantitative and fundamental investors seeking differentiated signal with institutional rigor.

Company details

Company typeSME
IndustryMarket Research
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

We are looking for an experienced, self-driven Data Engineer to take ownership of web scraping and web-based data acquisition within our data platform.

We already use web scraping across a number of ingestion processes, and we are looking to improve the maturity, scale and robustness of this capability. This includes integrating scraping workloads into our lakehouse architecture, improving scheduling, monitoring and storage, and establishing reusable patterns for building and operating scrapers reliably.

This is a senior, hands-on role. We are looking for someone who has previously built and operated production web-scraping systems at scale and can independently take a new data source from investigation through implementation, deployment and ongoing support.

Roles and Responsibilities

· Own the development and operation of web-scraping and web-data ingestion pipelines.

· Design and establish a scalable web-scraping framework, including common patterns for extraction, scheduling, storage, monitoring, validation and failure handling.

· Design, build and maintain reliable production scrapers for new and existing data sources.

· Evaluate and make build-versus-buy recommendations for scraping infrastructure and third-party services, considering capability, reliability, cost, operational complexity and risk.

· Investigate websites and determine the most appropriate acquisition approach, including APIs, direct HTTP requests, HTML parsing, browser automation or third-party tooling.

· Ensure scraping activities appropriately consider internal policies, website terms, robots.txt, access restrictions, privacy and intellectual-property constraints, escalating unclear cases where needed.

· Diagnose and resolve issues caused by changing websites, dynamic content, authentication, rate limits and other operational challenges.

· Use AI tooling to work more effectively while understanding, reviewing and being able to defend the code you ship.

Required Qualifications

· Demonstrated professional experience building and operating production web-scraping systems at scale.

· Proven ability to independently own substantial scraping projects from initial investigation through production operation.

· Strong production Python engineering skills, including building maintainable applications rather than standalone scripts.

· Strong practical experience with scraping technologies such as Requests/httpx, BeautifulSoup, Scrapy, Playwright or Selenium.

· Good understanding of HTTP, HTML, APIs, JavaScript-rendered websites and browser/network behaviour.

· Experience handling common scraping challenges such as pagination, authentication, sessions, retries, rate limiting, concurrency and proxies.

· Strong understanding of data pipelines, data quality and how collected data should be validated, stored and consumed downstream.

· Experience deploying, monitoring and supporting production workloads in a cloud environment.

· Strong debugging and problem-solving skills, with the ability to work independently and make sensible engineering decisions.

· Experience working effectively within a remote engineering team, including code review, documentation and ticket-based workflows.

Desirable Skills

· Experience with AWS.

· Experience with lakehouse or data-lake architectures, particularly Iceberg.

· Experience with PySpark or other distributed data-processing technologies.

· Experience with Docker and containerised workloads.

· Experience with Terraform or other infrastructure-as-code tooling.

· Experience with Grafana or similar observability platforms.

· Experience operating high-volume or distributed crawling systems.

· Experience evaluating or operating commercial scraping, proxy or browser-infrastructure services.

· Experience implementing automated scraper testing, canary runs or source-drift detection.

· Experience working with legal, compliance, privacy or data-governance teams on web-data acquisition.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Engineer Related jobs

Other jobs at Oxford Data Plan

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.