Logo for NaNLABS

Senior Data Product Engineer

Role overview

Qualifications

  • Strong hands-on experience with advanced web scraping at scale
  • Strong professional experience with Python and data processing
  • Experience building reliable data pipelines using Airflow or an equivalent orchestrator
  • English level B2 or higher

Responsibilities

  • Architect and scale distributed web scraping systems to reliably collect data from thousands of internet sources
  • Build strategies to navigate anti-scraping mechanisms such as Cloudflare, CAPTCHAs, rate limiting, IP banning, and browser fingerprinting
  • Refactor and productionize existing Python data collection pipelines, improving reliability, observability, error handling, retries, and alerting
  • Collaborate with Data Science and SaaS teams to integrate existing ML components and define reliable API contracts

Key facts

Hard skills

About the company

NaNLABS logo

NaNLABS

Software Development

We empower Automotive & EV, SaaS, Cybersecurity and other tech driven companies to scale, optimize, and innovate with cloud-native architectures, real-time data processing, and AI-driven applications. More than a typical dev shop, we are your tech sidekick—fully embedded in your team with deep ownership and technical excellence, co-creating tailored, high-impact solutions that drive real results. What We Do For over 12 years, we’ve enabled US-based companies to leverage cloud-native technologies to overcome complex data challenges, and future-proof their architecture for growth. Cloud Data Engineering Solutions - that drive scalable growth and cost-efficient data solutions while future-proofing your infrastructure. Real-Time Data Solutions - that deliver instant insights and support fast, adaptive business growth. AI & Machine Learning Solutions- that boost innovation while ensuring security and privacy. Why NaNLABS? Because every hero needs a sidekick who amplifies their vision, we bring: 🔑 Ownership: Fully embedded in your team, we treat your project as our own. 💡 Proactive Problem-Solving: We anticipate roadblocks to keep the project on track. 🚀 Impact-Driven Innovation: Beyond deliverables, we drive measurable results. 👨🏻💻 Advanced Tech Stack: We work with AWS, Databricks, and many others cutting-edge tools to build scalable, high-performance solutions. 🏃 Reliable Agility: We move at your pace, delivering responsive, dependable support. 🛠️ Craftsmanship: Thoughtfully crafted solutions that align with your vision and goals. Proven Results - CyberCube: Increased code quality by 90% code quality boost, strengthening engineering foundations. - HyreCar: Scaled user base 5x user growth, leading to its acquisition by Getaround. - Netskope / WootCloud: Achieved 3x faster deployments, accelerating product delivery. Ready to level up? As your tech sidekick, we're here to help you scale smarter with AI, cloud, and data solutions. Let's talk!

Company details

IndustrySoftware Development
Company size51-200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Senior Data Product Engineer

Remote – LATAM 

About the Project

We’re looking for a Senior Data Product Engineer to join an initiative focused on transforming an existing research-driven data collection platform into a scalable, production-grade data product for cyber risk analytics.

The project involves collecting and processing large volumes of data from thousands of internet sources, building resilient web scraping systems, and making the resulting data reliably available to SaaS teams through a structured data access layer.

You’ll work at the intersection of Data Collection, Data Science, and Data Engineering, owning the path from raw web data collection and ingestion through to a reliable, documented data product.

What you’ll do

  • Architect and scale distributed web scraping systems to reliably collect data from thousands of internet sources.

  • Build strategies to navigate anti-scraping mechanisms such as Cloudflare, CAPTCHAs, rate limiting, IP banning, and browser fingerprinting.

  • Work with proxy pools, headless browsers, and adaptive crawling techniques to ensure reliable data collection at scale.

  • Process and extract data from HTML and JSON using tools such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, or CSS selectors.

  • Refactor and productionize existing Python data collection pipelines, improving reliability, observability, error handling, retries, and alerting.

  • Build schedulable and containerized ingestion workflows using Airflow or equivalent orchestration tools.

  • Design PostgreSQL schemas, views, partitioning, and data access patterns for processed web data.

  • Work with AWS services including S3, EKS, IAM/IRSA, Parameter Store, and ECR.

  • Build and maintain the data access layer using GraphQL and Hasura.

  • Collaborate with Data Science and SaaS teams to integrate existing ML components and define reliable API contracts.

  • Take end-to-end ownership of the data product, contribute to code reviews, and drive technical improvements across the project.

Tech Stack You’ll Use

  • Web Scraping: Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml

  • Core & Data: Python, PostgreSQL, Airflow, Redis, Elasticsearch

  • AWS: boto3, S3, EKS, IAM/IRSA, Parameter Store, ECR

  • Data Access: GraphQL, Hasura

  • Infrastructure & DevOps: Docker, Kubernetes, Helm, GitHub Actions

  • Additional technologies: Liquibase/Flyway, SQLAlchemy

What we’re looking for

  • Strong hands-on experience with advanced web scraping at scale.

  • Experience overcoming anti-scraping and bot-protection mechanisms, including proxy rotation, headless browsers, rate limiting, IP blocking, fingerprinting, or similar challenges.

  • Strong experience with web scraping and DOM parsing tools such as Scrapy, Playwright, Selenium, Puppeteer, BeautifulSoup, lxml, XPath, or CSS selectors.

  • Strong professional experience with Python and data processing.

  • Hands-on experience with AWS, particularly S3 and boto3; experience with EKS, IAM/IRSA, Parameter Store, or ECR is highly valuable.

  • Experience building reliable data pipelines using Airflow or an equivalent orchestrator.

  • Strong knowledge of PostgreSQL and relational database fundamentals.

  • Ability to take end-to-end ownership of technical solutions and work autonomously.

  • Experience contributing to code reviews and maintaining high software engineering standards.

  • Experience with GraphQL, Hasura, Redis, Elasticsearch, Docker, or Kubernetes is a plus.

  • Experience with cybersecurity data, NLP/ML pipelines, or data-as-a-product environments is a plus.

  • English level B2 or higher, with the ability to communicate effectively with technical and cross-functional teams.

What you’ll get

  • Time off & well-being: vacations fully flexible and self-managed, sick leave and personal days, public holidays, paternity and maternity leave, study leave, and moving days.

  • Learning & growth: training in best practices and tech, books and light talks, in-house English classes, continuous feedback, and 1:1 career development sessions.

  • Work experience: flexible working hours, equipment and work materials provided, internal events and team activities, and a day off on your birthday.

  • Contract & setup: 100% remote positions across LATAM, under a contractor model with payment in USD.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Engineer Related jobs

Other jobs at NaNLABS

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.