Logo for Wynd Labs

Data Engineer

Role overview

Qualifications

  • Bachelor’s degree or equivalent work experience
  • Python (advanced) — strong grasp of async programming, multiprocessing
  • Web scraping at scale — hands-on experience with high-volume scraping
  • Distributed data pipelines — experience designing and operating pipelines

Responsibilities

  • Maintain, optimize, and troubleshoot database queries and related data systems
  • Assist in creating, maintaining, and improving data pipelines used to collect and process datasets
  • Support web scraping and data collection initiatives, including developing scripts or tools
  • Monitor and troubleshoot data pipeline issues, identify data quality concerns, and implement fixes

Key facts

Hard skills

About the company

Wynd Labs logo

Wynd Labs

IT Services & IT Consulting

Making AI Data Accessible. Building a suite of products powered by https://www.getgrass.io/

Company details

Company typeStartup
IndustryIT Services & IT Consulting
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

Who You Are:

  • Bachelor’s degree or equivalent work experience

  • Python (advanced) — strong grasp of async programming, multiprocessing, and writing production-grade code for long-running data jobs

  • Web scraping at scale — hands-on experience with high-volume scraping (proxies, rate limiting, anti-bot evasion). Experience with platform APIs and large media/metadata datasets (video platforms, social media)

  • Distributed data pipelines — experience designing and operating pipelines across many workers/servers using task queues (Celery, Kafka, RabbitMQ, or similar)

  • Data warehousing — practical experience with columnar/analytical warehouses; Databend, ClickHouse, or BigQuery strongly preferred; comfortable with complex analytical queries, partitioning strategies, cost-aware querying on cloud warehouses

  • Docker & Kubernetes — containerizing workloads, writing Helm charts/manifests, managing deployments, autoscaling scraping/processing workloads

  • Linux & bare-metal ops — comfortable managing services on Linux servers, debugging performance issues (disk I/O, network, memory) without managed-cloud abstractions

  • CI/CD for data workflows (GitHub Actions, ArgoCD)

  • Writing Scalable API

What You'll Be Doing:

  • Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability.

  • Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets.

  • Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data in accordance with Company requirements.

  • Monitor and troubleshoot data pipeline issues, identify data quality concerns, and assist in implementing timely fixes to maintain data accuracy and operational continuity.

  • Document engineering work, including database queries, pipeline processes, scraping workflows, technical decisions, issues encountered, and resolutions implemented.

  • Participate in research and development projects to improve the Company’s data products and workflows.

Why Work With Us:

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.

  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.
    We prioritize low ego and high output. This is a fully remote team.

  • Compensation. You’ll receive a competitive salary, benefits and equity package.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Engineer Related jobs

Other jobs at Wynd Labs

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.