Logo for Ryz Labs

Data Engineer

Role overview

Qualifications

  • 5+ years of experience building and operating production-grade data pipelines
  • Strong experience working with both batch data processing and event-driven architectures
  • Hands-on experience with AWS Glue, Athena, and S3-based data lakes
  • Strong proficiency in Python and SQL for data transformation, processing, and analysis

Responsibilities

  • Own and evolve data ingestion pipelines, including source ingestion/scraping, data cleaning and curation using AWS Glue
  • Design and maintain reliable batch and event-driven data workflows across our data platform
  • Diagnose and resolve identity-resolution issues across multiple data sources
  • Develop and maintain data quality and validation frameworks, including freshness monitoring and schema-drift detection

Key facts

Hard skills

Other skills

  • Communication
  • Collaboration
  • Troubleshooting (Problem Solving)

About the company

Ryz Labs logo

Ryz Labs

IT Services & IT Consulting

🤩 Where startups soar! We're not just building startups - we're crafting the future. From ideation to execution, we provide the blueprint for startup success. ✅ 💡 We specialize in nurturing startups from the ground up, and supercharging existing ones with our top-tier technical talent solutions. 🌐 Join us at Ryz Labs, where we turn promising ideas into thriving businesses. Let's shape the future of innovation together!

Company details

Company typeStartup
IndustryIT Services & IT Consulting
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Remote position - only for professionals based in LATAM.     We are looking for a Senior Data Engineer for one of our client's teams. You will take ownership of the systems responsible for ingesting, transforming, validating, and publishing data across our platform.   In this role, you will work at the data ingestion boundary, where data from multiple external sources enters our systems. You will be responsible for building reliable pipelines, resolving identity and entity-matching challenges, detecting data-quality issues, and ensuring data flows correctly into downstream products and services.   This is a hands-on engineering role that combines AWS data engineering, Python/SQL development, event-driven architectures, data quality, and identity resolution. You will collaborate closely with engineering, product, analytics, and downstream platform teams to troubleshoot issues and improve the reliability of our data ecosystem.     What You'll DoOwn and evolve data ingestion pipelines, including source ingestion/scraping, data cleaning and curation using AWS Glue, entity resolution, and event-driven data publishing. • Design and maintain reliable batch and event-driven data workflows across our data platform. • Diagnose and resolve identity-resolution issues across multiple data sources, including deduplication and entity-matching challenges. • Develop and maintain data quality and validation frameworks, including freshness monitoring, null-rate checks, consistency validation, and schema-drift detection. • Monitor data at the ingestion boundary and proactively identify issues before they impact downstream systems and products. • Partner with engineering teams responsible for the events bridge and downstream identity services to trace data and events end-to-end. • Collaborate with Product, Assessments, Analytics, and other stakeholders to understand how data flows into downstream products and ensure those requirements are reflected in the data pipelines. • Automate infrastructure and pipeline changes using Infrastructure as Code, primarily Terraform or AWS CDK. • Work with AWS services including Glue, Athena, S3, and event-driven services to build and operate scalable data infrastructure. • Participate in production incident response, quickly identifying the scope and impact of data-quality issues and implementing remediation. • Continuously improve pipeline reliability, observability, maintainability, and operational processes.   What You'll Bring5+ years of experience building and operating production-grade data pipelines. • Strong experience working with both batch data processing and event-driven architectures. • Hands-on experience with AWS Glue, Athena, and S3-based data lakes, including layered/medallion-style data transformations. • Strong proficiency in Python and SQL for data transformation, processing, and analysis. • Experience with identity resolution, entity matching, or deduplication, including exact and fuzzy matching approaches. • Understanding of the challenges and tradeoffs involved in maintaining durable identifiers and first-seen/locked identity mappings. • Experience working with event schemas and schema-registry-backed contracts, such as Protobuf, and an understanding of the risks associated with schema changes. • Strong troubleshooting and production incident-response skills, including the ability to determine impact, identify root causes, and implement fixes quickly. • Ability to understand and debug code written in a functional or concurrent programming language, such as Elixir. • Strong communication and collaboration skills, with the ability to work effectively across engineering, product, analytics, and other technical teams.   Nice to Have • Experience with Elixir/Phoenix or another BEAM-based concurrent processing framework. • Experience with DynamoDB-backed identity, lookup, or matching services. • Experience with the Snowflake ecosystem, including data modeling, Snowpipe, Streams, and Tasks. • Experience working with sports data providers, such as MLB.com or Sportradar, or with other licensed-content/data provider ecosystems. • Experience implementing data observability, including freshness/staleness alerts, null-rate monitoring, schema-drift detection, and data-quality dashboards. • Experience working with Infrastructure as Code using Terraform or AWS CDK.   Technologies Languages: Python, SQL, familiarity with Elixir AWS: Glue, Athena, S3, DynamoDB Data & Streaming: Data Lakes, Event-Driven Architecture, Protobuf, Schema Registries Infrastructure: Terraform, AWS CDK Data Platforms: Snowflake Engineering Practices: Data Quality, Data Observability, Identity Resolution, Entity Matching, Incident Response

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Engineer Related jobs

Other jobs at Ryz Labs

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.