Logo for Marvik

Senior Data Engineer

Role overview

Qualifications

  • 5+ years of hands-on data engineering experience building and operating production data pipelines at scale
  • Strong proficiency in Python, SQL, and PySpark / Apache Spark
  • Proven experience in end-to-end data modeling, schema design, and layered lakehouse architectures
  • Experience with cloud-native lakehouse platforms; familiarity with Microsoft Fabric is highly preferred

Responsibilities

  • Build and operate scalable ingestion, ELT/ETL, and orchestration pipelines within Microsoft Fabric and cloud lakehouse environments
  • Design and implement low-latency, real-time data ingestion flows
  • Implement layered architectures using PySpark/SQL with idempotent, backfillable jobs
  • Apply deduplication, normalization, schema validation, and lineage tracking to ensure high-quality data

Key facts

Hard skills

Other skills

  • Problem Solving
  • Teamwork

About the company

Marvik logo

Marvik

Artificial Intelligence & Machine Learning Services

The AI revolution isn't coming, it's already here. Companies that move now will define the market. At Marvik, we're building the AI that powers this transformation. Our team combines deep technical expertise with the speed and agility to turn bold ideas into production-ready solutions. We thrive where the challenges are complex, the stakes are high, and the opportunities are massive.

Company details

IndustryArtificial Intelligence & Machine Learning Services
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

This is a hands-on building role: you turn raw, messy fabrication data into the clean, well-modeled, AI-ready datasets that our AI/ML and analytics workloads run on πŸš€

πŸ§‘πŸ»β€πŸ’» Responsibilities: 

  • Pipeline Development & Operation: Build and operate scalable ingestion, ELT/ETL, and orchestration pipelines (batch and real-time streaming) within Microsoft Fabric and cloud lakehouse environments.

  • Real-Time Data Ingestion: Design and implement low-latency, real-time data ingestion flows to support live operational analytics and streaming workloads.

  • Data Modeling & Layering: Implement layered (medallion-style: Bronze/Silver/Gold) architectures using PySpark/SQL with idempotent, backfillable, and incrementally loaded jobs.

  • Data Quality & Governance: Apply deduplication, normalization, schema validation, and lineage tracking to ensure downstream data is high-quality, trustworthy, and audit-ready.

  • AI & Analytics Readiness: Deliver feature-ready, curated datasets to support business intelligence, analytics, vector search, and AI/ML agentic workloads.

  • Observability & Reliability: Establish testing, monitoring, and pipeline observability (freshness, volume, schema drift) with clear alerting to resolve failures proactively.

  • Tooling & AI Development: Utilize AI-assisted development tools (Claude Code, Copilot, Cursor) as a force multiplier for writing pipelines, query tuning, and data transformation scripts.

🀝 If you have:

  • Experience: 5+ years of hands-on data engineering experience building and operating production data pipelines at scale.

  • Core Technical Stack: Strong proficiency in Python, SQL, and PySpark / Apache Spark, backed by solid software engineering fundamentals (Git, CI/CD, unit/integration testing).

  • Real-Time Data Processing: Demonstrated hands-on experience implementing real-time data ingestion and streaming pipelines (not limited to batch processing).

  • Data Architecture & Modeling: Proven experience in end-to-end data modeling, schema design, and layered lakehouse architectures (Medallion architecture).

  • Platform Experience: Experience with cloud-native lakehouse platforms; hands-on experience or familiarity with Microsoft Fabric is highly preferred.

  • Data Quality & Observability: Strong grasp of data testing frameworks, pipeline monitoring, and data quality enforcement.

  • AI Tooling: Active experience leveraging AI-assisted development tools (Cursor, Copilot, Claude) to accelerate engineering velocity.

🦾 It’s a plus:

  • Hands-on experience with Microsoft Fabric (Fabric Lakehouse, Data Factory, Synapse Analytics).

  • Experience extracting data from document stores / NoSQL databases (specifically MongoDB / MongoDB Atlas and Change Streams / CDC).

  • Streaming frameworks experience (Event Hubs, Kafka, Spark Structured Streaming).

  • Exposure to vector embeddings, RAG-ready datasets, or feature stores for AI/ML workloads.

  • AEC / Construction / MEP domain experience.

This call is made within the framework of Law 19.691 on the Promotion of Employment for Persons with Disabilities, including individuals registered in the National Registry of Persons with Disabilities of the Ministry of Social Development

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
Β·

Data Engineer Related jobs

Other jobs at Marvik

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.