Job Title: Data Engineer
Location: Remote/Hybrid - USA
Reports to: Senior Director, Data
Department Name: Information Technology
Job Types: Full Time, Exempt
Position Summary
The Data Engineer is responsible for building and operating production data pipelines on Databricks and across the surrounding data stack. Reporting to the Senior Director of Data, you will play a crucial role in supporting analytics and data initiatives across departments.
You will collaborate with Account Management, Finance, Finance Ops, and other teams to support Rubicon’s analytics, business, and AI teams. Your goal will be to own production pipelines, end to end (including on call), and to deliver a new pipeline from requirements through production with data contracts, quality checks, and documentation.
Essential Duties & Key Responsibilities
- Build batch and streaming ingestion pipelines from databases, APIs, SaaS applications, event streams, and file drops using Lakeflow Declarative Pipelines for CDC and SCD Type 2 patterns.
- Build and maintain medallion architecture (bronze, silver, gold) with clear contracts at each layer, applied as part of every pipeline you deliver.
- Write and tune complex SQL and PySpark transformations against large Delta Lake tables, and model gold-layer datasets for analytics.
- Own operational health of your pipelines, including data quality checks, monitoring and alerting, Unity Catalog governance, Continuous Integration/Continuous Development (CI/CD) promotion, and on-call.
- Ability to travel and/or work onsite as needed.
- Performs other duties as assigned or apparent.
Supervisory Responsibilities:
- This job has no direct supervisory responsibilities.
- May assist with onboarding and training new hires on Information Technology processes, systems, and operational procedures.
- Serves as a resource for internal team members regarding IT/Data operational processes and standards.
- Supports knowledge sharing, continuous improvement initiatives, and best practices across the Information Technology and Data Engineering functions.
Experience & Qualifications:
- High School Diploma required. Bachelor’s degree in Business, Information Systems, Data Science, or related field preferred.
- 5+ years of experience in Data Engineering or adjacent roles with a track record of owning and fixing pipelines in production.
- Strong knowledge of Production Apache Spark (PySpark) and Python (modular, testable code with unit tests).
- Demonstrated understanding of data modeling and pipeline mechanics, including dimensional modeling, grain and keys, slowly changing dimensions, CDC, schema drift, and formats such as Parquet and Avro.
- Proficiency with Delivery fundamentals (orchestration with Lakeflow Jobs, Airflow, or Dagster; Git and code review; Continuous Integration/Continuous Deployment (CI/CD); cloud object storage on a major cloud provider).
- Advanced SQL skills, including window functions, CTEs, complex joins, aggregations over large datasets, and identifying why queries are slow through query plan analysis.
- An in-depth understanding of incremental processing, idempotency, and late-arriving data.
- Experience with debugging correctness issues by reasoning through joins, grain, and duplication.
- Hands on Databricks experience required, including:
- Delta Lake at scale: MERGE-based upserts, OPTIMIZE, time travel, schema evolution.
- Lakeflow Declarative Pipelines and Lakeflow Jobs. Prior Delta Live Tables and Workflows experience counts.
- Unity Catalog for governance, access control, and lineage.
- Databricks SQL warehouses and compute selection, including cost awareness.
- Databricks Certified Data Engineer Associate or Professional certifications are highly valued.
- Ability to communicate with stakeholders at all levels, ensuring action can be taken from written reports without follow-up meetings.
- Works well with analysts, product owners, and source system teams to turn ambiguous requirements into workable specifications.
- Scopes work into shippable increments.
- Strong organizational and project management skills with the ability to manage multiple priorities and deadlines.
- A proactive, can-do attitude with a willingness to take ownership of tasks and drive them to completion.
- Ability to work independently while also being a team player who thrives in a collaborative environment.
- Discretion and trustworthiness in handling sensitive information and supporting high-level strategic initiatives.
Preferred Experience:
- Lakeflow Connect or another managed ingestion or CDC tool, plus Databricks Asset Bundles, Terraform, or Docker.
- Data quality and observability tooling, BI tools and reverse ETL, and PII masking and retention practices.
- RAG, vector search, or document processing workloads, and MLflow tracing and evaluation.
- No prior AI experience is required, but GenAI Engineering experience (document ingestion pipelines, vector search indexes, and MLflow tracing and evaluation) is valued.
Physical Demands and Working Environment:
The physical demands described here are representative of those that must be met by an employee to successfully perform the essential functions of this position. Reasonable accommodation may be made to enable individuals with disabilities to perform the functions.
- While performing the duties of this job in a home office setting, the employee is regularly required to work on a computer for extended periods of time.
- Frequent use of a computer requires fine motor skills and hand-eye coordination.
- Ability to sit for extended periods while working from home or a designated workspace.
- Ability to perform tasks that require sustained attention and focus.
- Occasional lifting of materials up to 25 pounds.
- Travel to attend team meetings may be required.
- To facilitate working from home, and as a requirement for this role, the employee must provide reliable internet connection with sufficient bandwidth to execute all job functions and technology setup conducive to remote work. The company laptop will be provided.
- A quiet, distraction-free workspace is required for maintaining productivity.
- Collaboration with team members may occur through virtual meetings and communication platforms.