LeadStack Inc. is an award-winning, one of the nation's fastest-growing, certified minority-owned (MBE) staffing services provider of contingent workforce. As a recognized industry leader in contingent workforce solutions and Certified as a Great Place to Work, we're proud to partner with some of the most admired Fortune 500 brands in the world.
Job Title: Data Engineer
Job Duration: 6 Month(s)
Location: Remote, EST/CST only, Cincinnati or Chicago preferred
PR: $60/hr - $78/hr on W2
Manager would like each candidate to submit a short writeup on their actual AI experience in regards to machine learning platforms, not just AI coding assistants. Please include this in the additional information section of each submittal.
During the behavioral interview, we were looking for clear, structured communication about prior work—particularly the ability to explain the business context, design decisions, and tradeoffs behind a pipeline they had built. While the candidate demonstrated familiarity with technical concepts and optimization considerations, their initial responses were often highly detailed without directly addressing the question or providing the requested high-level context.
Information
Data Engineer profile but preference towards those with feature/machine learning experience.
Location: Remote, EST/CST only, Cincinnati or Chicago preferred
Project: Building a machine learning layer using Databricks or Vertex AI platform
Target start date: As soon as possible
Interview Process: 1 Round, 90 minutes (Behavioral and Tech screening)
Job Description
Your role will involve collaboration with cross-functional teams, leveraging cutting-edge technologies, and ensuring scalable, efficient, and secure data engineering practices. You'll be working with data across 2 cloud platforms (Azure and GCP) and thus will be working with a wide variety of technologies. A strong emphasis will be placed on expertise in python, Vertex AI, and advanced feature engineering techniques.
Responsibilities Take ownership of systems, processes, and the tech stack while driving features to completion through all phases of the entire SDLC. This includes internal and external facing applications as well as process improvement activities:
· Build and Maintain Data Pipelines: Design, build, and maintain scalable, efficient, and reliable data pipelines to support data ingestion, transformation, and integration across diverse sources and destinations, using tools such as Kafka, Databricks, and similar toolsets.
· Drive Digital Innovation: Leverage innovative technologies and approaches to modernize and extend core data assets, including SQL-based, NoSQL-based, cloud-based, and real-time streaming data platforms.
· Implement Feature Engineering: Develop and manage feature engineering pipelines for machine learning workflows, utilizing tools like Vertex AI, BigQuery ML, and custom Python libraries.
· Implement Automated Testing: Design and implement automated unit, integration, and performance testing frameworks to ensure data quality, reliability, and compliance with organizational standards.
· Optimize Data Workflows: Optimize data workflows for performance, cost efficiency, and scalability across large datasets and complex environments.
· Draft and Review Documentation: Draft and review architectural diagrams, interface specifications, and other design documents to ensure clear communication of data solutions and technical requirements.
Requirements:
- Bachelor's degree typically in Computer Science, Management Information Systems, Mathematics, Business Analytics or another STEM degree.
- 4+ years of professional Data Development experience.
- 4+ years of experience with SQL and NoSQL technologies.
- 3+ years of experience building and maintaining data pipelines and workflows.
- 2+ years of experience developing with Python.
- Experience in feature engineering for machine learning pipelines.
- Experience with CI/CD pipelines and processes.
- Experience with automated unit, integration, and performance testing.
- Experience with version control software such as Git.
- Full understanding of ETL and Data Warehousing concepts.
- Strong understanding of Agile principles (Scrum).
Preferred Qualifications
· Knowledge of Structured Streaming (Spark, Kafka, EventHub, or similar technologies).
· Experience with GCP services (or Databricks equivalent) such as BigQuery, Vertex AI Platform, Cloud Storage, AutoMLOps, and Dataflow.
· Experience with GitHub SaaS/GitHub Actions.
· Experience understanding Databricks concepts.
· Experience with PySpark and Spark development.
· Experience with Service Oriented Architecture.
To know more about current opportunities at LeadStack, please visit us at https://leadstackinc.com/careers/
Should you have any questions, feel free to call me on or send an email on _____________________