Logo for Cognitive Medical Systems, Inc.

Lead Data Engineer - Remote (WFH)

Role overview

Qualifications

  • Bachelor's degree in Computer Science, Data Engineering, or related field
  • 8 or more years in data engineering, including 4 or more years building and operating high-volume batch ETL systems
  • Expert-level Python and PySpark
  • Production experience with Apache Airflow (AWS MWAA) for pipeline orchestration at scale

Responsibilities

  • Own summary records of prescription drug transactions, also known as pharmacy drug events (PDE), ingestion, validation, editing, and storage pipelines
  • Engineer and maintain all vendor reference data edits from a variety of sources for drug event record, pharmacy data and files
  • Build, maintain, and optimize Apache Airflow and PySpark transformations across various technologies, including S3, Redshift Serverless, Snowflake, and Databricks
  • Monitor PDE submissions for adjustment trends and informational edit rates

About the company

Cognitive Medical Systems, Inc. logo

Cognitive Medical Systems, Inc.

Digital Health & Health Tech

Cognitive Medical Systems’ purpose is to empower people and organizations to optimize healthcare delivery through innovative technology solutions. Since our software development company’s founding in 2010, consistent focus on innovative and solutions-based HealthIT applications drives our progress forward. Our rapid expansion is fueled by our continued improvement and innovative approach to solving IT-related challenges. The Cognitive Way powers our results-driven approach to make our customers successful. Today our government contracting efforts serve the Department of Defense, Department of Veterans Affairs, the Department of Health and Human Services, and several commercial customers. Cognitive specializes in an area called Clinical Decision Support (CDS). CDS technology enables precise access to relevant point-of-care data to doctors, nurses, and patients that encompasses a patient’s entire healthcare record. Cognitive’s provides “real-time” CDS that helps simplify the complex clinical workflows that define modern healthcare. Simply, our innovations lead to better healthcare.

Company details

Company typeSME
IndustryDigital Health & Health Tech
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Cognitive:

Cognitive is an IT and software engineering services company dedicated to elevating the quality, speed and delivery of today’s US government healthcare programs. With a wealth of clinical expertise and hands-on experience, our team understands the significant challenges our government clients—and their customers—face every day. This real-world experience equips us to develop IT solutions that seamlessly connect all facets of healthcare delivery. At Cognitive, we’re guiding government agencies to the forefront of technology innovation in healthcare delivery.

Position Overview:

*This position is contingent upon contract award*

The Lead Data Engineer is responsible for PDE processing pipelines, data quality, and IDR integration for the Drug Data Processing System (DDPS) and Payment Reconciliation System (PRS) O&M contract. This individual owns the end-to-end PDE data pipeline — receiving, validating, editing, and storing approximately 9 to 10 million Prescription Drug Event records daily — and maintains all vendor reference data integrations including FDA, NCPDP, FDB, MediSpan, NPPES, and OIG data sources.

This is a remote position; however, Cognitive hires only in the following designated U.S. states based on contract and business requirements: VA, DC, MD, TN, FL, AZ, CO, OR, and TX.

Key Responsibilities:

  • Own summary records of prescription drug transactions, also known as pharmacy drug events (PDE), ingestion, validation, editing, and storage pipelines. 
  • Apply and maintain CMS-defined business rule validations, returning edit results to plan sponsors, and minimizing PDEs requiring manual analysis.
  • Engineer and maintain all vendor reference data edits from a variety of sources for drug event record, pharmacy data and files, exclusion and preclusion list; provide health plan and CMS status updates on PDE submissions.
  • Build, maintain, and optimize Apache Airflow and PySpark transformations across variety of technologies, including S3, Redshift Serverless, Snowflake, and Databricks.
  • Maintain system compatibility with other CMS data model changes.
  • Monitor PDE submissions for adjustment trends and informational edit rates; provide analytical capability to link rejected PDEs with IDR records; provide beneficiary and plan-level cost aggregations to PRS for year-end reconciliation.
  • Support data transfer projects and new data share integrations with CMS and downstream stakeholders including API and multi-cloud approaches as directed by the CMS.
  • Maintain target unit test coverage for all new pipeline code; support BDD and TDD practices; track and surface real-time PDE processing metrics including volumes, error rates, and reconciliation accuracy.
  • Operate within the CMS Lean-Agile Release Train adhering to the Iteration Schedule; deliver code through GitHub, Jenkins, JFrog Artifactory/XRay, SonarQube, and Snyk CI/CD pipeline.

Qualifications:

  • Bachelor's degree in Computer Science, Data Engineering, or related field; 8 or more years in data engineering, including 4 or more years building and operating high-volume batch ETL systems.
  • Expert-level Python andPySpark
  • Production experience with Apache Airflow (AWS MWAA) for pipeline orchestration at scale.
  • Production experience with AWS Redshift Serverless, Snowflake, and Amazon Athena for data warehousing and query at scale.
  • Hands-on experience with Databricks (Notebooks, Jobs) for large-scale data processing.
  • Strong SQL skills for multi-source data transformation, editing logic, and reconciliation of data pipelines.
  • Experience with AWS S3, SQS, Lambda, and Event Bridge for data ingestion and event-driven orchestration.
  • Experience with Linux environments including RHEL, CentOS, and Amazon Linux 2, and shell scripting in Bash.
  • Demonstrated experience operating data pipelines under strict data quality, accuracy, and timeliness SLAs in a federal or healthcare environment.
  • Experience with version control and CI/CD integration using GitHub, Jenkins, and JFrog Artifactory.
  • Ability to pass CMS and internal required background checks for public trust.

Why Join Us? 

  • Be part of a mission-driven organization making a difference in healthcare IT.
  • Collaborate with innovative and passionate professionals that are there to support you at every turn
  • Enjoy a supportive work/life balance with the flexibility of a 100% remote company.
  • Benefit from opportunities for growth and development in a dynamic environment.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Engineer Related jobs

Other jobs at Cognitive Medical Systems, Inc.

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.