Logo for Cherokee Federal

Healthcare Data Scientist

Role overview

Qualifications

  • Hands-on experience writing SQL queries for data transformation and analysis.
  • Experience using Python (e.g., pandas) for data processing.
  • Hands-on experience using PySpark for distributed data processing.
  • Experience working within cloud-based data platforms, preferably Azure Synapse or similar.

Responsibilities

  • Develop, maintain, and optimize data pipelines using SQL, Python (pandas), and PySpark.
  • Execute and manage notebook-based workflows within Azure Synapse, including debugging and documentation.
  • Perform structured data validation, including row counts, null checks, duplicate detection, schema validation, and allowed value enforcement.
  • Conduct exploratory data analysis to identify patterns, anomalies, and data quality concerns.

About the company

Cherokee Federal logo

Cherokee Federal

Government Administration

Cherokee Federal – a division of Cherokee Nation Businesses – is a team of tribally owned federal contracting companies focused on building solutions, solving complex challenges, and serving the nation’s mission around the globe for more than 60 federal clients. With our heritage of ingenuity coupled with modern business practices, we serve as a trusted partner that can innovate and implement solutions. Our team of companies, with more than 9,000+ employees, manages nearly 2,000 projects of all sizes across the construction, engineering and manufacturing, and mission solutions portfolios — ranging from advanced data analytics and telehealth to cybersecurity, cloud and logistics. Cherokee Federal’s team of small disadvantaged business entities, many of which are 8(a) and/or HUBZone certified, offer attractive contract vehicles with unique advantages – resulting in a streamlined, responsive contract management process.

Company details

Company typeXLarge
IndustryGovernment Administration
Company size5001 - 10000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

ATA is seeking a Data Scientist to support data pipeline development, validation, and analysis within a cloud-based Health IT data platform. This role is hands-on and delivery-focused, with an emphasis on building reliable, reproducible data workflows using SQL, Python, and PySpark in an Azure Synapse environment. 

A core expectation of this role is the ability to work across the full data lifecycle, from ingestion through transformation to final dataset delivery, while maintaining data quality and traceability. The ideal candidate is comfortable debugging data issues end-to-end, understands how data structure and join logic impact outputs, and applies disciplined validation and documentation practices. This role also supports exploratory data analysis and the development of derived datasets to enable analytics and downstream use cases. The position will work extensively with healthcare data originating from EHR systems and interface feeds, including HL7 v2 and FHIR data, clinical terminology, and source-to-target data mappings. 

Key Responsibilities: 
Data Pipeline Development 

• Develop, maintain, and optimize data pipelines using SQL, Python (pandas), and 

PySpark. 

ATA, LLC | 752 Walker Road | Suite D | Great Falls, VA 22066 | www.ata-llc.com 

Advanced Technology Applications 

• Execute and manage notebook-based workflows within Azure Synapse, including 

debugging and documentation. 

• Process and transform structured and semi-structured data in formats such as CSV, 

JSON/NDJSON, and Parquet. 

• Work within ETL/ELT pipelines across raw, curated, and production data layers. 

• Ingest, profile, map, and transform healthcare data from EHR systems and interface 

feeds while preserving source lineage and clinical context 

Data Quality and Validation 

• Perform structured data validation, including row counts, null checks, duplicate 

detection, schema validation, and allowed value enforcement.

• Identify and resolve data quality issues such as schema drift, inconsistencies, and 

transformation errors across pipeline stages. 

• Apply repeatable testing and validation practices, including reproducing issues, verifying 

fixes, and ensuring data reliability prior to downstream use. 

• Validate source-to-target mappings and reconcile records across source and destination 

systems during data conversion and migration activities. 

Data Analysis and Dataset Development 

• Conduct exploratory data analysis to identify patterns, anomalies, and data quality 

concerns. 

• Develop derived datasets to support reporting, analytics, and downstream data use 

cases. 

• Collaborate with stakeholders to translate data requirements into usable datasets and 

metrics. 

Data Modeling and Structure Awareness 

• Interpret and work with data schemas, including column definitions, data types, primary 

and composite keys, and table relationships. 

• Manage dataset grain and understand how join strategies (e.g., one-to-one vs. one-to

many) impact row counts and outputs. 

Debugging and Troubleshooting 

• Trace data issues from source ingestion through transformation logic to final outputs. 

• Use logs and debugging approaches to diagnose and resolve pipeline issues. 

• Document data transformations, assumptions, mappings, and validation results in a 

clear and consistent manner. 

• Collaborate with engineers, analysts, and stakeholders to ensure data usability, 

integrity, and alignment with requirements. 

ATA, LLC | 752 Walker Road | Suite D | Great Falls, VA 22066 | www.ata-llc.com 

Advanced Technology Applications 

• Communicate data issues, findings, and workflow updates with technical team 

members. 

Minimum Qualifications: 

• Hands-on experience writing SQL queries for data transformation and analysis. 

• Experience using Python (e.g., pandas) for data processing. 

• Hands-on experience using PySpark for distributed data processing. 

• Experience working within cloud-based data platforms, preferably Azure Synapse or 

similar. 

Understanding ETL/ELT concepts and data pipeline architecture. 

• Experience working with structured and semi-structured data formats (CSV, JSON, 

Parquet). 

• Familiarity with Git and collaborative development workflows. 

• Strong problem-solving and debugging skills across data pipelines. 

• Ability to validate and ensure data quality through structured checks and testing 

practices. 

• Strong written and verbal communication skills. 

Preferred Qualifications: 

• Experience working with healthcare interoperability standards and message formats, 

particularly HL7 v2 and FHIR. 

• Familiarity with clinical terminology and code systems such as SNOMED CT, ICD-10, 

LOINC, RxNorm, or CPT. 

• Experience supporting large-scale EHR data conversion or migration efforts, particularly 

between RPMS and Oracle Health/Cerner—or a comparable legacy-to-modern EHR 

migration. 

• Understanding of clinical data mapping, terminology normalization, provenance, and 

validation across source and target systems 

General personal traits we know will connect well with the team: 

• Dependable, self-directed, and able to meet commitments. 

• Comfortable learning new subject areas and solving unfamiliar problems. 

• Enjoys collaborating across technical and subject-matter disciplines. 

• Pragmatic and able to select the appropriate tool for the requirement. 

ATA, LLC | 752 Walker Road | Suite D | Great Falls, VA 22066 | 

Advanced Technology Applications 

About ATA: A leading provider of full-stack data and AI solutions with deep mission experience supporting the Department of Defense, Department of Homeland Security, IC, and federal agencies. Founded in 2008 and headquartered in Virginia, ATA specializes in secure, scalable, and operationally ready technologies that transform data into actionable insight. With a proven track record in advanced analytics, software development, and technology integration, ATA is uniquely positioned to accelerate federal organizations in a number of ways by leveraging powerful contracting options and delivering cutting-edge, automation-enhanced, AI-assisted capabilities—tailored to need and mission assurance imperatives. We believe our diversity of infrastructure, data, and application experience is valuable and is one of the attributes that sets ATA apart. 

Summary of Benefits: We expect each member of our team to fully engage creatively and work collaboratively and perform each day to the best of their ability.  To support this, we have created a benefit package focused on professional growth, achieving a healthy work-life balance, and participation in the long-term success of the company.  Our benefits include generous paid time-off; an employee incentive program; continuous learning culture, Internal

Investment Projects (IIP), virtual brown-bags/level-ups, and other professional development activities; recruiting bonuses; 3% 401k Safe Harbor contributions; Medical/Dental/Vision, Long & Short-term Disability, AD&D insurance, and Life Insurance.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Scientist Related jobs

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.