Logo for Ciklum

Data Engineer

Role overview

Qualifications

  • 3+ years of commercial Data Engineering experience building and maintaining production data pipelines
  • Solid proficiency in Python, SQL, and standard data manipulation frameworks
  • Direct working knowledge of graph databases (Neo4j), vector search engines (Qdrant), or static code parsing frameworks (Tree-sitter)
  • Proficiency with Git version control, Docker containerization, unit/integration testing for data pipelines, and CI/CD workflows

Responsibilities

  • Develop, test, and maintain robust data pipelines that process AST outputs, knowledge graph structures, and vector embeddings
  • Write data processing scripts to ingest, normalize, and update code dependency graphs in Neo4j and vector indexes in Qdrant
  • Parse raw source code structures and raw metadata into structured Markdown assets and standardized JSON/RAG inputs
  • Monitor execution throughput, resolve batch processing errors, and ensure data state consistency across pipeline runs

About the company

Ciklum logo

Ciklum

Ciklum is a global Digital Solutions Company for Fortune 500 and fast-growing organisations alike around the world. The company is headquartered in London and has software development centres and branch offices in the United States, Spain, Switzerland, Denmark, Israel, Poland, Ukraine, Czech Republic, Slovakia, Romania, UAE and Pakistan.Ciklum builds tailored digital solutions that leverage emerging technologies for such clients as Just Eat, Flixbus, Metro Markets, EFG International, Zurich Insurance, Lottoland and others.For more information about us visit www.ciklum.com

Company details

Company typeStartup
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Ciklum is looking for a Data Engineer to join our team full-time in Ukraine.

We are a custom product engineering company that supports both multinational organizations and scaling startups to solve their most complex business challenges. With a global team of over 4,000 highly skilled developers, consultants, analysts and product owners, we engineer technology that redefines industries and shapes the way people live.

About the role:

As a Data Engineer, become a part of a cross-functional development team engineering experiences of tomorrow.

The Legacy Code Semantic Documentation Project is a 26-week enterprise initiative for a global industrial automation leader. The primary objective is to engineer an automated, AI-assisted pipeline to generate structured, system-level Markdown documentation directly from an undocumented ~400K LOC codebase (spanning IEC 61131-3 languages and ANSI C/C++).

Operating within a dedicated, zero-data-egress secure tenant, the technical architecture combines Tree-sitter AST parsing, Neo4j knowledge graphs, Qdrant vector search, and self-hosted open-weight LLMs (Llama 3.1 / Mixtral family) with NLI-based validation. Delivery is structured across 7 work packages executing over a 6-month period, governed by strict contractual KPIs: ≥95% code coverage, ≥92% NLI-verified factual precision, and ≥90% SME validation acceptance. All outputs align with EU Cyber Resilience Act (EU CRA) requirements for SBOM and source-level traceability.

Responsibilities:

  • Data Pipeline Execution: Develop, test, and maintain robust data pipelines that process AST outputs, knowledge graph structures, and vector embeddings
  • Knowledge Base Ingestion: Write data processing scripts to ingest, normalize, and update code dependency graphs in Neo4j and vector indexes in Qdrant
  • Data Cleansing & Transformation: Parse raw source code structures and raw metadata into structured Markdown assets and standardized JSON/RAG inputs
  • Pipeline Monitoring & Debugging: Monitor execution throughput, resolve batch processing errors, and ensure data state consistency across pipeline runs
  • Collaboration: Work closely with Senior Data Engineers and AI/ML Engineers to optimize data retrieval speeds and pipeline efficiency

Requirements:

  • Professional Experience: 3+ years of commercial Data Engineering experience building and maintaining production data pipelines
  • Core Technical Proficiency: Solid proficiency in Python, SQL, and standard data manipulation frameworks
  • Hands-on Stack Exposure: Direct working knowledge of graph databases (Neo4j), vector search engines (Qdrant), or static code parsing frameworks (Tree-sitter)
  • Engineering Best Practices: Proficiency with Git version control, Docker containerization, unit/integration testing for data pipelines, and CI/CD workflows
  • Problem-Solving & Detail Orientation: Strong analytical skills with a focus on data accuracy, schema consistency, and output validation
  • Language: Professional proficiency in English (B2+/C1)

What’s in it for you?

  • Strong community: Work alongside top professionals in a friendly, open-door environment
  • Growth focus: Take on large-scale projects with a global impact and expand your expertise
  • Tailored learning: Boost your skills with internal events (meetups, conferences, workshops), Udemy access, language courses, and company-paid certifications
  • Endless opportunities: Explore diverse domains through internal mobility, finding the best fit to gain hands-on experience with cutting-edge technologies
  • Flexibility: Enjoy radical flexibility – work remotely or from an office, your choice
  • Care: We’ve got you covered with company-paid medical insurance, mental health support, and financial & legal consultations

About us:

At Ciklum, we are always exploring innovations, empowering each other to achieve more, and engineering solutions that matter. With us, you’ll work with cutting-edge technologies, contribute to impactful projects, and be part of a One Team culture that values collaboration and progress.

As one of Ukraine’s largest IT companies and a top employer recognized by Forbes, we’ve spent over 20 years delivering meaningful tech solutions. We proudly support diverse talent and military veterans, recognizing their unique skills and perspectives they bring to shaping the future.

Explore, empower, engineer with Ciklum!

Interested already? We would love to get to know you! Submit your application. We can’t wait to see you at Ciklum.

#LI-NV1

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Data Engineer Related jobs

Other jobs at Ciklum

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.