Principal33
IT Services & IT Consulting
See how your profile stacks up against this role.
We compared the job requirements to your profile to show where you're strong and where you fall short.
Data pipelines and integration
• Design, build and operate scalable batch, micro-batch and streaming data pipelines using Python, PySpark and Azure Databricks.
• Integrate internal and external data sources, including REST APIs, GraphQL, WebSocket and gRPC interfaces, databases, files, event streams and third-party data feeds.
• Develop robust ingestion solutions for structured, semi-structured and unstructured data, including JSON, CSV, Parquet, Delta and API-based payloads.
• Build reliable web scraping and data acquisition components where APIs or managed integration mechanisms are unavailable.
• Implement pagination, throttling, retries, exponential backoff, checkpointing, schema evolution and recovery patterns for external integrations.
• Design pipelines that support idempotent processing, reprocessing, controlled backfills and graceful recovery from partial failures.
• Develop and maintain batch and streaming patterns for market, weather, fundamental and time-series data.
Software engineering
• Develop modular, reusable and testable Python and PySpark components rather than relying on monolithic notebooks.
• Apply object-oriented and functional design principles appropriately to data-focused development.
• Structure solutions as maintainable software projects with clear separation between source code, configuration, tests, deployment assets and notebooks.
• Write clean, readable and well-documented code using type hints, meaningful interfaces and appropriate design patterns.
• Build automated unit, integration, contract and data-quality tests and incorporate them into delivery pipelines.
• Conduct code reviews and promote engineering standards covering readability, testability, security, performance and maintainability.
• Package reusable functionality as Python modules or wheels where appropriate.
• Troubleshoot complex issues across source systems, APIs, processing logic, infrastructure and production runtime environments.
Data architecture and modelling
• Design maintainable data models that support analysts, traders, reporting solutions and downstream data products.
• Implement Lakehouse and Medallion architecture patterns across Bronze, Silver and Gold layers.
• Preserve raw data appropriately while applying cleansing, validation, standardisation and business transformations in downstream layers.
• Design solutions for schema evolution, data retention, lineage and reproducible processing.
• Apply sound data architecture principles across operational, analytical, event-based and time-series workloads.
• Optimise data layouts, partitioning, joins, file sizes, caching and Spark execution plans for performance and cost.
Orchestration and DataOps
• Design, schedule and operate workflows using Databricks Workflows and Astronomer.
• Implement dependency management, parameterisation, environment-specific configuration and controlled promotion across development, test and production environments.
• Define operational runbooks and support effective diagnosis, recovery and problem management.
• Monitor pipeline health, freshness, completeness, performance and data-quality indicators.
• Use production-safe release patterns, including controlled rollouts, rollback and validation where appropriate.
GitOps, CI/CD and Infrastructure as Code
• Manage all production code through Git using clear branching, pull-request and review practices.
• Build and maintain automated CI/CD pipelines using GitHub Actions and/or Azure DevOps.
• Deploy Databricks jobs, pipelines and application artefacts using Databricks Asset Bundles or equivalent approved mechanisms.
• Provision and configure relevant cloud and Databricks resources through Terraform.
• Treat application code, infrastructure, data pipeline definitions and operational configuration as version-controlled artefacts.
• Apply automated validation, security scanning and testing before production deployment.
• Contribute to reusable pipeline templates, engineering standards and platform automation.
Reliability, performance and cost efficiency
• Engineer solutions for availability, recoverability, scalability and predictable operational behaviour.
• Optimise Spark workloads through appropriate partitioning, built-in Spark functions, efficient joins, adaptive execution and avoidance of unnecessary shuffles or UDFs.
• Select suitable compute models and cluster configurations based on workload characteristics.
• Apply cost-awareness to pipeline design, compute sizing, scheduling, storage and data-retention decisions.
• Monitor resource consumption and identify opportunities to reduce processing times and cloud costs without compromising reliability or data quality.
• Balance immediate delivery requirements with sustainable architecture and long-term maintainability.
Data quality, governance and security
• Implement automated data validation, schema checks, null checks, referential-integrity controls and business quality rules.
• Detect and manage schema drift and unexpected changes in source data.
• Use Delta Lake and Unity Catalog capabilities to support data lineage, access control, metadata and governance.
• Ensure secrets and credentials are handled securely using approved secret-management mechanisms and managed identities.
• Maintain technical documentation, metadata and operational information for assigned data products.
• Collaborate with data governance, architecture, security and platform teams to ensure alignment with enterprise standards.
Collaboration and delivery
• Work closely with traders, analysts, data scientists, software engineers, product owners and platform teams to translate business requirements into robust technical solutions.
• Communicate design decisions, risks, dependencies and technical trade-offs clearly to technical and non-technical stakeholders.
• Contribute reusable components, templates, documentation and engineering guidelines for the wider data community.
• Work effectively in a distributed, international and cross-functional environment.
An experienced data engineer with at elast 5 years of experience, ideally in the energy sector and/or trading.
Be able to operate fundamental power-price forecasting models for short- to mid-term trading and large amounts of data.
Be able to work fully remote in a collaborative environment with interdisciplinary teams.
Good communication skills and professional behaviour.
After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.
Marcus Rivera
Chief Revenue Officer

Ciklum

PRECISIONeffect

Precision For Medicine

Precision Medicine Group

2am.tech

Principal33

Principal33

Principal33