Growth through diversity, equity, and inclusion. As an ethical business, we do what is right — including ensuring equal opportunities and fostering a safe, respectful workplace for each of us. We believe diversity fuels both personal and business growth. We're committed to building an inclusive community where all our people thrive regardless of their backgrounds, identities, or other personal characteristics.
What You’ll Be Doing
Platform Design, Development & Evolution: Architect, build, and continuously evolve the core M&O platform and services, leveraging modern technologies and best practices to provide comprehensive observability and data quality functions.
Ensuring System Reliability & Performance: Actively maintain and enhance the stability, availability, and performance of critical applications and data infrastructure by integrating Site Reliability Engineering (SRE) principles directly into the platform's design and operation.
Proactive Issue Detection & Resolution: Develop and integrate intelligent systems within the platform to proactively identify, diagnose, and trigger automated or semi-automated resolution for technical issues, performance bottlenecks, and anomalies across all operational and data systems.
Advanced Data Quality Platform Implementation: Build and integrate capabilities within the M&O platform for continuously measuring, monitoring, and reporting on all critical data quality dimensions (Timeliness, Consistency, Completeness, Accuracy, Validity, Uniqueness) across diverse supply chain pipelines, including Warehousing, Transformation, SAP, Manufacturing, and other critical data sources.
Centralized Insights & Dashboarding: Develop and manage a unified dashboarding interface (e.g., Grafana, Power BI, Databricks) within the platform to visualize key performance indicators, system health, operational metrics, financial insights (FinOps), and granular data quality metrics for various stakeholders.
Automating Infrastructure & Operations: Drive the platform's automation capabilities through Infrastructure as Code (IaC), Continuous Integration/Continuous Delivery (CI/CD) pipelines, and extensive scripting (Python, Shell) for provisioning, deployment, and operational workflows.
Managing Core Data & Cloud Technologies: Integrate and optimize essential technologies like Fivetran, AKS, Kafka, Azure SQL Server, Databricks, and Flink into the M&O platform, ensuring seamless operation and data flow within the Azure cloud environment.
Optimizing Cloud Resources & Costs: Embed FinOps practices and reporting into the platform to monitor cloud resource utilization and spending, identifying opportunities for cost reduction and ensuring efficient allocation of infrastructure investments.
Fostering a Culture of Continuous Improvement: Leverage platform data and SRE practices to continuously analyze operational incidents, performance trends, and data quality issues, driving ongoing enhancements to both the platform and the systems it monitors.
Enabling Data-Driven Decision Making & Compliance: Ensure the M&O platform provides accurate, timely data and insights, derived from both monitoring and high-quality data across all domains, to support informed business decisions, while also guaranteeing compliance with regulatory requirements for data integrity and auditability.
What We’re Looking For
At least 4 years of experience in a similar position.
Hands-on experience with Microsoft Azure, including both PaaS and IaaS services.
Experience designing, building, and maintaining CI/CD pipelines using GitHub Actions for Azure-based environments.
Practical knowledge of administering Linux and Windows virtual machines.
Experience working with containerisation and orchestration technologies, including Docker and Kubernetes.
Ability to create and maintain automation scripts using Python, Bash, and PowerShell.
Experience with code quality and security analysis tools, particularly SonarQube.
Hands-on experience with Infrastructure as Code tools, especially Terraform and Ansible.
Good understanding of networking principles and practices in cloud and hybrid environments.
Experience troubleshooting and resolving issues related to system integrations.
Familiarity with monitoring and observability solutions, such as Prometheus, Grafana, or comparable technology stacks.
Strong analytical and problem-solving skills, with the ability to work effectively in a collaborative environment.
English at least B2.