We are sharing a full-time opportunity for an experienced Director of Infrastructure Engineering with deep expertise in AWS, GCP, infrastructure as code, CI/CD, platform engineering, reliability, security, and technical leadership to build and scale the infrastructure supporting production AI systems.
The role combines hands-on infrastructure engineering with strategic leadership across cloud architecture, developer platforms, observability, reliability, security, and engineering operations. The successful candidate will own both long-term infrastructure strategy and day-to-day execution, including production resilience, incident response, automation, and team development.
Key Responsibilities
Cloud Infrastructure Strategy
-
Own the architecture and evolution of multi-cloud infrastructure across AWS and GCP
-
Define long-term infrastructure strategy aligned with scalability, reliability, security, and cost objectives
-
Evaluate architecture trade-offs across cloud platforms and deployment models
-
Ensure infrastructure can support growing production workloads
-
Maintain strong operational standards as systems and teams scale
Infrastructure as Code & Platform Engineering
-
Build and maintain infrastructure as code using Terraform or comparable tooling
-
Ensure environments are reproducible, auditable, and version controlled
-
Develop reusable infrastructure patterns and platform abstractions
-
Reduce operational friction through automation and internal tooling
-
Improve developer productivity across engineering teams
CI/CD & Software Delivery
-
Design and improve CI/CD systems
-
Enable fast, reliable, and secure software delivery
-
Standardise deployment practices across engineering teams
-
Improve release safety through automation and validation
-
Reduce deployment risk while supporting rapid iteration
Reliability & Production Operations
-
Define and drive reliability practices across production systems
-
Establish SLOs, error budgets, incident-response processes, and on-call operations
-
Lead disaster-recovery planning and resilience initiatives
-
Improve operational maturity through blameless postmortems
-
Ensure infrastructure remains highly available and recoverable
Observability & Incident Management
-
Build comprehensive observability across metrics, logs, traces, and alerting
-
Improve detection of production issues before they affect users
-
Define meaningful operational signals and alert thresholds
-
Support structured troubleshooting and incident response
-
Use production data to guide reliability improvements
Security & Compliance
-
Partner with Security and Engineering leadership to embed security by default
-
Strengthen infrastructure controls and operational risk management
-
Support compliance with frameworks such as ISO 27001, SOC 2, and CMMC
-
Evaluate security risks across cloud architecture and platform tooling
-
Maintain secure infrastructure practices throughout the software lifecycle
Team Leadership & Engineering Culture
-
Lead and grow a high-performing Infrastructure or Platform Engineering team
-
Establish a culture of ownership, operational excellence, and continuous improvement
-
Mentor experienced engineers and support technical development
-
Influence infrastructure strategy across engineering and executive stakeholders
-
Balance strategic leadership with hands-on technical involvement
Ideal Profile
-
8+ years of experience building and operating production infrastructure, platform engineering, DevOps, or SRE systems
-
3+ years of experience leading engineering teams
-
Deep expertise with AWS, GCP, or multi-cloud production environments
-
Strong experience with Terraform or comparable infrastructure-as-code tooling
-
Strong knowledge of Kubernetes and containerised infrastructure
-
Experience building and operating modern CI/CD platforms
-
Demonstrated success designing highly available, observable, and resilient systems
-
Strong understanding of infrastructure security, compliance, and operational risk
-
Experience scaling infrastructure and engineering organisations in fast-moving environments
-
Excellent written and verbal communication skills
-
Ability to influence technical strategy across engineering and executive stakeholders
-
Experience supporting AI/ML infrastructure or large-scale data platforms is highly valuable
-
Familiarity with model training, inference, evaluation, or data-pipeline infrastructure is advantageous
-
Experience with regulated cloud environments such as FedRAMP, GovCloud, or CMMC Level 2 is beneficial
-
Contributions to platform engineering, open-source infrastructure, or developer-productivity initiatives are a plus
Engagement Details
-
Full-time engagement
-
Fully remote
-
Base compensation: $350,000–$500,000/year
-
Work will involve AWS, GCP, infrastructure as code, CI/CD, observability, reliability engineering, platform tooling, security, and technical leadership
-
Strong hands-on infrastructure expertise and leadership experience are central to this role
-
Responsibilities will span strategic architecture, production operations, developer experience, team leadership, and incident management
-
The role may support AI/ML infrastructure, large-scale data platforms, or regulated cloud environments
-
Infrastructure priorities, compliance requirements, and platform architecture may evolve as production systems scale
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy