We are sharing a specialised full-time opportunity for experienced Cloud and DevOps Engineers with strong hands-on expertise in Kubernetes operations, AWS infrastructure, Infrastructure-as-Code, and production CI/CD environments.
This role supports advanced infrastructure-focused technical work involving cloud systems, cluster operations, infrastructure automation, and deployment pipelines. Selected engineers will design and evaluate challenging technical tasks, diagnose production-style infrastructure issues, develop rigorous solutions, and contribute expert feedback across Kubernetes, AWS, IaC, and DevOps workflows.
Key Responsibilities
Kubernetes Operations & Troubleshooting
-
Design and evaluate technically challenging Kubernetes infrastructure tasks
-
Diagnose cluster failures and identify underlying configuration, networking, scheduling, resource, or service issues
-
Develop technically accurate solutions to complex Kubernetes operational problems
-
Review proposed remediation approaches for correctness and production feasibility
-
Apply hands-on operational judgement beyond basic manifest authoring or managed-control-plane usage
AWS Cloud Infrastructure
-
Design and review infrastructure scenarios involving production AWS environments
-
Evaluate integrations involving AWS Lambda, API Gateway, and DynamoDB
-
Assess architecture, service configuration, dependencies, and operational considerations
-
Identify technically incorrect assumptions or weak infrastructure designs
-
Review solutions for scalability, reliability, and practical implementation quality
Infrastructure-as-Code
-
Design and evaluate infrastructure automation using Terraform and/or AWS CDK
-
Review IaC architecture, resource definitions, dependencies, and deployment approaches
-
Assess maintainability, correctness, and operational safety of infrastructure definitions
-
Identify configuration weaknesses and infrastructure automation risks
-
Develop clear reference solutions for complex IaC scenarios
CI/CD Engineering
-
Evaluate CI/CD pipeline designs and deployment workflows
-
Review build, test, release, and deployment automation for technical soundness
-
Assess pipeline dependencies, failure handling, and operational reliability
-
Identify gaps in deployment logic, validation, or rollback planning
-
Evaluate approaches for maintaining reliable delivery across production environments
Technical Task Design
-
Design challenging, domain-relevant problems involving cloud infrastructure, Kubernetes, and automation
-
Write accurate and well-structured solutions to infrastructure engineering tasks
-
Develop scenarios that test practical diagnostic and engineering judgement
-
Ensure tasks reflect realistic production environments and engineering constraints
-
Calibrate technical difficulty and expected solution quality
Evaluation & Technical Review
-
Evaluate infrastructure engineering tasks and proposed solutions
-
Assess technical work for correctness, completeness, and production readiness
-
Identify infrastructure, architectural, automation, and reasoning errors
-
Provide clear written technical feedback
-
Distinguish substantive engineering failures from minor implementation differences
Rubrics & Quality Standards
-
Develop detailed evaluation frameworks for infrastructure engineering tasks
-
Define criteria for cluster failure diagnosis, IaC design quality, and CI/CD reasoning
-
Establish clear standards for technically correct and production-appropriate solutions
-
Apply evaluation criteria consistently across complex engineering scenarios
-
Collaborate with other technical experts to maintain accuracy and consistency
Technical Knowledge Development
-
Help identify gaps in cloud infrastructure and DevOps reasoning
-
Contribute expert knowledge across Kubernetes operations, AWS services, Infrastructure-as-Code, and deployment engineering
-
Explain complex technical decisions clearly and precisely
-
Evaluate unfamiliar infrastructure scenarios using first-principles engineering judgement
-
Support improvement of technical training and evaluation materials
Ideal Profile
-
4+ years of dedicated professional experience in cloud infrastructure, DevOps, Site Reliability Engineering, platform engineering, or a closely related field
-
Experience as a DevOps Engineer, Cloud Engineer, Infrastructure Engineer, Platform Engineer, Site Reliability Engineer (SRE), or similar professional
-
Strong hands-on production experience operating Kubernetes
-
Demonstrated ability to diagnose and repair Kubernetes cluster failures, beyond simply authoring manifests or using managed control planes
-
Production experience with Terraform and/or AWS CDK
-
Direct production experience integrating AWS Lambda, API Gateway, and DynamoDB
-
Experience building and owning CI/CD pipelines
-
Strong understanding of cloud infrastructure, deployment automation, production reliability, and operational troubleshooting
-
Demonstrable professional progression within infrastructure, DevOps, SRE, or platform engineering
-
Strong written communication skills and ability to explain complex technical decisions clearly
-
Professional experience within recognised, technically rigorous organisations is preferred
Engagement Details
-
Full-time position
-
Fully remote within the United States
-
40 hours per week
-
Weekday availability required
-
Compensation: $65–$100/hour
-
Work focuses on Kubernetes operations, AWS infrastructure, Infrastructure-as-Code, CI/CD pipelines, and related technical evaluation
-
Collaboration with research, engineering, and other subject-matter experts
-
Opportunity to contribute to advanced AI infrastructure and technical evaluation workstreams
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.