Logo for 24-MAG

Remote | Kubernetes Engineer — $60–$80/hour

Role overview

Qualifications

  • 3+ years of hands-on production Kubernetes experience
  • Experience operating EKS, GKE, AKS, or self-managed Kubernetes clusters
  • Deep understanding of CNI, DNS, ingress, PV/PVC storage, RBAC, scheduling, and cluster failure modes
  • Strong experience authoring and reviewing Kubernetes manifests and Helm charts

Responsibilities

  • Review production-oriented Kubernetes scenarios across managed or self-hosted environments
  • Assess cluster configuration, workload behaviour, and operational readiness
  • Evaluate whether proposed approaches reflect sound Kubernetes practices
  • Identify configuration errors, unsafe assumptions, or production risks

Key facts

Hard skills

Other skills

  • Communication

About the company

24-MAG logo

24-MAG

Business Consulting & Services

Company details

IndustryBusiness Consulting & Services
Company size2 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

We are sharing a specialised part-time consulting opportunity for experienced Kubernetes Engineers with strong hands-on expertise in production cluster operations, infrastructure troubleshooting, manifests, Helm, networking, storage, and platform reliability.

This role focuses on evaluating Kubernetes engineering tasks for technical correctness, production readiness, and operational realism. Selected engineers will review cluster scenarios, manifests, troubleshooting workflows, and failure modes while providing clear, rubric-based technical feedback.

Key Responsibilities

Kubernetes Cluster Operations

  • Review production-oriented Kubernetes scenarios across managed or self-hosted environments
  • Assess cluster configuration, workload behaviour, and operational readiness
  • Evaluate whether proposed approaches reflect sound Kubernetes practices
  • Identify configuration errors, unsafe assumptions, or production risks
  • Apply practical judgement grounded in real-world cluster operations

Cluster Internals & Networking

  • Evaluate Kubernetes networking and service-discovery configurations
  • Review CNI, DNS, ingress, and service connectivity
  • Diagnose connectivity and routing issues across workloads
  • Identify incorrect networking assumptions or configuration problems
  • Assess whether proposed fixes address the underlying cluster issue

Storage & Persistence

  • Review PersistentVolume and PersistentVolumeClaim (PV/PVC) configurations
  • Assess storage classes, volume lifecycle behaviour, and workload dependencies
  • Identify provisioning, mounting, or persistence issues
  • Evaluate storage-related failure scenarios
  • Review proposed remediation for technical correctness

RBAC & Access Control

  • Evaluate Kubernetes RBAC configurations
  • Review roles, cluster roles, bindings, and service-account permissions
  • Identify missing or excessive access
  • Assess whether permissions appropriately support workload requirements
  • Apply least-privilege principles when evaluating access-control decisions

Failure-Mode Troubleshooting

  • Diagnose common Kubernetes failures such as CrashLoopBackOff, OOMKilled, scheduling failures, and pod eviction
  • Review logs, events, workload configuration, and resource behaviour
  • Evaluate troubleshooting sequences for efficiency and technical accuracy
  • Identify root causes rather than superficial symptoms
  • Assess whether proposed corrective actions are production appropriate

Manifests & Helm

  • Author and review Kubernetes manifests
  • Evaluate Helm charts for correctness, maintainability, and deployment safety
  • Review resource specifications, probes, requests, limits, selectors, and dependencies
  • Identify templating or configuration problems
  • Assess whether manifests accurately represent the intended workload

Production Incident Review

  • Evaluate live-cluster troubleshooting scenarios
  • Review diagnostic reasoning and incident-response decisions
  • Assess prioritisation, escalation, and remediation approaches
  • Identify missed evidence or ineffective debugging paths
  • Determine whether proposed solutions reduce recurrence risk

Platform Reliability

  • Evaluate scaling, availability, resilience, and operational-readiness considerations
  • Review workloads involving autoscaling or service-mesh technologies where relevant
  • Assess observability and operational visibility
  • Identify reliability weaknesses in workload or cluster design
  • Evaluate recommendations for production readiness

Observability & Performance

  • Review monitoring and troubleshooting workflows using tools such as Prometheus and Grafana
  • Evaluate alerting, metrics, logs, and operational signals
  • Assess HPA configurations and scaling behaviour where relevant
  • Review service-mesh scenarios involving technologies such as Istio
  • Identify gaps affecting diagnosis or performance management

Rubric-Based Technical Evaluation

  • Assess Kubernetes tasks against structured technical criteria
  • Provide clear written explanations supporting evaluation decisions
  • Reference specific configuration, runtime behaviour, or diagnostic evidence
  • Apply grading standards consistently across assignments
  • Distinguish valid alternative approaches from genuinely incorrect solutions

Ideal Profile

  • 3+ years of hands-on production Kubernetes experience
  • Experience operating EKS, GKE, AKS, or self-managed Kubernetes clusters
  • Deep understanding of CNI, DNS, ingress, PV/PVC storage, RBAC, scheduling, and cluster failure modes
  • Strong experience authoring and reviewing Kubernetes manifests and Helm charts
  • Demonstrated experience debugging live production cluster incidents
  • Proficiency in Go, Python, or TypeScript
  • Strong understanding of workload resource management and production reliability
  • CKA or CKAD certification is preferred
  • Experience with Istio, HPA, Prometheus, Grafana, or comparable technologies is advantageous
  • Previous SRE, platform engineering, code-review, or technical task-grading experience is advantageous
  • Strong written communication and ability to provide precise technical feedback

Engagement Details

  • Part-time independent contractor engagement
  • Fully remote within the United States
  • Flexible scheduling based on project requirements
  • Compensation: $60–$80/hour
  • Work focuses on Kubernetes cluster operations, manifests, troubleshooting, reliability, and production-readiness evaluation
  • Projects may be extended, shortened, or concluded based on project needs and performance
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
  • H1-B and STEM OPT support is unavailable for this engagement

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at 24-MAG

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.