Logo for Gigster

Senior Platform Engineer, Cloud Infrastructure

Role overview

Qualifications

  • 6+ years of professional experience in software, platform, infrastructure, or site reliability engineering
  • Demonstrated experience building and operating production Kubernetes platforms
  • Production experience writing Go, Python or Java
  • Degree in Computer Science, Engineering, or a related field, or equivalent practical experience

Responsibilities

  • Design, build, and operate production Kubernetes clusters
  • Write, refactor, and maintain production services, controllers, and middleware
  • Lead incident response for platform-level issues and troubleshoot production systems
  • Own infrastructure as code across the platform and build/improve CI/CD workflows

Key facts

Hard skills

Other skills

  • Analytical Skills
  • Problem Solving
  • Communication

About the company

Gigster logo

Gigster

Software Development

Gigster builds top-tier software development teams. Our AI-powered platform ensures tailor-fit talent matching, accelerated delivery, cost efficiency, and guaranteed project outcomes. Perfect-Fit Talent Matching: Receive tailor-fit talent matching with Gigster’s AI-powered platform based on skillset, past performance, and personality type collected from 10+ years of project data. Guaranteed Outcome: Guaranteed pricing and outcomes for any fully-managed project based on 5,000+ deliverables. Flexibility: Get a fully-managed team with a project manager to orchestrate end-to-end project execution for big projects or elicit on-demand talent to add to your existing team. Elite Talent: Get access to our global talent pool of 50,000+ rigorously-vetted developers, designers, and project managers. Fast Start & Confident Delivery: Proven delivery workflow model for accelerated start and streamlined development for maximum efficiency in execution and project delivery. Founded in 2014, Gigster has completed over 5,000 projects with some of the largest companies in the world. Third party research firm, Constellation Research, completed a study that showed Gigster’s model results in 30% more efficiency in staffing, 60% lower delivery risk, and a 3.6x higher customer satisfaction score than other software development firms.

Company details

Company typeStartup
IndustrySoftware Development
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Senior Platform Engineer, Cloud Infrastructure

Type: Remote
Coverage: Pacific Hours (8:00 AM – 5:00 PM PST)

Job Description:

We are looking for a senior engineer to build and operate the cloud-native platform that our products run on. This is a hands-on infrastructure engineering role: you will design and run production Kubernetes platforms, write and maintain the Go services and controllers that extend them, and own the reliability of systems that other engineering teams depend on.

The work spans platform architecture, networking, observability, and production operations. You will write real code: Go services, Kubernetes controllers, custom middleware, but the value you create is measured in platform capability and reliability, not lines shipped. We are looking for someone who is comfortable owning ambiguous, multi-quarter initiatives and driving them to production.

Key Responsibilities:

Platform and Kubernetes Engineering:

  • Design, build, and operate production Kubernetes clusters, including cluster networking, workload isolation, and multi-region topologies.

  • Work with Kubernetes internals: resource quota management, scheduling and cluster behaviour, NetworkPolicy enforcement, and custom controllers or operators.

  • Implement and operate service mesh capabilities: mTLS between services, service-account-level authentication and authorisation, and traffic management across internal and external request paths.

  • Optimise containerised workloads for performance, cost, and resource efficiency.

Software Development in Go, Python or Java:

  • Write, refactor, and maintain production services, controllers, and middleware that extend the platform.

  • Read and contribute to complex existing codebases, including open-source projects that we customise or extend.

  • Build HTTP, REST, and gRPC service interfaces used by internal engineering teams.

  • Write meaningful unit and integration tests, and treat testability as a design property rather than an afterthought.

Reliability and Production Operations:

  • Lead incident response for platform-level issues; investigate root causes and author postmortems that result in durable fixes.

  • Troubleshoot production systems using logs, metrics, traces, and profiling tools.

  • Diagnose and resolve performance and reliability problems across distributed systems.

  • Define and drive SLOs, and build the alerting that makes them actionable.

Infrastructure as Code and Delivery:

  • Own infrastructure as code across the platform, authoring reusable modules and maintaining them as the platform evolves.

  • Build and improve CI/CD and GitOps delivery workflows so that teams can ship safely and frequently.

  • Balance developer velocity against reliability, security, and compliance requirements.

  • Plan and execute cloud migration initiatives, including moving production workloads between cloud providers or environments while maintaining reliability and minimizing downtime.

Observability:

  • Build and maintain metrics, dashboards, alerting policies, and distributed tracing.

  • Instrument services so that failures are diagnosable without a code change.

Collaboration:

  • Partner with product, security, and infrastructure teams to gather requirements and align on architecture.

  • Contribute to design reviews and help set technical direction.

  • Mentor other engineers and raise the standard of engineering practice around you.

Qualifications:

Education and Experience:

  • 6+ years of professional experience in software, platform, infrastructure, or site reliability engineering, including significant time operating production distributed systems.

  • Demonstrated experience building and operating production Kubernetes platforms, not only deploying onto them.

  • Production experience writing Go, Python or Java.

  • Experience designing systems from an ambiguous starting point and carrying them to production.

  • Experience migrating cloud services, including planning and executing moves of production workloads between providers or environments.

  • Degree in Computer Science, Engineering, or a related field, or equivalent practical experience.

Technical Skills:

  • Strong understanding of Kubernetes internals: networking (CNI), NetworkPolicy, resource management, and cluster behaviour under load.

  • Hands-on experience with a service mesh (Istio, Envoy, Linkerd, or similar) and with mTLS and workload identity.

  • Solid Linux fundamentals, including cgroups and resource management.

  • Infrastructure as code at scale, Terraform or an equivalent.

  • Production experience with at least one major cloud platform (GCP, AWS, or Azure); multi-cloud experience is a strong plus.

  • Observability tooling: Prometheus, Grafana, OpenTelemetry, and query languages such as PromQL.

  • Production experience with relational databases, including PostgreSQL or managed Postgres-compatible services, and an understanding of their replication and failover characteristics.

  • Docker and container tooling as part of the delivery lifecycle.

  • Strong debugging and performance profiling skills.

Preferred:

  • Experience with Go testing frameworks such as Ginkgo and Gomega.

  • Experience building Kubernetes controllers, operators, or other API-server extensions.

  • Additional strength in Python.

  • Experience with identity and access management: SSO, Keycloak, OIDC, SAML, or secret management with Vault or a cloud equivalent.

  • Experience designing for high availability and disaster recovery across regions or providers.

  • Experience working in a monorepo, and with build systems such as Bazel.

  • Ability to read Java.

  • Exposure to compliance frameworks such as SOC 2 or GDPR.

  • Interest in the reliability and safety of LLM-backed systems running in cloud-native environments.

  • Experience with Alibaba Cloud (AliCloud) is highly preferred.

  • Experience with large-scale Data Platform technologies such as Apache Spark and Apache Flink is highly preferred.

Soft Skills:

  • Strong analytical and problem-solving ability.

  • Clear written and verbal communication, particularly in design documents, postmortems, and cross-team requirements gathering.

  • Able to work independently while contributing effectively within a distributed team.

  • Comfortable in a fast-moving, highly technical environment.

Please Note:

  • This is a platform engineering role. We are not looking for candidates whose experience has been limited to consuming Kubernetes or cloud services. The ideal candidate understands how the underlying platforms work, has built and operated production Kubernetes infrastructure, and can explain the architecture, implementation, and operational trade-offs behind the tools they use.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Platform Engineer Related jobs

Other jobs at Gigster

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.