Logo for Resilient Co.

Sr Platform/ Infrastructure Engineer

Role overview

Qualifications

  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles
  • Proven experience deploying and operating Kubernetes in production
  • Strong Python skills for automation, tooling, and operational scripts
  • Cloud experience with AWS and Azure (designing, deploying, and operating services)

Responsibilities

  • Design, deploy, and maintain production Kubernetes clusters and related services
  • Build and maintain automation and tooling using Python to support platform operations
  • Integrate and operate Prometheus for monitoring, alerting, and observability
  • Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers

About the company

Resilient Co. logo

Resilient Co.

Staffing & Recruiting

Unknown

Company details

IndustryStaffing & Recruiting
Company sizeUnknown

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

We are seeking a senior Sr Platform/Infrastructure Engineer to strengthen our platform team and drive cloud-native infrastructure initiatives. This role focuses on deploying and maintaining Kubernetes services, integrating monitoring and storage platforms, and troubleshooting distributed systems to ensure resilient, scalable operations.

You will work with Python-driven tooling, Prometheus-based monitoring, Ceph-backed storage, and public cloud environments (AWS and Azure) to modernize and operate our platform. This is an opportunity to shape platform reliability and performance in a hands-on engineering role.

Responsibilities

  • Design, deploy, and maintain production Kubernetes clusters and related services.
  • Build and maintain automation and tooling using Python to support platform operations.
  • Integrate and operate Prometheus for monitoring, alerting, and observability.
  • Deploy and manage Ceph storage solutions for distributed workloads.
  • Support platform modernization initiatives and migrate services to cloud-native patterns.
  • Troubleshoot and resolve issues in distributed systems across compute, storage, and network layers.
  • Collaborate with development, SRE, and operations teams to define platform requirements and SLAs.
  • Document platform designs, runbooks, and operational procedures.
  • Participate in on-call rotations and incident response to maintain platform availability.

Requirements

  • 5+ years of experience in platform, infrastructure, or site reliability engineering roles.
  • Proven experience deploying and operating Kubernetes in production.
  • Strong Python skills for automation, tooling, and operational scripts.
  • Experience implementing and operating Prometheus-based monitoring and alerting.
  • Hands-on experience with Ceph or similar distributed storage systems.
  • Cloud experience with AWS and Azure (designing, deploying, and operating services).
  • Demonstrated ability to troubleshoot distributed systems and resolve production incidents.
  • Experience collaborating across teams to deliver platform improvements and migrations.

Nice to Have

  • Experience with OpenSearch.
  • Proficiency with Bash scripting.
  • Familiarity with Java-based services.
  • Experience with Fluent Bit for log collection.
  • Experience working with PostgreSQL.

Engagement & Logistics

  • Engagement Length: 12 months or more.
  • Time Zone: PST - 8:00 AM - 5:00 PM
  • Holiday Calendar: Client Holidays (USA – Mandatory)
  • Laptop: BYOD.
  • Overtime Required: No.



Selection process

  1. Meeting with Resilient Co. team.
  2. Technical interview
  3. Client (2 interviews - Manager + Technical panel)

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Infrastructure Engineer Related jobs

Other jobs at Resilient Co.

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.