Logo for ICU Medical

Senior Site Reliability Engineer, CloudOps

Role overview

Qualifications

  • 7+ years of hands-on experience in AWS Cloud Engineering, DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering
  • Bachelor’s degree from an accredited college or university
  • Practical background supporting production workloads in Linux/AWS environments
  • Prior experience in the healthcare industry maintaining HIPAA/HiTrust-compliant infrastructure

Responsibilities

  • Manage, maintain, and troubleshoot a multi-account AWS Organization environment and core services
  • Support production deployments, CI/CD pipelines, and infrastructure automation using Python, Bash, and AWS CloudFormation
  • Monitor system health and performance, investigate alerts, and perform root-cause analysis
  • Participate in a shared on-call rotation, managing incident response and failover/recovery validation for production applications

Key facts

  • Remote from: North Carolina (USA)
  • Full time
  • Senior (5-10 years)
  • Site Reliability Engineer (SRE)
  • English

Hard skills

Other skills

  • Analytical Thinking
  • Problem Solving
  • Collaboration

About the company

ICU Medical logo

ICU Medical

Medical Devices & Equipment

ICU Medical connects patients and caregivers through safe, life-saving, life-enhancing IV therapy systems, software, solutions, and consumables. Since IV therapy is our only business, meeting your IV needs with quality products and consistent supply is our only concern. We are 100% focused on bringing you intuitive, patient-centric IV products and services that provide meaningful clinical differentiation, consistent innovation, and superior value. We design our products to work within your existing workflows to minimize disruption and maximize the time you spend with patients. Together, we help forge the human and emotional connections that enhance clinical experience and are the essence of outstanding quality of care. For more than three decades, we have been dedicated to a singular purpose—improving the safety and efficiency of IV therapy. With the acquisition of Hospira Infusion Systems from Pfizer in 2017, we became the only company to focus exclusively on IV therapy across the continuum of care. Our focus allows us to bring you: > Dedicated and non-dedicated IV sets and needlefree connectors clinically proven to provide an effective barrier against bacterial transfer and colonization. > The industry’s broadest IV smart pump offering covering large volume, pain management, and ambulatory needs. > IV medication safety software providing full IV-EHR interoperability with the highest customer satisfaction and compatibility with more EHR systems than any other company. > Significant US IV solutions manufacturing and supply capabilities.

Company details

Company typeXLarge
IndustryMedical Devices & Equipment
Company size5001 - 10000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Position Summary 

We are seeking a Senior Site Reliability Engineer, CloudOps to support, scale, and optimize a multi-account AWS environment hosting healthcare-oriented applications and analytics platforms. In this role, you will bridge infrastructure engineering, operational reliability, production support, and cloud modernization initiatives across complex microservices architectures. The ideal candidate brings strong AWS expertise, solid Linux administration skills, and a proven track record of managing production systems in HIPAA/HiTrust regulated environments. You will participate in incident response, on-call rotations, and continuous deployment workflows while helping drive our transition toward containerized and Kubernetes-based platforms. This collaborative position is built for an analytical engineer who excels at resolving production incidents, partnering with developers, and continuously elevating operational excellence.

Essential Duties & Responsibilities

  • Manage, maintain, and troubleshoot a multi-account AWS Organization environment (35+ accounts) and core services, including EC2, ECS/Fargate, Lambda, S3, CloudFront, API Gateway, and Aurora/RDS databases.
  • Support production deployments, CI/CD pipelines (Jenkins, AWS CodePipeline), and infrastructure automation using Python, Bash, and AWS CloudFormation.
  • Monitor system health and performance using Datadog, CloudWatch, and Zabbix; investigate alerts, execute root-cause analysis, and refine monitoring coverage to reduce operational noise.
  • Participate in a shared on-call rotation, managing incident response and performing failover/recovery validation for production applications and data stores.
  • Maintain HIPAA/HiTrust compliance and security posture by managing tools like Prisma/Cortex Cloud, Security Hub, and GuardDuty, while enforcing proper IAM policies and network segmentation.
  • Support Java (Spring Boot) and Python applications running in containers, assisting developers during investigations and preparing for future Kubernetes (EKS) modernization initiatives.

Knowledge & Skills

  • Deep hands-on expertise with AWS core services (networking, compute, serverless, and database technologies) and CloudFormation IaC automation.
  • Strong Linux administration skills (primarily Ubuntu) along with proficiency in Python and Bash scripting for operational automation.
  • Experience with containerization technologies (Docker, ECS/Fargate) and familiarity with modern Kubernetes ecosystems (EKS, Helm, ArgoCD).
  • Solid understanding of observability tools (Datadog, CloudWatch, Zabbix) and CI/CD pipelines (Jenkins, CodePipeline, Git workflows).
  • Knowledge of cloud security best practices, access management (IAM), and compliance frameworks within regulated sectors (HIPAA/HiTrust).
  • Proven diagnostic, incident-management, and analytical troubleshooting skills for complex microservices architectures.

Minimum Qualifications, Education & Experience 

  • Must be at least 18 years of age.
  • High School Diploma required.
  • Bachelor’s degree from an accredited college or university is required.
  • 7+ years of hands-on experience in AWS Cloud Engineering, DevOps, Site Reliability Engineering (SRE), or Infrastructure Engineering.
  • Practical background supporting production workloads in Linux/AWS environments, reading application logs, and making minor code fixes.
  • Direct experience participating in on-call rotations and incident response protocols.
  • Prior experience in the healthcare industry maintaining HIPAA/HiTrust-compliant infrastructure.

Work Environment

  • This is largely a sedentary role. 
  • This job operates in a professional office environment and routinely uses standard office equipment.
  • Typically requires travel less than 5% of the time

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at ICU Medical

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.