Senior Software Engineer, DevOps (Remote – U.S.)
About the Opportunity
We are seeking an experienced Senior Software Engineer, DevOps to help design, build, and support modern cloud infrastructure and platform capabilities that power highly scalable applications. This role partners closely with software engineering teams to deliver secure, reliable, and automated systems while driving improvements in cloud infrastructure, CI/CD, observability, security, and operational excellence.
The ideal candidate is a hands-on technical leader with deep DevOps expertise, strong problem-solving skills, and a passion for building scalable platforms that enable engineering teams to move quickly and safely.
Key Responsibilities
- Design, implement, and maintain secure, scalable cloud infrastructure solutions.
- Develop and manage Infrastructure as Code (IaC) to ensure consistent, repeatable, and auditable deployments.
- Enhance CI/CD pipelines and infrastructure automation to improve software delivery and operational efficiency.
- Monitor, optimize, and support cloud infrastructure with a focus on performance, availability, reliability, and cost management.
- Establish and improve observability practices, including monitoring, logging, alerting, and incident response.
- Lead root cause analysis efforts and drive continuous service improvements.
- Implement DevSecOps best practices, including identity and access management, secrets management, infrastructure hardening, and vulnerability remediation.
- Troubleshoot complex platform and infrastructure issues across development and production environments.
- Identify and address technical debt while leading infrastructure modernization initiatives.
- Collaborate with engineering, security, and operations teams to align platform architecture and operational goals.
- Create and maintain technical documentation, standards, and operational playbooks.
- Mentor engineers and promote best practices in cloud technologies, automation, infrastructure, and DevOps methodologies.
- Participate in on-call rotations and support production incidents as needed.
Required Qualifications
- 5+ years of experience in DevOps, Platform Engineering, Cloud Engineering, Site Reliability Engineering (SRE), or Software Engineering.
- Proven experience designing, implementing, and supporting cloud-native infrastructure solutions.
- Strong experience with Infrastructure as Code (IaC) tools and practices.
- Demonstrated success improving platform reliability, standardization, automation, and operational efficiency.
- Experience supporting highly available production environments in public cloud platforms.
- Strong background in incident management, root cause analysis, and service reliability improvements.
- Experience implementing security best practices across cloud environments and infrastructure.
- Proficiency with Linux systems, command-line tools, and Git-based source control workflows.
- Experience developing scripts and automation solutions using Python, Bash, or similar scripting languages.
- Strong communication skills with the ability to work cross-functionally and mentor technical teams.
Preferred Qualifications
- Bachelor's degree in Computer Science, Software Engineering, or a related field, or equivalent practical experience.
- Experience with modern cloud technologies, containerization, orchestration, observability platforms, and CI/CD tooling.
- Experience establishing engineering standards, operational best practices, and platform governance.
Why Join Us?
This is an opportunity to work on large-scale cloud infrastructure challenges, influence platform strategy, and contribute to a high-performing engineering organization focused on innovation, reliability, security, and continuous improvement.
#LI-RS1
#LI-TL1