Logo for HavocAI

Head of Cloud Operations

Role overview

Qualifications

  • 8+ years of relevant experience across change management, release management, incident management, technical program management, service management, SRE, DevOps, platform engineering, or related disciplines.
  • Demonstrated experience leading or managing SRE, DevOps, or Platform Engineering teams.
  • Strong understanding of modern cloud operations, software delivery, infrastructure automation, and production reliability.
  • Ability to create scalable processes that provide appropriate control without unnecessarily slowing engineering teams.

Responsibilities

  • Lead and manage the SRE and DevOps teams, setting technical and operational direction across reliability, automation, and safe delivery.
  • Define and own change classes with clear approval paths and requirements.
  • Own HavocAI’s incident management framework, including incident declaration and escalation paths.
  • Maintain the authoritative record of deployed systems, including versions and dependencies.

Key facts

  • Remote from: United States
  • Full time
  • Senior (5-10 years)
  • English

Hard skills

Other skills

  • Program Management
  • Technical Acumen
  • Leadership
  • Communication
  • Decision Making
  • Problem Solving
  • Team Management

About the company

HavocAI logo

HavocAI

Defense Technology

HavocAI is revolutionizing maritime autonomy. Founded in 2024, HavocAI is bringing scalable maritime autonomy solutions and our ultra-low cost, high-rate production ASVs (Autonomous Surface Vessels) to the defense and commercial markets at the speed of relevance. HavocAI is poised to field thousands of autonomous multi-mission maritime assets in domains of pressing need, providing genuinely affordable mass efficiently and expeditiously.

Company details

IndustryDefense Technology
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Us:

Havoc is a leader in all-domain collaborative autonomy. Its software-defined hardware approach powers military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together in complex and contested environments. Havoc connects assets, enabling them to share information, adapt in real time, and continue operating even when communications are disrupted or denied. Havoc optimizes mission performance and minimizes human risk.

Havoc was founded in 2024 and headquartered in Providence, Rhode Island. Learn more at Havoc: All-Domain Collaborative Autonomy .

About the Role

HavocAI is seeking a Head of Cloud Operations to own the change, release, incident, and reliability practices that keep our systems dependable, auditable, and compliant as we scale.

Reporting to the Director of Cloud and partnering closely with the ISSO and engineering teams, you will define how production changes are approved and deployed, how releases are coordinated, how incidents are managed, and how we maintain a trustworthy record of what is running across our environments.

You will also lead our SRE and DevOps teams, setting direction across reliability, infrastructure automation, CI/CD, observability, and safe delivery. This role requires someone who can build disciplined processes without creating unnecessary bureaucracy—using automation and engineering practices wherever possible to make the right way of working the easiest way of working.

The ideal candidate combines strong operational leadership with enough technical depth to challenge assumptions, make decisions under pressure, and translate security and compliance requirements into practical engineering processes.

What You’ll Do

SRE & DevOps Leadership
  • Lead and manage the SRE and DevOps teams, setting technical and operational direction across reliability, automation, and safe delivery.

  • Hire, coach, develop, and manage performance for engineers across both functions.

  • Own reliability and delivery practices including SLIs, SLOs, error budgets, on-call health, CI/CD, and infrastructure automation.

  • Establish clear ownership and operating expectations across cloud reliability and delivery.

  • Partner with engineering leaders to identify systemic reliability risks and prioritize improvements.

  • Build an engineering culture that balances speed, reliability, security, and operational discipline.

Change & Release Management
  • Define and own change classes—including standard, normal, and emergency changes—with clear approval paths and requirements.

  • Establish and operate an appropriate change approval process, including impact assessments, rollback plans, and approval records.

  • Integrate change management with GitOps workflows, using merged, signed, peer-reviewed pull requests as the foundation of the change record.

  • Own maintenance windows, freeze periods, and emergency-change processes, including retroactive approvals where appropriate.

  • Ensure production changes are traceable to an approved request, approver, and rollback decision.

  • Own the release calendar, versioning strategy, and promotion across environments and tenants.

  • Establish pre-deployment verification requirements covering CI status, security scans, migrations, feature flags, and predefined rollback triggers.

  • Coordinate releases across Cloud Platform, Backend, Autonomy, and Frontend teams to prevent conflicts and manage dependencies.

  • Maintain complete, audit-ready deployment and release records.

Incident & Problem Management
  • Own HavocAI’s incident management framework, including incident declaration, severity levels, escalation paths, and incident command.

  • Establish clear authority and expectations for declaring and managing incidents.

  • Run incident command during significant events, coordinating roles, communications, escalation, and stakeholder or customer notifications.

  • Own on-call health, alert quality, escalation practices, and operational readiness.

  • Lead blameless post-incident reviews and ensure remediation actions are assigned, tracked, and completed.

  • Manage government-sponsor notification obligations and timelines for incidents affecting authorized systems, with company-wide scope beyond the IATT boundary.

  • Establish clear distinctions between incidents and problems and drive analysis of recurring issues.

  • Translate recurring operational issues into technical debt, reliability, and remediation priorities.

Configuration, Baselines & Compliance
  • Maintain the authoritative record of deployed systems, including versions, digests, and dependencies.

  • Keep deployment records synchronized with the ISSO’s system inventory.

  • Establish and maintain system baselines and processes for detecting configuration drift.

  • Produce audit-ready evidence for the ISSO, including change records, deployment logs, incident reports, and post-incident remediation actions.

  • Own execution tracking against POA&M commitments, partnering with the ISSO to ensure remediation dates and engineering commitments are met.

  • Translate security and compliance control language into practical engineering processes and clearly communicate engineering implementation back to security stakeholders.

  • Build automation wherever possible to reduce manual compliance work and improve the reliability of operational evidence.

What We’re Looking For

  • 8+ years of relevant experience across change management, release management, incident management, technical program management, service management, SRE, DevOps, platform engineering, or related disciplines.

  • Demonstrated experience leading or managing SRE, DevOps, or Platform Engineering teams, including hiring, coaching, and performance management.

  • Strong understanding of modern cloud operations, software delivery, infrastructure automation, and production reliability.

  • Experience establishing and operating change, release, and incident management processes in complex technical environments.

  • Proven ability to coordinate complex initiatives across engineering teams and stakeholders you do not directly manage.

  • Technical fluency sufficient to evaluate and challenge engineering impact assessments, deployment strategies, rollback plans, and root-cause analyses.

  • Ability to remain calm, decisive, and directive during active incidents and make sound decisions under pressure.

  • Strong written communication skills, particularly for incident communications, post-incident reports, operational documentation, and executive updates.

  • Ability to create scalable processes that provide appropriate control without unnecessarily slowing engineering teams.

  • Strong ownership, judgment, and comfort operating in a fast-moving and ambiguous environment.

  • Must be a U.S. Citizen and able to obtain and maintain a U.S. Government security clearance.

Nice to Have

  • Prior incident command experience in a regulated, defense, government, or safety-relevant environment.

  • Experience operating cloud systems subject to U.S. Government authorization or compliance requirements.

  • Familiarity with POA&Ms, security authorization processes, configuration baselines, and audit evidence management.

  • Experience implementing GitOps-based change and release processes.

  • Knowledge of ITIL practices or equivalent hands-on experience developing effective service management processes.

  • Experience with tools such as Jira, PagerDuty, status pages, runbook platforms, and incident management systems.

  • Experience with Kubernetes, infrastructure as code, CI/CD platforms, observability systems, and modern cloud infrastructure.

What Success Looks Like

Within your first 12 months, you will have:

  • Established an enforced, practical change and release management process that engineering teams consistently follow.

  • Ensured every production change is traceable to the appropriate approval, deployment record, and rollback decision.

  • Created a consistent incident management framework with clear severity levels, ownership, command structures, escalation paths, and communication standards.

  • Established effective post-incident practices with remediation actions tracked through completion.

  • Improved the health and effectiveness of SRE, DevOps, on-call, and reliability practices.

  • Created an automated and trustworthy record of what is deployed across authorized environments.

  • Kept deployment records synchronized with the ISSO’s inventory and maintained audit-ready operational evidence.

  • Established effective tracking and execution against POA&M remediation commitments.

  • Built operational processes that strengthen reliability and compliance without introducing unnecessary friction for engineering teams.

Benefits:

  • 100% Employer paid Health, Dental and Vision Insurance for you and your families

  • Life Insurance (Employer Paid)

  • Ability to participate in the companies 401k program (Matching)

  • Unlimited PTO policy with an enforced 2 week minimum

  • Equity Package

  • Work / Home Office Stipend

  • Global Entry

  • 16 Week Paid Parental Leave

  • Monthly Health and Wellness Stipend


Our Values:

  • Innovation: We are driven to break new ground. Every day presents an opportunity to challenge the status quo, think boldly, and deliver advanced solutions that transform the future of defense technology.

  • Integrity: We hold ourselves to the highest ethical standards, ensuring transparency, accountability, and trust in all our actions and partnerships.

  • Mission-Driven: We are focused on achieving impactful outcomes that align with our core mission—protecting lives through innovation.

  • Forward-Leaning: We continuously seek out new opportunities and remain at the forefront of technological advancements. We embrace change and anticipate the challenges of tomorrow with confidence and creativity.

  • Ownership of All Tasks: At HavocAI, no problem is too complex or too trivial. We believe that greatness comes from tackling the hardest challenges, but also in handling the smallest, sometimes thankless, tasks with the same level of commitment and care.

  • Servant Leadership: We lead by serving others, whether it’s supporting our employees, partners, or the broader community. Empowering those around us is key to achieving long-term success and making a lasting impact.

HavocAI is an Equal Opportunity Employer and is committed to creating an inclusive and diverse workplace. We welcome applicants from all backgrounds and do not discriminate based on race, color, religion, gender, sexual orientation, age, national origin, disability, veteran status, or any other legally protected status.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at HavocAI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.