Logo for Expedite Technology Solutions LLC

Senior Manager SRE

Role overview

Qualifications

  • 10+ years in engineering, operations, or SRE roles
  • 5+ years leading SRE, platform, or reliability-focused teams
  • Proven experience implementing SRE practices at scale (SLIs, SLOs, error budgets)
  • Strong background in cloud environments (AWS, Azure, GCP)

Responsibilities

  • Drive adoption of the SRE operating model across application teams
  • Define and enforce SLIs, SLOs, and Error Budgets
  • Partner with application product teams and engineering operations
  • Lead adoption of centralized observability standards

Key facts

  • Remote from: Anywhere
  • Full time
  • Senior (5-10 years)
  • English

Other skills

  • Strategic Thinking
  • Communication
  • Team Effectiveness

About the company

Expedite Technology Solutions LLC logo

Expedite Technology Solutions LLC

IT Services & IT Consulting

Expedite Technology Solutions has an extensive background in technology outsourcing, professional services, and staffing. We offer the expertise of a global company combined with ‘one to one’ customized service of a local business partner. Whether you want an implementation specialist for a short-term project or a software developer to add to your team, Expedite will work with you to provide the top-notch IT professionals you need. We can help you identify the right technical resources and manage the hiring process, or provide you with a complete project solution.

Company details

Company typeSME
IndustryIT Services & IT Consulting
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Key Responsibilities
SRE Activation & Operating Model
∙ Drive adoption of the SRE operating model across application teams
∙ Establish clarity in roles between:
o SRE
o Production Support Engineering (PSE)
o Application teams
∙ Ensure SRE practices are embedded into the development lifecycle, not treated as post-production activities

Reliability Standards & Governance
∙ Define and enforce:
o SLIs, SLOs, and Error Budgets
o Production readiness criteria
o Reliability best practices
∙ Lead SLO adoption and compliance reviews across the organization
∙ Establish governance frameworks to ensure consistent application of standards

Cross-Team Coordination & Enablement
∙ Partner with:
o Application product teams
o Production Support Engineering (MG team)
o Platform / Infrastructure / Observability teams
∙ Drive alignment and reduce friction between engineering and operations
∙ Ensure clear handoffs, escalation models, and operational ownership

Observability & Monitoring Strategy
∙ Lead adoption of centralized observability standards across:
o Metrics
o Logging
o Tracing
∙ Align tooling (AppDynamics, Splunk, Prometheus, etc.)
∙ Ensure monitoring and alerting are SLO-driven and actionable, not noise-based

Incident Management & Continuous Improvement
∙ Partner with PSE to strengthen:
o Incident management processes
o RCA (Root Cause Analysis) standards
∙ Drive identification of patterns and systemic issues
∙ Ensure learnings translate into engineering improvements and automation

Automation & Reliability Engineering
∙ Identify opportunities to:
o Reduce manual operational work
o Improve system resilience
o Enable self-healing capabilities
∙ Promote a culture of engineering over reaction

Reporting & Organizational Insight
∙ Define and track reliability metrics across FS&I
∙ Build reporting that provides visibility into:
o System health
o Incident trends
o SLO performance
∙ Translate technical data into actionable business insights

Required Qualifications
∙ 10+ years in engineering, operations, or SRE roles
∙ 5+ years leading SRE, platform, or reliability-focused teams
∙ Proven experience implementing SRE practices at scale (SLIs, SLOs, error budgets)
∙ Strong background in cloud environments (AWS, Azure, GCP)
∙ Hands-on experience with observability tools (Splunk, AppDynamics, Prometheus, etc.)
∙ Experience in incident management and production operations at scale
∙ Ability to operate effectively in high-pressure and complex enterprise environments

Preferred Qualifications
∙ Experience driving organizational transformation (not just technical implementation)
∙ Strong understanding of CI/CD, DevOps, and automation practices
∙ Experience working in regulated or large enterprise environments
∙ Familiarity with AIOps or advanced automation strategies

Key Success Indicators
∙ Increased adoption of SLOs and reliability standards
∙ Reduction in high-severity incidents over time
∙ Improved MTTR and operational efficiency
∙ Increased adoption of standardized observability practices
∙ Reduction in reactive, ticket-driven work across teams
∙ Clear alignment between SRE, PSE, and application teams

Core Competencies
∙ Strategic thinking with strong execution focus
∙ Ability to drive alignment across multiple teams and stakeholders
∙ Strong communication and influence skills
∙ Bias toward structure, clarity, and accountability
∙ Ability to operate with urgency and discipline in complex environments

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Expedite Technology Solutions LLC

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.