Logo for Sherweb

Site Reliability Specialist, IT Operations

Role overview

Qualifications

  • College or university degree in computer science, software development, information technology, or engineering
  • 3 to 5 years of experience in systems administration, IT operations, or similar technical role
  • Strong scripting or development skills with at least one language such as PowerShell, Python, or JavaScript

Responsibilities

  • Develop, maintain, and improve scripts, automation workflows, and operational tooling
  • Apply SRE principles to improve the reliability, availability, performance, and resilience of platforms
  • Provide advanced operational support and resolve incidents affecting production services
  • Improve monitoring, alerting, logging, metrics, and operational visibility

About the company

Sherweb logo

Sherweb

IT Infrastructure & Managed Services

Sherweb is more than a cloud marketplace. We go beyond distribution with value-added services that help IT professionals offload technical operations and extend their cloud expertise. From presales to helpdesk, you can count on our team of passionate experts to champion your growth.

Company details

Company typeLarge
IndustryIT Infrastructure & Managed Services
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Location : Work from Home - Province of Quebec 

The Site Reliability Specialist on the IT Operations team contributes to the reliability, availability, performance, and resilience of Sherweb's platforms and services.

This is a highly technical individual contributor role that applies Site Reliability Engineering (SRE) principles to production environments. The role combines systems administration, software development, automation, observability, and operational excellence to improve service reliability, reduce operational toil, and increase platform scalability.

Working closely with Infrastructure, Development, DevOps, Platform, Security, and Product teams, the Site Reliability Specialist helps ensure production systems remain stable, supportable, and continuously improving through engineering and automation practices.

 Here's how you will contribute to the success of the company

·        Develop, maintain, and improve scripts, automation workflows, and operational tooling to reduce manual intervention, improve reliability, and lower operational toil.

·        Apply SRE principles to improve the reliability, availability, performance, and resilience of Sherweb’s hosted platforms and production services.

·        Implement and support reliability standards, service level objectives (SLOs), service level indicators (SLIs), and operational practices established for platforms and services.

·        Use a developer mindset to transform repetitive operational tasks into scalable, reusable, documented, and supportable automation.

·        Provide advanced operational support and resolve incidents affecting production services while collaborating closely with SRE, Infrastructure, Development, DevOps, and Product teams.

·        Investigate recurring issues and perform root cause analysis to identify short-term corrective actions and long-term reliability improvements.

·        Build, support, and maintain production systems and hosted service technologies while following operational procedures, security best practices, and compliance requirements.

·        Improve monitoring, alerting, logging, metrics, and operational visibility to help the team detect issues earlier, understand system behavior, and prevent incidents.

·        Contribute to improving end-to-end observability and system understanding through metrics, logs, traces, telemetry, and operational diagnostics.

·        Contribute to observability-as-code, infrastructure-as-code, configuration-as-code, and automation practices where applicable.

·        Explore and leverage Azure AI Foundry, Power Automate, and AI agent capabilities to improve operational efficiency, automate repetitive workflows, and accelerate incident response or service reliability improvements.

·        Participate in platform lifecycle activities, deployments, maintenance windows, migrations, and continuous service improvement initiatives.

·        Collaborate with developers, architects, subject matter experts, DevOps, and infrastructure teams to support implementation, optimization, troubleshooting, and operational readiness of services.

·        Create and maintain operational documentation, including SOPs, runbooks, troubleshooting guides, automation documentation, maintenance procedures, and knowledge-sharing materials.

·        Track, organize, and manage incidents, requests, and service tickets to respect SLAs and ensure clear communication through the ITSM process.

·        Participate in rotational on-call duty and perform maintenance work outside normal business hours when required.

·        Carry out all other related tasks per the job’s evolution and departmental needs.

 

Here's what you need to have and master to get the job

Education

·        College or university degree in computer science, software development, information technology, engineering, or a combination of equivalent training and experience.

 Experience

·        3 to 5 years of experience in systems administration, IT operations, infrastructure support, software development, DevOps, automation, or a similar technical role.

·        Experience supporting production systems in business-critical and customer-facing environments.

·        Proven experience improving operational efficiency through automation and engineering practices.

 Core Skills

·        Strong scripting or development skills with at least one language such as PowerShell, Python, Bash, JavaScript, TypeScript, or C#.

·        Proven ability to design, write, test, troubleshoot, document, and maintain scripts or small applications used to automate operational tasks.

·        Proven experience supporting Microsoft and/or Linux server environments, including troubleshooting, maintenance, operational support, and automation.

·        Strong diagnostic, investigation, and problem-solving skills with the ability to analyze incidents, identify root causes, and implement sustainable improvements through automation or engineering practices.

·        Good understanding of distributed systems, networking concepts, system dependencies, availability, performance, reliability, and service operations in production environments.

·        Experience with monitoring, alerting, observability, log management, telemetry and operational data analysis to detect, troubleshoot, and prevent issues.

·        Familiarity with version control, Git-based workflows, CI/CD pipelines, code review practices, and deployment automation.

·        Experience with infrastructure as code, configuration management, or automation tools such as Terraform, Ansible, DSC, Azure DevOps, GitHub Actions, Docker, or Kubernetes is an asset.

·        Familiarity with Azure AI Foundry, Power Automate, Copilot/AI agents, or agent-based automation concepts is an asset.

·        Knowledge of high availability environments, virtualization, cloud services, backup and restore practices, and production support models is an asset.

 Professional Attributes

·        Autonomous, reliable, and motivated, with a continuous learning mindset and a strong interest in improving reliability through software engineering and automation practices.

·        Strong communication and collaboration skills, with the ability to work effectively with technical and non-technical stakeholders.

·        Excellent English skills, both spoken and written, are essential; fluency in French is an asset.

·        Relevant industry certifications such as Microsoft Azure, Red Hat, Linux Foundation, Kubernetes, DevOps, or observability platforms are considered assets.

 Additional Requirements

*Availability for rotation on the on-call schedule in a 24/7 environment.

 

Benefits of working at Sherweb 

Sherweb is first and foremost a culture where our customers’ needs are at the heart of everything we do, supported by Sherwebers committed to living our values of passion, teamwork, and integrity.  

Dynamic and flexible work environment   

  • Beautiful offices designed for collaboration (Sherbrooke and downtown Montreal)  
  • Access to cutting-edge tools and technologies  
  • A results-driven culture where ideas, actions, and expertise are recognized  
  • Supportive colleagues from diverse professional and cultural backgrounds  
  • A generous employee referral program  

Flexible total compensation package   

  • Base salary between $71,330.00 and $101,900.00 per year  
  • Vacation allowance based on prior experience  
  • Paid time off starting on day one (vacation and personal days)  
  • Employee benefits program with savings at local merchants  

Unlimited growth opportunities   

  • Close relationship with your direct manager and open communication to support your development  
  • Multiple learning and development opportunities  
  • Tools to track your progress and support career growth  

A vibrant social life (Sherweblife)  

A rich calendar of virtual and in-person activities designed to foster connection throughout the year.  

At Sherweb, we believe in transparency and pay equity. The salary range provided is intended to give an indication of what you can expect for this role. However, we recognize that each candidate brings a unique set of skills and experiences. The final compensation package will be tailored to reflect the selected candidate’s qualifications and expertise, ensuring we remain competitive and fair in our offers.

Reasons for the requirement of English: Sherweb has international customers and fluency in English is the only way to ensure proper service delivery to them. The main tasks related to this position require written and oral communication with an English-speaking clientele at all times.

#LI-Remote

#LI-SG1

 

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Sherweb

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.