Logo for Coderio Software Company

Ssr Monitoring and Observability Analyst

Role overview

Qualifications

  • 3+ years in Monitoring, IT Operations, SRE, or Systems Administration roles.
  • Advanced expertise with platforms such as Prometheus, Grafana, ELK Stack, New Relic, or Datadog.
  • Hands-on experience monitoring Cloud environments (AWS, Azure, or GCP) and containerized workloads (Docker, Kubernetes).
  • Strong Linux administration skills combined with deep root-cause analysis (RCA) and event correlation capabilities.

Responsibilities

  • Help define and execute the company’s observability strategy following SRE and DevOps best practices.
  • Configure alert thresholds based on real business impact while actively working to reduce notification noise.
  • Develop and maintain intuitive, real-time dashboards (e.g., Grafana, Kibana) for operational visibility across multiple teams.
  • Implement monitoring automation from agent deployment to automated incident response.

About the company

Coderio Software Company logo

Coderio Software Company

Software Development

Company details

IndustrySoftware Development

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Coderio

Coderio designs and delivers scalable digital solutions for global businesses. With a strong technical foundation and a product-driven mindset, our engineering teams lead complex software projects from architecture to execution. We value autonomy, clear communication, and technical excellence.

🌍 Learn more: coderio.com

Role Overview

As an Observability & Monitoring Analyst, you will design, implement, and maintain proactive monitoring and alerting systems to ensure high availability, performance, and health across our global clients' IT infrastructure, applications, and services.

You will build end-to-end monitoring solutions utilizing metrics, logs, and distributed traces, establishing business-impact-driven alert thresholds (SLIs/SLOs), and supporting incident resolution through detailed Root Cause Analysis (RCA). Working closely with Operations and DevOps teams, your goal will be to minimize MTTR (Mean Time to Recovery) and drive continuous ecosystem reliability.

Key Responsibilities

  • Strategy & Design: Help define and execute the company’s observability strategy following SRE and DevOps best practices. Design robust, end-to-end monitoring architectures.

  • SLO/SLI Management: Configure alert thresholds based on real business impact while actively working to reduce notification noise.

  • Visualization: Develop and maintain intuitive, real-time dashboards (e.g., Grafana, Kibana) for operational visibility across multiple teams.

  • Automation & AIOps: Implement monitoring automation from agent deployment to automated incident response (basic/intermediate AIOps).

  • Platform Administration: Administer, patch, and maintain monitoring platforms while optimizing resource and infrastructure costs.

  • Documentation & Operations: Author and maintain service maps, monitoring runbooks, and troubleshooting procedures to streamline incident response.

What We Are Looking For

  • Experience: 3+ years in Monitoring, IT Operations, SRE, or Systems Administration roles.

  • Observability Tooling: Advanced expertise with platforms such as Prometheus, Grafana, ELK Stack, New Relic, or Datadog.

  • Cloud & Containers: Hands-on experience monitoring Cloud environments (AWS, Azure, or GCP) and containerized workloads (Docker, Kubernetes).

  • Logs & Tracing: Solid knowledge of log aggregation (Fluentd, Logstash, Loki) and Distributed Tracing (Jaeger, Zipkin, OpenTelemetry).

  • Automation & Scripting: Practical proficiency in Python or Bash for custom checker creation and operational automation.

  • OS & Troubleshooting: Strong Linux administration skills combined with deep root-cause analysis (RCA) and event correlation capabilities.

  • Mindset & Soft Skills: Proactive problem solver, strong communicator, and a team player accustomed to collaborative DevOps cultures.

  • Education: Bachelor’s degree in Computer Science, Systems Engineering, or equivalent practical experience.

Nice to Have

  • Official Cloud Certifications (AWS, Azure, or GCP).

  • Tooling Certifications (Datadog, Dynatrace, Elastic, Prometheus).

  • SRE or DevOps certifications/foundational knowledge.

  • Solid grasp of core networking concepts (TCP/IP, DNS, Load Balancing).

Benefits

  • 100% remote Long-term commitment, with autonomy and impact

  • Strategic and high-visibility role in a modern engineering culture

  • Collaborative international team and strong technical leadership

  • Clear path to growth and leadership within Coderio 

 

Why join Coderio?

At Coderio, we value talent regardless of location. We are a remote-first company, passionate about technology, collaborative work, and fair compensation.We offer an inclusive, challenging environment with real opportunities for growth.If you are motivated to build solutions with impact, we are waiting for you. 
Apply now.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Related jobs

Other jobs at Coderio Software Company

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.