Logo for Pozent Corporation

Lead Site Reliability Engineer - USA (Remote Role)

Role overview

Qualifications

  • Strong experience in observability and monitoring, including hands‐on expertise with Dynatrace and OpenTelemetry
  • Proven experience designing and executing automated regression testing frameworks
  • Strong proficiency in Infrastructure as Code (IaC) using Terraform
  • Experience with CI/CD pipelines and deployment automation

Responsibilities

  • Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure
  • Develop comprehensive observability strategies leveraging Dynatrace and OpenTelemetry
  • Lead incident response activities for high‐severity production events
  • Drive operational excellence through automation of repetitive tasks and workflows

Key facts

Other skills

  • Analytical Skills
  • Troubleshooting (Problem Solving)
  • Organizational Skills
  • Time Management
  • Communication
  • Collaboration
  • Leadership

About the company

Pozent Corporation logo

Pozent Corporation

IT Services & IT Consulting

POZENT is a digital technology and workforce solutions company with expertise in Digital Artificial Intelligence & Block Chain technologies. Our Center of Excellence is focused on building solutions and applications of cutting-edge technology in the development of Digital Ecosystems that will propel the efficiency and profit margin of your organization. Our workforce solutions engage the most highly qualified and motivated business, functional and technical professionals to provide your business a custom fit solution that will propel growth and presence. We work for your success!

Company details

IndustryIT Services & IT Consulting
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Project Description

Responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. Serves as a senior individual contributor responsible for designing, implementing, and improving Site Reliability Engineering (SRE) practices across the software development lifecycle. Works closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management while coaching and influencing others.

Responsibilities

Design & Reliability Engineering

  • Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.
  • Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.
  • Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.

Observability & Monitoring

  • Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.
  • Design and maintain end‐to‐end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
  • Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.

Incident & Problem Management

  • Lead incident response activities for high‐severity production events, coordinating cross‐functional teams to restore services and minimize customer impact.
  • Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.

Automation & Operational Excellence

  • Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
  • Partner with development teams to build reliable and observable services throughout the SDLC.

Testing & Validation

  • Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
  • Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.

Infrastructure & Cloud Engineering

  • Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization.
  • Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.
  • Utilize Azure‐native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability.

Resiliency & Readiness

  • Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains.
  • Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.
  • Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization.

Capacity & Performance

  • Lead capacity planning, performance tuning, and workload optimization efforts across production environments.

Documentation & Knowledge Management

  • Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
  • Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries.

Communication & Leadership

  • Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
  • Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.
  • Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.

Risk & Compliance

  • Understand and adhere to the company's risk and regulatory standards, policies, and controls in accordance with the company's Risk Appetite.
  • Identify reliability, operational, and technology risks requiring escalation to management.
  • Promote an environment that supports a culture of belonging and reflects the client brand.
  • Maintain internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.
  • Complete other related duties as assigned.

Skills - Must Have

  • Strong experience in observability and monitoring, including hands‐on expertise with Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics collection and analysis, centralized logging, alerting, and dashboard development.
  • Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
  • Strong proficiency in Infrastructure as Code (IaC) using Terraform.
  • Experience with CI/CD pipelines, deployment automation, and operational tooling.
  • Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
  • Strong understanding of application performance management, distributed systems, and modern cloud‐native architectures.

Cloud & Platform Expertise

  • Strong experience with Microsoft Azure, including Azure App Services, Resource Groups, Azure networking concepts, scaling and performance optimization, deployment and release management, and application lifecycle management.
  • Experience leveraging Azure‐native operational tooling such as Azure Monitor, Application Insights, Log Analytics, Azure dashboards, and alerting.
  • Experience supporting cloud‐native and hybrid infrastructure environments.

Reliability & Engineering Practices

  • Demonstrated experience implementing and operating SRE practices, including SLOs, SLIs, error budgets, incident management, problem management, RCA, and reliability automation.
  • Ability to improve system reliability through performance tuning, capacity planning, observability‐driven insights, proactive issue detection, and reliability engineering initiatives.
  • Experience developing automated recovery mechanisms and self‐healing solutions.
  • Knowledge of resiliency engineering patterns, disaster recovery planning, and high‐availability architectures.

Nice to Have

  • Experience supporting large‐scale enterprise applications in regulated environments.
  • Strong analytical and troubleshooting skills related to production systems and distributed architectures.
  • Experience working in Agile and DevOps operating models.
  • Ability to work autonomously and lead complex reliability initiatives.
  • Strong organizational and time management skills.
  • Advanced verbal and written communication skills.
  • Experience driving project milestones and delivery commitments.
  • Proven experience leading major incident response and post‐incident improvement efforts.
  • Experience partnering with architecture, infrastructure, cybersecurity, and application development teams.
  • Experience with scripting and automation using PowerShell, Python, Bash, or similar technologies.
  • Industry certifications in Azure, Terraform, Cloud Engineering, or Site Reliability Engineering preferred.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Pozent Corporation

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.