Logo for Fidelity Investments

Principal Site Reliability Engineer

Role overview

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field
  • Five (5) years of experience as a Principal Site Reliability Engineer or related occupation
  • Demonstrated expertise in software performance benchmarking and engineering
  • Experience with AWS services and DevOps practices

Responsibilities

  • Defines and leads enterprise-level reliability strategies
  • Architects resilient systems and infrastructure
  • Creates and publishes performance test results report with recommendations on quality improvement
  • Implements advanced observability practices and techniques at scale

Key facts

Hard skills

Other skills

  • Microsoft Windows
  • Problem Solving
  • Mentorship
  • Communication

About the company

Fidelity Investments logo

Fidelity Investments

Financial Services

Fidelity’s mission is to strengthen the financial well-being of our customers and deliver better outcomes for the clients and businesses we serve. Fidelity’s strength comes from the scale of our diversified, market-leading financial services businesses that serve individuals, families, employers, wealth management firms, and institutions. With assets under administration of $15.0 trillion, including discretionary assets of $5.9 trillion as of March 31, 2025, we focus on meeting the unique needs of a broad and growing customer base. Privately held for 78 years, Fidelity employs more than 77,000 associates across the United States, Ireland, and India. For our Terms and Conditions, please visit http://go.fidelity.com/LIterms

Company details

IndustryFinancial Services
Company size10,001+

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Job Description:

Note: Fidelity will not provide immigration sponsorship for this position.

Position Description:

Deploys and supports distributed, multi-tiered systems at scale while ensuring high availability and fault tolerance across multiple environments. Builds and operates resilient platforms in Amazon Web Services (AWS) using Elastic Compute Cloud (EC2), Simple Storage Service (S3), and Auto Scaling Groups for dynamic resource management. Designs, develops, and executes performance tests using Java-based frameworks, Apache JMeter, k6, and Rush-hour to validate system behavior under day-to-day traffic patterns. Defines and implements observability practices to monitor system health, latency, and error rates through metrics, logs, and distributed tracing using Datadog, Grafana, Splunk, and the Elasticsearch, Logstash, and Kibana (ELK) stack. Automates operational workflows with Python and Shell scripting to enhance efficiency and reduce manual tasks. Supports consistent build, deployment, and orchestration processes using cloud computing and DevOps technologies -- Continuous Integration and Continuous Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets, and implementing proactive monitoring and incident response strategies. Builds and refines methodologies for performance, load, stress, and chaos testing and develops analytics and reports aligned with business needs to improve system resilience and optimization.

Primary Responsibilities: 

  • Defines and leads enterprise-level reliability strategies.
  • Architects resilient systems and infrastructure.
  • Creates and publishes performance test results report with recommendations on quality improvement.
  • Maintains scalability and resiliency of complex environment.
  • Implements advanced observability practices and techniques at scale.
  • Manages and interprets large datasets using query languages and visualization tools.
  • Advises senior leadership on reliability engineering best practices.
  • Mentors junior engineers.
  • Performs independent and complex technical and functional analysis for multiple divisional initiatives.
  • Develops innovative solutions to improve system availability, scalability, and performance.
  • Designs, implements, and maintains performance test frameworks.

Education and Experience:

Bachelor’s degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment.

Or, alternatively, Master’s degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and three (3) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment.

Skills and Knowledge:

Candidate must also possess:

  • Demonstrated Expertise (“DE”) performing software performance benchmarking and engineering for online financial web applications, Application Programming Interfaces (APIs), and mobile transactions according to DevOps practices, using performance benchmarking tools Rushhour, Locust, K6, and JMeter; and configuring CI/CD and test automation, using Jenkins, Sonar, Ant, Maven, Artifactory, and Terraform in AWS.
  • DE solutioning, designing, architecting, and building scalable and resilient enterprise-grade software platforms using cloud-based architecture and AWS services (EC2, Elastic Container Service (ECS), Lambda, Elastic MapReduce (EMR), and CloudFormation); developing microservices on Elastic Kubernetes Service (EKS), implementing CI/CD pipelines using DevOps tools (Bitbucket, GitHub, Artifactory, Sonar, Veracode, and Helm), and adhering to DevOps practices along with leveraging Java, Python, Spring Boot, Docker, EKS, and AWS.
  • DE analyzing and monitoring system and application performance across Apache, NGINX, Java, and Node.js platforms, and Linux and Windows environments, using Splunk, Datadog, Kibana, Grafana, and AWS CloudWatch; diagnosing performance bottlenecks, recommending tuning strategies, reducing Mean Time to Detect (MTTD) and Mean Time to Repair (MTTR), using Application Performance Monitoring (APM) tools -- Dynatrace, New Relic, Splunk, and Datadog; and performing capacity planning to optimize Central Processing Unit (CPU), memory, and process configurations.
  • DE instrumenting advanced observability practices at scale across cloud-native and hybrid environments; defining and tracking SLOs and SLIs to ensure reliability and performance metrics, using Python automation, Infrastructure as Code (IaC) methodologies, and observability tools (Datadog, Splunk, Dynatrace, Grafana, and the ELK stack; developing custom dashboards, alerting rules, and automated incident response workflows to proactively detect and resolve performance degradations, using Datadog, Catchpoint, Grafana, ELK stack, and Cloudwatch; and enabling actionable insights through trace-level correlation of end-to-end (E2E) user journeys and system behaviors, using Dynatrace, Splunk, Draw.io, and Miro.

#PE1M2

#LI-DNI

Fidelity’s Onsite Working Model
Fidelity is transitioning to a full-time onsite working model through a phased rollout across regions and roles. Currently, some roles and locations require 100% onsite presence, while others require less. Onsite expectations are likely to evolve as the rollout continues. This transition does not apply to fully remote roles.

Certifications:

Category:

Information Technology

Please be advised that Fidelity’s business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement-related financial activities and the rules and regulations of numerous self-regulatory organizations, including FINRA, among others. Those laws and regulations may restrict Fidelity from hiring and/or associating with individuals with certain Criminal Histories.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer Related jobs

Other jobs at Fidelity Investments

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.