Logo for Nortal

(EG0069) Senior Site Reliability Engineer (SRE) - Cassandra & AWS - Talent Connection

Role overview

Qualifications

  • Bachelor's Degree in Computer Science, Engineering, or a related field.
  • 5+ years of experience working in Site Reliability Engineering, DevOps, platform engineering or a similar production infrastructure role.
  • Apache Cassandra experience, including data modeling, replication, consistency, tuning and multi-datacenter deployments.
  • Hands-on AWS experience and familiarity with cloud-native architecture.

Responsibilities

  • Design, build and operate highly available, fault-tolerant distributed systems and caching platforms.
  • Evaluate the existing architecture and identify opportunities to scale the platform toward 6x current traffic while maintaining performance and resiliency.
  • Improve caching strategies, including TTLs, refresh patterns, cache placement, performance and downstream dependency management.
  • Help evolve the platform toward active-active resiliency and validate failure, recovery and capacity scenarios.

Key facts

Hard skills

Other skills

  • Communication
  • Problem Solving

About the company

Nortal logo

Nortal

Digital Transformation Consulting

Nortal is a strategic innovation and technology company with an unparalleled track-record of delivering successful transformation projects for over 20 years. As a valued partner for governments, healthcare institutions, leading businesses, and Fortune 500 companies we deliver value by challenging the status quo and succeeding where legacy firms fail.

Company details

Company typeLarge
IndustryDigital Transformation Consulting
Company size1001 - 5000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Senior Site Reliability Engineer (SRE) - Talent Connection 

Real ownership from day one, not check-ins and micromanagement. That's what the Senior Site Reliability Engineer (SRE) role at Nortal looks like. 

What working with us looks like 

  • Remote work with a LATAM team. Coffee breaks, tech talks, and games keep it human, even at a distance. 
  • No micromanagement. We hire for autonomy and expect you to use it. 
  • Support beyond the job. Our People Care team helps with time off, wellness, and anything in between. 
  • Our accounts team handles client relationships, so you focus on the work. 

What you get 

  • Competitive USD salary. 
  • Full remote work, with coworking spaces across LATAM if you want to meet the team in person. 
  • Paid time off under your country's rules, at full salary. 
  • National holidays. 
  • Sick leave, no stress attached. 
  • A yearly refundable credit. Spend it on anything related to your health and well-being. 
  • A day off for your birthday. 

About this search 

Great talent doesn't wait for job postings, so we don't either. We're building a network of skilled professionals for roles that come up regularly with our clients. 

Join our Future Talent network and you'll be one of the first people we reach when the right opening appears. We make around 160 hires a year, so it happens often. 

The role 

As a Senior Site Reliability Engineer, I help modernize and scale production platforms, most recently a customer-data caching platform that serves 20–25+ consuming applications at about 3,000 requests per second. I work across SRE, DevOps, engineering and architecture to improve resiliency, automation, observability, deployment practices and overall platform reliability. I work independently, set technical direction and do well in small engineering teams. 

Your day-to-day: 

  • Design, build and operate highly available, fault-tolerant distributed systems and caching platforms.  
  • Evaluate the existing architecture and identify opportunities to scale the platform toward 6x current traffic while maintaining performance and resiliency.  
  • Improve caching strategies, including TTLs, refresh patterns, cache placement, performance and downstream dependency management.  
  • Help evolve the platform toward active-active resiliency and validate failure, recovery and capacity scenarios.  
  • Design and implement automated CI/CD pipelines, including rolling, blue/green or canary deployment strategies, automated checks and quality gates.  
  • Establish infrastructure, configuration, secrets and application deployment practices using an everything-as-code approach.  
  • Build performance and load-testing capabilities and establish meaningful performance gates for releases.  
  • Develop monitoring, alerting and observability for cache latency, throughput, availability, errors and other key reliability indicators.  
  • Define and improve SLIs, SLOs, error budgets and operational KPIs where appropriate.  
  • Automate operational tasks and reduce manual intervention through scripting, tooling and infrastructure automation.  
  • Troubleshoot production issues, perform root-cause analysis and drive reliability improvements through blameless post-incident reviews.  
  • Provide technical guidance and establish engineering guardrails for SRE and development teams.  
  • Collaborate with Java/Spring Boot engineers, architects, client FTEs and other technical teams to deliver platform improvements. 

What we're looking for 

  • Bachelor's Degree in Computer Science, Engineering, or a related field. 
  • 5+ years of experience working in Site Reliability Engineering, DevOps, platform engineering or a similar production infrastructure role.  
  • Apache Cassandra experience, including data modeling, replication, consistency, tuning and multi-datacenter deployments.  
  • Experience operating and scaling distributed systems in production environments.  
  • Strong understanding of caching architectures, performance optimization and high-availability systems.  
  • Hands-on AWS experience and familiarity with cloud-native architecture.  
  • Experience designing and implementing automated CI/CD pipelines and deployment strategies for highly available systems.  
  • Experience with infrastructure/configuration/secrets as code and a strong preference for automated, repeatable processes.  
  • Strong understanding of observability, monitoring, alerting, performance metrics and production troubleshooting.  
  • Experience with performance testing, capacity planning and identifying system bottlenecks.  
  • Strong programming or scripting experience for automation and tooling; Java/Spring Boot experience strongly preferred.  
  • Ability to work independently, take ownership of ambiguous technical problems and move work forward with limited direction.  
  • Ability to provide technical leadership and establish practical engineering standards and guardrails within a small team.
  • Advanced English Level is required for this role, as you will work with US clients. Effective communication in English is essential to deliver the best solutions to our clients and expand your horizons.  

Strongly Preferred  

  • Experience with GraphQL and API gateway architectures.  
  • Kubernetes and containerized application experience.  
  • GitLab CI/CD experience.  
  • Terraform or other infrastructure-as-code tools.  
  • Prometheus, Grafana, Datadog or similar observability platforms.  
  • Kafka, Amazon MSK or other event-streaming technologies.  
  • Experience designing active-active or multi-region architectures.  
  • Experience with enterprise customer-data platforms or high-volume caching systems. 

How the process works 

  1. Apply and upload your resume. 
  2. We review your profile and add it to our talent database. 
  3. We reach out for an interview to learn about your experience and share more about what we're looking for. 
  4. When a role opens that fits, we contact you to start the process. Because we already know you, it moves faster. 
  5. No intermediaries. You deal with us directly. 

About Nortal 

For over 25 years, we've built the systems that global enterprises and public institutions run on. Today, that means using AI and data to help organizations make better decisions, not just build what they asked for. 

We believe good solutions don't have to be complicated. That's why our clients trust us with the initiatives that matter most. 

Real ownership, real impact, no micromanagement. If that's what you're after, apply now. 👇 

By applying to this position, you authorize Nortal to collect, store, transfer, and process your personal data in accordance with our Privacy Policy. For more information, please review our Privacy Policy.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer Related jobs

Other jobs at Nortal

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.