Logo for Akamai Technologies

Senior Site Reliability Engineer

Role overview

Qualifications

  • Expert level experience in SysAdmin (Linux/Unix Administration), DevOps or SRE role
  • Demonstrated expertise in Kubernetes and large-scale containerization systems
  • Proficiency in at least one programming language (Python/Golang)
  • Experience with observing tools like Prometheus, Grafana, and distributed tracing

Responsibilities

  • Provide support and mentorship for other engineers within the department
  • Develop and maintain automated tools and scripts to enhance system reliability, deployment processes, and incident response efficiency
  • Participate in on-call rotations, guiding restoration and repair of service-impacting issues
  • Contribute to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

Key facts

Hard skills

Other skills

  • Collaboration
  • Mentorship
  • Problem Solving

About the company

Akamai Technologies logo

Akamai Technologies

Cloud Computing & Infrastructure (IaaS/PaaS)

At Akamai, we make life better for billions of people, billions of times a day. Every day, billions of people around the world connect with their favorite brands to shop online, play the latest video games, log into mobile banking apps, learn remotely, share videos with friends, and so much more. They may not know it, but Akamai is there, powering and protecting life online. Over 20 years ago, we set out to solve the toughest challenge of the early internet: the “World Wide Wait.” And we’ve been solving the internet’s toughest challenges ever since, working toward our vision of a safer and more connected world. With the world’s most distributed compute platform — from cloud to edge — we make it easy for businesses to develop and run applications, while we keep experiences closer to users and threats farther away. That’s why innovative companies worldwide choose Akamai to build, deliver, and secure their digital experiences. Our leading security, compute, and delivery solutions are helping global companies make life better for billions of people, billions of times a day. Devoted, determined problem-solvers who share a passion for technology, we’re always pushing ground-breaking ideas and driving innovation. Want to power and protect life online, by solving the toughest challenges? Be part of an amazing team. Let’s connect: LinkedIn: https://www.linkedin.com/company/akamai-technologies Twitter: https://twitter.com/Akamai Blog: https://www.akamai.com/blog

Company details

Company typeXLarge
IndustryCloud Computing & Infrastructure (IaaS/PaaS)
Company size5001 - 10000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Are you passionate about cutting edge technology?

Do solving some of the Internet's most difficult content delivery challenges interest you?

Join our highly skilled Site Reliability team

Our team designs, develops, and manages applications and infrastructure that support Akamai's Compute products and services. We do this while maintaining Akamai's mission at the forefront of what we do. Make life better for billions of people, billions of times a day.

Partner with the best

The Senior Engineer creates solutions to improve automation and efficiency for systems and teams. Responsibilities include optimizing workflows, infrastructure, and applications. Expertise in Linux administration, configuration management, and performance tuning is essential. Collaborate on deployment, monitoring, and resolving incidents. Focus on reliability, scalability, and efficiency through automation and resource optimization. Promote continuous improvement and operational excellence across all systems.

As a Senior Site Reliability Engineer, you will be:

  • Providing support and mentorship for other engineers within the department

  • Developing and maintaining automated tools and scripts to enhance system reliability, deployment processes, and incident response efficiency.

  • Improving our system monitoring to speed error detection and remediation, enhancing performance and reliability of virtualization platform

  • Participating in on-call rotations, guiding restoration and repair of service-impacting issues

  • Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response

  • Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

Do what you love

To be successful in this role you will:

  • Possess expert level experience in a SysAdmin (Linux/Unix Administration), DevOps or SRE role, working with large scale distributed systems

  • Demonstrate expertise in Kubernetes and large-scale containerization systems.

  • Possess at least one programming language (Python/Golang) and configuration management with Terraform/SaltStack/Ansible

  • Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.

  • Have experience with architecting software and infrastructure at scale

  • Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.

Build your career at Akamai

Our ability to shape digital life today relies on developing exceptional people like you. The kind that can turn impossible into possible. We’re doing everything we can to make Akamai a great place to work. A place where you can learn, grow and have a meaningful impact.

With our company moving so fast, it’s important that you’re able to build new skills, explore new roles, and try out different opportunities. There are so many different ways to build your career at Akamai, and we want to support you as much as possible. We have all kinds of development opportunities available, from programs such as GROW and Mentoring, to internal events like the APEX Expo and tools such as Linkedin Learning, all to help you expand your knowledge and experience here.

Learn more

Not sure if this job is the right match for you or want to learn more about the job before you apply? Schedule a 15-minute exploratory call with the Recruiter and they would be happy to share more details.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer Related jobs

Other jobs at Akamai Technologies

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.