Logo for Printify

Senior Site Reliability Engineer (remote within EMEA)

Role overview

Qualifications

  • Solid Linux systems administration background and comfort scripting in Python.
  • Strong AWS knowledge: EKS, IAM, VPC networking, RDS, S3, SQS.
  • Hands-on experience operating and troubleshooting Kubernetes (EKS) at production scale.
  • Proficiency with Terraform and ideally Terragrunt for multi-environment management.

Responsibilities

  • Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts.
  • Drive large-scale automation projects and set standards for using Terraform and GitOps.
  • Be the go-to person for solving complex, cross-service infrastructure problems.
  • Mentor mid-level SREs, provide detailed feedback, and support onboarding of new team members.

About the company

Printify logo

Printify

E-commerce & Online Marketplaces

Printify is a transparent print-on-demand and dropshipping platform designed to help online merchants make more money in a simple and easy way. Our platform lets anyone start a business with as little investment and risk as possible, connecting entrepreneurs with more than 90 print providers across the globe. Our mission is to empower entrepreneurs to build their own businesses, independently. We created Printify to break down the needlessly complex print and merch industry so that creative minds, time-strapped entrepreneurs and everyone in between can do what they do best: create great products, drive sales, and run their own business. Our print-on-demand platform has been on a steep growth curve, more than doubling our employee count during the 2020 due to high demand. Financial Times ranked Printify as the 15th fastest-growing technology company in the USA in 2020, our key market. Printify is headquartered in Riga with 280 employees and counting, backed by leading angel investors from Silicon Valley and Europe.

Company details

Company typeSME
IndustryE-commerce & Online Marketplaces
Company size501 - 1000

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About the team: 

Platform Infrastructure builds, operates, and continuously evolves FYUL's container platform and cloud foundation. We foster a DevOps culture through self-service tooling, enabling product engineering teams to ship reliable, secure, and cost-efficient services as the business scales. The team owns our AWS cloud accounts, Kubernetes platform, cloud networking, observability stack, core databases, CI/CD pipelines, and infrastructure-as-code, and acts as the go-to partner for engineering teams on cloud and DevOps topics.


About the role:

We're hiring a Senior SRE II to join Platform Infrastructure as one of the team's senior individual contributors. At this level, you're the go-to person for our most complex infrastructure problems: you architect and drive large-scale automation and reliability initiatives, set standards other engineers follow, and mentor Associate and mid-level SREs. You'll split your time between hands-on platform work - Kubernetes, AWS, GCP, CI/CD, observability - and technical leadership: proposing designs, reviewing others' work, and helping the team make good build-vs-buy and cost/reliability trade-offs.

Your daily tasks will include:

  • Infrastructure & reliability: Architect and manage highly available, secure, and scalable infrastructure across multiple AWS accounts and environments using infrastructure as code.

  • Design and operate our Amazon EKS clusters, including networking policies, persistent storage, and scaling strategies for containerized workloads.

  • Own and evolve core platform services: cloud networking, Kubernetes, and the databases and messaging systems engineering teams depend on.

  • Automation & infrastructure as code: Drive large-scale automation projects and set standards for using Terraform / Terragrunt and GitOps (ArgoCD) across teams.

  • Lead adoption of automation to reduce manual operational work and keep environments consistent and repeatable.

  • Observability & incident response: Be the go-to person for solving complex, cross-service infrastructure problems.

  • Drive initiatives that improve reliability and observability (Grafana, Prometheus, Loki, Tempo, Mimir) so systems scale with minimal manual intervention.

  • Participate in on-call rotation, lead incident response for production issues, and write clear runbooks, ADRs, and postmortems.

  • Security & cost efficiency: Lead security efforts within the team - IAM, encryption, secure logging - and mentor others on secure infrastructure practices.

  • Audit infrastructure spend regularly and drive cost optimization across the platform (rightsizing, autoscaling, FinOps practices).

  • Collaboration & mentorship: Mentor mid-level SREs, provide detailed feedback, and support onboarding of new team members.

  • Communicate complex technical concepts clearly to both engineers and non-technical stakeholders.

  • Partner with product engineering squads to understand their needs and represent Platform Infrastructure in cross-team initiatives.

Your qualifications:

These reflect the technical bar we hold Senior SRE II's to internally, based on our SRE competency framework and current stack.

  1. Core technical experience:

  • Solid Linux systems administration background and comfort scripting in Python.

  • Strong AWS knowledge: EKS, IAM (roles, policies, IRSA), VPC networking, RDS, S3, SQS, and familiarity with the Well-Architected Framework; experience in multi-account AWS environments is a strong plus.

  • Hands-on experience operating and troubleshooting Kubernetes (EKS) at production scale, including Helm chart development, CNI networking (we run Cilium), pod networking/IPAM concepts, and container security (ECR, image scanning).

  • Proficiency with Terraform (modules, state management) and ideally Terragrunt for multi-environment management; GitOps experience with ArgoCD.

  • Experience with Postgres, MySQL and/or MongoDB in production scale, including Aurora.

  • CI/CD experience with Jenkins (Jenkinsfile, shared libraries) and/or GitHub Actions, and familiarity with deployment strategies such as blue-green and canary.

  • Experience with the Grafana observability stack (Grafana, Prometheus, Loki, Tempo, Mimir) - metrics design, dashboarding, alerting, log aggregation, and distributed tracing. Not only using but also maintaining it.

  • Practical incident management experience: on-call rotations, structured incident response, and writing runbooks/postmortems.

  • Working knowledge of 12-Factor App principles and cost optimization / FinOps awareness.

    2. How you work:

  • A methodical, data-driven approach to troubleshooting rather than guessing.

  • Strong written communication - you write runbooks, ADRs, and postmortems that others can actually follow.

  • Comfortable driving initiatives with ambiguous ownership, and taking accountability for outcomes rather than waiting to be asked.

  • Track record of mentoring less senior engineers and giving direct, constructive feedback.

  • Several years of hands-on production infrastructure/SRE experience, with demonstrated ownership of initiatives at a senior individual-contributor level (leading design work, setting standards, being the escalation point for hard problems).

    3. Nice to have:

  • GCP Experience.

  • Experience with Kafka / AWS MSK.

  • Prior experience in regulated or compliance-sensitive environments (security best practices, access reviews).

  • Experience contributing to a platform/DevEx roadmap that other engineering teams consume as a self-service product.

Our tech stack:

  • Languages: PHP (Symfony), Node.js (TypeScript), Angular (TypeScript).

  • Data: PostgreSQL, Redis, MongoDB.

  • Infra: AWS, Kubernetes, Terraform, Helm, Atlantis

  • Engineering Tools: Postman, Git, GitHub Copilot, PhpStorm, Grafana, Kibana, Prometheus.

  • Remote work Tools: Jira, Miro, Google Workspace, Slack.

  • Development Practices: Pair Programming, Code Reviews, Continuous Integration/Deployment.

What we offer:

  • A global, inclusive team that’s as supportive as it is ambitious and serious about getting things done

  • An opportunity to work remotely or in a modern and welcoming office in Riga

  • Flexible working hours (start your day as late as 11 AM)

  • Private health insurance

  • 2 extra paid days off to focus on your mental or physical well-being

  • 1 extra paid day off to celebrate a Birthday or any other celebration of your choice

  • Internal and external learning opportunities

  • Access to mentorship, internal meetups, and hackathons, both on-site and online

  • Free and healthy lunch if you work from the Rīga office

  • Design and order your own merch using our platforms with an employee discount

  • Exciting team-building events and parties you’ll never forget!



FYUL is the engine that powers on-demand commerce at global scale.

Formed in 2024 through the merger of Printful, Printify, and Snow Commerce, we bring together tech, talent, and infrastructure to help people turn ideas into beautiful products.

From solo creators to entertainment giants, FYUL powers merch that connects with millions, backed by advanced tech, premium production, and global reach.

We're a fast-growing global company working toward powering great brands, great experiences, and great people.


We are an equal-opportunity workplace. We’re committed to diversity and inclusion and make hiring decisions based solely on qualifications, merit, and work experience.

If you think you’d excel in this role, send us your resume in English, showing us why you are the right person for the job.

Interested, but don’t think this is the right fit for you? Feel free to share it with friends and check out other open positions at our career site. We’re always looking for creative and driven minds to join our ever-growing team!

AS Printful Latvia (Reģ. Nr. 40203050078)

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Printify

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.