Logo for Writer

Infrastructure engineer (UK)

Role overview

Qualifications

  • 7+ years of experience in infrastructure engineering, DevOps, or a similar role building and operating large-scale, high-availability production systems.
  • Deep expertise with cloud platforms (AWS strongly preferred) and containerization technologies (Docker and Kubernetes), plus Infrastructure-as-Code tools like Terraform.
  • Proficiency in programming languages such as Python, Java, or Go for automation and monitoring.
  • Knowledge of monitoring and logging tooling (Prometheus, Grafana, ELK Stack) to maintain system health and performance.

Responsibilities

  • Automate operational tasks and infrastructure management by developing robust tools and platforms using Python, Go, or similar languages to reduce manual toil across production.
  • Design and implement scalable, fault-tolerant infrastructure on public cloud providers (AWS, GCP, Azure) to support the rapidly expanding, high-traffic AI platform.
  • Own the reliability, performance, and efficiency of core services by defining and upholding SLOs and error budgets.
  • Own the observability stack for monitoring, logging, and alerting to ensure rapid detection of issues across distributed systems and lead incident response, post-mortems, and root cause analyses.

About the company

Writer logo

Writer

Artificial Intelligence & Machine Learning Services

Writer is the generative AI platform for enterprises. We empower your people β€” product, operations, support, marketing, HR, and more β€” to maximize creativity and 10x productivity by transforming the way they work.Our secure platform snaps easily into your business data sources and delivers accurate answers and content that are fine-tuned on your own data and follow your own AI guardrails. We put generative AI in people’s hands right where they work so they can create, analyze, and govern with ease.Our platform is enterprise-grade, doesn’t use or share your data, and features open and transparent LLMs that are Writer-built and deployable in a variety of ways, including self-hosted. We're compliant with SOC 2 Type II, GDPR, HIPAA, and PCI, and are deployed at leading enterprises, including Intuit, Spotify, L’Oreal, Uber, and Deloitte. Visit us at writer.com.

Company details

Company typeScaleup
IndustryArtificial Intelligence & Machine Learning Services
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

πŸš€ About WRITER

WRITER is where the world's leading enterprises orchestrate AI-powered work. Our vision is to expand human capacity through superintelligence. And we're proving it's possible – through powerful, trustworthy AI that unites IT and business teams together to unlock enterprise-wide transformation. With WRITER's end-to-end platform, hundreds of companies like Mars, Marriott, Uber, and Vanguard are building and deploying AI agents that are grounded in their company's data and fueled by WRITER's enterprise-grade LLMs. Valued at $1.9B and backed by industry-leading investors including Premji Invest, Radical Ventures, and ICONIQ Growth, WRITER is rapidly cementing its position as the leader in enterprise generative AI.

Founded in 2020 with office hubs in San Francisco, New York City, Austin, Chicago, and London, our team thinks big and moves fast, and we're looking for smart, hardworking builders and scalers to join us on our journey to create a better future of work with AI.

πŸ“ About the role

At WRITER, our mission to expand human capacity with superintelligence relies on a foundational truth: our platform must be available, performant, and reliable, 24/7. As an Infrastructure engineer, you'll be at the heart of making this a reality, impacting every enterprise customer who trusts us with their AI-powered workflows. This isn't just about keeping the lights on; it's about pushing the boundaries of what's possible, proactively identifying and solving complex systemic challenges, and laying the groundwork for our rapid growth and the evolving demands of enterprise generative AI. You'll build resilient systems, automate across the stack, and champion reliability best practices, directly enabling our ambitious product roadmap and ensuring our customers always have access to the powerful tools they need.

This is a hybrid position, based out of our New York City or London hubs. You'll report to our director of engineering.

πŸ¦ΈπŸ»β€β™€οΈ What you'll do

  • Automate operational tasks and infrastructure management by developing robust tools and platforms using Python, Go, or similar languages, significantly reducing manual toil across our production environment

  • Design and implement scalable, fault-tolerant infrastructure solutions on public cloud providers (AWS, GCP, Azure) to support WRITER's rapidly expanding, high-traffic AI platform

  • Own the reliability, performance, and efficiency of WRITER’s core services, defining and upholding stringent Service Level Objectives (SLOs) and Error Budgets

  • Own the observability stack for monitoring, logging, and alerting systems to ensure rapid detection of issues across our complex distributed systems

  • Lead incident response, post-mortems, and root cause analyses, applying learnings to proactively prevent future outages and build a more resilient system architecture

  • Collaborate closely with product and engineering teams, providing expert guidance on system design for reliability, performance, and scalability from conception through launch

⭐️ What you need

  • A solid 7+ years of experience in infrastructure engineering, DevOps, or a similar role focused on building and operating large-scale, high-availability production systems

  • Deep expertise with cloud platforms (AWS strongly preferred), containerization technologies like Docker and Kubernetes, and Infrastructure-as-Code tools such as Terraform

  • Strong proficiency in programming languages such as Python, Java, Go for automation and monitoring

  • Knowledge of monitoring and logging tools (e.g., Prometheus, Grafana, ELK Stack) to maintain system health and performance

  • Demonstrated ability to Challenge the status quo, proactively identify systemic weaknesses, and propose innovative solutions to complex reliability problems

  • Excellent communication, collaboration, and problem-solving skills, with a talent for building strong relationships and Connecting with cross-functional teams

  • A strong sense of ownership and accountability, eager to Own mission-critical systems and drive them toward peak performance and unparalleled reliability

🍩 Benefits & perks (UK full-time employees):

  • Generous PTO, plus company holidays

  • Comprehensive medical and dental insurance

  • Paid parental leave for all parents (16 weeks)

  • Fertility and family planning support

  • Early-detection cancer testing through Galleri

  • Competitive pension scheme and company contribution

  • Annual work-life stipends for:

    • Wellness stipend for gym, massage/chiropractor, personal training, etc.

    • Learning and development stipend

  • Company-wide off-sites and team off-sites

  • Competitive compensation and company stock options

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
Β·

Infrastructure Engineer Related jobs

Other jobs at Writer

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.