Logo for Filevine

Senior Site Reliability engineer

Role overview

Qualifications

  • 8+ years of hands-on experience in software engineering, cloud infrastructure, platform engineering, DevOps, or related technical roles
  • Strong knowledge of distributed systems and experience operating Kubernetes workloads and cloud infrastructure in AWS or a comparable platform
  • Strong proficiency with Python, Go, Bash, or similar languages
  • Demonstrated experience applying AI and machine learning to operational data and engineering workflows

Responsibilities

  • Design and improve monitoring, logging, distributed tracing, dashboards, alerting, SLIs, and SLOs
  • Build and maintain automation, internal tools, and CI/CD systems for reliable deployments
  • Drive implementation and continuous improvement of reliable systems for Filevine products
  • Lead significant technical initiatives and mentor other Site Reliability Engineers

About the company

Filevine logo

Filevine

Computer Software / SaaS

Filevine is changing the way legal work gets done for law practitioners and their clients. As the leading legal work platform, Filevine is dedicated to empowering all organizations with tools to simplify and elevate complex, high-stakes legal work. Powering everything from document management and client communication to contract lifecycle management and business analytics, over 25,000 legal professionals use Filevine daily to deliver excellence in every contract, deadline, and result. Filevine is recognized on the Deloitte Fast 500, has been named one of the Utah Business Fast 50 and is among the top 50 fastest-growing privately-owned software companies according to the 2021, 2022, and 2023 Inc. 5000 list.Filevine believes in a brighter future where the intersection of legal work and business is made more seamless, transparent, and effortless for all legal professionals and everyone they interact with through the power of legal technology.

Company details

Company typeScaleup
IndustryComputer Software / SaaS
Company size201 - 500

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Filevine is a Legal AI company delivering Legal Operating Intelligence for the future of legal work. Grounded in a singular system of truth, Filevine brings together data, documents, workflows, and teams into one unified platform—where modern legal work happens with clarity and consistency.
 
Powered by LOIS, the Legal Operating Intelligence System, Filevine connects context across every matter to transform legal operations from reactive to proactive. LOIS reads, understands, and reasons across your data to surface insight, automate complexity, and give professionals the clarity and confidence to see more, know more, and do more. Fueled by a team of exceptional collaborators and innovators, Filevine’s rapid growth has earned AI awards and recognition from Deloitte and Inc. as one of the most innovative and fastest-growing technology companies in the country.

Responsibilities

• Design and improve the monitoring, logging, distributed tracing, dashboards, alerting, SLIs,
and SLOs that give teams meaningful visibility into production health and customer impact.
• Build and maintain automation, internal tools, and CI/CD systems that increase engineering
efficiency, reduce toil, and support reliable deployments at scale. Take responsibility for the
quality and reliability of tools and services you support.
• Drive the implementation and continuous improvement of reliable systems for building,
deploying, testing, and operating Filevine products, proactively identifying and resolving
reliability, performance, scalability, and security risks before they impact customers.
• Own complex production incidents through detection, triage, communication, resolution,
and follow-up. Turn incident learning into durable corrective actions, stronger runbooks and
operating practices, and improvements that reduce recurring incidents and operational
burden.

Lead significant technical initiatives from problem definition and design through
implementation and adoption. Coordinate work across engineers and teams, communicate
tradeoffs and risks, and help ensure the work delivers the intended results.
• Mentor other Site Reliability Engineers through design reviews, incident follow-ups, paired
problem-solving, and meaningful delegation. Help engineers develop stronger technical
judgment and become increasingly capable of handling complex production work
independently.
• Participate in the shared on-call rotation and help ensure production systems are prepared
to operate reliably at scale through capacity planning, operational readiness, and
continuous improvements to resilience and recovery.
• Apply AI and machine learning to analyze operational signals, identify patterns, forecast
reliability and capacity risks, and implement improvements that make systems more
reliable, efficient, and easier to operate.


Qualifications
• 8+ years of hands-on experience in software engineering, cloud infrastructure, platform
engineering, DevOps, or related technical roles, including at least 5 years in a Site Reliability
Engineering or reliability-focused role.
• Strong knowledge of distributed systems and hands-on experience operating Kubernetes
workloads and cloud infrastructure in AWS or a comparable platform, with proficiency in
Infrastructure as Code, monitoring, logging, alerting, distributed tracing, SLIs, and SLOs.
• Strong proficiency with Python, Go, Bash, or a similar language, with demonstrated
experience building and maintaining production tooling, automation, CI/CD pipelines, and
deployment systems that reduce toil, improve reliability, and simplify ongoing operations.
• Demonstrated ability to lead troubleshooting, incident response, root cause analysis, and
long-term reliability improvements for complex production systems, including the
elimination of recurring incidents and operational work.
• Proven ability to mentor Site Reliability Engineers, help others build stronger technical
judgment, communicate clearly with technical and business stakeholders, and lead complex
initiatives from planning through delivery.
• Demonstrated experience applying AI and machine learning to operational data and
engineering workflows to identify patterns, forecast reliability or capacity risks, and
implement measurable improvements with appropriate safeguards.

Cool Company Benefits:
- A dynamic, rapidly growing company, focused on helping organizations thrive 
- Medical, Dental, & Vision Insurance (for full-time employees)
- Competitive & Fair Pay
- Maternity & paternity leave (for full-time employees)
- Short & long-term disability
- Opportunity to learn from a dedicated leadership team
- Top-of-the-line company swag
 
Privacy Policy Notice
Filevine will handle your personal information according to what’s outlined in our Privacy Policy.
 
Communication about this opportunity, or any open role at Filevine, will only come from representatives with email addresses using "filevine.com". Other addresses reaching out are not affiliated with Filevine and should not be responded to.
 

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Filevine

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.