Logo for Titan AI

VP of Site Reliability

Role overview

Qualifications

  • Ten or more years in engineering
  • At least five years building SRE or platform operations functions
  • Experience with companies shipping software to customers
  • Knowledge of managing multi-tenant and multi-deployment-model infrastructure

Responsibilities

  • Build the SRE practice and operate it yourself
  • Define severity tiers and SLA commitments for production support
  • Set the operating system across all engineering lanes
  • Write SLOs and lead incident response at live bank customers

About the company

Titan AI logo

Titan AI

Artificial Intelligence & Machine Learning Services

Titan is the first banking-native AI platform built by a founding team with deep experience in AI, bank operations, and regulatory compliance. Titan is designed to help Banks, FinTechs and Credit Unions deliver outsized impact. We allow these customers to: 🏦 Adopt AI safely through a secure, private interface that provides access to multiple foundational models, explainability tools, and bank-grade security. 📊 Reason with their own data using Titan’s own banking models designed to think like seasoned bank operators, leaders, and regulators. 🤖 Automate intelligently with function-specific banking agents that handle critical workflows.

Company details

IndustryArtificial Intelligence & Machine Learning Services
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About Titan

Titan builds AI software for banks: purpose-built small language models, a banking ontology, and AI bankers that financial institutions can trust. Our models outperform general-purpose LLMs by 30 to 80 percent on banking tasks. Customers include community banks, credit unions, and large regional and super-regional institutions. We are backed by leading fintech investors and operate under the compliance, audit, and model-risk standards that banking requires.

Why This Role Exists

Titan is scaling from a handful of live banking customers to thirty, then to hundreds. Each bank deploys differently: Azure, private cloud, or the bank's existing infrastructure. The core problem this role solves is making the platform work consistently and reliably across all of them, managing the last-mile deployment complexity that grows with every new customer.

This is a hands-on, principal-level role. You are not coming in to build an org chart. You are coming in to do the work: write the runbooks, stand up the on-call rotation, own incident command when a bank has an outage, and build the deployment playbook that takes us from client 10 to client 350. The practices get built before the teams do.

What You Own

Site Reliability Engineering. You build the SRE practice and operate it yourself first: SLO framework, on-call rotation, and incident command process. You write the SLOs, run the rotation, lead incident response at live bank customers, and produce the postmortems. Once the practice is stable and documented, you bring in an SRE Lead to own it and grow the function.

Production Support. Before the first support hire, you define severity tiers, SLA commitments per customer tier, and escalation paths, and you route alerts into a real queue. You are the technical accountable owner when a bank has a production incident. Once the structure works, you hire into it cost-efficiently and hand off the day-to-day to a Support Lead as customer volume justifies it.

Engineering Operations. You set the operating system across all four engineering lanes: sprint discipline, release rituals, code review standards, change management evidence, and the metrics the CEO and board read monthly. You own the SOC 2 artifacts, model risk review documentation, and the change traceability that bank examiners scrutinize.

What You Will Not Own

• Technical direction and architecture. Owned by the CTO and Chief Architect.

• Quality engineering. QE is being built as a separate function with dedicated QE engineers and a QE Lead hired independently.

• People management of the AI Toolbelt, Product Engineering, and Banking Models lanes. Lane leads manage their own teams. You influence through process, not reporting lines.

Who You Are

Ten or more years in engineering, with at least five years personally building SRE or platform operations functions at a software company selling into enterprise or regulated markets. You have not spent your career in internal bank IT. You come from companies that ship software to customers and operate it at scale: ServiceNow, MongoDB, AWS, GCP, or comparable. You have managed multi-tenant and multi-deployment-model infrastructure and know the last-mile complexity that comes with it.

You have written SLOs that people actually use. You have stood up an on-call rotation from nothing. You have been the technical owner during a production incident and know what it costs to not have a process. You earn trust from senior engineers without leaning on title. You see process as leverage, not overhead. You are not here to manage. You are here to build.

What Success Looks Like

In your first 90 days: a diagnostic of engineering operations shared with the CEO and CTO, written SLOs on customer-facing services with a clear baseline of where performance stands, and the support triage structure defined before the first support hire. In your first six months: the on-call rotation is running, incident command has been tested in production, and the deployment playbook covers the three deployment models we operate. At one year: the platform runs reliably from client 10 to client 30 and the foundation is in place to reach 100. The operating system runs without you prompting it.

Compensation and Structure

• Competitive base and meaningful equity.

• Atlanta, GA strongly preferred. West Coast considered on a case-by-case basis. Remote-friendly for the right candidate.

• Reports to the CEO. Peer to the CTO and the Chief Customer and Growth Officer.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Site Reliability Engineer (SRE) Related jobs

Other jobs at Titan AI

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.