Logo for Wizdaa

Platform Architect (AI/ML Infrastructure, GCP-focused)

Role overview

Qualifications

  • Experience in ML infrastructure and SRE
  • Expertise in Google Cloud Platform
  • Proficiency in managing inference services and pipelines
  • Strong background in DevOps principles

Responsibilities

  • Build and operate model and inference serving infrastructure
  • Own the ML deployment lifecycle including model registry and rollout strategies
  • Operate agentic and LLM workloads in production
  • Drive ML cost efficiency and manage resource allocation

About the company

Wizdaa logo

Wizdaa

Staffing & Recruiting

Wizdaa, formerly known as PrideLogic, excels in creating world-class development teams and extending operational runways for innovative startups. We source, place, and pay top-tier developers who operate within U.S. time zones at highly competitive rates. Recognizing the uniqueness of each startup, we provide tailored solutions to meet specific needs, ensuring every client receives bespoke services that drive success.

Company details

Company typeStartup
IndustryStaffing & Recruiting
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

We're looking for a Platform Architect who can set the standard for how we build, ship, and operate ML and AI systems at scale. You sit at the intersection of ML infrastructure and SRE. You'll own the path from model and pipeline to reliable production service, and you'll bring DevOps rigor to systems that are historically under-engineered. The immediate focus is AI/ML infrastructure on Google Cloud.

This is not a ticket-processing role, and it's not a research role. You'll tackle hard problems: model serving reliability, inference cost and latency, reproducible pipelines, and agentic workload operations. You'll have the scope to solve them properly. Senior professionals here identify problems before they're asked and raise the ceiling on what the platform can do.

WHAT YOU'LL WORK ON

  • Build and operate model and inference serving infrastructure, managing latency, throughput, autoscaling, and reliability for real-time and batch inference across multiple tenants.

  • Own the ML deployment lifecycle: model registry, versioning, promotion workflows, rollout strategies (canary, shadow, A/B), and safe rollback.

  • Operate agentic and LLM workloads in production, managing inference providers and gateways, quota and throttling behavior (TPS/TUPS limits), guardrails, prompt/version management, and graceful degradation under load.

  • Build reproducible, automated ML pipelines: training, evaluation, and deployment pipelines as code, with lineage and reproducibility built in.

  • Extend infrastructure-as-code to ML systems, using Terraform patterns and multi-project design that bring ML infrastructure under the same standards as the rest of the platform.

  • Operate GitOps for ML workloads, owning ArgoCD configuration and promotion workflows across environments and tenants.

  • Run ML and AI workloads on multi-tenant Kubernetes (GKE), managing GPU/accelerator scheduling, workload placement, tenant isolation, and cost-aware capacity.

  • Own ML reliability and observability: SLOs for inference services, model and data drift detection, performance regression monitoring, alert quality, on-call ergonomics, and runbook culture.

  • Drive ML cost efficiency by right-sizing accelerators, managing committed-use and Spot VM capacity, and attributing inference cost across tenants and workloads.

  • Use agentic coding tools for infrastructure and pipeline work: scaffolding environments, generating and reviewing IaC and pipeline code, and accelerating automation.

WHAT YOU WON'T FIND HERE

A platform team that maintains the status quo. We're actively building: new scale requirements, new architectural domains, and an ML/AI footprint that's growing fast. Senior engineers here shape how the platform evolves, and the tools available to do it are better than they've ever been.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Solutions Architect Related jobs

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.