Logo for Nonstop Administration & Insurance Services

Senior ML / AI Engineer

Role overview

Qualifications

  • 5+ years building and shipping ML/AI systems in production
  • Strong Python and solid software-engineering fundamentals
  • Practical depth in LLM application patterns
  • Cloud + MLOps experience on AWS

Responsibilities

  • Design and ship agentic systems
  • Build the RAG layer with strict access controls
  • Operate self-hosted inference at scale
  • Own the MLOps / governance plane

Key facts

  • Remote from: United States
  • Full time
  • Senior (5-10 years)
  • AI/ML Engineer
  • English

Hard skills

Other skills

  • Communication
  • Problem Solving
  • Collaboration

About the company

Nonstop Administration & Insurance Services logo

Nonstop Administration & Insurance Services

Health Insurance (Payers)

With Nonstop Health, organizations with over 50 members on healthcare benefits save an average of 12.5% on their healthcare spend – savings for both the employer and employee. Headquartered in the San Francisco Bay Area and Portland, Oregon, Nonstop was founded with the mission to help nonprofits be more sustainable businesses by offering the best benefits and administrative services to hire and retain top talent.

Company details

IndustryHealth Insurance (Payers)
Company size51 - 200

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Job Type Full-time Description

About The Role

We're building the AI/ML enablement backbone for Nonstop Health - a shared platform on which every AI initiative ships as a governed, production-grade service. The flagship work today is an agentic platform for healthcare claims (EOB parsing, claim substantiation against member policy + IRS/MERP rules) built on self-hosted LLMs/VLMs, a multi-agent framework, and a drag-and-drop Agent Studio - all under HIPAA / SOC 2 / ISO 27001 controls. You'll own hard parts of this platform and stand up new initiatives on top of the same foundation.

This is a builder role for someone equally comfortable reasoning about an agent's failure modes, tuning a retrieval pipeline, and making a service survive an RDS failover. You care that AI systems handling sensitive, money-moving decisions are correct, grounded, observable, and defensible - not just demo-able.

What You'll Do

  • Design and ship agentic systems - dynamic tool-calling agents (LangGraph + structured Pydantic outputs) that reason at runtime over a governed tool catalog, with per-step verification, grounding checks, LLM-as-judge, human-in-the-loop, and bounded, auditable loops.Build the RAG layer - document ingestion ? chunking ? embeddings ? hybrid retrieval ? reranking, with strict per-member/per-client access scoping and injection/poisoning defenses.
  • Operate self-hosted inference at scale - vLLM (chat/vision) and an embeddings/rerank service on EKS GPU nodes; optimize throughput (continuous batching, prefix caching, quantization, KV-cache tuning) and enforce fair-share concurrency across tenants.
  • Own the MLOps / governance plane - offline eval harnesses and golden sets, quality scoring, canary/gray-release with auto-rollback, cost/budget governance, and end-to-end observability (OpenTelemetry traces + metrics, immutable audit).
  • Make it production-safe - durable state, circuit breakers, retries, PHI-safe logging/redaction, RBAC, fail-closed defaults, and horizontal scalability on AWS (EKS, RDS/pgvector, SQS, S3/KMS, Cognito) via Terraform.
  • Extend the platform to new domains - take a new AI initiative from problem framing to a governed, evaluated, deployed service on the shared foundation, and mentor engineers on doing the same.
  • Raise the bar - clean, typed, tested code (the platform holds a ruff + mypy + full-test-green bar on every change); thoughtful design reviews; pragmatic trade-offs between autonomy and determinism for high-stakes tasks.
Requirements

What we're looking for (Required)

  • 5+ years building and shipping ML/AI systems in production (not just notebooks or POCs), including hands-on LLM application work in the last 1–2 years.
  • Strong Python (typed, tested, production-grade) and solid software-engineering fundamentals; comfortable across an async web service, a data layer, and infra.
  • Practical depth in LLM application patterns: prompting, structured/ function-calling outputs, RAG (embeddings, vector search, retrieval quality), agent/tool-use loops, and - critically - how to evaluate and de-risk them (grounding, hallucination control, eval sets, guardrails).
  • Cloud + MLOps experience on AWS (or equivalent): containers + Kubernetes, IaC (Terraform), CI/CD, observability, and cost/perf tuning of model-serving.
  • Track record of owning reliability: state durability, failure handling, scaling, and debugging production incidents.
  • Clear written communication and the judgment to make sensible calls under ambiguity.
  • Self-hosting/optimizing open-weight models (vLLM, TGI, or similar) on GPUs; embeddings/rerank serving (e.g., bge / Infinity).
  • LangGraph / LangChain or comparable agent frameworks; multi-agent orchestration.
  • Regulated-data experience - HIPAA / SOC 2 / ISO 27001, PHI/PII handling, RBAC, audit trails; healthcare/insurance/claims domain (X12 835, EOB, benefits) a strong plus.
  • pgvector / Postgres, MongoDB/DocumentDB, SQS, Cognito/OIDC.
  • OpenTelemetry, Prometheus/Grafana; frontend comfort (React/TypeScript) to extend an internal builder/console.
  • Amazon Bedrock or a multi-provider abstraction; on-prem/air-gapped deployment.

Tech You'll Work With

Python · FastAPI · LangGraph · Pydantic · vLLM · pgvector · MongoDB · Postgres · Redis · SQS · React/TypeScript · AWS (EKS + GPU, Cognito, RDS, S3/KMS, SQS) · Terraform · OpenTelemetry · Docker/Kubernetes

Why Join?

You'll work on genuinely hard, high-impact AI - systems that make sensitive decisions and therefore have to be right - with the mandate and the platform to do it well: real governance, real evals, real observability, and a team that treats "governed autonomy" as the goal, not an afterthought.

Compensation

The base salary range for this role is $150,000 - $210,000 depending on experience.

Great benefits aren't just for our clients! We cover 100% of medical, dental and vision benefits for employees and dependents. We also contribute to your retirement goals, with a 401k match up to 4%.

**Must be authorized to work for any employer in the United States. This job does not offer visa sponsorship.

Who We Are

Our main offering is a MERP that, when paired with a HDHP, offers first dollar coverage to employees at a lower premium cost. This means people have access to early care without the worry of a copay, coinsurance, or deductible. Ultimately, early care drives down overall cost and improves health and happiness. We think this is the way healthcare should work! Our primary sales channel is selling into brokers. We have a national presence with a growing book of business and expanded product offerings.

Our team at Nonstop is where our magic really comes to life. We are driven, curious, collaborative, passionate, mission oriented, and most of all - we are real people who love what we do and believe in the real impact we are driving for people all over the US.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

AI/ML Engineer Related jobs

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.