Logo for Clera

Founding AI Engineer

Role overview

Qualifications

  • Demonstrable track record shipping production multimodal and computer vision systems
  • Bachelor's degree in Computer Science, Machine Learning, Engineering, or equivalent
  • Experience with applied agentic AI or model orchestration in production settings
  • Strong foundation in CS, ML, or engineering

Responsibilities

  • Build and ship the production agentic Vision-Language Model (VLM) pipeline
  • Own model orchestration and runtime optimization for edge inference
  • Design and build the evaluation harness and data flywheel from scratch
  • Ship real-time voice and video AI interfaces tailored to different end-user profiles

Key facts

Other skills

  • Problem Solving

About the company

Clera logo

Clera

Staffing & Recruiting

Clera is the first AI talent agent: a personal headhunter that acts on behalf of top talent. We connect you to your dream job.

Company details

IndustryStaffing & Recruiting
Company size1 - 10

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

About the Role

We're a seed-stage wearable AI company building an AI co-pilot for skilled field technicians — delivered through industrial smart glasses — that helps workers in high-stakes industries like data centers and energy infrastructure operate at an expert level. Our stack spans edge inference, real-time voice/video, and agentic visual reasoning running on real hardware in demanding environments.

We're looking for a Founding AI Engineer with 1–5 years of experience who has shipped production multimodal and agentic AI systems to real users. This is a hands-on, end-to-end ownership role at the core of our product — not a research or prototyping position.

What You'll Do

  • Build and ship the production agentic Vision-Language Model (VLM) pipeline running on industrial smart glasses — multi-step, tool-using visual-reasoning loops against real customer workflows (SOPs, inspection, field service).

  • Own model orchestration and runtime optimization for edge inference, balancing model quality against latency with graceful degradation across variable connectivity conditions.

  • Design and build the evaluation harness and data flywheel from scratch — failure-mode capture, customer-data fine-tuning loops, and measurable quality improvements each release cycle.

  • Ship real-time voice and video AI interfaces tailored to different end-user profiles: video-heavy, conversational speech, and proactive alerting.

  • Build RAG pipelines for efficient creation and querying of enterprise knowledge bases from field operator data.

  • Drive multimodal model training for on-premise deployments: open-source model SFT, RL post-training, and quantization.

What We're Looking For

Required (Dealbreakers):

  • Demonstrable track record shipping production multimodal and computer vision systems in the VLM era, owning the model layer end-to-end — with hands-on expertise in visual-language and/or video-language VLMs/VLAs.

  • Bachelor's degree in Computer Science, Machine Learning, Engineering, or equivalent — graduated 2018 or later.

  • Willingness to work on-site 5 days/week in San Francisco, CA.

Also Required:

  • Experience with applied agentic AI or model orchestration in production settings.

  • Experience building production AI products at a startup or high-ownership AI team, or relevant big-company experience (AR/smart glasses, real-time video/streaming, on-device/edge ML) paired with a strong builder signal (e.g., early startup, side projects, open-source contributions).

  • Experience with rigorous evaluation methodologies — ground-truth evals, trajectory evals, tool-call accuracy, and regression testing for comparing models and orchestration stacks.

  • Strong foundation in CS, ML, or engineering, or a demonstrated equivalent through shipping history.

Nice to Have:

  • Experience with production AR or wearable AI (e.g., AR headsets, mixed reality platforms) or autonomous driving computer vision.

  • On-prem / self-hosted model deployment, including serving and optimizing open-weight models on customer hardware, or hands-on fine-tuning and deployment of vLLMs.

  • Exposure to industrial domains such as data centers, energy grids, aerospace, or manufacturing.

  • Experience with in-context grounding or RAG against a knowledge base, including tool and knowledge base wiring.

  • Master's degree with a vision or multimodal research component.

Tech Stack Includes: Python, PyTorch, vLLM, Triton, Ray Serve, ONNX, TensorRT, Hugging Face Transformers, LangChain, RAG, RLHF, SFT, Quantization (GPTQ, AWQ), Edge AI, Multimodal LLMs, Vision-Language Models, Docker.

Compensation & Benefits

  • Salary: $180,000 – $240,000 USD annually

  • Early-stage equity as a founding team member

  • Visa sponsorship: Not available

Location

On-site, 5 days/week in San Francisco, CA. This is not a remote or hybrid role.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Artificial Intelligence Engineer Related jobs

Other jobs at Clera

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.