Logo for Spark Tek Inc

Machine Learning Operations Engineer, AWS Stack

Role overview

Qualifications

  • Bachelor's degree in computer science, engineering, data science, information systems, or a related technical field
  • 3+ years of experience in machine learning engineering, MLOps, cloud engineering, data engineering, DevOps, or production analytics support
  • Practical experience working with AWS services used for machine learning or data workflows
  • Strong Python skills

Responsibilities

  • Support the deployment and day-to-day operation of machine learning and computer vision models
  • Partner with data scientists and machine learning engineers to package approved models for production use
  • Build and maintain practical AWS-based workflows for data movement, model execution, batch inference, and output delivery
  • Help create repeatable deployment processes so models can move from development to production

About the company

Spark Tek Inc logo

Spark Tek Inc

IT Services & IT Consulting

We are a team of experts with deep knowledge of various business functions. We aim to provide customized solutions that can be aligned with our customer demands. We provide business consulting services to help you build high performing organizations which are scalable and sustainable. Our reliable solutions have proven success across varied industries along with different growth stages. We are fast expanding industry and domain experts with in depth experience in leadership and global organizations.

Company details

Company typeStartup
IndustryIT Services & IT Consulting
Company size11 - 50

Your match analysis

See how your profile stacks up against this role.

We compared the job requirements to your profile to show where you're strong and where you fall short.

Job description

Role: Machine Learning Operations Engineer, AWS Stack

Duration: Long Term Contract

Location: Bay Area, CA – Remote

Department Overview

The computer vision team develops machine learning solutions that convert aerial inspection imagery into actionable intelligence for PG&E. The team works cross-functionally across product, inspection, data science, machine learning engineering, cloud platform, and business stakeholders to deliver scalable analytics products that support safer operations, better asset visibility, and more informed decisions. The team combines practical model development, AWS-based deployment, structured change management, and production support discipline to move models from concept to operational use.

Position Summary

PG&E is seeking a Machine Learning Operations Engineer with practical AWS experience to help deploy, monitor, and support machine learning and computer vision solutions in production. In this role, you will work with data scientists, machine learning engineers, product teams, and business stakeholders to turn approved models into reliable, repeatable, and well-documented production workflows. The ideal candidate understands the basics of machine learning, enjoys building dependable cloud-based processes, and can communicate clearly with both technical and non-technical teams.

What You'll Do

  • Support the deployment and day-to-day operation of machine learning and computer vision models used for inspection and asset intelligence use cases, including overhead equipment inspection and unauthorized attachment detection.
  • Partner with data scientists and machine learning engineers to package approved models for production use and make sure model handoffs are clear, tested, and documented.
  • Build and maintain practical AWS-based workflows for data movement, model execution, batch inference, and output delivery using services such as Amazon S3, SageMaker, Lambda, Step Functions, CloudWatch, and related AWS tools.
  • Help create repeatable deployment processes so models can move from development to testing to production in a controlled and consistent way.
  • Support CI/CD practices for machine learning workflows, including code versioning, automated checks, deployment readiness steps, and release coordination.
  • Monitor production model runs for job completion, data issues, system errors, performance changes, and operational readiness.
  • Assist with troubleshooting production inference issues by reviewing logs, validating inputs and outputs, coordinating fixes, and communicating status to stakeholders.
  • Maintain clear runbooks, deployment notes, monitoring summaries, and support documentation so production workflows can be operated consistently by the broader team.
  • Work with product managers, SMEs, data teams, cloud platform teams, and business stakeholders to align on production requirements, release timing, support needs, and success measures.
  • Help improve reliability, scalability, security, and cost awareness for machine learning workloads without over-engineering the solution.

What You Bring

  • Bachelor's degree in computer science, engineering, data science, information systems, or a related technical field, or equivalent combination of education and relevant experience.
  • 3+ years of experience in machine learning engineering, MLOps, cloud engineering, data engineering, DevOps, or production analytics support.
  • Practical experience working with AWS services used for machine learning or data workflows, such as Amazon S3, SageMaker, Lambda, Step Functions, CloudWatch, IAM, ECR, ECS, or related services.
  • Strong Python skills and comfort working with scripts, APIs, logs, configuration files, and version-controlled repositories.
  • Understanding of how machine learning models move from development into production, including model packaging, testing, deployment, monitoring, and support.
  • Experience supporting batch processing, inference pipelines, data validation, or production data workflows.
  • Familiarity with CI/CD concepts, source control, deployment coordination, and basic release management practices.
  • Ability to troubleshoot issues across data, code, cloud services, permissions, and operational workflows.
  • Ability to work across cross-functional teams and explain technical issues clearly to technical and business stakeholders.
  • Strong analytical, problem-solving, documentation, and communication skills.

Desired Qualifications

  • Experience with computer vision, image-based analytics, inspection workflows, or large-scale image datasets.
  • Experience with Docker, container-based deployments, or model packaging for production use.
  • Exposure to infrastructure-as-code tools such as Terraform, CloudFormation, or AWS CDK.
  • Experience with model monitoring, data quality checks, operational dashboards, or alerting workflows.
  • Familiarity with ML lifecycle tools such as model registries, experiment tracking, or workflow orchestration.
  • Experience in utility, infrastructure, industrial inspection, or similar analytics environments using image-based data for decision-making is a strong advantage.

Apply once. Then go straight to the hiring manager.

After you apply, unlock the direct contact details of the people who actually make the call. A quick follow-up makes you 5x more likely to land an interview.

MR

Marcus Rivera

Chief Revenue Officer

m.rivera@company.com
linkedin.com/in/marcusrivera
Unlocked after you apply
·

Machine Learning Engineer Related jobs

Other jobs at Spark Tek Inc

Premium

Reach out to the hiring manager directly.

Gain access to the contact details of the hiring managers who actually decide, and reach out to network with them directly. That, plus more when you upgrade:

  • Full match report with fit score and gaps
  • Career diagnostics on how recruiters read you
  • Curated company matches and warm intros
  • 48h early access to new roles

Cancel anytime.