Skip to content
ConfidoConfido

Senior ML Ops Engineer

Be the first dedicated owner of Confido's ML platform, owning end-to-end ML pipelines, infrastructure for training/inference/agentic workloads, and providing reproducible environments for the AI/ML team in a fast-growing CPG AI startup.

About the job

What you'll do

  • Own ML pipelines end to end — experimentation to production — and the infrastructure behind training, inference, and agentic workloads
  • Give the AI/ML team a paved road: reproducible environments and fast paths from prototype to production, so they can try new models and agents without fighting the infra
  • Stand up the cloud foundation as Infrastructure as Code and the CI/CD that ships ML safely
  • Serve and optimize inference and forecasting workloads — latency, throughput, and cost — and the data streams feeding them (e.g. turning a heavy synchronous model call into an async, parallelized one)
  • Own the data interface with data engineering: serve the right data to models and agents, and write their outputs back into the platform's data systems for the rest of Confido to use
  • Make reliability, observability, security, and privacy the default — and keep model and agent quality measurable in production through online evals and human-in-the-loop review, not just uptime

Requirements

  • 5+ years in MLOps, ML platform, AI infrastructure, or platform engineering — on production ML systems, not pipelines on paper
  • Live at the seam of software and infrastructure: equally at home writing production code and standing up cloud infra
  • Driven a real pipeline end to end and can walk through it: the architecture, the security and cost trade-offs, and what you'd change
  • Deep cloud infrastructure understanding, distributed data systems, and IaC — you can boot an environment from scratch, wire CI/CD, and run containerized workloads in production without hand-holding
  • Strong Python and comfort in a production app codebase (Ruby, Java)
  • Monitoring, security, and cost are instincts, not afterthoughts
  • High ownership in a fast-moving startup, and experience productionizing what research/AI teams build

Nice to have

  • LLMOps tooling — tracing, prompt/version management, eval harnesses
  • Inference optimization (vLLM, ONNX, TensorRT) and GPU / spot-instance economics
  • ML platform and orchestration tooling (MLflow, BentoML, Ray, Airflow)
  • Large-scale data systems (Snowflake, Kafka) and vector databases
  • Managed ML services (Bedrock, SageMaker, Vertex AI)
  • Multimodal or generative AI in production

Stack

Python · Ruby/Rails · AWS · Terraform · Kubernetes · GitHub Actions · Snowflake · Aurora/RDS · Redis · Kafka

Perks + Benefits

  • Equity — own a piece of what you're building
  • Fully paid health coverage with Aetna (we cover 100% of premiums)
  • Top-tier dental and vision through Guardian
  • 12 weeks paid parental leave
  • Unlimited PTO, plus regular 4-day holiday weekends we actually take
  • 401(k) through Vestwell
  • Paid relocation — we'll get you here
  • Full desk setup on day one (laptop, monitor, keyboard) + a $200 stipend to make it yours
  • Catered Friday lunches, team dinners on us, and unlimited coffee + snacks featuring our own brands

Skills

MLOps, Python, AWS, Terraform, Kubernetes, MLflow, Bentoml, Ray, Airflow, Snowflake, Kafka, vLLM, Onnx, TensorRT, SageMaker

Fetch

Fetch

United States

Senior Machine Learning Engineer II
$211k+/yrRemote6+ YOEML Engineering

Build and operate low-latency machine learning systems for ad ranking, relevance, and optimization, including feature pipelines, experimentation, evaluation, and production inference. The role requires 6+ years of software engineering experience, strong Python skills, AWS experience, and practical LLM application experience.

Checkr

Checkr

San Francisco, CA

Senior Machine Learning Engineer
$207k+/yrOn-site6+ YOEML Engineering

Build and operate production ML and AI services using Python, LLM APIs, and robust software engineering practices. The role requires 6+ years of professional software experience, including production ML systems, and partners closely with product and engineering teams.

Front

Front

San Francisco, CA

Senior Applied AI Engineer
$205k+/yrHybrid5+ YOEML Engineering

Build and deploy generative AI and LLM-powered agentic applications at Front to automate customer support inquiries, enhance product capabilities, and drive operational insights. Requires 5+ years software engineering experience with strong production AI/ML focus, agentic/RAG expertise, and proficiency in Node.js, TS, and Python.

Reddit

Reddit

Ontario, Canada

Senior Machine Learning Engineer, Ads
$217k+/yrRemote5+ YOEML Engineering

Design, build, and deploy production ML systems for recommendations, search, ranking, and advertising at internet scale. Own the full ML lifecycle from modeling to monitoring with strong cross-functional collaboration.

Instacart

Instacart

United States
Senior Machine Learning Engineer, Digital Twin Platform
$201k+/yrRemote5+ YOEML Engineering

Develop and deploy production machine learning models for real-time inventory and shelf-stocking intelligence at scale. The role requires 5+ years of production ML experience, strong Python and ML framework skills, cloud and data pipeline expertise, and a bachelor's degree or equivalent experience.