Skip to content
AttainAttain

Machine Learning Engineer

Build and operate infrastructure-first production ML systems, including pipelines, model serving, CI/CD, monitoring, and automated retraining. The role requires 5+ years of production ML experience, strong Python and platform engineering skills, and hands-on expertise with cloud, Kubernetes, and MLOps tooling.

About the job

Responsibilities

  • Build, deploy, and operate production ML systems supporting financial services and real-time decisioning.
  • Develop pipelines and serving infrastructure for predictive models used in consumer decisioning, fraud, churn, transaction intelligence, and related use cases.
  • Own production model lifecycle infrastructure, including feature pipelines, deployment, CI/CD, monitoring, and automated retraining.
  • Build reusable modeling pipelines, feature engineering systems, model-serving infrastructure, and production-quality tooling in GCP and Kubernetes environments using Terraform and CI/CD.
  • Define monitoring, alerting, dashboards, and metrics to detect model drift, degradation, and system issues.
  • Automate repetitive ML lifecycle workflows and enable data scientists to deploy, iterate on, and retrain models safely.
  • Build low-latency online model-serving services for real-time decisioning.
  • Implement model versioning, reproducibility, and progressive rollout strategies such as shadow, canary, and champion-challenger deployments.
  • Use AI coding agents to write, test, ship, operate, and debug infrastructure and pipeline code, verifying their output with sound engineering judgment.
  • Collaborate with data scientists, analysts, platform engineers, product managers, and business stakeholders.

Requirements

  • 5+ years of experience building and operating production ML systems as a Machine Learning Engineer, ML Platform Engineer, MLOps Engineer, Applied Scientist, or similar.
  • Strong experience deploying, serving, monitoring, and operating ML models in production, including feature engineering, training/serving parity, retraining, and model diagnostics.
  • Hands-on MLOps experience with pipelines, ML CI/CD, Docker, Kubernetes, Terraform, and workflow schedulers such as Airflow.
  • Strong software and platform engineering fundamentals.
  • Strong Python skills for building pipelines, services, and production tooling; Go or Rust experience is a plus.
  • Experience with distributed computing and GPU-accelerated workloads, including Spark, Ray, Dask, or distributed training and inference.
  • Strong SQL skills and experience with cloud data warehouses and operational databases such as BigQuery and Spanner.
  • Experience with observability tools such as Prometheus, Grafana, or Datadog.
  • Experience with cloud computing platforms; GCP is preferred.
  • Strong written and verbal communication skills.
  • Ability to apply critical thinking, abstract reasoning, and engineering judgment to complex technical and business problems.
  • Experience replacing manual ML workflows with durable automation.
  • Willingness to work across engineering, infrastructure, and ML execution as business needs require.

Nice-to-haves

  • STEM degree in Computer Science, Statistics, Economics, Mathematics, Engineering, Physics, Operations Research, or a related quantitative field.
  • Experience with service meshes such as Istio and low-latency gRPC or microservice serving.
  • Experience directing AI coding agents such as Claude Code or Cursor.
  • Experience supporting credit decisioning, risk modeling, fraud, churn, or consumer behavior modeling.
  • Familiarity with model explainability, auditability, and compliance for regulated decisioning and fintech.
  • Experience with Go or Rust.

Skills

Python, MLOps, Kubernetes, Docker, Terraform, Airflow, GCP, SQL, BigQuery, Spanner, Prometheus, Grafana, Spark, Ray

Thinking Machines Lab

Thinking Machines Lab

San Francisco, CA

Research Software Engineer, Post Training
$350k+/yrHybridML Engineering

Build and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, AI for Chip Design
$266k+/yrHybridML Engineering

Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.

Rollstack

Rollstack

United States
AI Software Engineer
No salary listedRemote3+ YOEML Engineering

Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.

ClickUp

ClickUp

United States

Machine Learning Engineer, Ranking & Retrieval
$200k+/yrRemote5+ YOEML Engineering

Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.

PathAI

PathAI

Boston, MA
Machine Learning Engineer III
$131k+/yrOn-site5+ YOEML Engineering

Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.