# Machine Learning Engineer

**Company:** [Attain](https://hotfix.jobs/companies/attain)
**Location:** Chicago, IL, New York, NY, Redwood City, CA
**Role:** ML Engineering
**Experience:** 5+ years
**Skills:** Python, MLOps, Kubernetes, Docker, Terraform, Airflow, GCP, SQL, BigQuery, Spanner, Prometheus, Grafana, Spark, Ray
**Posted:** 2026-09-02

> Build and operate infrastructure-first production ML systems, including pipelines, model serving, CI/CD, monitoring, and automated retraining. The role requires 5+ years of production ML experience, strong Python and platform engineering skills, and hands-on expertise with cloud, Kubernetes, and MLOps tooling.

## Job Description

## Responsibilities
- Build, deploy, and operate production ML systems supporting financial services and real-time decisioning.
- Develop pipelines and serving infrastructure for predictive models used in consumer decisioning, fraud, churn, transaction intelligence, and related use cases.
- Own production model lifecycle infrastructure, including feature pipelines, deployment, CI/CD, monitoring, and automated retraining.
- Build reusable modeling pipelines, feature engineering systems, model-serving infrastructure, and production-quality tooling in GCP and Kubernetes environments using Terraform and CI/CD.
- Define monitoring, alerting, dashboards, and metrics to detect model drift, degradation, and system issues.
- Automate repetitive ML lifecycle workflows and enable data scientists to deploy, iterate on, and retrain models safely.
- Build low-latency online model-serving services for real-time decisioning.
- Implement model versioning, reproducibility, and progressive rollout strategies such as shadow, canary, and champion-challenger deployments.
- Use AI coding agents to write, test, ship, operate, and debug infrastructure and pipeline code, verifying their output with sound engineering judgment.
- Collaborate with data scientists, analysts, platform engineers, product managers, and business stakeholders.

## Requirements
- 5+ years of experience building and operating production ML systems as a Machine Learning Engineer, ML Platform Engineer, MLOps Engineer, Applied Scientist, or similar.
- Strong experience deploying, serving, monitoring, and operating ML models in production, including feature engineering, training/serving parity, retraining, and model diagnostics.
- Hands-on MLOps experience with pipelines, ML CI/CD, Docker, Kubernetes, Terraform, and workflow schedulers such as Airflow.
- Strong software and platform engineering fundamentals.
- Strong Python skills for building pipelines, services, and production tooling; Go or Rust experience is a plus.
- Experience with distributed computing and GPU-accelerated workloads, including Spark, Ray, Dask, or distributed training and inference.
- Strong SQL skills and experience with cloud data warehouses and operational databases such as BigQuery and Spanner.
- Experience with observability tools such as Prometheus, Grafana, or Datadog.
- Experience with cloud computing platforms; GCP is preferred.
- Strong written and verbal communication skills.
- Ability to apply critical thinking, abstract reasoning, and engineering judgment to complex technical and business problems.
- Experience replacing manual ML workflows with durable automation.
- Willingness to work across engineering, infrastructure, and ML execution as business needs require.

## Nice-to-haves
- STEM degree in Computer Science, Statistics, Economics, Mathematics, Engineering, Physics, Operations Research, or a related quantitative field.
- Experience with service meshes such as Istio and low-latency gRPC or microservice serving.
- Experience directing AI coding agents such as Claude Code or Cursor.
- Experience supporting credit decisioning, risk modeling, fraud, churn, or consumer behavior modeling.
- Familiarity with model explainability, auditability, and compliance for regulated decisioning and fintech.
- Experience with Go or Rust.

## Similar jobs

- [Research Software Engineer, Post Training](https://hotfix.jobs/jobs/168e3c8f-8577-4482-bf94-91b3d11744ba) - Thinking Machines Lab - San Francisco, CA - $350k – $475k/yr
- [Software Engineer, AI for Chip Design](https://hotfix.jobs/jobs/abfb017d-1ce2-4d24-901d-16084eb7b3bc) - OpenAI - San Francisco, CA - $266k – $468k/yr
- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Machine Learning Engineer, Ranking & Retrieval](https://hotfix.jobs/jobs/adb5411d-8bfc-48b8-a8e7-15cbf5ea21b2) - ClickUp - Remote - $200k – $250k/yr
- [Machine Learning Engineer III](https://hotfix.jobs/jobs/2359be26-5005-4fa7-94c9-8a86066a6bb5) - PathAI - Boston, MA - $131k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/25640abd-aa1f-49d9-b4cb-74804bb13328
**Canonical:** https://hotfix.jobs/jobs/25640abd-aa1f-49d9-b4cb-74804bb13328