Machine Learning Engineer
Build and operate infrastructure-first production ML systems, including pipelines, model serving, CI/CD, monitoring, and automated retraining. The role requires 5+ years of production ML experience, strong Python and platform engineering skills, and hands-on expertise with cloud, Kubernetes, and MLOps tooling.
About the job
Responsibilities
- Build, deploy, and operate production ML systems supporting financial services and real-time decisioning.
- Develop pipelines and serving infrastructure for predictive models used in consumer decisioning, fraud, churn, transaction intelligence, and related use cases.
- Own production model lifecycle infrastructure, including feature pipelines, deployment, CI/CD, monitoring, and automated retraining.
- Build reusable modeling pipelines, feature engineering systems, model-serving infrastructure, and production-quality tooling in GCP and Kubernetes environments using Terraform and CI/CD.
- Define monitoring, alerting, dashboards, and metrics to detect model drift, degradation, and system issues.
- Automate repetitive ML lifecycle workflows and enable data scientists to deploy, iterate on, and retrain models safely.
- Build low-latency online model-serving services for real-time decisioning.
- Implement model versioning, reproducibility, and progressive rollout strategies such as shadow, canary, and champion-challenger deployments.
- Use AI coding agents to write, test, ship, operate, and debug infrastructure and pipeline code, verifying their output with sound engineering judgment.
- Collaborate with data scientists, analysts, platform engineers, product managers, and business stakeholders.
Requirements
- 5+ years of experience building and operating production ML systems as a Machine Learning Engineer, ML Platform Engineer, MLOps Engineer, Applied Scientist, or similar.
- Strong experience deploying, serving, monitoring, and operating ML models in production, including feature engineering, training/serving parity, retraining, and model diagnostics.
- Hands-on MLOps experience with pipelines, ML CI/CD, Docker, Kubernetes, Terraform, and workflow schedulers such as Airflow.
- Strong software and platform engineering fundamentals.
- Strong Python skills for building pipelines, services, and production tooling; Go or Rust experience is a plus.
- Experience with distributed computing and GPU-accelerated workloads, including Spark, Ray, Dask, or distributed training and inference.
- Strong SQL skills and experience with cloud data warehouses and operational databases such as BigQuery and Spanner.
- Experience with observability tools such as Prometheus, Grafana, or Datadog.
- Experience with cloud computing platforms; GCP is preferred.
- Strong written and verbal communication skills.
- Ability to apply critical thinking, abstract reasoning, and engineering judgment to complex technical and business problems.
- Experience replacing manual ML workflows with durable automation.
- Willingness to work across engineering, infrastructure, and ML execution as business needs require.
Nice-to-haves
- STEM degree in Computer Science, Statistics, Economics, Mathematics, Engineering, Physics, Operations Research, or a related quantitative field.
- Experience with service meshes such as Istio and low-latency gRPC or microservice serving.
- Experience directing AI coding agents such as Claude Code or Cursor.
- Experience supporting credit decisioning, risk modeling, fraud, churn, or consumer behavior modeling.
- Familiarity with model explainability, auditability, and compliance for regulated decisioning and fintech.
- Experience with Go or Rust.
Skills
Python, MLOps, Kubernetes, Docker, Terraform, Airflow, GCP, SQL, BigQuery, Spanner, Prometheus, Grafana, Spark, Ray
Similar jobs
ML Engineering jobsBuild and operate the engineering systems that support post-training research, including reinforcement learning infrastructure, sandboxed execution, data pipelines, and agent scaffolding. The role requires strong Python and systems engineering skills, project ownership, and a relevant bachelor’s degree or equivalent experience.
Build research infrastructure and tooling that enables AI models to design silicon, including reinforcement learning environments, EDA integrations, evaluations, and experiment workflows. The role requires strong software engineering fundamentals and comfort working across research, tooling, and chip-design systems.
Build production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Develop and deploy machine learning models for biomedical research and AI products, collaborating with scientific, engineering, and product teams. Requires an advanced quantitative degree, substantial ML experience, Python proficiency, and experience bringing models into production or research applications.