MLOps Engineer
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
About the job
Responsibilities
- Build and evolve the MLOps platform and CI/CD workflows, including experiment tracking, model registries, packaging, automated training and retraining, deployment, and safe rollout and rollback.
- Operate reliable, scalable batch, streaming, and real-time model-serving infrastructure for models and digital twins.
- Build ML data and feature pipelines for telemetry, process and knowledge graphs, images, time-series data, conversations, and other production sources.
- Maintain feature-store capabilities and a lakehouse foundation using Apache Iceberg on Amazon S3, with data quality, lineage, versioning, and reproducibility.
- Build model observability and human-in-the-loop systems that capture production signals and expert corrections as ground truth for evaluation and retraining.
- Develop standardized tooling and workflows for reliable, reproducible ML development.
- Drive infrastructure, reliability, scalability, architecture, and technical direction from design through production.
- Collaborate with Process, Chemical, Materials, Simulation, Software, Data, and AI engineers.
Requirements
- 5+ years of relevant industry experience building production software, infrastructure, data, or machine learning systems.
- Experience building and operating machine learning systems in production with an MLOps or DevOps orientation.
- Strong DevOps fundamentals, including CI/CD, containers, Kubernetes, cloud infrastructure, and infrastructure as code.
- Proficiency in Python and SQL.
- Hands-on experience with MLflow or similar experiment-tracking, model-registry, and model-lifecycle tooling.
- Experience with Amazon S3, lakehouse technologies such as Apache Iceberg, and workflow orchestration tools such as Airflow or Dagster.
- Experience building pipelines for multimodal ML data, including images, time-series, structured, and semi-structured data.
- Familiarity with manufacturing systems, sensors, process automation, or other physical-world data systems.
- Strong problem-solving, data-quality, reliability, and technical communication skills.
- Bachelor's or master's degree in Computer Science, Data Engineering, Data Science, a related STEM field, or equivalent practical experience.
Nice-to-Haves
- Feature stores, human-in-the-loop systems, active learning, or data-labeling infrastructure.
- Robotics or robotic automation, including sensors, vision systems, or robotics data.
- Operating ML systems in manufacturing or other physical-world environments.
- Internal tools for expert feedback, labeling, model evaluation, or model interaction.
- Shared ML infrastructure or platforms used across multiple teams or applications.
Compensation
- Salary range: $200,000–$250,000 USD.
- Compensation also includes equity and benefits.
Skills
MLOps, DevOps, CI/CD, Kubernetes, Docker, Python, SQL, MLflow, Amazon S3, Apache Iceberg, Airflow, Dagster, Feature Stores, Infrastructure As Code
Similar jobs
ML Engineering jobsBuild and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.
Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.
Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.
Build trustworthy infrastructure for production LLM agents, closed-loop evaluation, and autonomous research workflows. The role requires strong Python and distributed-systems experience, hands-on LLM post-training and inference knowledge, and experience operating agent systems at scale.