ML Operations Engineer
Build and operate scalable MLOps infrastructure for distributed training and inference across on-premises and GPU environments. The role partners with data science and AI teams and requires production experience with Python, Spark, Docker, Kubernetes, and open-source MLOps tooling.
About the job
Responsibilities
- Partner with Data Science and AI Engineering teams to adopt MLOps best practices and migrate training and inference workloads onto the platform.
- Build and scale a high-performance machine learning and AI platform across on-premises data centers and GPU resources.
- Implement and maintain model tracking, versioning, experiment management, and model-performance/drift observability.
- Build CI/CD pipelines for ML/AI artifacts, including model registries, container-image pipelines, and automated retraining triggers.
- Improve platform reliability, cost efficiency, and codebase maintainability.
- Establish monitoring and observability for GPU utilization, model latency, throughput, and custom ML metrics.
- Design and operate ML/AI deployment infrastructure, including GPU cluster architecture, model serving, and tooling for training and inference.
- Build and maintain infrastructure for LLM and generative AI workloads, including fine-tuning pipelines, vector databases, RAG architectures, and inference optimization.
- Collaborate with business units and Product on ML/AI-driven feature development.
- Manage project priorities, deadlines, and deliverables.
Requirements
- Bachelor's degree in Computer Science or a similar technical field, or equivalent practical experience.
- Strong software engineering skills in complex, distributed, multi-language systems; Python preferred.
- Production experience with Spark, Docker, and Kubernetes.
- Experience building and operating end-to-end distributed systems.
- Experience developing and maintaining ML systems using open-source MLOps tools such as MLflow, Argo, Metaflow, Airflow, or Kubeflow.
- Strong understanding of software testing, benchmarking, and CI/CD practices.
- Solid understanding of Linux systems administration.
- Ability to work closely with data scientists and understand tools and workflows such as Jupyter, notebooks, and experiment tracking.
- Enthusiasm for learning and keeping up with the evolving ML/AI infrastructure landscape.
Nice-to-Haves
- Familiarity with LLM and generative AI tooling and concepts, including model-serving frameworks, embeddings, vector stores, and RAG pipelines.
- Familiarity with GPU infrastructure, including CUDA fundamentals, GPU scheduling and orchestration in Kubernetes, and NVIDIA driver, CUDA toolkit, and cuDNN management.
Compensation and Benefits
- Competitive base salary plus performance-based bonus.
- Comprehensive medical insurance, flexible PTO, professional development reimbursement, WiFi reimbursement, and parental leave for Europe-based employees.
- Hybrid-friendly culture with flexible work options.
Skills
Python, Spark, Docker, Kubernetes, MLflow, Argo Workflows, Airflow, Kubeflow, Linux, CI/CD, Prometheus, Grafana, CUDA, Nvidia Drivers, RAG
Similar jobs
ML Engineering jobsBuild the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build and operate machine-learning systems that improve search ranking quality across retrieval and later-stage ranking. The role requires deep search or recommender-systems expertise, production ranking ownership, and at least five years of relevant industry experience.
Develop and deploy scalable machine learning and AI systems for user-facing products, including forecasting and AutoML capabilities. The role requires strong production ML engineering, modeling, software engineering, statistics, and systems knowledge.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.