Skip to content

ML Operations Engineer

Build and operate scalable MLOps infrastructure for distributed training and inference across on-premises and GPU environments. The role partners with data science and AI teams and requires production experience with Python, Spark, Docker, Kubernetes, and open-source MLOps tooling.

About the job

Responsibilities

  • Partner with Data Science and AI Engineering teams to adopt MLOps best practices and migrate training and inference workloads onto the platform.
  • Build and scale a high-performance machine learning and AI platform across on-premises data centers and GPU resources.
  • Implement and maintain model tracking, versioning, experiment management, and model-performance/drift observability.
  • Build CI/CD pipelines for ML/AI artifacts, including model registries, container-image pipelines, and automated retraining triggers.
  • Improve platform reliability, cost efficiency, and codebase maintainability.
  • Establish monitoring and observability for GPU utilization, model latency, throughput, and custom ML metrics.
  • Design and operate ML/AI deployment infrastructure, including GPU cluster architecture, model serving, and tooling for training and inference.
  • Build and maintain infrastructure for LLM and generative AI workloads, including fine-tuning pipelines, vector databases, RAG architectures, and inference optimization.
  • Collaborate with business units and Product on ML/AI-driven feature development.
  • Manage project priorities, deadlines, and deliverables.

Requirements

  • Bachelor's degree in Computer Science or a similar technical field, or equivalent practical experience.
  • Strong software engineering skills in complex, distributed, multi-language systems; Python preferred.
  • Production experience with Spark, Docker, and Kubernetes.
  • Experience building and operating end-to-end distributed systems.
  • Experience developing and maintaining ML systems using open-source MLOps tools such as MLflow, Argo, Metaflow, Airflow, or Kubeflow.
  • Strong understanding of software testing, benchmarking, and CI/CD practices.
  • Solid understanding of Linux systems administration.
  • Ability to work closely with data scientists and understand tools and workflows such as Jupyter, notebooks, and experiment tracking.
  • Enthusiasm for learning and keeping up with the evolving ML/AI infrastructure landscape.

Nice-to-Haves

  • Familiarity with LLM and generative AI tooling and concepts, including model-serving frameworks, embeddings, vector stores, and RAG pipelines.
  • Familiarity with GPU infrastructure, including CUDA fundamentals, GPU scheduling and orchestration in Kubernetes, and NVIDIA driver, CUDA toolkit, and cuDNN management.

Compensation and Benefits

  • Competitive base salary plus performance-based bonus.
  • Comprehensive medical insurance, flexible PTO, professional development reimbursement, WiFi reimbursement, and parental leave for Europe-based employees.
  • Hybrid-friendly culture with flexible work options.

Skills

Python, Spark, Docker, Kubernetes, MLflow, Argo Workflows, Airflow, Kubeflow, Linux, CI/CD, Prometheus, Grafana, CUDA, Nvidia Drivers, RAG

Protege

Protege

Remote

AI Engineer - New Verticals
No salary listedRemote3+ YOEML Engineering

Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.

Perplexity

Perplexity

Belgrade, Serbia
Member Of Technical Staff
No salary listedOn-site5+ YOEML Engineering

Build and operate machine-learning systems that improve search ranking quality across retrieval and later-stage ranking. The role requires deep search or recommender-systems expertise, production ranking ownership, and at least five years of relevant industry experience.

Databricks

Databricks

Belgrade, Serbia

Senior Applied AI Engineer
No salary listedOn-site5+ YOEML Engineering

Develop and deploy scalable machine learning and AI systems for user-facing products, including forecasting and AutoML capabilities. The role requires strong production ML engineering, modeling, software engineering, statistics, and systems knowledge.

Payabli

Payabli

Remote

Staff Machine Learning Engineer
No salary listedRemote8+ YOEML Engineering

Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.