Skip to content

MLOps Engineer

Build and manage scalable MLOps infrastructure, automated ML pipelines, model serving, and monitoring for a Large Tabular Model (LTM) at an enterprise AI company. Requires 5+ years MLOps/DevOps experience with Kubernetes, PyTorch/TensorFlow, and cloud infrastructure.

About the job

Key Responsibilities

  • Develop and manage scalable, automated machine learning pipelines, CI/CD workflows, and orchestration frameworks.
  • Design and implement robust model serving infrastructure using platforms like TorchServe, TensorFlow, Triton etc.
  • Develop scalable inference architectures optimized for ultra-low latency and high throughput.
  • Ensure seamless model deployment by implementing A/B testing, canary releases, and rollback capabilities.
  • Develop logging, alerting, and monitoring solutions to track model development and reliability.
  • Improve GPU usage, enable autoscaling, and streamline resource allocation to boost efficiency.
  • Design, implement, and maintain feature stores, robust data pipelines, and scalable storage solutions to efficiently handle large volumes of data.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
  • 5+ years of experience as MLOps engineer or DevOps roles, working with MLOps platforms (MLflow, WandB etc.) and frameworks (PyTorch, TensorFlow etc.).
  • Experience building and designing MLOps infrastructure from the ground up.
  • Experience with model serving frameworks (TorchServe, TensorFlow Serving, Triton, KServe etc.) for high scalability and low latency inference.
  • Experience in building and managing data pipelines to support both model training and inference.
  • Experience with Kubernetes on a major cloud provider (AWS, GCP, or Azure) and with infrastructure as code (e.g. Terraform, Helm, GitOps).
  • Strong software engineering skills in Python, Bash, and Go, with a focus on writing clean, maintainable, and scalable code.
  • Experience in AI/ML systems security, compliance, and model governance.
  • Proficient with observability and monitoring tools, such as Prometheus, Grafana, Datadog, and OpenTelemetry.

Nice-to-Haves

  • Experience with ML workflow tooling (MLflow, Kubeflow, or similar).
  • Experience with FastAPI and Backend applications.
  • Familiarity with data platforms like Databricks or Snowflake.
  • Exposure to SRE practices or cloud security certifications.
  • Hands-on experience with Prometheus, Grafana, or Datadog.

Skills

MLOps, Kubernetes, PyTorch, TensorFlow, Python, Terraform, Helm, Prometheus, Grafana, Datadog, MLflow, Torchserve, Triton, AWS, GCP

Protege

Protege

Remote

AI Engineer - New Verticals
No salary listedRemote3+ YOEML Engineering

Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.

Payabli

Payabli

Remote

Staff Machine Learning Engineer
No salary listedRemote8+ YOEML Engineering

Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.