# MLOps Engineer

**Company:** [Fundamental Research Labs](https://hotfix.jobs/companies/fundamental-research-labs)
**Location:** Remote
**Role:** ML Engineering
**Experience:** 5+ years
**Skills:** MLOps, Kubernetes, PyTorch, TensorFlow, Python, Terraform, Helm, Prometheus, Grafana, Datadog, MLflow, torchserve, triton, AWS, GCP
**Posted:** 2026-07-20

> Build and manage scalable MLOps infrastructure, automated ML pipelines, model serving, and monitoring for a Large Tabular Model (LTM) at an enterprise AI company. Requires 5+ years MLOps/DevOps experience with Kubernetes, PyTorch/TensorFlow, and cloud infrastructure.

## Job Description

## Key Responsibilities
- Develop and manage scalable, automated machine learning pipelines, CI/CD workflows, and orchestration frameworks.
- Design and implement robust model serving infrastructure using platforms like TorchServe, TensorFlow, Triton etc.
- Develop scalable inference architectures optimized for ultra-low latency and high throughput.
- Ensure seamless model deployment by implementing A/B testing, canary releases, and rollback capabilities.
- Develop logging, alerting, and monitoring solutions to track model development and reliability.
- Improve GPU usage, enable autoscaling, and streamline resource allocation to boost efficiency.
- Design, implement, and maintain feature stores, robust data pipelines, and scalable storage solutions to efficiently handle large volumes of data.

## Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
- 5+ years of experience as MLOps engineer or DevOps roles, working with MLOps platforms (MLflow, WandB etc.) and frameworks (PyTorch, TensorFlow etc.).
- Experience building and designing MLOps infrastructure from the ground up.
- Experience with model serving frameworks (TorchServe, TensorFlow Serving, Triton, KServe etc.) for high scalability and low latency inference.
- Experience in building and managing data pipelines to support both model training and inference.
- Experience with Kubernetes on a major cloud provider (AWS, GCP, or Azure) and with infrastructure as code (e.g. Terraform, Helm, GitOps).
- Strong software engineering skills in Python, Bash, and Go, with a focus on writing clean, maintainable, and scalable code.
- Experience in AI/ML systems security, compliance, and model governance.
- Proficient with observability and monitoring tools, such as Prometheus, Grafana, Datadog, and OpenTelemetry.

## Nice-to-Haves
- Experience with ML workflow tooling (MLflow, Kubeflow, or similar).
- Experience with FastAPI and Backend applications.
- Familiarity with data platforms like Databricks or Snowflake.
- Exposure to SRE practices or cloud security certifications.
- Hands-on experience with Prometheus, Grafana, or Datadog.

## Similar roles

- [Machine Learning Infrastructure Engineer, Safeguards Research](https://hotfix.jobs/jobs/1b6b27bd-8416-45da-b2b5-eddcaa1d8e2d) - Anthropic - San Francisco, CA - $350k – $500k/yr
- [Applied AI Inference Engineer](https://hotfix.jobs/jobs/c853e6ba-960e-4f07-a0f7-7a1920e5d1ad) - Crusoe - San Francisco, CA - $250k – $300k/yr
- [Perception Engineer](https://hotfix.jobs/jobs/4b0279a4-3c15-4f99-aed7-f216ca5c51ee) - Applied Intuition - Sunnyvale, CA - $180k – $255k/yr
- [Research Scientist / Engineer](https://hotfix.jobs/jobs/063941d9-e414-4af2-b6c8-dff6c2af103c) - Luma AI - Redwood City, CA - $188k – $395k/yr
- [Applied Scientist, AI](https://hotfix.jobs/jobs/5cafa37f-9dbc-4b42-b2e6-624c0e71a688) - Sprinter Health - San Francisco, CA - $180k – $260k/yr

**Apply:** https://hotfix.jobs/jobs/7451307e-56a5-49ac-a072-7fe9370478b0
**Canonical:** https://hotfix.jobs/jobs/7451307e-56a5-49ac-a072-7fe9370478b0