Skip to content
Fundamental Research LabsFundamental Research LabsUnited States

MLOps Engineer

Build and manage scalable MLOps infrastructure, automated ML pipelines, model serving, and monitoring for a Large Tabular Model (LTM) at an enterprise AI company. Requires 5+ years MLOps/DevOps experience with Kubernetes, PyTorch/TensorFlow, and cloud infrastructure.

Salary not listed
Remote5+ YOEML Engineering

About the role

Key Responsibilities

  • Develop and manage scalable, automated machine learning pipelines, CI/CD workflows, and orchestration frameworks.
  • Design and implement robust model serving infrastructure using platforms like TorchServe, TensorFlow, Triton etc.
  • Develop scalable inference architectures optimized for ultra-low latency and high throughput.
  • Ensure seamless model deployment by implementing A/B testing, canary releases, and rollback capabilities.
  • Develop logging, alerting, and monitoring solutions to track model development and reliability.
  • Improve GPU usage, enable autoscaling, and streamline resource allocation to boost efficiency.
  • Design, implement, and maintain feature stores, robust data pipelines, and scalable storage solutions to efficiently handle large volumes of data.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
  • 5+ years of experience as MLOps engineer or DevOps roles, working with MLOps platforms (MLflow, WandB etc.) and frameworks (PyTorch, TensorFlow etc.).
  • Experience building and designing MLOps infrastructure from the ground up.
  • Experience with model serving frameworks (TorchServe, TensorFlow Serving, Triton, KServe etc.) for high scalability and low latency inference.
  • Experience in building and managing data pipelines to support both model training and inference.
  • Experience with Kubernetes on a major cloud provider (AWS, GCP, or Azure) and with infrastructure as code (e.g. Terraform, Helm, GitOps).
  • Strong software engineering skills in Python, Bash, and Go, with a focus on writing clean, maintainable, and scalable code.
  • Experience in AI/ML systems security, compliance, and model governance.
  • Proficient with observability and monitoring tools, such as Prometheus, Grafana, Datadog, and OpenTelemetry.

Nice-to-Haves

  • Experience with ML workflow tooling (MLflow, Kubeflow, or similar).
  • Experience with FastAPI and Backend applications.
  • Familiarity with data platforms like Databricks or Snowflake.
  • Exposure to SRE practices or cloud security certifications.
  • Hands-on experience with Prometheus, Grafana, or Datadog.

Skills

MLOpsKubernetesPyTorchTensorFlowPythonTerraformHelmPrometheusGrafanaDatadogMLflowtorchservetritonAWSGCP

Similar roles

ML Engineering jobs
Anthropic

Machine Learning Infrastructure Engineer, Safeguards Research

AnthropicSan Francisco, CA +1

Build and own ML infrastructure, data pipelines, and tooling for Safeguards research at Anthropic. Focus on fast researcher iteration for training/evaluating lightweight detectors on model internals while ensuring correctness at scale. Requires strong Python, distributed systems, and production infrastructure experience.

350k – 500k/yr
Hybrid5+ YOEML Engineering
Crusoe

Applied AI Inference Engineer

CrusoeSan Francisco, CA +1

Build and optimize the end-to-end LLM inference stack for production deployments. Profile and tune serving frameworks (vLLM/SGLang) and CUDA kernels for latency, throughput and cost; partner directly with customer teams to take workloads from POC to monitored production.

250k – 300k/yr
On-site5+ YOEML Engineering
Applied Intuition

Perception Engineer

Applied IntuitionSunnyvale, CA

Perception Engineer owning outcomes for autonomous mining vehicles. Responsible for sensor selection, model adaptation to new sites/domains, diagnosing failures, data strategies, and translating customer needs into technical KPIs and solutions. Requires strong systems understanding of perception/full stack and real-world deployment experience.

180k – 255k/yr
On-site5+ YOEML Engineering
Luma AI

Research Scientist / Engineer

Luma AIRedwood City, CA

Build and scale distributed reinforcement learning infrastructure for post-training large multimodal foundation models, including rollout generation, environments, rewards, and evaluation systems for agentic tasks.

188k – 395k/yr
Hybrid5+ YOEML Engineering
Sprinter Health

Applied Scientist, AI

Sprinter HealthSan Francisco, CA

Build, evaluate, and productionize ML/AI models (including LLMs and NLP) that solve ambiguous healthcare, product, and operational problems at Sprinter Health. Requires strong experimentation, error analysis, stakeholder collaboration with clinicians, and focus on real-world impact, bias, and evaluation.

180k – 260k/yr
HybridML Engineering