Build and manage scalable MLOps infrastructure, automated ML pipelines, model serving, and monitoring for a Large Tabular Model (LTM) at an enterprise AI company. Requires 5+ years MLOps/DevOps experience with Kubernetes, PyTorch/TensorFlow, and cloud infrastructure.
Salary not listed
Remote5+ YOEML Engineering
About the role
Key Responsibilities
Develop and manage scalable, automated machine learning pipelines, CI/CD workflows, and orchestration frameworks.
Design and implement robust model serving infrastructure using platforms like TorchServe, TensorFlow, Triton etc.
Develop scalable inference architectures optimized for ultra-low latency and high throughput.
Ensure seamless model deployment by implementing A/B testing, canary releases, and rollback capabilities.
Develop logging, alerting, and monitoring solutions to track model development and reliability.
Improve GPU usage, enable autoscaling, and streamline resource allocation to boost efficiency.
Design, implement, and maintain feature stores, robust data pipelines, and scalable storage solutions to efficiently handle large volumes of data.
Requirements
Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
5+ years of experience as MLOps engineer or DevOps roles, working with MLOps platforms (MLflow, WandB etc.) and frameworks (PyTorch, TensorFlow etc.).
Experience building and designing MLOps infrastructure from the ground up.
Experience with model serving frameworks (TorchServe, TensorFlow Serving, Triton, KServe etc.) for high scalability and low latency inference.
Experience in building and managing data pipelines to support both model training and inference.
Experience with Kubernetes on a major cloud provider (AWS, GCP, or Azure) and with infrastructure as code (e.g. Terraform, Helm, GitOps).
Strong software engineering skills in Python, Bash, and Go, with a focus on writing clean, maintainable, and scalable code.
Experience in AI/ML systems security, compliance, and model governance.
Proficient with observability and monitoring tools, such as Prometheus, Grafana, Datadog, and OpenTelemetry.
Nice-to-Haves
Experience with ML workflow tooling (MLflow, Kubeflow, or similar).
Experience with FastAPI and Backend applications.
Familiarity with data platforms like Databricks or Snowflake.
Exposure to SRE practices or cloud security certifications.
Hands-on experience with Prometheus, Grafana, or Datadog.
Machine Learning Infrastructure Engineer, Safeguards Research
AnthropicSan Francisco, CA +1
Build and own ML infrastructure, data pipelines, and tooling for Safeguards research at Anthropic. Focus on fast researcher iteration for training/evaluating lightweight detectors on model internals while ensuring correctness at scale. Requires strong Python, distributed systems, and production infrastructure experience.
350k – 500k/yr
Hybrid5+ YOEML Engineering
Applied AI Inference Engineer
CrusoeSan Francisco, CA +1
Build and optimize the end-to-end LLM inference stack for production deployments. Profile and tune serving frameworks (vLLM/SGLang) and CUDA kernels for latency, throughput and cost; partner directly with customer teams to take workloads from POC to monitored production.
250k – 300k/yr
On-site5+ YOEML Engineering
Perception Engineer
Applied IntuitionSunnyvale, CA
Perception Engineer owning outcomes for autonomous mining vehicles. Responsible for sensor selection, model adaptation to new sites/domains, diagnosing failures, data strategies, and translating customer needs into technical KPIs and solutions. Requires strong systems understanding of perception/full stack and real-world deployment experience.
180k – 255k/yr
On-site5+ YOEML Engineering
Research Scientist / Engineer
Luma AIRedwood City, CA
Build and scale distributed reinforcement learning infrastructure for post-training large multimodal foundation models, including rollout generation, environments, rewards, and evaluation systems for agentic tasks.
188k – 395k/yr
Hybrid5+ YOEML Engineering
Applied Scientist, AI
Sprinter HealthSan Francisco, CA
Build, evaluate, and productionize ML/AI models (including LLMs and NLP) that solve ambiguous healthcare, product, and operational problems at Sprinter Health. Requires strong experimentation, error analysis, stakeholder collaboration with clinicians, and focus on real-world impact, bias, and evaluation.