MLOps Engineer
Build and manage scalable MLOps infrastructure, automated ML pipelines, model serving, and monitoring for a Large Tabular Model (LTM) at an enterprise AI company. Requires 5+ years MLOps/DevOps experience with Kubernetes, PyTorch/TensorFlow, and cloud infrastructure.
About the job
Key Responsibilities
- Develop and manage scalable, automated machine learning pipelines, CI/CD workflows, and orchestration frameworks.
- Design and implement robust model serving infrastructure using platforms like TorchServe, TensorFlow, Triton etc.
- Develop scalable inference architectures optimized for ultra-low latency and high throughput.
- Ensure seamless model deployment by implementing A/B testing, canary releases, and rollback capabilities.
- Develop logging, alerting, and monitoring solutions to track model development and reliability.
- Improve GPU usage, enable autoscaling, and streamline resource allocation to boost efficiency.
- Design, implement, and maintain feature stores, robust data pipelines, and scalable storage solutions to efficiently handle large volumes of data.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
- 5+ years of experience as MLOps engineer or DevOps roles, working with MLOps platforms (MLflow, WandB etc.) and frameworks (PyTorch, TensorFlow etc.).
- Experience building and designing MLOps infrastructure from the ground up.
- Experience with model serving frameworks (TorchServe, TensorFlow Serving, Triton, KServe etc.) for high scalability and low latency inference.
- Experience in building and managing data pipelines to support both model training and inference.
- Experience with Kubernetes on a major cloud provider (AWS, GCP, or Azure) and with infrastructure as code (e.g. Terraform, Helm, GitOps).
- Strong software engineering skills in Python, Bash, and Go, with a focus on writing clean, maintainable, and scalable code.
- Experience in AI/ML systems security, compliance, and model governance.
- Proficient with observability and monitoring tools, such as Prometheus, Grafana, Datadog, and OpenTelemetry.
Nice-to-Haves
- Experience with ML workflow tooling (MLflow, Kubeflow, or similar).
- Experience with FastAPI and Backend applications.
- Familiarity with data platforms like Databricks or Snowflake.
- Exposure to SRE practices or cloud security certifications.
- Hands-on experience with Prometheus, Grafana, or Datadog.
Skills
MLOps, Kubernetes, PyTorch, TensorFlow, Python, Terraform, Helm, Prometheus, Grafana, Datadog, MLflow, Torchserve, Triton, AWS, GCP
Similar jobs
ML Engineering jobsBuild the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.