Senior Software Engineer, ML Ops
Designs and scales reliable machine learning infrastructure, automation, and deployment workflows across AWS and on-premises environments. Requires 5+ years of software engineering experience, strong Python and backend skills, Kubernetes, cloud, observability, and infrastructure-as-code expertise.
About the job
Responsibilities
- Architect and build machine learning infrastructure and automation in AWS and on-premises environments.
- Lead system design and architectural discussions for MLOps platforms, addressing performance, security, and compliance requirements.
- Research, evaluate, and implement MLOps tools, frameworks, and best practices.
- Collaborate with machine learning engineers, data scientists, product engineering, and infrastructure teams to move research into production.
- Optimize machine learning workflows for efficient and reproducible model deployment and monitoring.
- Establish engineering standards, conduct design reviews, and mentor junior engineers.
- Automate machine learning operations, including model CI/CD, feature engineering pipelines, and deployment strategies using Kubernetes, Airflow, and other orchestration tools.
Requirements
- Bachelor's or master's degree in computer science, computer engineering, software engineering, or a closely related field.
- 5+ years of software engineering experience focused on production-grade frameworks or applications.
- Strong software engineering skills in complex, multi-language systems and scalable backend architecture.
- Experience with Kubernetes and cloud computing platforms, preferably AWS.
- Experience with observability and monitoring tools such as Prometheus, Grafana, or Datadog.
- Understanding of DevOps principles and infrastructure as code, including Helm and Terraform.
- Experience owning development platforms and serving internal customers.
- Proficiency in Python and exposure to additional programming languages.
Preferred Qualifications
- Experience with machine learning frameworks such as PyTorch or scikit-learn.
- Experience with data workflow orchestration frameworks such as Airflow or Kubeflow.
- Expertise in MLOps principles, including model lifecycle management, feature stores, model monitoring, and machine learning CI/CD.
- Experience with streaming data processing using Kafka, Flink, or Spark Streaming.
- Familiarity with security and compliance best practices for machine learning systems.
- Experience using AI development assistants such as Copilot or Cursor.
Compensation
- Expected annual salary range: $127,500–$195,500.
Skills
Python, Kubernetes, AWS, Prometheus, Grafana, Datadog, Helm, Terraform, PyTorch, scikit-learn, Airflow, Kubeflow, Kafka, Flink, Spark
Similar jobs
ML Engineering jobsBuild and deploy secure enterprise AI agents, integrations, and automated workflows that improve internal productivity and business operations. The role requires at least five years of enterprise, software, automation, or IT systems engineering experience, plus hands-on workflow and LLM application development.
Build and operate agentic AI systems for rare disease research, including typed workflows, evaluation, observability, and production deployment. The role requires a bachelor’s degree or equivalent experience, 5+ years building production systems, and 2+ years shipping LLM-powered applications.
Build customer-facing agentic AI experiences for loyalty programs, combining LLM workflows, product engineering, personalization, and safety controls. The role requires 6+ years of software engineering experience and production experience with AI or LLM-powered systems.
Build marketplace search and ranking features while supporting MLOps infrastructure, model deployment, feature stores, and real-time data pipelines. The role requires 5+ years of software engineering or MLOps experience, backend or full-stack expertise, and familiarity with cloud and machine learning tooling.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.