Principal ML System Engineer
Leads the technical vision, architecture, and engineering standards for a company-wide ML platform supporting model development, deployment, serving, and monitoring. The role requires principal-level expertise in Python and Java, scalable MLOps, cloud infrastructure, security, and technical leadership across teams.
About the job
Responsibilities
- Set the technical vision and strategy for the machine learning platform supporting traditional ML and generative AI development.
- Partner with product and engineering leadership to translate business objectives into a multi-quarter ML platform strategy and roadmap.
- Define reference architectures and standards for scalable data and ML pipelines covering model training, evaluation, deployment, and serving.
- Establish company-wide MLOps practices, including model CI/CD, model registries, feature stores, experiment tracking, and build-versus-buy decisions.
- Establish reliability, observability, and performance practices for production ML systems, including monitoring, alerting, and automated remediation.
- Define ML platform security architecture covering authentication, role-based access control, audit logging, and compliance monitoring.
- Define secure, cost-efficient integration and infrastructure patterns for existing systems, APIs, and data sources.
- Provide technical leadership and mentorship across engineering teams and guide the technical roadmap for ML infrastructure.
Requirements
- Expert-level Python and Java skills with strong software engineering fundamentals.
- Deep experience designing and building ML platforms and MLOps workflows at scale.
- Experience with MLflow, Kubeflow, Ray, and model-serving frameworks or equivalents.
- Extensive experience with cloud platforms such as AWS, Azure, and/or Google Cloud.
- Experience with Docker and Kubernetes.
- Demonstrated ability to set technical direction and drive organization-wide initiatives across multiple teams.
Nice-to-Haves
- Bachelor’s degree or higher in Computer Science, Machine Learning, or a related field.
- Familiarity with Azure Machine Learning, Databricks processing and serverless environments, and machine learning frameworks.
- Experience leading and sustaining critical cross-team systems.
- Experience implementing security at scale, including role-based access control, multifactor authentication, network security best practices, and compliance monitoring.
- Experience optimizing large-model training and inference, including LLM serving, for performance and cost.
Compensation
- Annual salary range: $195,000–$217,000.
Skills
Python, Java, MLflow, Kubeflow, Ray, AWS, Azure, GCP, Docker, Kubernetes, MLOps, Databricks, Azure Machine Learning, Llm Serving, Role-Based Access Control
Similar jobs
ML Engineering jobsPrincipal technical leader defining architecture and multi-year strategy for Pinterest’s Homefeed, Search, and AI Assistant experiences. The role requires 15+ years of large-scale systems or machine-learning experience, deep expertise in discovery and generative AI, and hands-on leadership across engineering and product organizations.
Own the architecture and delivery of production AI systems for patient-provider matching, search relevance, personalization, clinical workflows, and engagement. The role requires extensive software engineering, distributed systems, search or recommendation, ML applications, and foundation-model experience in a regulated healthcare setting.
Leads the technical strategy and engineering execution required to achieve driverless freeway operation for autonomous vehicles. Requires 10+ years of software experience, demonstrated freeway autonomy leadership, deep expertise in an autonomy domain, and strong executive communication skills.
Design and scale ML infrastructure and real-time learning systems powering personalization, search, ranking, and ad tech for millions of consumers. The role requires deep distributed-systems and data-pipeline expertise, strong architecture leadership, and experience delivering zero-to-one ML systems.
Develop and productize online mapping models for autonomous navigation using real-world sensor data. The role requires deep ML expertise, robotics or computer vision experience, strong Python and deep learning framework skills, and a staff-level ability to deliver practical solutions.