Principal ML System Engineer
Leads the technical vision, architecture, and roadmap for a company-wide machine learning platform supporting model training, deployment, serving, monitoring, and generative AI. The role requires expert Python and Java skills, large-scale MLOps experience, cloud and Kubernetes expertise, and organization-wide technical leadership.
About the job
Responsibilities
- Partner with product and engineering leadership to translate business and product objectives into a multi-quarter technical strategy and roadmap for the ML platform.
- Define reference architectures and standards for scalable data and ML pipelines spanning model training, evaluation, deployment, and serving.
- Set company-wide MLOps direction and best practices, including model CI/CD, model registries, feature stores, experiment tracking, and build-vs-buy decisions.
- Establish reliability, observability, and performance practices for production ML systems, including monitoring, alerting, and automated remediation.
- Establish ML platform security architecture, including authentication, role-based access control, audit logging, and compliance monitoring.
- Define secure, cost-efficient integration and infrastructure patterns for connecting the platform with existing systems, APIs, and data sources at scale.
- Provide technical leadership and mentorship across engineering teams and influence the organization-wide ML infrastructure roadmap.
Requirements
- Expert-level proficiency in Python and Java with strong software engineering fundamentals.
- Deep experience designing and building ML platforms and MLOps workflows at scale.
- Experience with cloud platforms, containerization, and orchestration.
- Demonstrated success setting technical direction and driving organization-wide technical initiatives across multiple teams.
Nice-to-haves
- Bachelor’s degree or higher in Computer Science, Machine Learning, or a related field.
- Familiarity with Azure Machine Learning, Databricks processing and serverless environments, and ML frameworks.
- Experience leading and sustaining critical cross-team systems.
- Experience implementing security at scale, including role-based access control, multifactor authentication, network security, and compliance monitoring.
- Experience optimizing large-model training and inference, including LLM serving, for performance and cost.
Compensation
- Annual salary: $176,000–$195,000.
Skills
Python, Java, Ml Platforms, MLOps, MLflow, Kubeflow, Ray, AWS, Azure, GCP, Docker, Kubernetes, Databricks, Azure Machine Learning, Llm Serving
Similar jobs
ML Engineering jobsLeads the architecture, development, and governance of agentic AI platforms and reusable patterns across SaaS products. The role requires principal-level experience with Python, Java, Azure and Anthropic AI technologies, RAG pipelines, model serving, evaluation frameworks, and production governance.
The Staff Machine Learning Engineer will architect and deploy scalable generative AI and machine learning systems, including retrieval, inference, evaluation, and agentic workflows. The role requires 7+ years of software development experience, strong Python skills, applied ML expertise, and deep familiarity with modern GenAI platforms and frameworks.
Builds and owns production multi-agent AI infrastructure, backend integrations, and workflow automation for marketing operations. Requires 8+ years of software engineering experience, strong Python and JavaScript/Node.js skills, production LLM experience, and deep Google Cloud expertise.
Sets the technical direction for production machine learning across a payments platform, building and scaling models for risk, authorization, disputes, and forecasting. Requires 8+ years of ML engineering experience, including production model ownership and strong technical leadership.
Own end-to-end development, evaluation, and production deployment of AI models serving high-volume real-time products. The role requires 5+ years of production Python experience, hands-on fine-tuning and ML operations, cloud infrastructure expertise, and strong technical ownership.