Staff Machine Learning Engineer
Develops, deploys, and optimizes state-of-the-art ML models for production at scale, handling terabyte datasets and neural networks. Requires 8+ years experience, expertise in PyTorch/TensorFlow/Python, and focus in CV/NLP.
About the job
Responsibilities
- Design, code, train, tune, deploy, and analyze ML models for production use cases at scale with high throughput and uptime.
- Write and maintain scalable, performant code shared across platforms.
- Contribute to product and core backend systems with improvements.
- Improve engineering standards, tooling, and processes.
- Develop novel, accurate, performant ML algorithms.
- Conduct metric-driven research experiments to improve model performance.
- Provide mentorship and onboarding to ML engineers.
- Lead cross-functional collaboration.
- Contribute to strategic direction and roadmap planning.
- Maintain awareness of industry best practices for data handling.
Minimum Requirements
- Bachelor's Degree in computer science or related field.
- 8+ years of experience building web applications.
- Experience implementing highly-available distributed systems/microservices.
- Delivered scalable backend APIs.
- Strong interpersonal and communication skills with bias towards action.
- Experience writing code and training across distributed systems.
- Ability to understand and make well-reasoned tradeoffs in feature design.
- Expert in machine learning frameworks such as PyTorch or TensorFlow.
- Expert in scripting languages such as Python and/or shell scripts for data analysis.
- Subject matter expert in at least one ML focus area (e.g., computer vision or natural language processing).
- Lead end-to-end development of new products.
Skills
PyTorch, TensorFlow, Python, Machine Learning, Deep Learning, Computer Vision, Natural Language Processing, Distributed Systems, Data Pipelines, Neural Networks
Similar jobs
ML Engineering jobsBuild and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.
Develop production C++ perception capabilities for autonomous systems, spanning algorithms, libraries, integration, validation, and release. The role requires deep expertise in at least one perception domain, strong systems debugging, and experience delivering maintainable software in complex robotics or real-time environments.
Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.
Staff-level engineer responsible for building AI agents and automation, evaluating developer AI tools, and driving adoption across the engineering organization. Requires 8+ years of software engineering experience plus production experience with LLMs, agentic systems, and applied machine learning.
Develop and productize online mapping models for autonomous navigation using real-world sensor data. The role requires deep ML expertise, robotics or computer vision experience, strong Python and deep learning framework skills, and a staff-level ability to deliver practical solutions.