Senior Software Engineer - ML Infrastructure
Builds distributed ML infrastructure including GPU training, end-to-end pipelines, and deployment platforms. Requires 3+ years experience in production ML systems, strong software engineering, and familiarity with open-source tools.
About the job
Responsibilities
- Design and implement distributed cloud GPU training approaches for deep learning model training and evaluation
- Build end-to-end machine learning pipelines and integrate them into core product workflows
- Encourage change, especially in support of ML engineering best practices, and maintain a high standard of excellence
- Collaborate with engineers across the entire company to solve complex data problems at scale
Requirements
- Bachelor's degree in Computer Science, Software Engineering, or equivalent
- 3+ years of professional experience
- Experience with building software components to address production, full-stack machine learning challenges
- Opinions about building a company-wide platform for ML training, evaluation, and deployment
- Knowledge of the open source landscape with judgment on when to choose open source versus build in-house
- Excellent analytical and problem-solving skills
Nice to Have
- Experience with developing, running, and managing orchestration systems like Airflow and Flyte that non-engineers can use to build data pipelines
- Experience with ML modeling frameworks (PyTorch, Tensorflow, etc.), and model serving platforms (TorchServe, TensorFlow Serving, NVIDIA Triton inference server, etc.)
Compensation
- Base salary range: $153,000 - $222,000 USD annually
- Equity, comprehensive health/dental/vision insurance, 401k with employer match, learning/wellness stipends, paid time off
Skills
PyTorch, TensorFlow, Airflow, Flyte, Kubernetes, GPU, Distributed Training, Machine Learning Pipelines, Torchserve, Nvidia Triton
Similar jobs
ML Engineering jobsBuild and deploy production machine-learning models and data systems that classify and enrich Internet telemetry for internal platforms and customer-facing products. The role requires 5+ years of applied ML, data science, or software engineering experience, plus strong Python or Go skills.
Designs and ships production multi-agent compliance systems, including LLM pipelines, model training, evaluation, monitoring, and explainability. Requires 5+ years of applied AI/ML engineering experience, strong Python, and experience deploying production ML systems.
Build and operate the platform that deploys, serves, observes, and retrains production machine-learning models for real-time fraud and financial-crime risk decisions. Requires 5+ years of ML engineering, backend, or MLOps experience, strong Python skills, and production model-serving expertise.
Leads hands-on development and deployment of production AI agents and workflows, sets technical standards, and mentors engineers. Requires 4+ years of software engineering experience, production LLM application experience, and strong Python or TypeScript skills.
Owns the full lifecycle of data and ML solutions, from ingestion and feature-ready datasets through production deployment and business-impact measurement. The role combines data engineering, applied machine learning, MLOps, and generative AI to build risk detection capabilities.