Senior Software Engineer - Machine Learning Platform
Build and operate backend infrastructure for machine learning model training, serving, feature management, and marketplace simulation. The role requires 6+ years of software engineering experience, distributed systems expertise, and experience with production ML platforms.
About the job
Responsibilities
- Build, maintain, and optimize a machine learning and simulation platform for scale, performance, and confidence in decisioning.
- Develop backend software applications and APIs that apply machine learning models to evolving business needs.
- Build self-service tooling for ML teams to register features and deploy models independently.
- Develop feature infrastructure covering feature definition, storage, serving, and offline-to-online parity.
- Design and contribute to simulation systems that reflect production environments, reduce simulation costs, and broaden usage.
- Collaborate with ML, Engineering, Product, and Data Engineering teams.
- Mentor engineers on distributed systems, MLOps, and scalable architecture.
Requirements
- 6+ years of software engineering experience.
- Experience building and maintaining backend software services and APIs.
- Experience with distributed systems or large-scale data processing using Spark, Databricks, Ray, or equivalent technologies.
- Experience with an ML platform or production ML workflows, such as training pipelines, model serving, feature pipelines, or training data platforms.
- Proficiency with some or many of Python, Kotlin, Databricks, and AWS.
- Ability to understand complex requirements and communicate them to technical and non-technical partners.
Nice-to-haves
- Metaflow, MLflow, gRPC, Spark/PySpark, dbt, Ray, or GPU experience.
- Knowledge of simulation, experimentation, or backtesting systems.
- Experience building self-service or configuration-driven internal tooling.
- Strong quantitative reasoning and interest in engineering and machine learning.
- Strong ownership, accountability, written communication, and verbal communication skills.
- Ability to work effectively in self-directed and collaborative environments.
Compensation and Benefits
- Anticipated base salary: $166,900–$230,000 USD, varying by geographic compensation region.
- Target bonuses, equity compensation, and annual equity grants.
- Medical, dental, vision, wellness, retirement, paid time off, sick leave, company holidays, family and parental leave, life insurance, and disability coverage.
- 401(k) or Group Retirement Savings Plan match and an Employee Stock Purchase Plan for eligible employees.
Skills
Python, Kotlin, Databricks, AWS, Spark, Ray, Distributed Systems, MLOps, MLflow, Metaflow, gRPC, dbt, Gpu Computing, Feature Engineering, Model Serving
Similar jobs
ML Engineering jobsBuild and optimize large language model training and post-training pipelines, improving model quality, distributed performance, evaluation, and production readiness. The role requires deep PyTorch and transformer experience, strong distributed-systems and software-engineering skills, and expertise in modern LLM optimization techniques.
The Senior Algorithm Engineer leads development and production deployment of machine and deep learning algorithms for biosignal and medical-device applications. The role requires 5+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated health or similar domains.
Develop and deploy real-time perception and sensor-fusion software for autonomous battery-electric rail vehicles. The role requires strong robotics, geometry-based computer vision, C/C++ and Rust experience, plus hands-on work with multimodal sensors and production systems.
Build and deploy machine learning systems that apply economic theory, econometrics, and causal inference to marketplace problems. The role requires advanced training in economics, strong Python and data skills, and production ML experience for senior-level hires.
Builds and operates AI platform capabilities including RAG pipelines, semantic retrieval, agentic orchestration, and LLM integrations to power legal tech products. Requires 4+ years in distributed cloud systems, AI/ML experience, and proficiency in modern programming languages.