Senior Software Engineer, ML Platform
Build and maintain scalable ML platform for model experimentation, training, evaluation, inference, and feature store to power underwriting products. Requires 5+ years experience with Python, ML stacks like Databricks/AWS, and MLOps systems.
About the job
What You'll Do
- Turn notebooks into software. Decompose data scientist training/inference notebooks into reusable, tested components (libraries, pipelines, templates) with clear interfaces and documentation.
- Create developer-friendly ML abstractions. Build SDKs, CLIs, and templates that make it simple to define features, train/evaluate models, and deploy to batch or real-time targets with minimal boilerplate.
- Build our real-time ML inference platform. Stand up and scale low-latency model serving.
- Expand batch ML inference. Improve scheduling, parallelism, cost controls, observability, and failure/rollback for large-scale batch scoring and post-processing.
- Own and expand the feature store. Design offline/online feature definitions, high read/write throughput, and consistent offline/online semantics.
- Platform reliability and observability. Instrument training/inference for latency, throughput, accuracy, drift, data quality, and cost; build alerting and dashboards; drive incident response and postmortems.
- Underwriting infrastructure partnership. Support production batch and real-time underwriting systems in collaboration with Data Science; collaborate on model interfaces, SLAs, safety checks, and product integrations.
What We Are Looking For
- 5+ years of software engineering experience, including experience on ML platform/MLOps systems (training, deployment, and/or feature pipelines).
- Strong Python; solid software design and testing fundamentals. Proficiency with SQL; hands-on Spark/PySpark experience.
- Knowledge of ML fundamentals—probability & statistics, supervised vs. unsupervised learning, bias/variance & regularization, feature engineering, model evaluation metrics, validation strategies, and production concerns like drift, stability, and monitoring.
- Expertise with modern data/ML stacks—AWS, Databricks (workflows, lakehouse, MLflow/registry, Model Serving), and Airflow (or equivalent orchestration).
- Experience building real-time systems (service design, caching, rate limiting, backpressure) and batch pipelines at scale.
- Practical knowledge of feature-store concepts (offline/online stores, backfills, point-in-time correctness), model registries, experiment tracking, and evaluation frameworks.
- Strong problem-solving skills and a proactive attitude toward ownership and platform health.
- Excellent communication and collaboration skills, especially in cross-functional settings.
Bonus Points
- Databricks experience (MLflow, Model Serving).
- Experience with feature stores (e.g., Tecton, Feast) and streaming (Kafka/Kinesis).
- Experience with fintech, risk, or underwriting systems; familiarity with model safety checks, rejection/override flows, and auditability.
- Background with A/B testing platforms, shadow/canary deployments, and automated rollback.
- Experience with low-latency inference systems.
What We Offer
Salary Range: $230k - $265k
Equity grant
Medical, dental & vision insurance
Work from home flexibility
Unlimited PTO
Commuter benefits
Free lunches
Paid parental leave
401(k)
Employee assistance program
Skills
Python, SQL, Pyspark, Spark, AWS, Databricks, Airflow, MLflow, Kubernetes, Kafka, Tecton, Feast
Similar jobs
ML Engineering jobsSenior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.
Leads the development and production deployment of large-scale ASR and TTS systems for conversational intelligence products. The role requires 5+ years of industry experience, deep speech-model expertise, and strong software engineering and ML operations capabilities.
Build and improve production AI systems for clinical products, owning evaluations, model behavior, agentic workflows, data flywheels, deployment, and observability. The role requires 5+ years of production ML or applied AI experience, strong Python and modern ML framework skills, and hands-on debugging expertise.
Leads development of speech models, decoders, and low-latency inference systems for next-generation voice agents. Requires 5+ years in speech ML or related audio AI, strong Python and PyTorch experience, and the ability to guide technical direction and mentor engineers.
Design, build, and deploy production ML systems for recommendations, search, ranking, and advertising at internet scale. Own the full ML lifecycle from modeling to monitoring with strong cross-functional collaboration.