Senior Software Engineer, ML/AI Platform
Senior Software Engineer on the ML Platform team building and operating data, tooling, serving, and inference layers for a PB-scale feature store and ML lifecycle at a high-growth AI marketing company.
About the job
What You’ll Accomplish
- Unlock offline & real-time access to trillions of data points for our ML and Data Science teams.
- Manage, expand, and optimize our feature store that enables feature engineering, multi-TB scale training jobs, and offline / real-time inferencing.
- Support PB scale data operations on the feature store using Apache Spark, Spark Structured Streaming, Kafka, and Ray.
- Partner with other teams and business stakeholders to deliver ML and AI initiatives.
Your Expertise
- 5+ years working in Data Engineering / MLOps, with experience building and maturing pipelines of a PB-scale feature store.
- Deep experience with Apache Spark, Spark Streaming, and Ray Data for building data pipelines for ML use cases.
- Understanding of the correlation between data cardinality, query plans, configuration settings, hardware, and their impact on data pipeline performance.
- Experience creating infrastructure for training ML models and fine-tuning LLMs.
- Understanding of key differences between online and offline ML inferences and critical elements for success with each.
What We Use
- Infrastructure runs primarily in Kubernetes hosted in AWS’s EKS.
- Infrastructure tooling includes Istio, Datadog, Terraform, CloudFlare, and Helm.
- Backend: Java / Spring Boot microservices, built with Gradle, coupled with DynamoDB, Kinesis, AirFlow, Postgres, Planetscale, and Redis, hosted via AWS.
- Frontend: React and TypeScript, using GraphQL, Storybook, Radix UI, Vite, esbuild, and Playwright.
- Automation: Python, Metaflow, HuggingFace, PyTorch, TensorFlow, and Pandas.
Skills
Spark, Spark Streaming, Ray Data, Kubernetes, Aws Eks, Python, PyTorch, TensorFlow, Metaflow, Kafka
Similar jobs
ML Engineering jobsDesign, build, and deploy production ML systems for recommendations, search, ranking, and advertising at internet scale. Own the full ML lifecycle from modeling to monitoring with strong cross-functional collaboration.
Leads development of speech models, decoders, and low-latency inference systems for next-generation voice agents. Requires 5+ years in speech ML or related audio AI, strong Python and PyTorch experience, and the ability to guide technical direction and mentor engineers.
Build and improve production AI systems for clinical products, owning evaluations, model behavior, agentic workflows, data flywheels, deployment, and observability. The role requires 5+ years of production ML or applied AI experience, strong Python and modern ML framework skills, and hands-on debugging expertise.
Build and operate low-latency machine learning systems for ad ranking, relevance, and optimization, including feature pipelines, experimentation, evaluation, and production inference. The role requires 6+ years of software engineering experience, strong Python skills, AWS experience, and practical LLM application experience.
Senior AI Engineer responsible for production LLM agents that enrich business identity data through web discovery, verification, classification, and risk scoring. The role requires strong asynchronous Python, agent and evaluation expertise, browser automation, and experience operating AI systems in production.