Staff Software Engineer, Machine Learning
Builds, scales, and operates production-grade ML systems for real-time personalization on Attentive's platform. Requires 6+ years experience with Python, PyTorch/TensorFlow, and scalable ML pipelines in a fast-paced environment.
About the job
What You’ll Accomplish
- You have a proven track record of building systems that maintain a high bar of quality
- You deeply loathe regressions and take proactive steps to protect against them through a variety of testing techniques
- You are a collaborator, technical leader, and a great communicator
- You are constantly improving the quality of the project you are working on, both via direct contributions as well as long-term advocacy for larger-scale changes
- You are enthusiastic about the high impact, fast-paced work environment of an late-stage startup
- 10+ years experience is ideal
Your Expertise
- You have worked professionally building systems for 6+ years with experience on a single system long enough to see the consequences of your decisions
- Experience with TensorFlow/PyTorch, xgboost, pandas, matplotlib, SQL, Spark or similar tools
- You have proficiency or experience with Python
- You have extensive experience using machine learning and data analysis, or similar, to build scalable systems and data-driven products, working with cross-functional teams
- You have a proven track record of building scalable, efficient, automated processes for large-scale data analyses, model development, model validation, and model implementation from modern research
- You have led cross-functional machine learning projects across teams
What We Use
- Our infrastructure runs primarily in Kubernetes hosted in AWS’s EKS
- Infrastructure tooling includes Istio, Datadog, Terraform, CloudFlare, and Helm
- Our backend is Java / Spring Boot microservices, built with Gradle, coupled with things like DynamoDB, Kinesis, AirFlow, Postgres, Planetscale, and Redis, hosted via AWS
- Our frontend is built with React and TypeScript, and uses best practices like GraphQL, Storybook, Radix UI, Vite, esbuild, and Playwright
- Our automation is driven by custom and open source machine learning models, lots of data and built with Python, Metaflow, HuggingFace, PyTorch, TensorFlow, and Pandas
Compensation
For US based applicants: The US base salary range for this full-time position is $320,000 - $360,000 annually + equity + benefits
Skills
Python, PyTorch, TensorFlow, Xgboost, pandas, Kubernetes, AWS, Spark, SQL, Huggingface, Metaflow
Similar jobs
ML Engineering jobsBuild and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.
Leads end-to-end development of production algorithmic systems for healthcare, spanning machine learning, optimization, and LLM applications. The player-coach role requires 6+ years of industry experience, strong problem-solving and metrics judgment, and technical leadership of a small team.
Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.
Leads technical strategy for Reddit’s Ads ML Platform, improving feature development, training-data generation, experimentation, and the path to production ML serving. The role requires 8+ years in infrastructure or distributed systems, production ML platform experience, and strong cross-team technical leadership.
Build and operate production machine learning systems for ranking, retrieval, recommendations, personalization, and customer intelligence. The role requires 12+ years of production software and ML experience, strong expertise in intelligent systems, and sound judgment around trustworthy customer-impacting signals.