Skip to content
StripeStripe

Staff Software Engineer, Machine Learning Platform

As a Staff Software Engineer, Machine Learning Platform, you will be a technical lead, defining strategy and leading the technical direction for the next generation of ML infrastructure at Stripe. You will take ownership of end-to-end architecture and system design for complex projects, working cross-functionally to drive significant business impact.

About the job

What you’ll do

You will serve as a technical lead across the ML Platform space and a key contributor to the evolution of the platforms that power Stripe's ML-driven products. As a Staff Engineer, you'll be empowered to make decisions with a large impact on Stripe. You will influence our investments and strategy while making our systems more reliable, secure, and a delight to use. You will work cross-functionally with other tech staff, data science, product, and senior leadership to land a bigger impact of ML at Stripe. You will help define the long-term strategy and lead the technical direction for the next generation of ML infrastructure that powers Stripe's ML-driven products.

Responsibilities

  • Take ownership of end-to-end architecture and system design for large, complex projects across ML Platform.
  • Define technical directions for projects with high ambiguity, transforming complex user needs into long-lasting platform strategy.
  • Design the system architecture and solutions for the most challenging problems in the ML Platform domain, including low-latency model inference, large-scale feature stores, real-time monitoring, and LLM/agent orchestration.
  • Turn high-leverage ideas into tangible, robust solutions that shape platform and product roadmap , combining technical excellence with creative problem-solving.
  • Scope and lead large projects with significant business impact, driving them from requirements through design, implementation, and production operation.
  • Work with ML engineers, data scientists, and product teams directly to translate their needs into functional requirements and scalable technical solutions.
  • Arbitrate critical decisions that balance competing priorities while meeting latency, reliability, cost, and security constraints.
  • Serve as a key engineering representative, engaging senior leaders across Stripe and advising the leadership team on key technical considerations related to the end-to-end ML lifecycle.
  • Drive cross-team technical initiatives that improve ML development velocity and MLOps maturity across the company.
  • Mentor and grow other engineers. Serve as a role model for designing, implementing, and operating great software systems.

Who you are

We’re looking for someone who meets the minimum requirements to be considered for the role. If you meet these requirements, you are encouraged to apply. The preferred qualifications are a bonus, not a requirement.

Minimum requirements

  • 10+ years of professional software development experience, or equivalent domain expertise, with a solid background in service-oriented architecture and large-scale distributed systems.
  • Track record of serving as a technical lead, with the ability to provide technical direction, lead multi-team initiatives, and mentor team members.
  • Experience working on production ML platform services.
  • Strong product instincts and a deep understanding of the business context in which you operate.
  • Strong communication skills with the ability to explain complex technical concepts to both technical and non-technical stakeholders.
  • Demonstrated ability to work cross-functionally, collaborating effectively with ML engineers, data scientists, software engineers, product managers, and business stakeholders.
  • The ability to thrive on a high level of autonomy and responsibility, and comfort operating in ambiguous environments.
  • Hands on experience using AI tools to accelerate how you work.

Preferred qualifications

  • Experience building large-scale serving or data infrastructure for machine learning use cases (e.g., model inference, feature stores, real-time feature computation, model registries).
  • Familiarity with LLMs, LLM frameworks, and agentic AI patterns (e.g., tool use, multi-agent orchestration, retrieval-augmented generation).
  • Experience rapidly developing prototypes and iterating based on user feedback.
  • Familiarity with cloud services (e.g., AWS) and cloud-based AI/ML services (e.g., SageMaker, Bedrock, Databricks, OpenAI).
  • Experience training and shipping machine learning models to production to solve critical business problems.
  • Ability to synthesize ideas across the organization while setting a compelling technical vision.
  • Comfortable working with geographically distributed teams.
  • Passion for side-projects, open source, or self-driven technical initiatives.

Skills

Machine Learning, Distributed Systems, Service-Oriented Architecture, MLOps, LLMs, AWS, SageMaker, Databricks, OpenAI, Python

Anthropic

Anthropic

San Francisco, CA

Staff+ Software Engineer, ML Inference Path
$320k+/yrHybrid7+ YOEML Engineering

Build and operate scalable ML inference infrastructure for Claude’s safety systems, translating safety research into reliable production deployments. The role requires deep production ML infrastructure experience, distributed systems expertise, and proficiency with Python and modern ML frameworks.

Shield AI

Shield AI

San Diego, CA

Staff Engineer, Perception Software
$200k+/yrOn-site7+ YOEML Engineering

Develop production C++ perception capabilities for autonomous systems, spanning algorithms, libraries, integration, validation, and release. The role requires deep expertise in at least one perception domain, strong systems debugging, and experience delivering maintainable software in complex robotics or real-time environments.

Garner Health

Garner Health

New York, NY

Staff Machine Learning Operations Engineer
$298k+/yrHybrid7+ YOEML Engineering

Leads the reliability, architecture, deployment automation, and monitoring of production machine learning systems. Requires 7+ years of software engineering experience, deep MLOps platform expertise, and strong Kubernetes, cloud, infrastructure-as-code, and observability fundamentals.

Talkiatry

Talkiatry

United States

Staff AI Enablement Engineer
$190k+/yrRemote8+ YOEML Engineering

Staff-level engineer responsible for building AI agents and automation, evaluating developer AI tools, and driving adoption across the engineering organization. Requires 8+ years of software engineering experience plus production experience with LLMs, agentic systems, and applied machine learning.

Nuro

Nuro

Mountain View, CA

Senior/Staff Engineer, Machine Learning - Online Mapping
$194k+/yrOn-site7+ YOEML Engineering

Develop and productize online mapping models for autonomous navigation using real-world sensor data. The role requires deep ML expertise, robotics or computer vision experience, strong Python and deep learning framework skills, and a staff-level ability to deliver practical solutions.