Skip to content
ConfluentConfluentMountain View, CA

Staff Software Engineer

Build and operate backend services for AI and model inference on Confluent's real-time streaming data platform. Own end-to-end feature delivery across model lifecycle, inference routing, and agent execution with strong distributed systems expertise.

236k – 277k/yr
Remote10+ YOEML Engineering

About the role

What You Will Do

  • Design and build the backend services (primarily Go, Java, and Python) that run AI and model inference on real-time data.
  • Own features end to end — drafting the design, aligning stakeholders inside and outside the team, and driving the decision to a conclusion.
  • Make the technical calls on systems that span teams: model lifecycle, inference routing, and agent execution.
  • Own the quality of what you ship — code, test coverage, documentation, operability, and rollout safety. This is production infrastructure serving live inference, so reliability isn't an afterthought.
  • Make the engineers around you better through code review, design feedback, and being someone the team trusts with ambiguous, cross-cutting work.
  • Participate in on-call for the services your team owns, and help keep the team's processes and rituals healthy.

What You Will Bring

  • 10+ years of significant experience designing, building, and operating distributed systems or cloud-native backend infrastructure in production.
  • Strong working knowledge of Kubernetes and distributed-systems patterns (control loops, API servers, high-scale control planes), plus the fundamentals — containerization, networking, resource isolation.
  • Proficiency in at least one of Go, Java, or Python, and the willingness to work across all three.
  • A track record of leading cross-team technical work: turning ambiguous requirements into designs others can rally behind.
  • Excellent written and verbal communication — you can write a design doc that aligns people who don't report to you.

What Gives You an Edge

  • Exposure to model serving, LLM/agent infrastructure, or streaming data systems.
  • You don't need a background in ML research or model training — this role is about building and operating the platform that serves AI reliably at scale, not inventing the models.

Skills

GoJavaPythonKubernetesDistributed Systemscloud nativecontainerizationNetworkingmodel servingllm infrastructurestreaming data systems

Similar roles

ML Engineering jobs
Snowflake

Senior/Staff Software Engineer - Machine Learning Platform (Inference)

SnowflakeMenlo Park, CA

Build and lead the Snowflake ML Platform for scalable, native machine learning and LLM inference workloads. Requires 7+ years experience with ML serving systems, inference engines (vLLM, TensorRT-LLM), and frameworks like PyTorch.

236k – 339k/yr
On-site7+ YOEML Engineering
Harvey

Staff Software Engineer, Model Infrastructure

HarveySan Francisco, CA

Lead design and development of Harvey's Model Infrastructure platform powering all AI requests, including unified model controller, intelligent routing, multi-provider integrations, observability, and capacity management for high reliability, low latency, and efficiency. Requires 7+ years building large-scale distributed systems with strong programming and leadership skills; AI/LLM infrastructure experience preferred.

236k – 290k/yr
On-site7+ YOEML Engineering
Snowflake

Staff AI Engineer - Cortex Code Quality

SnowflakeMenlo Park, CA

Staff AI Engineer builds and owns quality systems for Cortex Code AI coding agents, including agent strategy, experimentation pipelines, failure analysis, and cross-team alignment. Requires 8+ years shipping AI/ML software, proficiency in Python/TypeScript/Go, and expertise in LLM evaluation harnesses.

236k – 339k/yr
On-site8+ YOEML Engineering
Nuro

Senior/Staff Software Engineer, Labeling Platform

NuroMountain View, CA

Build and productionize scalable, highly available distributed systems and infrastructure for Nuro's data labeling platform that powers autonomous driving ML models. Requires 5+ years experience with reliable large-scale data systems, technical leadership, and strong programming skills in Python, C++, or Go.

235k – 352k/yr
On-site5+ YOEML Engineering
Datadog

Staff GenAI Engineer - Application Performance Monitoring

DatadogNew York, NY

Technical leader on Datadog's APM team building and deploying GenAI/ML models for agentic investigations, automated troubleshooting, and incident triaging. Requires 10+ years experience leading large-scale GenAI initiatives end-to-end in product environments.

234k – 234k/yr
Hybrid10+ YOEML Engineering