As a Software Engineer on the Inference Stack team, you will build the distributed runtime that powers large-scale LLM inference. This role involves working across the stack, from developer experience to low-level infrastructure, and owning systems in production.
180k – 360k/yrHybridML Engineering
Software Engineer - Voice AI (Inference Runtime)
BasetenSan Francisco, CA +1
Build and own high-performance inference runtime for Voice AI models including STT, TTS, and voice agents. Design real-time systems with low tail latency, collaborate cross-team, and optimize model serving for production workloads. Requires CS degree and real-time systems experience.
165k – 330k/yrHybridML Engineering
Post-Training Research Engineer
BasetenSan Francisco, CA
Build in-house tooling for post-training custom ML models using advanced techniques like RL and finetuning. Requires deep expertise in transformer training, PyTorch distributed systems, parallelism strategies, GPU performance optimization, and HPC platforms.
200k – 275k/yrHybridML Engineering
Software Engineer - AI Enablement
BasetenSan Francisco, CA
Builds and operates internal AI agents and LLM-powered workflows to boost engineering productivity across code writing, PR reviews, debugging, and documentation. Evaluates and deploys cutting-edge AI coding tools tailored to company needs.
150k – 230k/yrOn-siteML Engineering
Software Engineer, Model Performance Tooling
BasetenSan Francisco, CA
Builds performance benchmarking, diagnostic, and optimization tools for LLM inference on GPU clusters. Early-career role requiring Python proficiency, systems curiosity, and interest in AI hardware—no prior experience needed.
160k – 200k/yrOn-siteEntry levelML Engineering
Software Engineer - Training Infrastructure
BasetenSan Francisco, CA +1
Architects and leads development of scalable ML training infrastructure, including scheduling, storage, networking, and reinforcement learning systems. Requires proficiency in Go, Kubernetes expertise, distributed systems knowledge, and experience with cloud providers and ML workloads.
165k – 330k/yrHybridML Engineering
Software Engineer - Model Performance
BasetenSan Francisco, CA +1
Software Engineer optimizes ML model inference performance using techniques like quantization and speculative decoding. Requires backend experience with PyTorch, TensorRT, CUDA, and deep GPU knowledge for LLMs.
180k – 360k/yrHybridML Engineering
Search
Location
7 jobs
Job results
Software Engineer - BIS
BasetenSan Francisco, CA
As a Software Engineer on the Inference Stack team, you will build the distributed runtime that powers large-scale LLM inference. This role involves working across the stack, from developer experience to low-level infrastructure, and owning systems in production.
180k – 360k/yrHybridML Engineering
Software Engineer - Voice AI (Inference Runtime)
BasetenSan Francisco, CA +1
Build and own high-performance inference runtime for Voice AI models including STT, TTS, and voice agents. Design real-time systems with low tail latency, collaborate cross-team, and optimize model serving for production workloads. Requires CS degree and real-time systems experience.
165k – 330k/yrHybridML Engineering
Post-Training Research Engineer
BasetenSan Francisco, CA
Build in-house tooling for post-training custom ML models using advanced techniques like RL and finetuning. Requires deep expertise in transformer training, PyTorch distributed systems, parallelism strategies, GPU performance optimization, and HPC platforms.
200k – 275k/yrHybridML Engineering
Software Engineer - AI Enablement
BasetenSan Francisco, CA
Builds and operates internal AI agents and LLM-powered workflows to boost engineering productivity across code writing, PR reviews, debugging, and documentation. Evaluates and deploys cutting-edge AI coding tools tailored to company needs.
150k – 230k/yrOn-siteML Engineering
Software Engineer, Model Performance Tooling
BasetenSan Francisco, CA
Builds performance benchmarking, diagnostic, and optimization tools for LLM inference on GPU clusters. Early-career role requiring Python proficiency, systems curiosity, and interest in AI hardware—no prior experience needed.
160k – 200k/yrOn-siteEntry levelML Engineering
Software Engineer - Training Infrastructure
BasetenSan Francisco, CA +1
Architects and leads development of scalable ML training infrastructure, including scheduling, storage, networking, and reinforcement learning systems. Requires proficiency in Go, Kubernetes expertise, distributed systems knowledge, and experience with cloud providers and ML workloads.
165k – 330k/yrHybridML Engineering
Software Engineer - Model Performance
BasetenSan Francisco, CA +1
Software Engineer optimizes ML model inference performance using techniques like quantization and speculative decoding. Requires backend experience with PyTorch, TensorRT, CUDA, and deep GPU knowledge for LLMs.