Member of Technical Staff - Inference
Build and optimize large-scale LLM inference and serving infrastructure across cloud GPU fleets, integrating inference systems with RL training. Requires 3+ years operating ML/LLM services, strong distributed systems and GPU expertise, and hands-on experience with modern inference frameworks.
About the job
Responsibilities
- Build multi-tenant LLM serving across cloud GPU fleets.
- Design GPU-aware placement and scheduling algorithms for heterogeneous accelerators.
- Implement multi-region and multi-zone failover, traffic shifting, autoscaling, routing, and load balancing.
- Optimize model distribution and cold-start times across clusters.
- Integrate and contribute to inference frameworks such as vLLM, SGLang, and TensorRT-LLM.
- Tune tensor, pipeline, and expert parallelism; prefix caching; memory management; and related performance configurations.
- Profile kernels, memory bandwidth, and transport; apply quantization and speculative decoding.
- Develop reproducible performance suites covering latency, throughput, context length, batch size, and precision.
- Embed and optimize distributed inference within the RL stack.
- Establish CI/CD, artifact promotion, performance gates, and reproducible builds.
- Build observability with metrics, logs, and tracing; support incident response and SLO management.
- Document architectures, playbooks, and API contracts; mentor and collaborate cross-functionally.
Requirements
- 3+ years building and operating large-scale ML or LLM services with latency and availability SLOs.
- Hands-on experience with at least one of vLLM, SGLang, or TensorRT-LLM.
- Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo.
- Deep understanding of prefill and decode, KV-cache behavior, batching, sampling, speculative decoding, and parallelism strategies.
- Ability to debug CUDA/NCCL, drivers and kernels, containers, service mesh and networking, and storage end to end.
- Python systems tooling and backend-services experience.
- PyTorch experience in inference-engine development, integration, and deployment readiness.
- AWS or Google Cloud experience and cloud deployment patterns.
- Kubernetes experience running infrastructure at scale.
- Understanding of GPU architecture, CUDA runtime, NCCL, InfiniBand, and GPU-aware bin-packing and scheduling.
Nice to Have
- CUDA or Triton kernel development and Nsight Systems/Compute profiling.
- Rust or C++ experience.
- Kafka/Pub/Sub, Redis, gRPC/Protobuf, Prometheus/Grafana, or OpenTelemetry.
- Terraform or Ansible and infrastructure-as-code practices.
- Contributions to serving, inference, or RL infrastructure open-source projects.
Compensation and Benefits
- Cash compensation of $150,000-$300,000 with significant equity incentives.
- Flexible work arrangement, remote or San Francisco office.
- Visa sponsorship and relocation support.
- Professional development budget.
- Team off-sites and conference attendance.
Skills
Llm Serving, vLLM, Sglang, Tensorrt-Llm, Nvidia Dynamo, Python, PyTorch, AWS, GCP, Kubernetes, CUDA, Nccl, InfiniBand, Triton, Terraform
Similar jobs
ML Engineering jobsBuild and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.
Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.