# Member of Technical Staff - Inference

**Company:** [Prime Intellect](https://hotfix.jobs/companies/prime-intellect)
**Location:** San Francisco, CA
**Role:** ML Engineering
**Salary:** $150k – $300k/yr
**Experience:** 3+ years
**Skills:** Llm Serving, vLLM, Sglang, Tensorrt-Llm, Nvidia Dynamo, Python, PyTorch, AWS, GCP, Kubernetes, CUDA, Nccl, InfiniBand, Triton, Terraform
**Posted:** 2026-07-08

> Build and optimize large-scale LLM inference and serving infrastructure across cloud GPU fleets, integrating inference systems with RL training. Requires 3+ years operating ML/LLM services, strong distributed systems and GPU expertise, and hands-on experience with modern inference frameworks.

## Job Description

## Responsibilities
- Build multi-tenant LLM serving across cloud GPU fleets.
- Design GPU-aware placement and scheduling algorithms for heterogeneous accelerators.
- Implement multi-region and multi-zone failover, traffic shifting, autoscaling, routing, and load balancing.
- Optimize model distribution and cold-start times across clusters.
- Integrate and contribute to inference frameworks such as vLLM, SGLang, and TensorRT-LLM.
- Tune tensor, pipeline, and expert parallelism; prefix caching; memory management; and related performance configurations.
- Profile kernels, memory bandwidth, and transport; apply quantization and speculative decoding.
- Develop reproducible performance suites covering latency, throughput, context length, batch size, and precision.
- Embed and optimize distributed inference within the RL stack.
- Establish CI/CD, artifact promotion, performance gates, and reproducible builds.
- Build observability with metrics, logs, and tracing; support incident response and SLO management.
- Document architectures, playbooks, and API contracts; mentor and collaborate cross-functionally.

## Requirements
- 3+ years building and operating large-scale ML or LLM services with latency and availability SLOs.
- Hands-on experience with at least one of vLLM, SGLang, or TensorRT-LLM.
- Familiarity with distributed and disaggregated serving infrastructure such as NVIDIA Dynamo.
- Deep understanding of prefill and decode, KV-cache behavior, batching, sampling, speculative decoding, and parallelism strategies.
- Ability to debug CUDA/NCCL, drivers and kernels, containers, service mesh and networking, and storage end to end.
- Python systems tooling and backend-services experience.
- PyTorch experience in inference-engine development, integration, and deployment readiness.
- AWS or Google Cloud experience and cloud deployment patterns.
- Kubernetes experience running infrastructure at scale.
- Understanding of GPU architecture, CUDA runtime, NCCL, InfiniBand, and GPU-aware bin-packing and scheduling.

## Nice to Have
- CUDA or Triton kernel development and Nsight Systems/Compute profiling.
- Rust or C++ experience.
- Kafka/Pub/Sub, Redis, gRPC/Protobuf, Prometheus/Grafana, or OpenTelemetry.
- Terraform or Ansible and infrastructure-as-code practices.
- Contributions to serving, inference, or RL infrastructure open-source projects.

## Compensation and Benefits
- Cash compensation of $150,000-$300,000 with significant equity incentives.
- Flexible work arrangement, remote or San Francisco office.
- Visa sponsorship and relocation support.
- Professional development budget.
- Team off-sites and conference attendance.

## Similar jobs

- [AI Engineer, Enablement](https://hotfix.jobs/jobs/6ae315a1-a605-45dc-b1ee-88fa8e9dee24) - LangChain - New York, NY - $150k – $195k/yr
- [Member of Technical Staff — Frontier Data](https://hotfix.jobs/jobs/20efe17f-aa7c-401b-8add-7086d4571fba) - Roboflow - Remote - $150k – $300k/yr
- [Algorithm Engineer](https://hotfix.jobs/jobs/3bef67e8-d4e7-4866-a9b8-6c7863fe2e96) - Beacon Biosignals - Remote - $150k – $170k/yr
- [Software Engineer - Prediction and Planning ML](https://hotfix.jobs/jobs/21b9c778-e1ae-4695-b26d-fec68ea8a8cc) - Applied Intuition - Sunnyvale, CA - $151k – $258k/yr
- [Software Engineer, AI Platform](https://hotfix.jobs/jobs/7dcee5ac-38bb-4da0-b399-0b5896976722) - Fab2 - Austin, TX - $140k – $200k/yr

**Apply:** https://hotfix.jobs/jobs/b58356c5-b389-4d91-96be-ec31e4ed22ad
**Canonical:** https://hotfix.jobs/jobs/b58356c5-b389-4d91-96be-ec31e4ed22ad