# Member Of Technical Staff

**Company:** [Perplexity](https://hotfix.jobs/companies/perplexity)
**Location:** London, United Kingdom
**Role:** ML Engineering
**Experience:** 3+ years
**Skills:** Rust, Python, CUDA, Cute Dsl, Triton, Cutlass, PyTorch, JAX, TensorFlow, Kubernetes, Nccl, Nvlink, InfiniBand, Rdma, Quantization
**Posted:** 2026-04-13

> Build and optimize the production inference infrastructure powering large-scale language and multimodal models. The role requires 3+ years of software engineering experience, deep GPU performance expertise, distributed-systems experience, and proficiency across Rust, Python, and CUDA-based tooling.

## Job Description

## Responsibilities
- Support transformer-based retrieval, text-generation, and multimodal models in inference infrastructure, including weight loading, request scheduling, KV-cache management, and API Gateway support.
- Port in-house CUDA kernels to NVIDIA's CuTe DSL for current and future GPU architectures.
- Develop a Rust-based inference server to address Python runtime limitations and support growing traffic.
- Profile and resolve bottlenecks across network ingress, continuous batching, and GPU-kernel interleaving.
- Build dashboards, alerts, and automated remediation for reliability and observability.
- Respond to and learn from production incidents.

## Requirements
- 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
- Deep experience with GPU programming and performance optimization, such as CUDA, Triton, or CUTLASS.
- Understanding of modern LLM architectures and production deployment.
- Experience building and operating production distributed systems under real load, ideally performance-critical systems.
- Comfortable working across Rust, Python, and CUDA/CuTe DSL.
- Ability to work end-to-end from research implementation and kernel development through production debugging.
- Familiarity with at least one deep learning framework: PyTorch, JAX, or TensorFlow.
- Understanding of GPU architectures, including memory hierarchy, warp scheduling, and tensor cores.
- Understanding of LLM architectures and inference optimization techniques such as quantization, speculative decoding, and prefill-decode disaggregation.
- Self-directed and effective in fast-moving environments.

## Nice-to-haves
- ML compilers and framework internals, including PyTorch internals, `torch.compile`, and custom operators.
- Distributed GPU communication, including NCCL, NVLink, InfiniBand, RDMA libraries, and model or tensor parallelism.
- Low-precision inference, including INT8, FP8, FP4 quantization, and mixed-precision serving.
- Profiling and debugging tools such as Nsight Compute, Nsight Systems, CUDA-GDB, and PTX/SASS analysis.
- Container orchestration, including Kubernetes, GPU scheduling, and autoscaling inference workloads.

## Compensation
- Equity may be part of the total compensation package in addition to base salary.

## Similar jobs

- [AI Software Engineer](https://hotfix.jobs/jobs/c9a0e889-a36e-488f-87d0-e9c27ab63bf0) - Rollstack - Remote
- [Applied AI Engineer, Digital Natives](https://hotfix.jobs/jobs/74eb3e31-bab1-479e-90e1-2e172658343a) - OpenAI - London, United Kingdom
- [Agent Engineer](https://hotfix.jobs/jobs/627cc378-6ac6-467e-bbf0-2c07738ae0a5) - Elliptic - London, United Kingdom
- [AI Engineer - New Verticals](https://hotfix.jobs/jobs/4159c72e-536e-4211-969c-6bcb9ad805fd) - Protege - Remote
- [AI Engineer - Assistant Experience](https://hotfix.jobs/jobs/00d11204-96cf-4c85-a12f-9058f0581c15) - Build - New York, NY - $120k – $240k/yr

**Apply:** https://hotfix.jobs/jobs/d2ffffd9-8bfc-428d-8c84-481ec72294f6
**Canonical:** https://hotfix.jobs/jobs/d2ffffd9-8bfc-428d-8c84-481ec72294f6