Member Of Technical Staff
Build and optimize the production inference infrastructure powering large-scale language and multimodal models. The role requires 3+ years of software engineering experience, deep GPU performance expertise, distributed-systems experience, and proficiency across Rust, Python, and CUDA-based tooling.
About the job
Responsibilities
- Support transformer-based retrieval, text-generation, and multimodal models in inference infrastructure, including weight loading, request scheduling, KV-cache management, and API Gateway support.
- Port in-house CUDA kernels to NVIDIA's CuTe DSL for current and future GPU architectures.
- Develop a Rust-based inference server to address Python runtime limitations and support growing traffic.
- Profile and resolve bottlenecks across network ingress, continuous batching, and GPU-kernel interleaving.
- Build dashboards, alerts, and automated remediation for reliability and observability.
- Respond to and learn from production incidents.
Requirements
- 3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems.
- Deep experience with GPU programming and performance optimization, such as CUDA, Triton, or CUTLASS.
- Understanding of modern LLM architectures and production deployment.
- Experience building and operating production distributed systems under real load, ideally performance-critical systems.
- Comfortable working across Rust, Python, and CUDA/CuTe DSL.
- Ability to work end-to-end from research implementation and kernel development through production debugging.
- Familiarity with at least one deep learning framework: PyTorch, JAX, or TensorFlow.
- Understanding of GPU architectures, including memory hierarchy, warp scheduling, and tensor cores.
- Understanding of LLM architectures and inference optimization techniques such as quantization, speculative decoding, and prefill-decode disaggregation.
- Self-directed and effective in fast-moving environments.
Nice-to-haves
- ML compilers and framework internals, including PyTorch internals,
torch.compile, and custom operators. - Distributed GPU communication, including NCCL, NVLink, InfiniBand, RDMA libraries, and model or tensor parallelism.
- Low-precision inference, including INT8, FP8, FP4 quantization, and mixed-precision serving.
- Profiling and debugging tools such as Nsight Compute, Nsight Systems, CUDA-GDB, and PTX/SASS analysis.
- Container orchestration, including Kubernetes, GPU scheduling, and autoscaling inference workloads.
Compensation
- Equity may be part of the total compensation package in addition to base salary.
Skills
Rust, Python, CUDA, Cute Dsl, Triton, Cutlass, PyTorch, JAX, TensorFlow, Kubernetes, Nccl, Nvlink, InfiniBand, Rdma, Quantization
Similar jobs
ML Engineering jobsBuild production AI capabilities for automated slide and document generation, working across LLM applications, data analysis, and content generation. The role requires 3+ years in machine learning and NLP, advanced Python, and experience with LLM frameworks and production systems.
Build and deploy AI-powered products for digital-native customers, taking systems from experimentation through production and scale. The role requires strong Python skills, hands-on production engineering, systematic AI evaluation, and the ability to navigate reliability, security, governance, and customer impact.
Build full-stack AI agent fleets, APIs, workflows, and internal services that automate complex business processes. The role requires at least five years of engineering experience, hands-on LLM framework experience, production AWS expertise, Kubernetes, and strong API and database skills.
Build the technical foundation for a new business vertical, creating reusable infrastructure and leading early customer engagements from scoping through delivery. The role requires 3+ years of engineering experience, strong Python and SQL skills, backend/data expertise, and comfort operating in ambiguity.
Build and operate Dougie, an agentic AI system that executes workflows, evaluates its own performance, retains institutional context, and improves in production. The role requires experience deploying unattended agentic systems and engineering reliable memory, retrieval, orchestration, and feedback loops.