Software Engineer - GPU Kernels
Develops and optimizes high-performance GPU kernels for AI/ML operations like matrix multiplications and attention mechanisms using CUDA and C++. Requires deep GPU architecture knowledge, performance profiling, and low-level optimization expertise.
About the job
Responsibilities
Core Engineering Responsibilities
- Design and implement high-performance GPU kernels for key ML operations, including matrix multiplications, attention mechanisms, and mixture-of-experts routing
- Write and optimize code using CUDA, PTX assembly, and architecture-specific techniques
- Apply advanced performance optimization methods such as memory coalescing, warp-level programming, tensor core acceleration, and compute/memory overlap
Performance & Innovation
- Implement cutting-edge features like quantization (FP8/FP4), sparsity, and compute/communication overlap
- Identify and resolve performance bottlenecks using tools like Nsight Systems, Nsight Compute, and Torch Profiler
- Collaborate with research teams to productionize theoretical advancements
Impact & Collaboration
- Contribute to internal and open-source GPU libraries
- Present technical contributions at industry conferences (e.g., NVIDIA GTC, AWS re:Invent)
Requirements
- Strong understanding of GPU architecture and programming paradigms:
- Memory hierarchy (global, shared, registers, L1/L2 cache)
- Thread/block/grid organization
- Synchronization techniques and race condition mitigation
- Proficient in C++ and GPU performance profiling tools
- Knowledge of:
- CUDA C++ API
- Memory access patterns and bandwidth optimization
- Numerical precision and quantization strategies
- Modern GPU features (e.g., tensor cores, async operations)
Nice to Have
- Experience with Transformer models and attention optimization (e.g., Flash Attention)
- Familiarity with GPU kernel libraries: Cutlass, Triton, Thrust, CUB
- Background in GEMM tuning and distributed/multi-GPU compute
- Contributions to open-source GPU projects
- Research publications or conference presentations on GPU performance
Skills
CUDA, C++, Ptx, Nsight Systems, Nsight Compute, Torch Profiler, Tensor Cores, Flash Attention, Cutlass, Triton
Similar jobs
Backend Engineering jobsBuild and maintain high-performance backend trading infrastructure, including order-management systems and event-driven APIs. The role requires at least three years of backend engineering experience, concurrent programming expertise, relational database proficiency, AWS familiarity, and Go, Rust, or C++ skills.
Build and scale backend APIs, microservices, data pipelines, and enterprise integrations powering an AI automation platform. The role requires 3–5 years of backend experience, strong Python skills, and familiarity with databases, cloud infrastructure, security, and distributed systems.
Build and own backend services, APIs, data pipelines, and infrastructure powering an AI-native video platform. The role requires 5+ years of industry experience and strong expertise in distributed systems, production scaling, and generative AI integration.
Build and operate scalable backend systems for voice-agent products, spanning distributed services, data infrastructure, real-time audio, and ML-enabled workflows. The role requires expert programming ability, strong database and reliability engineering skills, Kubernetes experience, and production experience with systems at scale.
Build foundational, high-performance C++ systems that capture, store, retrieve, and transmit data from autonomous vehicle fleets. The role requires a bachelor’s degree, 3+ years of software development experience, and expertise in concurrency, networking, or inter-process communication.