Skip to content
AnthropicAnthropic

Performance Engineer, GPU

Architects and optimizes GPU systems for large-scale AI models, focusing on kernel development, distributed training, and performance breakthroughs to enhance inference efficiency and model capabilities.

About the job

You might be a good fit if you:

  • Have deep experience with GPU programming and optimization at scale
  • Are impact-driven, passionate about delivering measurable performance breakthroughs
  • Can navigate complex systems from hardware interfaces to high-level ML frameworks
  • Enjoy collaborative problem-solving and pair programming
  • Want to work on state-of-the-art language models with real-world impact
  • Care about the societal impacts of your work
  • Thrive in ambiguous environments where you define the path forward

Strong candidates may also have experience with:

  • GPU Kernel Development: CUDA, Triton, CUTLASS, Flash Attention, tensor core optimization
  • ML Compilers & Frameworks: PyTorch/JAX internals, torch.compile, XLA, custom operators
  • Performance Engineering: Kernel fusion, memory bandwidth optimization, profiling with Nsight
  • Distributed Systems: NCCL, NVLink, collective communication, model parallelism
  • Low-Precision: INT8/FP8 quantization, mixed-precision techniques
  • Production Systems: Large-scale training infrastructure, fault tolerance, cluster orchestration

Representative projects:

  • Co-design attention mechanisms and algorithms for next-generation hardware architectures
  • Develop custom kernels for emerging quantization formats and mixed-precision techniques
  • Design distributed communication strategies for multi-node GPU clusters
  • Optimize end-to-end training and inference pipelines for frontier language models
  • Build performance modeling frameworks to predict and optimize GPU utilization
  • Implement kernel fusion strategies to minimize memory bandwidth bottlenecks
  • Create resilient systems for planet-scale distributed training infrastructure
  • Profile and eliminate performance bottlenecks in production serving infrastructure
  • Partner with hardware vendors to influence future accelerator capabilities and software stacks

Logistics

Annual Salary: $280,000 — $850,000 USD

Education requirements: We require at least a Bachelor's degree in a related field or equivalent experience.

Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Skills

CUDA, Triton, Cutlass, Flash Attention, PyTorch, JAX, Nsight, Nccl, Nvlink, Xla

OpenAI

OpenAI

San Francisco, CA

ASIC Package SI/PI Engineer
$266k+/yrHybrid5+ YOEHardware Engineering

Owns signal-integrity and power-integrity architecture, modeling, optimization, and validation for advanced AI ASIC packages. The role requires at least five years of high-speed interconnect and package electrical-design experience, plus expertise with electromagnetic and SI/PI simulation tools.

Anthropic

Anthropic

United States

Product Engineer - Manufacturing Operations
$320k+/yrHybridHardware Engineering

Owns data center hardware manufacturing from early builds through mass production, driving process readiness, yield, quality, engineering changes, and failure resolution with contract manufacturers and ODMs. Requires hands-on hardware manufacturing experience, strong root-cause skills, and willingness to travel up to 50%.

OpenAI

OpenAI

San Francisco, CA

Lab Operations Manager, Systems Integration | Consumer Devices
$230k+/yrOn-siteHardware Engineering

Leads day-to-day operations for a large-scale consumer device test lab, managing device fleets, test rigs, infrastructure, provisioning, maintenance, and logistics. The role partners with engineering and QA teams to maintain reliable environments for validation and release testing.

OpenAI

OpenAI

San Francisco, CA

Electrical Engineer, Actuator Test Infrastructure
$225k+/yrHybridHardware Engineering

Designs, commissions, and operates the electrical infrastructure for robotic actuator dynamometers and test cells, including power systems, motor drives, instrumentation, DAQ, and safety circuits. The role requires hands-on experience with electromechanical test equipment, power electronics, measurement integrity, and electrical documentation.

OpenAI

OpenAI

San Francisco, CA

Product Manufacturing Engineer - PCB/PCBA
$207k+/yrHybridHardware Engineering

Drives PCBA manufacturing strategy, process development, production readiness, and quality for next-generation AI hardware from prototype through high-volume production. The role requires deep PCBA experience and cross-functional execution with engineering teams, suppliers, and manufacturing partners.