Hardware / Software CoDesign Engineer
Co-designs AI-optimized hardware with vendors, develops performance models, and influences architectures for efficient ML training and inference. Requires 4+ years experience in software/hardware co-design, GPU programming, and collaboration with ML/kernel teams.
About the job
In this role, you will:
- Co-design future hardware for programmability and performance with our hardware vendors
- Assist hardware vendors in developing optimal kernels and add support for it in our compiler
- Develop performance estimates for critical kernels for different hardware configurations and drive decisions on compute core and memory hierarchy features
- Build system performance models at different abstraction levels and carry out analysis to drive decisions on scale up, scale out, front end networking
- Work with machine learning engineers, kernel engineers and compiler developers to understand their vision and needs from high performance accelerators
- Manage communication and coordination with internal and external partners
- Influence the roadmap of hardware partners to optimize them for OpenAI’s workloads
- Evaluate potential partners’ accelerators and platforms
- As the scope of the role and team grows, understand and influence roadmaps for hardware partners for our datacenter networks, racks, and buildings
You might thrive in this role if you have:
- 4+ years of industry experience, including experience harnessing compute at scale and optimizing ML platform code to run efficiently on target hardware
- Strong experience in software/hardware co-design
- Deep understanding of GPU and/or other AI accelerators
- Experience with CUDA, Triton or a related accelerator programming language
- Experience driving Machine Learning accuracy with low precision formats
- Experience with system performance modeling and analysis to optimize ML model deployment
- Strong coding skills in C/C++ and Python
- Are familiar with the fundamentals of deep learning computing and chip architecture/microarchitecture
- Able to actively collaborate with ML engineers, kernel writers, compiler developers, system engineers, chip architects/microarchitects
Nice to have:
- PhD in Computer Science and Engineering with a specialization in Computer Architecture, Parallel Computing, Compilers or other Systems
- Strong understanding of LLMs and challenges related to their training and inference
Benefits and Perks
- Medical, dental, and vision insurance for you and your family
- Mental health and wellness support
- 401(k) plan with 4% matching
- Unlimited time off and 18+ company holidays per year
- Paid parental leave (20 weeks) and family-planning support
- Annual learning & development stipend ($1,500 per year)
Skills
CUDA, Triton, C++, Python, GPU, Ai Accelerators, System Performance Modeling, Compilers, Low Precision Formats, Chip Architecture
Similar jobs
Embedded Engineering jobsDesign and implement real-time control infrastructure for intelligent, reliable robots, spanning actuators, hardware interfaces, state estimation, and whole-robot behavior. The role requires strong robotics and control fundamentals, hands-on hardware experience, and production-quality C++ or Rust development.
Build the low-level runtime for a custom AI accelerator, including kernel scheduling, device memory management, synchronization, and hardware-software interfaces. The role requires strong systems programming experience and expertise in concurrency, memory semantics, simulation, and performance debugging.
Leads systems software bringup and validation for new AI silicon, developing diagnostics, automation, observability, and stress infrastructure from first power-on through integrated model execution. The role requires low-level programming, computer architecture knowledge, hardware-software debugging, and cross-functional technical leadership.
Build and harden OS foundations for AI consumer devices, spanning kernel, services, security, performance, and application interfaces. Requires strong systems programming in C/C++, kernel experience, and debugging skills.
Develops real-time perception, sensor-control, and fusion software for heterogeneous autonomous vehicle platforms, integrating ML algorithms across sensors and embedded systems. Requires advanced engineering education or 5+ years of relevant experience, multi-modal sensing expertise, Linux/Docker proficiency, and U.S. security-clearance eligibility.