Software Engineer, Accelerators
Develop and optimize low-level software kernels and systems for new AI accelerator platforms to enable efficient large-scale training and inference of models like LLMs. Requires 3+ years in AI infrastructure, experience with data center-scale accelerators like TPUs, and strong systems skills.
About the job
Responsibilities
- Prototype and enable OpenAI's AI software stack on new, exploratory accelerator platforms.
- Optimize large-scale model performance (LLMs, recommender systems, distributed AI workloads) for diverse hardware environments.
- Develop kernels, sharding mechanisms, and system scaling strategies tailored to emerging accelerators.
- Collaborate on optimizations at the model code level (e.g. PyTorch) and below to enhance performance on non-traditional hardware.
- Perform system-level performance modeling, debug bottlenecks, and drive end-to-end optimization.
- Work with hardware teams and vendors to evaluate alternatives to existing platforms and adapt the software stack to their architectures.
- Contribute to runtime improvements, compute/communication overlapping, and scaling efforts for frontier AI workloads.
Requirements
- 3+ years of experience working on AI infrastructure, including kernels, systems, or hardware-software co-design.
- Hands-on experience with accelerator platforms for AI at data center scale (e.g., TPUs, custom silicon, exploratory architectures).
- Strong understanding of kernels, sharding, runtime systems, or distributed scaling techniques.
- Familiarity with optimizing LLMs, CNNs, or recommender models for hardware efficiency.
- Experience with performance modeling, system debugging, and software stack adaptation for novel architectures.
- Exposure to mobile accelerators is welcome, but experience enabling data center-scale AI hardware is preferred.
- Ability to operate across multiple levels of the stack, rapidly prototype solutions, and navigate ambiguity in early hardware bring-up phases.
- Interest in shaping the future of AI compute through exploration of alternatives to mainstream accelerators.
Skills
Kernels, Sharding, PyTorch, Tpus, Runtime Systems, Distributed Systems, Performance Modeling, Hardware-Software Co-Design, LLMs, Cnns
Similar jobs
Fullstack Engineering jobsBuilds zero-to-one, AI-native financial products across the frontend and backend, including enterprise memory, APIs, integrations, and professional workflows. The role requires 5+ years of software engineering experience, TypeScript and React expertise, backend development skills, and strong product judgment.
Build and ship AI-powered web products end to end, translating emerging model capabilities into user experiences while owning architecture and collaborating with research, product, and design teams. Requires 5+ years of full-stack software engineering experience.
Build data pipelines, APIs, libraries, and web interfaces that help AI researchers manage, query, and analyze training and evaluation data. The role requires significant software engineering experience and close collaboration with technical users; experience with large-scale data systems and ML tooling is advantageous.
Build full-stack product systems for expert acquisition, onboarding, matching, activation, and retention. The role requires 3–5 years of software engineering experience, strong product judgment, and the ability to ship measurable improvements with cross-functional partners.
Build and ship net-new full-stack product experiences that turn emerging model capabilities into impactful user products. The role requires strong product intuition, autonomy, rapid prototyping, and effective cross-functional collaboration in ambiguous environments.