
Quadric
Burlingame, CA
AI accelerator IP for on-device inference
About
Quadric develops licensable General Purpose NPU (GPNPU) processor IP and toolchain for efficient on-device AI inference. It enables chip designers to integrate flexible AI acceleration into SoCs for edge devices, automotive, and enterprise applications. The programmable architecture supports current and future models without hardware changes, scaling up to 864 TOPS.
Tech stack
Python, C++, PyTorch, TensorFlow, CUDA, NumPy, ONNX, Docker, Linux, C
Perks & benefits
Competitive salary, Equity, Health insurance, Dental insurance, Vision insurance, 401k, Life insurance
More AI companies
AI companiesSan Francisco, CA
Redwood City, CA
Cambridge, MA
San Francisco, CA
New York, NY
San Francisco, CA
Open jobs
13Leads architecture and implementation of SoC-level verification environments for GPNPU-based ASICs, using UVM, SystemVerilog/Verilog, VCS, co-simulation, assembly, and Python. Requires at least five years of complex SoC/ASIC design verification experience.
Design and integrate SoC RTL subsystems, third-party IP, and architecture for next-generation edge AI processors. The role requires 5+ years of SoC/ASIC front-end experience, strong SystemVerilog or Verilog skills, and expertise in computer architecture, low-power design, and subsystem development.
Forward Deployed Engineer owning technical customer engagements from evaluation to design win on Quadric's novel Chimera GPNPU. Port workloads, write custom kernels, embed with customer teams, build agentic tooling/playbooks, and relay insights to shape product roadmap. Requires strong C++/Python, systems performance intuition, and high agency in ambiguous environments.
Build and lead marketing from the ground up for a growth-stage edge AI semiconductor IP company. Own strategy, brand narrative, product marketing, PR, and GTM initiatives while partnering closely with technical and executive teams.
Own the Chimera GPNPU IP hardware roadmap and PPA targets for Quadric. Drive architecture decisions, anchor-customer engagement, and feature gating for silicon shipping 2027-2028. Requires 5-8 years silicon PM experience and deep NPU/SoC expertise.
Develops and optimizes AI and LLM inference kernels for Quadric’s neural processing platform, profiling performance across hardware configurations and improving compiler and runtime components. Requires 5+ years of kernel optimization experience, strong C/C++ and Python skills, and familiarity with CUDA, DSP, NEON, or Triton.
Develops and optimizes AI/LLM kernels for efficient neural network inference on Quadric's GPNPU platform. Requires 5+ years in AI kernel development, C/C++ proficiency, and experience with compute frameworks like CUDA or DSP.
Leads program management for AI acceleration chips and embedded systems, overseeing software/hardware releases, customer requirements, cross-functional execution, and safety certifications. Requires 15+ years experience, master's in CS, and expertise in PM tools and AI/ML domains.
Develops and maintains web platform (DevStudio) for neural processing unit toolchain, implementing features, providing support, and troubleshooting issues in collaboration with engineering teams. Requires 5+ years full-stack experience with Golang backend and React frontend expertise.
Develops compiler algorithms and optimization passes to lower and optimize deep learning and high-performance computing workloads for Quadric’s edge-focused neural processing architecture. Requires advanced computer science education, eight or more years of industry experience, and expertise in optimization, graphs, and machine-learning algorithms.
Develops and optimizes deep neural networks for Quadric's GPNPU architecture, focusing on algorithmic lowering, graph-based execution, and performance extraction. Requires MS/PhD, 8+ years experience in optimization, ML algorithms, and graphs.
The senior compiler engineer will optimize code generation for Quadric’s neural processing unit, develop LLVM compiler passes, and co-design compiler strategies with hardware engineers. The role requires at least five years of compiler experience, strong C++ skills, and expertise with LLVM or GCC.
Designs and implements microarchitecture and RTL for innovative neural processing unit using SystemVerilog/SystemC. Optimizes PPA, contributes to timing closure, requires 5+ years CPU/GPU/ASIC experience and strong computer architecture background.