
Relace
San Francisco, CA
AI models and infrastructure for coding agents
About
Relace builds specialized AI models and infrastructure optimized for coding agents, enabling ultra-fast code retrieval, merging, and generation at speeds like 10k+ tokens/second across large repositories. They serve engineering teams and AI code generation companies, powering production workflows for partners like Figma to make autonomous code editing reliable and efficient. This infrastructure reduces errors and accelerates software development in AI-native environments.
Tech stack
React, Next.js, Python, C++, Rust, PyTorch, CUDA, JAX, AWS, GCP, Azure, Kubernetes, Docker, Terraform, Linux, JavaScript, TypeScript, GPU optimization
More Developer Tools companies
Developer Tools companiesSan Francisco, CA
New York, NY
San Francisco, CA
San Francisco, CA
San Francisco, CA
San Francisco, CA
Open jobs
6Builds polished, fast, intuitive frontend interfaces for developer tools interacting with AI models and infrastructure. Requires 2+ years experience with React/Next.js, strong design instincts, and collaboration with backend teams in a high-velocity environment.
Machine Learning Engineer optimizes ML models for speed and efficiency through low-level CUDA kernel tuning, GPU scheduling, and hardware-aware systems design. Requires 2+ years in ML infrastructure with Python/C++/Rust and distributed frameworks like PyTorch.
Advances small, high-performance language models for retrieval, application, and code generation. Requires 2+ years ML research/production experience, Python/PyTorch/JAX fluency, optimization expertise, and advanced degree.
Builds stories, tutorials, and demos to connect cutting-edge AI models and infrastructure with developer communities. Requires strong technical writing, code comfort, and familiarity with developer ecosystems in an onsite SF role.
Leads growth initiatives by running campaigns on social channels, designing user acquisition loops, owning the growth stack, and collaborating with engineering to drive developer adoption through rapid experimentation.
Designs and operates high-performance inference and training infrastructure for ML models, focusing on GPU scheduling, distributed systems, and cloud optimization. Requires 2+ years experience with cloud platforms like AWS/GCP/Azure.