Member of Technical Staff
Build and operate high-performance inference systems across kernels, runtimes, serving infrastructure, and heterogeneous clusters while owning customer workloads in production. The role requires exceptional production engineering, customer-facing technical judgment, and the ability to solve ambiguous problems independently.
About the job
Responsibilities
- Develop and optimize high-performance computing kernels across inference engine internals and serving infrastructure.
- Develop AI agents for autonomous inference engineering.
- Design, deploy, and operate heterogeneous clusters across hardware vendors.
- Own customer accounts end to end, including scoping workloads, building proofs, operating production systems, and managing technical relationships.
- Conduct technical evaluations on customers’ workloads and explain benchmark methodology and results to engineering teams.
- Monitor and maintain production workloads, diagnosing and resolving latency and error-rate issues.
- Inform the product roadmap based on real-world workload behavior.
Requirements
- Exceptional software engineering ability, with experience shipping and operating production systems.
- Ability to use profilers, read traces, and independently diagnose unfamiliar systems.
- Experience owning customer relationships and resolving technical problems for named accounts.
- Ability to communicate credibly with engineering and ML teams, defend benchmark methodology, and incorporate feedback.
- Ability to work independently on ambiguous, poorly specified problems and produce scoped solutions.
Nice-to-haves
- Inference optimization experience.
Compensation and Benefits
- $200,000–$300,000 base salary plus generous equity.
- Fully covered medical, dental, and vision insurance.
- Daily lunch and dinner.
- Unlimited paid time off and parental leave.
- $1,000/month post-tax housing stipend for employees living within 0.5 miles of the office.
- Covered Uber/Waymo transportation to and from the office.
- Visa sponsorship available.
Skills
C++, Python, Ai Inference, High-Performance Computing, Computing Kernels, Compilers, Inference Engines, Serving Infrastructure, AI Agents, Heterogeneous Clusters, Profiling, Benchmarking
Similar jobs
ML Engineering jobsBuild and operate large-scale ranking and retrieval systems that power search relevance, including hybrid lexical/vector search, embeddings, query understanding, and permission-aware retrieval. Requires a bachelor's degree and 5+ years of ML engineering experience in ranking or information retrieval.
Build and operate production ML infrastructure spanning training, deployment, serving, monitoring, data pipelines, and feedback-driven retraining. The role requires strong MLOps and DevOps experience, Python and SQL proficiency, and ownership of reliable cloud-based systems.
Build and scale post-training, reinforcement-learning, evaluation, and inference systems for long-horizon agents operating over complex enterprise software. The role requires strong Python and PyTorch or JAX skills, distributed GPU experience, empirical rigor, and the ability to take research results into production.
Build and operate production AI agents that transform enterprise processes, data, and code. The role focuses on tool layers, retrieval, context management, evaluations, monitoring, auditability, and guardrails, requiring strong Python and TypeScript plus experience with production LLM systems and traditional machine learning.
Build and productionize applied AI/ML systems for document understanding, agentic workflows, and demand forecasting using rich, messy enterprise data. The role requires 3+ years of production AI/ML experience, strong evaluation and monitoring practices, and a STEM master’s degree.