Software Engineer, Inference
Develops low-latency, high-throughput inference services for OCR and multimodal models, optimizing batching, kernels, and autoscaling while evaluating serving frameworks. Requires 3+ years in performance engineering or ML systems with strong Python and GPU experience.
About the job
Responsibilities
- Build inference services with smart batching and caching
- Optimize kernels, tokenization, and model graphs
- Evaluate vLLM, TensorRT LLM, and Triton tradeoffs
- Implement autoscaling and admission control with clear SLOs
- Own performance dashboards and capacity planning
Requirements
- 3+ years in performance engineering or ML systems
- Strong Python, plus C++ or CUDA exposure
- Experience with GPU profiling and model serving
Nice to Have
- Experience reducing p95 and cost in production ML systems
- Prior startup or founding experience
Compensation and Benefits
- Competitive base salary plus equity
- Performance-based bonus
- Relocation assistance for Bay Area moves
- Daily meal stipend
- Medical, vision, and dental coverage
Skills
Python, C++, CUDA, vLLM, Tensorrt Llm, Triton, Gpu Profiling, Model Serving, Autoscaling, SLOs
Similar jobs
ML Engineering jobsBuild and teach reliable AI agent systems through customer workshops, technical content, guidance, and reference implementations. The role requires strong Python and agent-development experience plus a background delivering customer-facing technical training.
Build reinforcement-learning environments, evaluations, datasets, and scalable infrastructure for frontier AI capabilities. The role suits a high-agency generalist engineer with experience in agents, evaluations, or RL workflows and strong communication skills.
Develop and productionize machine- and deep-learning algorithms for biosignal and EEG data used in medical devices, clinical development, and diagnostics. The role requires 4+ years of industry experience, DSP and statistics expertise, PyTorch proficiency, and familiarity with regulated environments and production ML practices.
Develop and deploy ML-first behavior prediction and planning systems for autonomous vehicles, forecasting the motion and interactions of road users. Requires a bachelor's degree, deep learning lifecycle expertise, and at least three years of production software experience with C++ or Python.
Build the AI platform behind fab2, including model infrastructure, agent systems, evaluations, and tools for engineering and fab operations. The role requires strong production software engineering skills and comfort working across frontend, backend, infrastructure, and data.