What You'll Be Working On
- Bring current inference techniques into production and refine them.
- Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches.
- Work down into the serving stack, from frameworks like vLLM and SGLang to the CUDA kernels underneath, profiling and running in-depth analysis to find and fix performance problems.
- Adapt and scale optimization methods across many kinds of ML models, with an emphasis on large language models.
- Profile and tune deployments against clear targets for latency, throughput, and cost, and keep them dependable under real traffic.
- Tailor deployments to each customer's models and constraints, partnering with their engineering teams to move a workload from an early proof of concept through to a live, well-monitored production service.
- Build and support the software and product features around the inference stack in a production setting, using one or more general-purpose languages, with Python preferred given how central it is to ML work.
- Experiment quickly: take fuzzy goals, shape them into clear specs and focused proofs of concept, run fast experiments to find what works, and ship well-tested results without delay.
- Own delivery end to end, from the first experiment through to the optimization running in production, keeping the underlying performance goals, clear specs, and follow-through front of mind, and drafting features and product requirement documents together with other engineering and product teams.
- Work through ambiguity and make sound calls on tradeoffs and tooling, steering away from complexity that is not needed.
- Take real pride and ownership in your work, hold yourself accountable, and look for the same from the people around you.
What You'll Bring to the Team
- A Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or a related field.
- Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python.
- Familiarity with methods for optimizing LLMs for high throughput / low latency inference.
- Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level.
- A firm grasp of how GPUs are built and how they behave.
- Clear interest and hands-on experience with large language models.
- A working knowledge of AI/ML pipelines and the full path of developing and deploying ML models.
- Strong communication skills, particularly when explaining hard technical topics to customers and teammates.
Bonus Points
- A track record of making software systems run faster, especially for large language models.
- Experience with CUDA or comparable technologies.
- A strong command of software engineering fundamentals, with a record of building and shipping AI/ML inference systems.
- Experience with Docker and Kubernetes.
- Prior work building or tuning AI/ML projects, particularly in a customer-facing setting.
Benefits
- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location
Compensation Range
Compensation will be paid in the range of up to $250,000 - $300,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.