Build and scale GPU compute infrastructure, training frameworks, data pipelines, and ML platforms to power large-scale training and inference at xAI. Requires strong distributed systems and ML infrastructure experience plus Python proficiency.
180k – 440k/yr
On-site7+ YOEML Engineering
About the role
Responsibilities
Designing, building, and scaling GPU compute infrastructure, training frameworks, and experimentation tools to enable rapid iteration on ML hypotheses
Developing data pipelines and integrating large-scale data, training, and inference systems
Collaborating with ML teams to productionize models and ensure seamless integration across the stack
Ensuring scalability, reliability, and efficiency of large-scale machine learning systems
Working across the full stack to solve complex problems independently
Mentoring junior engineers and contributing to the growth of the team
Basic Qualifications
Bachelor, Master, Post-graduate or PhD in computer science, machine learning, or other quantitative discipline; or equivalent work experience
2+ years of industry experience working with high traffic or large-scale production environments, distributed systems, GPU infrastructure, and/or deep learning applications
2+ years experience with ML platforms, training infrastructure, or close collaboration with modeling engineers and data scientists
Strong proficiency with Python and experience with compiled languages such as C++ or Rust
Preferred Skills and Experience
Deep familiarity with modern ML frameworks such as JAX or PyTorch
Low-level understanding of compute systems, including distributed storage, NVIDIA drivers, CUDA toolkits, and networking
Comfortable with Linux systems and orchestration tools
Experience with job schedulers (e.g., Slurm), configuration management (Puppet/Ansible), or related infrastructure tooling
Compensation and Benefits
$180,000 - $440,000 USD total compensation (base salary is just one part; package also includes equity, comprehensive medical, vision, and dental coverage, 401(k), disability insurance, life insurance, and perks)
Senior Software Engineer building automated evaluation pipelines, test infrastructure, and monitoring systems to validate quality of Deepgram's speech, audio, LLM, and multimodal AI models before production release. Requires 5+ years building test/evaluation frameworks, strong analytical skills, and backend experience in Python/Rust/Go.
180k – 240k/yrRemote5+ YOEML Engineering
Software Engineer, ML Data
LiftoffSan Francisco, CA +1
Software Engineer building scalable ML data platform infrastructure including data lakes, dataset generation, model training, analytics and monitoring systems at large scale (1.5B inferences/sec). Requires 6+ years experience with strong Python/Golang and CS fundamentals.
180k – 230k/yrHybrid6+ YOEML Engineering
Senior Software Engineer, AI
IroncladSan Francisco, CA
Lead design and delivery of high-priority AI initiatives across multiple codebases. Build and ship AI-powered features with strong backend fundamentals and product sense.
180k – 220k/yrHybrid5+ YOEML Engineering
Senior Machine Learning Engineer
TeleskopeNew York, NY
Leads team building and scaling production ML pipelines for entity extraction and data classification across text, documents, and OCR. Requires 5+ years deploying ML systems, strong NLP expertise, Python proficiency, and translating research to production.
180k – 210k/yrHybrid5+ YOEML Engineering
Senior Machine Learning Engineer
ArcadeSan Francisco, CA
Builds and fine-tunes generative ML models for image, video, and text to create brand-aligned content. Collaborates with design, engineering, and product teams to prototype, evaluate, and deploy production APIs. Requires 5+ years ML experience with PyTorch and modern tooling.