Member of Technical Staff, AI Training Infrastructure
Design and optimize distributed infrastructure and training pipelines for large-scale language and multimodal models. The role requires 3+ years of distributed systems or ML infrastructure experience, PyTorch, cloud platforms, container orchestration, and distributed training expertise.
200k – 350k/yr
Hybrid3+ YOEML Engineering
About the role
Responsibilities
Design and implement scalable infrastructure for large-scale model training workloads.
Develop and maintain distributed training pipelines for LLMs and multimodal models.
Optimize training performance across multiple GPUs, nodes, and data centers.
Implement monitoring, logging, and debugging tools for training operations.
Architect and maintain data storage solutions for large-scale training datasets.
Automate infrastructure provisioning, scaling, and orchestration for model training.
Collaborate with researchers to implement and optimize training methodologies.
Analyze and improve the efficiency, scalability, and cost-effectiveness of training systems.
Troubleshoot complex performance issues in distributed training environments.
Requirements
Bachelor's degree in Computer Science, Computer Engineering, or a related field, or equivalent practical experience.
3+ years of experience with distributed systems and ML infrastructure.
Experience with PyTorch.
Proficiency in cloud platforms.
Experience with containerization and orchestration.
Knowledge of distributed training techniques, including data parallelism, model parallelism, and FSDP.
Nice-to-haves
Master's or PhD in Computer Science or a related field.
Experience training large language models or multimodal AI systems.
Experience with ML workflow orchestration tools.
Background in optimizing high-performance distributed computing systems.
Familiarity with ML DevOps practices.
Contributions to open-source ML infrastructure or related projects.
Benefits
Solve hard problems at the forefront of AI infrastructure.
Build with bleeding-edge technology impacting how businesses and developers harness AI.
Own work that directly shapes the future of AI in a fast-growing team.
Collaborate with experienced engineers and AI researchers.
Build and optimize prompts, tools, skills, and memory systems to shape how Perplexity's AI models respond, use tools, and leverage context across products. Requires strong software engineering skills and experience with LLM behavior design.
200k – 330k/yrHybrid2+ YOEML Engineering
Member of Technical Staff
PhyloSouth San Francisco, CA
Build and evaluate production AI agent harness systems for biomedical discovery. Own multi-agent coordination, planning, tool use, memory, rigorous evaluations, and infrastructure for reliable scientific agents. Requires strong backend/ML engineering and quantitative judgment.
200k – 300k/yrOn-site5+ YOEML Engineering
Senior / Staff AI Platform Engineer
Clear StreetUnited States
Build and own the core AI platform powering an AI-native trading copilot. Develop high-performance Rust backend for streaming, tool execution, and safe trading actions; design robust APIs with observability and security. Requires 8+ years systems programming experience.
200k – 350k/yrRemote8+ YOEML Engineering
Senior / Staff AI Model Engineer
Clear StreetUnited States
Own reliability and quality for an AI copilot in a trading platform. Design evaluation systems, benchmarks, quality gates, model improvement loops, and AI monitoring for correctness, safety, and performance in market analysis and trading workflows. Requires 8+ years production software experience and strong ML eval expertise.
200k – 350k/yrRemote8+ YOEML Engineering
Staff Software Engineer, Agents
DecagonSan Francisco, CA
Build and own end-to-end AI agents for enterprise customers, integrating latest text/voice models and iterating based on real-world usage. Requires 8+ years of software engineering experience with Python and TypeScript.