Build and own software state machines and control planes that automate the full lifecycle of GPU infrastructure from bare metal provisioning to running AI inference clusters. Requires strong software engineering experience with orchestration, reconciliation loops, and event-driven systems.
240k – 280k/yrOn-site7+ YOEDevOps / SRE
Staff Platform Engineer, Service Infrastructure
Together AISan Francisco, CA
Staff Platform Engineer owning service infrastructure strategy for Together AI's Product Foundations (API, UI, Billing, IAM). Lead Kubernetes, AWS, Terraform, networking, and reusable primitives to improve reliability, consistency, and scalability across teams.
240k – 280k/yrOn-site7+ YOEDevOps / SRE
Senior Software Engineer, Observability
Together AISan Francisco, CA
Senior Software Engineer building scalable observability platforms (metrics, logs, traces) with Prometheus, Grafana, OpenTelemetry and related tools for Together AI's GPU cloud infrastructure. Requires strong distributed systems and infrastructure-as-code experience.
200k – 280k/yrOn-site5+ YOEDevOps / SRE
Senior Network Engineer
Together AISan Francisco, CA
Senior Network Engineer responsible for designing, implementing, and maintaining high-performance compute network infrastructure for AI systems. Requires 8+ years experience with large-scale data center networks, deep expertise in routing/switching protocols, automation, and multi-vendor hardware.
190k – 270k/yrOn-site8+ YOEDevOps / SRE
Platform Engineer, Model Shaping
Together AISan Francisco, CA
Build and operate backend services and infrastructure for model customization and evaluation at Together AI. Requires 3+ years building production infrastructure, strong Python/Go skills, and deep experience with Kubernetes, Linux, and cloud platforms.
200k – 290k/yrHybrid3+ YOEDevOps / SRE
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AISan Francisco, CA
Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.
250k – 300k/yrOn-site8+ YOEDevOps / SRE
AI Infrastructure Engineer
Together AISan Francisco, CA
Builds and maintains AI infrastructure using Ansible, Terraform, and Kubernetes, ensuring scalability, reliability, and high availability. Handles on-call incident response, monitoring, debugging, and infrastructure growth planning. Requires 5+ years experience and CS bachelor's.
190k – 270k/yrOn-site5+ YOEDevOps / SRE
Infrastructure Design Engineer
Together AISan Francisco, CA
Designs whitespace environments for large-scale AI GPU clusters, including rack layouts, power, cooling, and cabling. Reviews contractor designs, ensures compliance with standards, and owns capacity planning. Requires 7+ years in data center infrastructure.
210k – 250k/yrHybrid7+ YOEDevOps / SRE
Director, Data Center Operations
Together AISan Francisco, CA
Lead design, fit-out, and commissioning of data center sites focused on power, cooling, and IT infrastructure for high-density GPU workloads. Build and manage a 20-person operations team, oversee multi-site portfolio, and establish processes from scratch.
250k – 300k/yrOn-siteDevOps / SRE
Senior Software Engineer - Together Cloud Infrastructure
Together AISan Francisco, CA
Build and maintain highly available AI cloud infrastructure virtualizing ML hardware like GB200 GPUs and BlueField DPUs, enabling self-serve Kubernetes/Slurm clusters for internal and external customers. Requires 5+ years experience with distributed systems, backend development (Golang preferred), and cloud providers.
160k – 230k/yrRemote5+ YOEDevOps / SRE
Senior Developer Productivity Engineer
Together AISan Francisco, CA
Owns systems and tooling to optimize developer workflows, CI/CD pipelines, and local environments for faster software delivery. Requires 5+ years in DevOps/SRE, proficiency in Python/Go/JS, and CI/CD expertise.
Build and own software state machines and control planes that automate the full lifecycle of GPU infrastructure from bare metal provisioning to running AI inference clusters. Requires strong software engineering experience with orchestration, reconciliation loops, and event-driven systems.
240k – 280k/yrOn-site7+ YOEDevOps / SRE
Staff Platform Engineer, Service Infrastructure
Together AISan Francisco, CA
Staff Platform Engineer owning service infrastructure strategy for Together AI's Product Foundations (API, UI, Billing, IAM). Lead Kubernetes, AWS, Terraform, networking, and reusable primitives to improve reliability, consistency, and scalability across teams.
240k – 280k/yrOn-site7+ YOEDevOps / SRE
Senior Software Engineer, Observability
Together AISan Francisco, CA
Senior Software Engineer building scalable observability platforms (metrics, logs, traces) with Prometheus, Grafana, OpenTelemetry and related tools for Together AI's GPU cloud infrastructure. Requires strong distributed systems and infrastructure-as-code experience.
200k – 280k/yrOn-site5+ YOEDevOps / SRE
Senior Network Engineer
Together AISan Francisco, CA
Senior Network Engineer responsible for designing, implementing, and maintaining high-performance compute network infrastructure for AI systems. Requires 8+ years experience with large-scale data center networks, deep expertise in routing/switching protocols, automation, and multi-vendor hardware.
190k – 270k/yrOn-site8+ YOEDevOps / SRE
Platform Engineer, Model Shaping
Together AISan Francisco, CA
Build and operate backend services and infrastructure for model customization and evaluation at Together AI. Requires 3+ years building production infrastructure, strong Python/Go skills, and deep experience with Kubernetes, Linux, and cloud platforms.
200k – 290k/yrHybrid3+ YOEDevOps / SRE
Staff Engineer, Distributed Storage and HPC & AI Infrastructure
Together AISan Francisco, CA
Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.
250k – 300k/yrOn-site8+ YOEDevOps / SRE
AI Infrastructure Engineer
Together AISan Francisco, CA
Builds and maintains AI infrastructure using Ansible, Terraform, and Kubernetes, ensuring scalability, reliability, and high availability. Handles on-call incident response, monitoring, debugging, and infrastructure growth planning. Requires 5+ years experience and CS bachelor's.
190k – 270k/yrOn-site5+ YOEDevOps / SRE
Infrastructure Design Engineer
Together AISan Francisco, CA
Designs whitespace environments for large-scale AI GPU clusters, including rack layouts, power, cooling, and cabling. Reviews contractor designs, ensures compliance with standards, and owns capacity planning. Requires 7+ years in data center infrastructure.
210k – 250k/yrHybrid7+ YOEDevOps / SRE
Get new-job notifications on iOS
Hotfix on iOS
Get a push summary when new jobs match your saved alerts.
Director, Data Center Operations
Together AISan Francisco, CA
Lead design, fit-out, and commissioning of data center sites focused on power, cooling, and IT infrastructure for high-density GPU workloads. Build and manage a 20-person operations team, oversee multi-site portfolio, and establish processes from scratch.
250k – 300k/yrOn-siteDevOps / SRE
Senior Software Engineer - Together Cloud Infrastructure
Together AISan Francisco, CA
Build and maintain highly available AI cloud infrastructure virtualizing ML hardware like GB200 GPUs and BlueField DPUs, enabling self-serve Kubernetes/Slurm clusters for internal and external customers. Requires 5+ years experience with distributed systems, backend development (Golang preferred), and cloud providers.
160k – 230k/yrRemote5+ YOEDevOps / SRE
Senior Developer Productivity Engineer
Together AISan Francisco, CA
Owns systems and tooling to optimize developer workflows, CI/CD pipelines, and local environments for faster software delivery. Requires 5+ years in DevOps/SRE, proficiency in Python/Go/JS, and CI/CD expertise.