Skip to content

Latest DevOps / SRE jobs at Together AI

Search
Location
11 jobs

Job results

Together AI

Staff Software Engineer, GPU Infrastructure Lifecycle Management

Together AISan Francisco, CA

Build and own software state machines and control planes that automate the full lifecycle of GPU infrastructure from bare metal provisioning to running AI inference clusters. Requires strong software engineering experience with orchestration, reconciliation loops, and event-driven systems.

240k – 280k/yrOn-site7+ YOEDevOps / SRE
Together AI

Staff Platform Engineer, Service Infrastructure

Together AISan Francisco, CA

Staff Platform Engineer owning service infrastructure strategy for Together AI's Product Foundations (API, UI, Billing, IAM). Lead Kubernetes, AWS, Terraform, networking, and reusable primitives to improve reliability, consistency, and scalability across teams.

240k – 280k/yrOn-site7+ YOEDevOps / SRE
Together AI

Senior Software Engineer, Observability

Together AISan Francisco, CA

Senior Software Engineer building scalable observability platforms (metrics, logs, traces) with Prometheus, Grafana, OpenTelemetry and related tools for Together AI's GPU cloud infrastructure. Requires strong distributed systems and infrastructure-as-code experience.

200k – 280k/yrOn-site5+ YOEDevOps / SRE
Together AI

Senior Network Engineer

Together AISan Francisco, CA

Senior Network Engineer responsible for designing, implementing, and maintaining high-performance compute network infrastructure for AI systems. Requires 8+ years experience with large-scale data center networks, deep expertise in routing/switching protocols, automation, and multi-vendor hardware.

190k – 270k/yrOn-site8+ YOEDevOps / SRE
Together AI

Platform Engineer, Model Shaping

Together AISan Francisco, CA

Build and operate backend services and infrastructure for model customization and evaluation at Together AI. Requires 3+ years building production infrastructure, strong Python/Go skills, and deep experience with Kubernetes, Linux, and cloud platforms.

200k – 290k/yrHybrid3+ YOEDevOps / SRE
Together AI

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AISan Francisco, CA

Design and operate multi-petabyte distributed storage systems for large-scale AI training and inference, integrating parallel filesystems and building Kubernetes-native storage platforms.

250k – 300k/yrOn-site8+ YOEDevOps / SRE
Together AI

AI Infrastructure Engineer

Together AISan Francisco, CA

Builds and maintains AI infrastructure using Ansible, Terraform, and Kubernetes, ensuring scalability, reliability, and high availability. Handles on-call incident response, monitoring, debugging, and infrastructure growth planning. Requires 5+ years experience and CS bachelor's.

190k – 270k/yrOn-site5+ YOEDevOps / SRE
Together AI

Infrastructure Design Engineer

Together AISan Francisco, CA

Designs whitespace environments for large-scale AI GPU clusters, including rack layouts, power, cooling, and cabling. Reviews contractor designs, ensures compliance with standards, and owns capacity planning. Requires 7+ years in data center infrastructure.

210k – 250k/yrHybrid7+ YOEDevOps / SRE
Together AI

Director, Data Center Operations

Together AISan Francisco, CA

Lead design, fit-out, and commissioning of data center sites focused on power, cooling, and IT infrastructure for high-density GPU workloads. Build and manage a 20-person operations team, oversee multi-site portfolio, and establish processes from scratch.

250k – 300k/yrOn-siteDevOps / SRE
Together AI

Senior Software Engineer - Together Cloud Infrastructure

Together AISan Francisco, CA

Build and maintain highly available AI cloud infrastructure virtualizing ML hardware like GB200 GPUs and BlueField DPUs, enabling self-serve Kubernetes/Slurm clusters for internal and external customers. Requires 5+ years experience with distributed systems, backend development (Golang preferred), and cloud providers.

160k – 230k/yrRemote5+ YOEDevOps / SRE
Together AI

Senior Developer Productivity Engineer

Together AISan Francisco, CA

Owns systems and tooling to optimize developer workflows, CI/CD pipelines, and local environments for faster software delivery. Requires 5+ years in DevOps/SRE, proficiency in Python/Go/JS, and CI/CD expertise.

150k – 230k/yrOn-site5+ YOEDevOps / SRE