Skip to content
TavilyTavily

DevOps Engineer

Owns and manages Kubernetes clusters, infrastructure as code, CI/CD pipelines, real-time data pipelines, monitoring, and production debugging for large-scale AI infrastructure. Requires 3+ years DevOps experience with distributed systems and cloud environments.

About the job

Responsibilities

  • Managing Kubernetes clusters across multiple environments and regions
  • Owning infrastructure as code for all resources
  • Maintaining and improving CI/CD pipelines and GitOps-based deployments
  • Maintaining and optimizing real-time data pipelines that process billions of events per day across distributed queues and stream processors
  • Building out monitoring, alerting, and observability
  • Debugging production issues across services
  • Managing cloud costs and capacity planning
  • Working closely with a small engineering team — owning infra end-to-end

Requirements

  • ~3+ years in a DevOps or platform engineering role, working in production environments
  • Proven experience designing and operating large-scale, distributed systems, with a solid understanding of API design, reliability, and performance at scale
  • Strong Kubernetes experience in a managed cloud environment
  • Proficiency with infrastructure as code (Terraform or similar)
  • Experience with GitOps-based deployment workflows
  • Built or maintained observability stacks (logging, metrics, alerting)
  • Experience handling production incidents calmly and methodically

Nice to Have

  • Multi-region deployments
  • Search infrastructure
  • Data pipeline experience (streaming, warehousing)
  • Proxy/networking infrastructure at scale

Skills

Kubernetes, Terraform, GitOps, CI/CD, Observability, Data Pipelines, Monitoring, Alerting, Infrastructure As Code

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.

Beacon AI

Beacon AI

San Carlos, CA

Software Engineer, Cloud Infrastructure
$135k+/yrHybridDevOps / SRE

Build and operate AWS cloud, ML, LLM, RAG, and IoT infrastructure, including deployment platforms, data pipelines, vector search, observability, security, and cost controls. The role requires deep AWS experience and production experience with LLM-powered applications.