Skip to content
CohereCohere

Site Reliability Engineer, Inference Infrastructure

Operates and scales reliable infrastructure for serving Cohere’s language models, building Kubernetes automation, observability, resilience, and customized production deployments. Requires 5+ years running large-scale production infrastructure and experience with distributed systems, cloud platforms, Linux, GPU workloads, and high-performance server development.

About the job

Responsibilities

  • Build self-service systems that automate managing, deploying, and operating services, including custom Kubernetes operators supporting language model deployments.
  • Automate environment observability and resilience so developers can troubleshoot and resolve problems.
  • Help ensure defined SLOs are met, including participation in an on-call rotation.
  • Build strong relationships with internal developers and influence the Infrastructure team’s roadmap based on their feedback.
  • Develop the team through knowledge sharing and an active review process.
  • Deploy optimized NLP models to production in low-latency, high-throughput, and high-availability environments.
  • Interface with customers and create customized deployments for their specific needs.

Requirements

  • 5+ years of engineering experience running production infrastructure at a large scale.
  • Experience designing large, highly available distributed systems with Kubernetes and GPU workloads on those clusters.
  • Experience with Kubernetes development, production coding, and support.
  • Experience with GCP, Azure, AWS, OCI, and multi-cloud, on-premises, or hybrid serving.
  • Experience designing, deploying, supporting, and troubleshooting complex Linux-based computing environments.
  • Experience with compute, storage, network resource, and cost management.
  • Strong collaboration and troubleshooting skills for building mission-critical systems and ensuring smooth operations.
  • Grit and adaptability to solve evolving complex technical challenges.
  • Familiarity with computational characteristics of GPUs, TPUs, and/or custom accelerators, especially their influence on inference latency and throughput.
  • Strong understanding of or working experience with distributed systems.
  • Experience with Golang, C++, or other languages designed for high-performance, scalable servers.

Benefits and Compensation

  • Weekly lunch stipend of $75/£75 or equivalent in local currency.
  • Full health and dental benefits, including a separate mental health budget.
  • RRSP matching, 401(k), or pension scheme.
  • 100% parental leave top-up for up to 6 months for either parent.
  • Annual enrichment benefits for arts and culture, fitness and wellness, quality time, and workspace improvements.
  • Education and learning stipend for conferences, courses, and coaching.
  • Six weeks of paid vacation (30 working days).
  • Travel budget for remote employees to visit other offices and an annual company offsite.
  • Coworking benefit for employees not near an office.
  • $500 home office stipend.

Skills

Kubernetes, Gpu Workloads, GCP, Azure, AWS, Oracle Cloud Infrastructure, Linux, Distributed Systems, Go, C++, Observability, SLOs, On-Call Operations, Nlp Models, Multi-Cloud

Cloudflare

Cloudflare

London, United Kingdom

Software Engineer: Resiliency - Deploy at Scale
No salary listedHybrid4+ YOEDevOps / SRE

Build and maintain Cloudflare’s deployment platform, enabling progressive rollouts, health-mediated releases, and automated workflows at scale. The role requires at least four years of software development experience, backend and frontend experience, and comfort with rapid delivery and on-call support.

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Baseten

Baseten

San Francisco, CA

Capacity Ops Engineer
$170k+/yrHybrid5+ YOEDevOps / SRE

Leads global GPU capacity management across acquisition, orchestration, infrastructure automation, and incident response. The role requires 5+ years of experience, deep Kubernetes expertise, production Go or Python skills, and the ability to balance reliability with unit economics.

Teleport

Teleport

United States

IT Security and Automation Engineer
$149k+/yrRemoteDevOps / SRE

Build IT workflow automation and security tooling using Go, Temporal, Kubernetes, Terraform, and shell scripting. The role also supports internal IT systems and integrations across endpoint management, identity, access, and security administration.

Crusoe

Crusoe

United States

Electrical Field Engineer - Data Center
$196k+/yrRemote5+ YOEDevOps / SRE

Electrical Field Engineer supports on-site installation, testing, and commissioning of data center power systems like switchgear, transformers, UPS, and generators. Requires 5+ years experience, Bachelor's in Electrical Engineering, and 50%+ travel to sites.