Latest DevOps / SRE jobs
Job results
Leads the design, automation, and reliability of large-scale, multi-cloud infrastructure supporting search, NoSQL, and AI-driven workloads. Requires 7+ years in infrastructure, DevOps, or SRE, plus deep Kubernetes, Terraform, Linux, and distributed data-systems expertise.
Operates and improves cloud infrastructure and Kubernetes-based systems for Astronomer customers, handling incidents, observability, automation, and production guidance. Requires 5+ years of cloud infrastructure experience, 3+ years with Kubernetes, and strong Linux, networking, distributed-systems, and customer-facing troubleshooting skills.
Build scalable infrastructure, automation, and network-resilience tools for a global cloud network. The role requires 6+ years in SRE, DevOps, or software engineering, strong Linux and distributed-systems expertise, programming skills in Python or Go, and experience with observability and IaC.
Build and operate internal developer platforms that improve engineering velocity, reliability, and security. The role spans developer tooling, CI/CD, GitOps, AI-assisted development, and automated engineering guardrails.
Build and operate distributed cloud infrastructure and platform services that support product teams globally. The role requires 5+ years of software development experience, strong distributed-systems expertise, and experience with cloud infrastructure, reliability, and observability.
The DevOps Engineer will design secure CI/CD pipelines, automate application delivery, and manage cloud infrastructure across development, testing, and production environments. The role requires 4+ years of DevOps experience, strong networking knowledge, and expertise with Kubernetes, Jenkins, scripting, and major cloud platforms.
Build and improve cloud infrastructure, developer workflows, and internal tooling that make software development, testing, and releases more efficient and reliable. The role requires cloud architecture knowledge, CI/CD experience, Terraform and Bazel proficiency, and software development skills in Go, Python, or C++.
Build and operate scalable build systems, CI pipelines, and developer infrastructure for consumer-device software. The role requires 5+ years of engineering experience, expertise with Bazel or comparable build systems, and experience improving CI reliability and performance at scale.
Build and operate production software, APIs, and automation for large-scale bare-metal and GPU infrastructure. The role requires 8+ years of software or infrastructure engineering experience, strong Python and Linux skills, and expertise in provisioning, lifecycle management, and reliability.
Leads the design, operation, and evolution of Flexport’s cloud infrastructure, platform tooling, observability, and incident response systems. Requires 10+ years of software, SRE, or infrastructure engineering experience, deep AWS expertise, and strong Terraform and automation skills.
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Provides technical leadership for reliability, scalability, and operational excellence across Reddit’s advertising systems. The role requires 8+ years operating large-scale distributed systems, strong software engineering skills, and expertise in cloud-native architectures, observability, and incident response.
Own and evolve AWS cloud infrastructure, deployment, reliability, observability, and security for a growing financial and hospitality technology platform. The hands-on role requires 8+ years operating production cloud infrastructure, strong AWS and container orchestration expertise, and experience with migrations and incident response.
Leads reliability engineering for critical AI serving systems, spanning SLOs, observability, high availability, and incident response. Requires strong distributed-systems or infrastructure experience, with model-serving, accelerator, networking, and resilience-testing expertise valued.
Production Engineer responsible for building and operating scalable infrastructure, driving reliability and architecture initiatives, and enabling product teams through platform tooling. Requires software engineering experience, distributed-systems expertise, cloud experience, and cross-team technical leadership.
Site Reliability Engineer drives end-to-end reliability for AI fine-tuning platform Tinker, including SLOs, monitoring, incident response, and multi-tenant GPU scheduling. Requires distributed systems experience, software proficiency for reliability, and production incident handling.
Develops scalable observability tooling and infrastructure for large-scale distributed systems, including logging, metrics, tracing, dashboards, and alerting. Requires 7+ years of production software experience and a bachelor’s degree or higher in computer science or a related field.
Senior Platform Engineer responsible for architecting scalable, secure infrastructure and improving reliability, observability, and production operations. The role requires strong AWS, Infrastructure as Code, and Kubernetes experience, along with technical leadership and mentoring skills.
Winter infrastructure and site reliability internship focused on building and operating minimal, on-premises backend infrastructure for a semiconductor fabrication facility. The role requires systems programming, Linux, networking, distributed systems, and hands-on infrastructure or automation experience.
Own and improve the CI/CD, testing, and deployment infrastructure that enables fast, safe, observable releases at scale. The role requires strong distributed-systems expertise, hands-on Kubernetes and infrastructure-as-code experience, and a track record of measurable cross-team improvements.
Designs and delivers large-scale data center connectivity, structured cabling, ISP circuits, and network upgrades for high-bandwidth infrastructure. Requires a bachelor’s degree or equivalent experience, at least five years in data center connectivity, and hands-on expertise with fiber, copper, testing, and low-voltage systems.
Build and operate self-service datastore infrastructure, embedding provisioning, observability, disaster recovery, compliance, and cost controls into a platform used by product engineering teams. Requires 3+ years in SRE or infrastructure-focused work, production software delivery, and AWS and Kubernetes experience.
Senior DevOps Engineer responsible for building and operating Kubernetes-based infrastructure, AWS cloud systems, deployment workflows, and observability for reliable services at scale. Requires 5+ years of DevOps or platform engineering experience and strong production Kubernetes expertise.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.
Build and operate secure, highly available Kubernetes platforms on AWS, including cluster creation, scaling, service mesh, automation, and incident response. The Staff-level role requires deep experience with Kubernetes, Terraform, AWS, Helm, Karpenter, and Istio.
Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.
Supports and evolves Azure cloud services for NLM, maintaining production environments, automating administration, and assisting developers with cloud operations. Requires strong Linux and Azure administration experience, networking knowledge, scripting, and troubleshooting skills.
Leads cloud infrastructure, platform strategy, deployment pipelines, and infrastructure automation for a growing consumer platform. Requires 5+ years in infrastructure, DevOps, platform engineering, or SRE, plus deep AWS, coding, containerization, and infrastructure-as-code experience.
Build and own production-grade AI agent infrastructure across multiple clouds, with responsibility for Kubernetes, Terraform, observability, security, reliability, and automation. Requires 5+ years of cloud infrastructure experience and strong CI/CD, networking, and production operations expertise.
Leads platform observability, monitoring, and incident response efforts, designing alerting and automation that improve reliability and customer experience. Requires 6+ years in SRE, DevOps, production engineering, or a similar role, along with cloud, container orchestration, and monitoring expertise.
Senior software engineer responsible for operating and evolving Voltus’s infrastructure platform across AWS, Kubernetes, Nomad, observability, stateful systems, and developer tooling. The role requires 6+ years of engineering experience, deep production Kubernetes and AWS expertise, and strong Go or Python skills.
Build and operate deployment platforms, automation, and developer tooling that make software releases safer, more reliable, and self-service. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience with production systems and cloud or distributed infrastructure.
Build and operate developer platform systems for continuous integration, Kubernetes-based ephemeral environments, automated testing, and internal tooling. The role requires a bachelor’s degree or equivalent, three years of software engineering experience, and experience operating production software or infrastructure.
Operates and evolves high-throughput MariaDB infrastructure, improving reliability, automation, security, observability, and disaster recovery. Requires 5+ years of production MariaDB/MySQL experience plus expertise in distributed databases, Kubernetes, infrastructure as code, and incident readiness.
Staff Engineer responsible for deploying, integrating, maintaining, and developing an AI training factory across isolated environments. The role requires 7+ years of related experience, cloud and Kubernetes expertise, Linux networking knowledge, application support skills, and automation experience.
Build and operate scalable control-plane and data-plane infrastructure for distributed AI workloads, including Ray cluster orchestration, scheduling, observability, and accelerator integration. Requires a bachelor's degree or equivalent experience, 3+ years of production coding, cloud-native expertise, Kubernetes, and Go/Python proficiency.
Leads infrastructure and platform strategy for a production healthcare AI platform, owning AWS, reliability, disaster recovery, compliance, CI/CD, and secure AI-agent operations. Requires deep cloud and Terraform expertise, audit-cycle experience, and prior technical leadership.
Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.
Lead the technical direction and evolution of Fetch’s CI/CD platform, improving build, test, deployment, and developer workflows at scale. The role requires 8+ years of software engineering experience, Staff-level cross-team leadership, and strong expertise in DevOps, cloud infrastructure, automation, and software delivery systems.
Own and scale Nango’s cloud platform, customer-controlled deployments, infrastructure automation, reliability, and data layer. The role requires 10+ years in platform, infrastructure, DevOps, or SRE work, with deep Kubernetes, AWS, Terraform, database, and compliance experience.
Own and scale the company’s cloud platform, BYOC deployments, infrastructure automation, reliability, data layer, and infrastructure security. Requires 10+ years in platform, infrastructure, DevOps, or SRE roles, with deep Kubernetes, AWS, Terraform, and database expertise.
The DevOps Engineer will build and operate reliable infrastructure, deployment workflows, and observability for data pipelines and AI/ML systems. The role requires at least three years of DevOps, SRE, or infrastructure experience plus strong cloud, Terraform, containerization, and MLOps expertise.
Associate Platform Engineer role for a 2027 college graduate supporting internal developer tooling, system monitoring, outage troubleshooting, and reliability improvements. The role offers hands-on learning in Linux, cloud infrastructure, Kubernetes, and systems analysis while working in person from the Redwood City office.
Build developer-experience tooling and release systems within Benchling’s Platform team, helping engineering teams develop, test, package, and ship high-quality software rapidly. The role requires 4+ years of software engineering experience, web framework expertise, strong problem-solving, and effective cross-functional communication.
Operate and scale Kong’s multi-region SaaS platform across major cloud providers, Kubernetes, and distributed data systems. The role requires strong infrastructure automation, observability, CI/CD, and production reliability experience, with participation in a global on-call rotation.
Leads reliability and networking for highly available, secure cloud services in Okta’s Federal SRE organization. The role requires active TS/SCI clearance with full-scope polygraph, Federal/DoD compliance experience, and deep expertise in AWS networking, Terraform, observability, and automation.
The Senior Cloud Operations Engineer designs, operates, and improves highly available multi-cloud infrastructure, observability, automation, and distributed systems across AWS and GCP. The role requires 5–7 years of DevOps or SRE experience, strong Kubernetes and infrastructure-as-code expertise, and participation in on-call operations.
Build and operate Kubernetes-based infrastructure, cloud systems, CI/CD, observability, and reliability tooling for a high-scale prediction market platform. The role requires 8+ years of infrastructure or platform engineering experience and strong production expertise with Kubernetes, cloud providers, infrastructure as code, and software development.
Senior security-focused DevOps engineer responsible for improving platform security through architecture, automation, secure defaults, and observability. The role partners with engineering teams across Google Cloud, GKE, Kubernetes, identity, networking, CI/CD, and software supply chain security.
Automate, manage, and optimize large-scale ClickHouse clusters handling trillions of events and 100+ PB data. Build provisioning systems with Terraform, Ansible, Kubernetes; focus on performance, scaling, and bleeding-edge features.