Latest DevOps / SRE jobs
Job results
Own the cloud platform, deployment architecture, container infrastructure, networking, autoscaling, cost controls, and Python runtime health for a high-scale healthcare technology platform. The role requires 8+ years in infrastructure, platform, or SRE work, deep AWS expertise, Terraform experience, and Staff-level cross-team influence.
Leads Ireland-based Platform Developer Enablement and SRE teams, defining platform strategy, developer self-service, reliability objectives, and observability standards. Requires senior software, SRE, or platform engineering experience, management leadership, and expertise in cloud infrastructure, Kubernetes, Terraform, CI/CD, and distributed systems.
Build and operate software systems that manage the GPU fleet powering ChatGPT inference, including fleet health, capacity planning, resource utilization, and operational automation. The role requires 5+ years of production infrastructure experience and strong programming and distributed-systems skills.
The Senior Site Reliability Engineer will operate and improve Kubernetes-based distributed infrastructure for AI application workloads, focusing on scalability, observability, reliability, and tenant isolation. The role requires 6+ years of distributed-systems experience, production Kubernetes expertise, cloud infrastructure knowledge, and strong programming skills.
Provides technical leadership for the reliability architecture and operational foundations of a multi-region, multi-cloud platform for AI applications. The role requires deep Kubernetes and distributed-systems expertise, infrastructure programming, cloud knowledge, and experience setting SRE standards and mentoring engineers.
Build and operate a self-service internal development platform that helps engineering teams deploy and run production services reliably. The role requires strong backend programming, production Kubernetes operations, cloud infrastructure, observability, networking, and distributed-systems experience.
Staff Engineer responsible for architecting and operating MongoDB’s large-scale observability collection and ingestion infrastructure. The role requires 10+ years of experience with distributed or highly concurrent systems, expert programming skills, and strong database, performance, and systems fundamentals.
Leads Cloud Operations Engineering activities, combining technical leadership, incident response, production troubleshooting, automation, and team coaching. The role requires expertise in Linux, networking, cloud infrastructure, monitoring, distributed systems, and at least two programming languages.
This SRE will operate and improve MongoDB Atlas’s multi-tenant distributed storage infrastructure, focusing on reliability, performance, observability, automation, and incident response. The role requires 6+ years of distributed-systems experience plus expertise in storage or databases, Kubernetes, cloud platforms, Linux, and networking.
Build and operate AI-powered developer tools, internal MCP integrations, and platform capabilities across the engineering organization. The role requires strong coding and debugging skills, Kubernetes operations experience, and the ability to lead projects, improve developer experience, and mentor teammates.
Operates and scales shared telemetry infrastructure spanning metrics, logs, traces, alerting, dashboards, and profiling. The role requires at least three years of production engineering experience, distributed-systems troubleshooting, Infrastructure as Code, container orchestration, incident response, and on-call participation.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.
Build and evolve the developer platform that enables reliable, efficient software delivery across the company. The role requires 5+ years of software engineering experience, strong programming and system-design fundamentals, and expertise in build systems, CI/CD, testing, and deployment automation.
Build and operate continuous delivery infrastructure for Kubernetes deployments across global regions, including progressive rollouts, automated health evaluation, and rollback systems. The role requires strong Go or Python skills, large-scale Kubernetes experience, and familiarity with GitOps tooling.
Builds and operates resilient, observable cloud infrastructure and automation for internal IT services. The role requires 5+ years of production software engineering experience, strong Python skills, infrastructure-as-code expertise, and hands-on cloud and container experience.
Develops large-scale observability infrastructure for Databricks products and platform systems, including logging, metrics, tracing, dashboards, and alerting. The role requires 5+ years of production software engineering experience and familiarity with distributed systems and observability practices.
Owns secure, scalable Azure infrastructure for healthcare applications, including cloud migrations, Terraform-based automation, CI/CD pipelines, monitoring, and compliance. Requires 3–5+ years of Azure experience and strong DevOps and cloud-security expertise.
The Senior SRE will design, automate, and operate high-throughput AWS and Kubernetes infrastructure, improving reliability, observability, CI/CD, and cost efficiency. The role requires 4–6 years of production SRE, DevOps, or systems engineering experience and strong Terraform, Linux, scripting, and incident-response skills.
Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.
Site Reliability Engineers build and operate scalable production infrastructure, automate operational workflows, and improve observability, incident response, and service reliability. The role spans Intermediate through Senior Staff levels and requires experience with Kubernetes, infrastructure as code, cloud platforms, and software engineering.
Leads production reliability for Grafana Cloud’s multi-tenant database products, partnering with product engineering teams to improve SLOs, scalability, observability, automation, and incident response. Requires 8+ years of engineering experience, including substantial SRE or production engineering work, plus strong Kubernetes and cloud expertise.
Builds and operates core platform services, Kafka-based event streaming, and developer frameworks for high-scale distributed systems in fintech. Requires 7+ years experience with AWS, Kubernetes, and cloud-native infrastructure.
Leads the architecture, reliability, observability, and operational excellence of secure cloud and hybrid infrastructure for a mission-critical collaboration platform. Requires 5+ years in SRE, DevOps, or cloud infrastructure, with expertise in Kubernetes, Terraform, AWS, and regulated environments.
Build and operate Zoox’s large-scale HPC platform for distributed compute, storage, scheduling, and developer workflows. The Staff Engineer will lead platform strategy, reliability and scalability initiatives, and cross-functional infrastructure improvements while mentoring engineers.
Build and operate scalable cloud infrastructure and application platforms, enabling frequent deployments, resilient systems, observability, and self-healing capabilities. The role requires strong troubleshooting, Linux and cloud experience, networking knowledge, and expertise in one or more platform engineering focus areas.
Operates and scales Kraken’s core infrastructure platforms, with a focus on OpenStack, Ceph, Linux, distributed systems, and automation. The role requires 3+ years of infrastructure or software engineering experience and supports reliable compute and storage services across cloud and on-premises environments.
Build and operate highly reliable continuous integration infrastructure, intelligent test-selection systems, and incident-response automation at scale. The role requires 10+ years of experience with large-scale CI/CD systems, container orchestration, and developer productivity tooling.
Leads reliability engineering for large-scale, customer-facing cloud services, including incident response, observability, automation, and platform improvements. Requires deep Kubernetes, cloud infrastructure, infrastructure-as-code, software engineering, and technical leadership experience.
Improves facility operations through process standardization, maintenance and construction workflow optimization, operational dashboards, data analysis, and automation. Requires an engineering bachelor's degree, at least one year of operations or process-improvement experience, and Excel, SQL, or Python skills.
Own end-to-end releases for vehicle software components while coordinating cross-functional readiness and building automation for branch management, CI/CD, visibility, and deployment. The role requires 3–5 years of release or build engineering experience, Git and scripting expertise, and a bachelor's or master's degree or equivalent experience.
Build and operate scalable asynchronous platform infrastructure, developing automation and tooling while improving reliability, performance, and operational efficiency. The role requires 5+ years of software development, automation, and systems engineering experience, plus cloud and distributed-systems expertise.
Owns and evolves Webflow’s highly available, multi-cloud infrastructure, including Kubernetes, networking, infrastructure as code, observability, and AI-powered automation. The role requires 5+ years operating customer-facing cloud infrastructure and deep AWS experience.
Leads Webflow’s deployment strategy and GitOps platform, improving CI/CD reliability, progressive delivery, and developer productivity. The role requires 7+ years in DevOps, SRE, or infrastructure engineering and deep experience with Kubernetes, AWS, Docker, and infrastructure as code.
Leads developer experience and AI tooling across the engineering organization, building internal agents, ephemeral environments, and faster CI/CD workflows. Requires 5+ years in cloud infrastructure, platform, or developer tooling, plus hands-on AI assistant and LLM workflow experience.
Deploys and expands global datacenter and physical network infrastructure, coordinating contractors, vendors, installations, and operational processes. Requires at least two years of datacenter or Linux systems administration experience plus networking, configuration management, scripting, and project coordination skills.
Hardens and operates large-scale GCP and AWS infrastructure, automating security remediation, incident response, IAM, and reliability improvements. The role requires deep cloud security and DevSecOps experience, strong infrastructure-as-code skills, and expertise in Kubernetes, Linux, and security automation.
Manages data center technicians and critical infrastructure supporting AI compute systems, including power, cooling, networking, hardware deployments, incidents, vendors, and capacity expansion. Requires 5+ years in data center operations and 3+ years managing technical teams.
Supervise data center technicians while overseeing server and network infrastructure installation, maintenance, troubleshooting, and operational improvement. The role requires 5+ years of relevant hardware and repair experience, technical leadership, Linux proficiency, and scripting experience.
Own end-to-end cost optimization across AWS infrastructure and AI/LLM usage, building tooling, dashboards, and guardrails while balancing cost, performance, and reliability. Requires 4+ years of AWS infrastructure experience, cost-management expertise, and proficiency with Terraform and a programming language.
The engineer will evolve and operate Notion’s async task runner and configuration management platform, supporting reliability and scalability for more than 100 million users. The role requires at least four years of software development experience and knowledge of distributed systems, production operations, and infrastructure tradeoffs.
Deploys and expands Cloudflare’s global data center and network infrastructure, coordinating vendors and contractors while automating provisioning and operational workflows. The role requires at least five years of relevant infrastructure experience, strong networking and Linux skills, and proficiency with Python or Bash automation.
Owns production reliability (SLOs, monitoring, incident response) and platform engineering (CI/CD, infrastructure as code) for AI developer tools used by hundreds of thousands. Requires deep production systems experience, strong coding skills, and cloud proficiency.
Builds and owns distributed training infrastructure, experiment orchestration, data pipelines, and performance optimizations for large-scale AI research on GPU clusters. Requires deep systems expertise, Python/C++/PyTorch proficiency, and ML understanding to accelerate frontier research.
Build and operate the compute, orchestration, networking, developer platform, and reliability systems underlying AI agents and developer tools. The role requires large-scale infrastructure experience, Kubernetes and cloud expertise, Python proficiency, and a strong security and observability mindset.
Leads the design and validation of resilient regional infrastructure for globally scaled payment systems, including launches, failovers, migrations, and CI/CD readiness gates. Requires 12+ years of software or infrastructure engineering experience, distributed-systems expertise, cloud experience, and strong technical leadership.
Site Reliability Engineer responsible for operating and scaling a large multi-region, multi-account AWS + Kubernetes platform. Focus on automation, IaC with Terraform/Terragrunt, reducing operational toil, and owning production stateful systems end-to-end including on-call.
Owns and automates production infrastructure on multi-region AWS with EKS clusters, focusing on scaling, reliability, and self-healing systems. Requires deep Kubernetes, Terraform, and Linux expertise for large-scale stateful workloads.
The DevOps Engineer designs, automates, and supports secure, scalable AWS infrastructure, Kubernetes environments, CI/CD pipelines, and monitoring systems. The role requires 2–4 years of DevOps experience plus expertise in infrastructure as code, testing, troubleshooting, and production support.
Build and improve CI/CD, deployment, build, and validation infrastructure for embedded automotive software across SIL/HIL and cloud environments. The role requires 2+ years of experience, a bachelor's degree, and familiarity with automotive toolchains and embedded systems.
The Senior DevOps Engineer will design and operate secure, scalable cloud infrastructure, lead Kubernetes operations, automate environments, and improve CI/CD and observability. The role requires 5+ years of DevOps or infrastructure experience and expertise with Kubernetes, Terraform, Docker, and cloud platforms.