Latest DevOps / SRE jobs
Job results
Staff Infrastructure Engineer responsible for designing and operating scalable infrastructure for growth systems, including onboarding, referrals, and user acquisition. The role requires 7+ years of production infrastructure experience, strong reliability instincts, and independent judgment in a high-autonomy environment.
Senior Site Reliability Engineer responsible for designing and operating reliable, scalable production infrastructure, leading incident response, and improving observability and resilience. Requires 5+ years of reliability-focused engineering experience and expertise across cloud, infrastructure as code, Kubernetes, monitoring, and application development.
Leads the architecture, development, and operation of cloud, Kubernetes, on-premises, and hybrid infrastructure, while building developer platforms and CI/CD automation. Requires at least six years of infrastructure or related engineering experience, deep Kubernetes expertise, strong programming skills, and technical leadership.
Senior platform engineer responsible for reliable, secure, and scalable infrastructure, developer tooling, observability, and AI enablement. The role requires 6+ years in platform engineering, SRE, or DevOps, with strong AWS and incident leadership experience.
Build and operate scalable GPU/TPU HPC infrastructure for training and serving frontier AI models. The role partners with AI researchers, optimizes distributed workloads across clouds, and requires expertise in Kubernetes, Python, Go, Linux, and high-performance networking.
Supports cloud infrastructure, automation, CI/CD, monitoring, and service reliability while learning alongside a global DevOps team. The entry-level role requires a bachelor’s degree, foundational systems knowledge, and exposure to cloud and DevOps tools.
Own the network architecture and standards for a multi-cloud enterprise AI platform deployed across Kubernetes environments and customer-controlled networks. The role requires deep cloud and Kubernetes networking expertise, strong security fundamentals, and the judgment to establish scalable, supportable connectivity patterns.
Designs, operates, and troubleshoots large-scale multi-vendor networks supporting high-performance AI compute infrastructure. Requires 8+ years of production networking experience, deep routing expertise, automation skills, and hands-on RoCE or InfiniBand experience.
Build and operate a Kubernetes-native control plane for provisioning, scheduling, self-healing, and optimizing GPU inference infrastructure. The role requires strong software engineering, durable workflow orchestration, reconciliation systems, event-driven architecture, and platform API experience.
Staff Platform Engineer will build and improve automated delivery pipelines, developer environments, infrastructure, and release systems across the engineering organization. The role requires 6+ years of engineering experience, a bachelor’s degree, and expertise with CI/CD, cloud infrastructure, containers, and infrastructure as code.
Build and operate autonomous infrastructure systems for large-scale GPU fleets, including cluster lifecycle automation, fleet intelligence, validation, and remediation. The role requires 3+ years of distributed systems or infrastructure engineering experience and strong Python, Go, or Rust skills.
Leads technical cloud operations by standardizing production changes, building service procedures, and improving automation, documentation, and self-service. The role requires production cloud or SaaS operations experience, strong operational judgment, and familiarity with cloud-native systems such as AWS and Kubernetes.
Senior Release Engineer responsible for building reliable CI/CD pipelines and release automation for enterprise SaaS platforms such as Salesforce and Zuora. The role requires 7+ years of release engineering or DevOps experience, strong Python skills, and hands-on use of approved AI-assisted tools.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.
Staff DevSecOps Engineer designing and automating security controls across AWS infrastructure, containers, CI/CD, and platform services. Requires 7+ years of related experience plus expertise in cloud security, infrastructure as code, hardened images, vulnerability scanning, identity, and secrets management.
Designs, automates, and operates AWS infrastructure, shared development environments, and container platforms. The role requires strong experience with Kubernetes, infrastructure as code, environment lifecycle automation, cloud security, compliance, and cost optimization.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services, improving observability, resilience, incident response, and operational tooling. Requires 7+ years of experience, major-cloud infrastructure expertise, infrastructure as code, distributed systems, and strong technical leadership.
Leads the establishment and maturation of SRE practices across cloud infrastructure and platform services. This hands-on technical role focuses on reliability targets, observability, incident response, resilience, automation, and mentoring engineering teams.
Operates and improves large-scale NVIDIA GPU infrastructure supporting AI model training, spanning provisioning, scheduling, networking, storage, automation, performance tuning, and hardware operations. Requires production Linux or GPU environment experience and strong systems and infrastructure skills.
Own reliability, scalability, security, observability, and incident response for mission-critical applications across Kubernetes, AWS, and on-premise DoD environments. Requires an active Top Secret clearance and at least five years of infrastructure-focused SRE, DevOps, or platform engineering experience.
Designs and operates highly available OT infrastructure for AI supercomputer campuses, including power, cooling, facility controls, and industrial software platforms. Requires at least three years of OT or ICS administration experience, strong systems and networking skills, and onsite availability in the Memphis/Southaven area.
Designs, deploys, and operates high-performance networks powering AI supercomputer campuses, including training fabrics, storage, OT, and site networks. Requires substantial data-center networking experience, automation expertise, and readiness for on-call, hands-on infrastructure work.
Leads cross-functional technical initiatives and builds scalable business operations and customer-facing systems. Requires Python, system design, production engineering experience, and strong stakeholder collaboration; platform, AWS, SaaS, and analytics experience are preferred.
Own and improve the Linux production infrastructure layer, from performance tuning and incident response to configuration management, orchestration, networking, virtualization, secrets, and observability. The role requires 6+ years of infrastructure or SRE experience and deep Linux expertise.
Own the reliability, resilience, observability, and automation of AWS and Kubernetes infrastructure supporting production products and AI/ML workloads. The role requires 4+ years of cloud infrastructure experience, strong Kubernetes and Terraform expertise, and senior-level incident response and software engineering skills.
Leads campus-scale site reliability for a data center environment, owning observability, incident command, postmortems, runbooks, and cross-functional reliability initiatives across infrastructure and facilities. Requires a bachelor's degree or equivalent experience and at least five years in SRE, systems engineering, or large-scale operations.
Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.
Build and operate cloud-agnostic deployment and observability infrastructure across public clouds and on-premises environments. The role requires 5+ years of infrastructure experience, strong networking and IaC expertise, and ownership of production systems and cross-functional projects.
Build and operate Tempo’s blockchain infrastructure, improving reliability, developer velocity, observability, and enterprise validator onboarding. The role requires production experience with bare-metal and cloud systems, Kubernetes, infrastructure as code, scripting, Linux, and networking.
Build and operate foundational observability infrastructure spanning telemetry pipelines, profiling, tracing, and diagnostic tooling across large-scale compute clusters. The role requires deep systems-level experience and 10+ years of relevant industry experience.
Senior site reliability engineer who will build and operate observability, anomaly detection, reconciliation, and reliability tooling for GitLab’s monetization systems. The role requires Ruby on Rails and observability experience, with knowledge of monitoring platforms, data pipelines, and business-critical billing systems.
Supports production operations by monitoring batch processes, troubleshooting issues, documenting procedures, and collaborating with technical teams in a 24/7 environment. The role is suited to a student pursuing a bachelor’s degree with an interest in technical operations, enterprise software, and financial services.
Designs and operates secure, highly available cloud infrastructure supporting engineering teams, with a focus on GCP, GKE, Terraform, Kubernetes, observability, and developer self-service. Requires 5–8 years of production infrastructure experience and strong cloud, automation, and Linux expertise.
Leads the design and deployment of AI-enabled manufacturing systems, MES, connected-factory infrastructure, and automation for aircraft production. Requires a bachelor’s degree and 8+ years of experience in digital manufacturing, industrial automation, or software-enabled operations.
Infrastructure Engineer responsible for securing, scaling, and operating cloud infrastructure and machine-learning platforms across a distributed startup. The role requires Kubernetes, infrastructure-as-code, cloud operations, CI/CD, programming, observability, and security experience.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting mission-critical payment systems. Requires 8+ years of distributed-systems experience and deep expertise in infrastructure as code, Kubernetes, automation, and cloud networking.
Leads the architecture, automation, observability, and reliability of multi-region AWS infrastructure supporting high-throughput payments. Requires 10+ years of distributed-systems experience and deep expertise in cloud infrastructure, Kubernetes, infrastructure as code, and modern SRE practices.
Leads development of Coinbase’s CI, build, and deployment infrastructure used by engineers across the organization. The role requires 8+ years building production distributed systems, strong Go or systems-language expertise, and demonstrated technical leadership across complex platform initiatives.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Supports and evolves the networking, compute, Kubernetes, and ingress infrastructure powering PagerDuty’s real-time platform. Requires 0–1+ years of relevant experience, Linux production operations, cloud infrastructure knowledge, programming proficiency, and Infrastructure as Code experience.
Operates and evolves foundational networking, compute, Kubernetes, and ingress infrastructure for PagerDuty’s real-time platform. Requires 3+ years in SRE, DevOps, or platform engineering, with Linux production operations, cloud infrastructure, programming, and Infrastructure as Code experience.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
Leads reliability engineering for Reddit’s critical user-facing systems, improving availability, scalability, performance, automation, and incident response at internet scale. Requires 8+ years operating distributed systems and strong expertise in programming, observability, high availability, and production troubleshooting.
Leads end-to-end infrastructure for a scientific imaging platform, covering Linux administration, GPU/HPC systems, storage, upgrades, and vendor coordination. The role supports AI-enabled imaging workflows and requires extensive production Linux, Image Artist, GPU, HPC, and enterprise storage experience.
Staff Cloud Infrastructure Engineer will architect secure, scalable AWS and data-center infrastructure while building identity platforms, automation, and security controls for AI-powered systems. The role requires 7+ years of relevant experience, strong AWS and Kubernetes expertise, and proficiency in Python or Go and Terraform.
Leads the design and operation of CI/CD, build infrastructure, cloud systems, and developer tooling that improve engineering productivity. Requires 7+ years of infrastructure or software engineering experience, strong automation and troubleshooting skills, and technical leadership across teams.
Junior Site Reliability Engineer supporting production operations, observability, incident response, and automation for critical services. The role suits candidates with 0–2 years of experience, programming or scripting skills, and an interest in cloud infrastructure and distributed systems.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.
The Senior DevOps Engineer will evolve multi-cloud infrastructure, production Kubernetes platforms, AI workloads, databases, observability, networking, and automation. The role requires 7+ years in infrastructure, DevOps, or SRE, strong Terraform and Kubernetes expertise, and proficiency in Python or Go.