Latest DevOps / SRE jobs
Job results
Leads technical direction for Airbnb’s service developer tooling platform, spanning AI-assisted development, JVM build infrastructure, testing, modernization, and observability. Requires 9+ years of industry experience, strong backend and distributed-systems expertise, and the ability to influence organizations and deliver multi-quarter infrastructure initiatives.
Build and Release Engineer responsible for release orchestration, CI/CD pipelines, artifact lifecycle management, and an internal release portal. The role requires strong software development skills, Git expertise, and graduation by December 2026.
Build and operate reliable, secure cloud infrastructure across AWS and Kubernetes while automating delivery, observability, disaster recovery, and cost optimization. The role requires 8+ years in DevOps, SRE, or infrastructure engineering and strong hands-on experience with Terraform, Kubernetes, AWS, and CI/CD.
Leads Firefox performance engineering by writing code, profiling bottlenecks, improving benchmarks, and guiding cross-functional teams. Requires 7+ years of experience, strong C++ and JavaScript skills, and expertise in performance-critical software, profiling, concurrency, and systems analysis.
The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.
The Senior Site Reliability Engineer will build and operate secure, scalable infrastructure and Snowflake data systems, automate deployments and operational processes, and lead incident response. The role requires strong coding, Terraform, Kubernetes, CI/CD, and data-platform experience, plus U.S. Person status.
Leads reliability engineering for highly available, FedRAMP-compliant cloud services, including infrastructure architecture, automation, observability, incident response, and operational standards. Requires extensive Kubernetes, cloud, software engineering, and cross-team technical leadership experience, plus US-person eligibility and residence on US soil.
Build and operate Hebbia’s AWS infrastructure and developer platform entirely through code. The role focuses on multi-account architecture, CI/CD, container orchestration, cloud cost controls, security compliance, and scalable platform foundations, requiring 5+ years of production cloud infrastructure experience.
Build and own Webflow’s corporate cloud foundation, including landing zones, networking, security, Infrastructure as Code, GitHub delivery pipelines, self-service deployment patterns, and observability. The role requires 5+ years of platform or cloud engineering experience and strong AWS, Azure, or GCP expertise.
Owns the multi-region AI runtime and platform reliability for customer-facing products, including Bedrock infrastructure, sandboxed execution, observability, cost controls, Terraform, and delivery pipelines. Requires 5+ years in DevOps, SRE, platform, or infrastructure engineering plus production experience operating LLM-backed workloads.
The Senior Infrastructure Engineer designs and operates internal data platforms and production web-service environments, develops cloud and Linux integrations, and ensures capacity and security. The role requires strong Terraform, Kubernetes, Python, Linux, networking, and cloud-provider experience.
Leads on-site deployment of data center physical infrastructure, managing contractors, performing QA/QC on fiber optics and cabling, and ensuring compliance with standards. Requires 5+ years experience, SME-level fiber optic expertise, bachelor's degree, and 40% travel readiness.
Build and operate robust infrastructure, support enterprise deployments, and improve on-premises delivery for a rapidly scaling AI code review platform. The role requires networking expertise, cloud and container experience, and at least one year of infrastructure or software engineering experience.
Own reliability, deployments, observability, compliance, and AI infrastructure across AWS and Kubernetes for a fintech platform. The role requires strong DevOps/SRE depth, backend software engineering experience, and hands-on ownership of SOC 2 and PCI-DSS controls.
The Senior Infrastructure Engineer will build and operate high-scale infrastructure while leading Kubernetes adoption, an AWS-to-GCP migration, and PostgreSQL modernization. The role requires 5+ years of infrastructure or SRE experience, production Go or Python skills, and demonstrated expertise in scalable systems and cloud efficiency.
Leads operations outcomes for partner-operated data center sites, directing vendors, defining operational standards, and ensuring deployment velocity, availability, repair performance, and incident response. Requires 8+ years in data center or infrastructure operations, vendor oversight experience, and hands-on server, network, and rack-level expertise.
Build Mercury’s secure, observable infrastructure platform across AWS, networking, containers, and developer tooling. The role requires strong Linux fundamentals, cloud-native experience, technical writing ability, and software development skills, with opportunities to support AI-agent infrastructure.
Owns enterprise DevSecOps architecture across Salesforce, NetSuite, Workday, AEM, and modern web platforms. The role requires 8+ years of DevSecOps, SRE, or security engineering experience, strong CI/CD and edge-security expertise, and leadership in secure automation, observability, identity, and compliance.
Build and operate Mercor’s enterprise agent platform across security, routing, isolated execution, orchestration, deployment, and production scalability. The role requires 5+ years building high-scale platforms, architectural ownership, and experience with core infrastructure primitives across multiple clouds.
Build and maintain scalable software, infrastructure, and services for network automation, improving resilience and reducing operational toil. The role requires 3+ years of software development experience, Python or Go proficiency, Linux expertise, database and observability experience, and familiarity with networking systems.
Leads the design, development, and operation of Stripe’s large-scale CI and developer productivity systems. Requires 10+ years of hands-on software development, distributed-systems expertise, technical leadership, and mentoring experience.
Own and evolve a platform domain supporting reliable, secure, and cost-effective multi-tenant infrastructure. The role requires strong Kubernetes, cloud, Terraform, Helm, GitOps, Python, security, and agentic coding expertise, along with excellent technical judgment and communication.
Build and operate the Kubernetes-based platform infrastructure, developer tooling, CI/CD systems, and agent infrastructure used across the engineering organization. The role requires production cloud experience, strong infrastructure-as-code skills, Python proficiency, and sound engineering judgment.
Own and scale secure cloud infrastructure, deployments, observability, compliance, and incident response for a hardware collaboration platform. The role requires substantial cloud or security engineering experience, AWS and Linux expertise, and the ability to lead cross-functional infrastructure initiatives.
Build and operate the Kubernetes-based cloud and on-premises infrastructure powering large-scale crawling, search, and ML workloads. The role requires 5+ years in DevOps, platform engineering, or cloud infrastructure, with strong Kubernetes, cloud, Docker, Terraform, and distributed-systems experience.
Own the infrastructure, deployment, and operational tooling for Coinbase’s latency-sensitive institutional trading platform across cloud and colocated environments. The role requires 8+ years of infrastructure, platform, or SRE experience, strong Linux and networking fundamentals, and experience operating regulated, low-latency systems.
Build and operate highly scalable, reliable cloud infrastructure and the platforms, tools, and automation that support Snowflake’s globally distributed services. The role requires software engineering expertise, cloud experience, and strong skills in at least one infrastructure domain.
The DevOps Engineer will design and operate AWS and hybrid infrastructure, improve CI/CD reliability, and strengthen disaster recovery and business continuity. The role requires 5+ years of cloud infrastructure experience, strong AWS proficiency, and hands-on infrastructure-as-code expertise.
Build and operate portable infrastructure that enables Claude to run reliably across multiple cloud providers and accelerator platforms. The role requires 8+ years of distributed-systems experience, multi-cloud architecture expertise, production programming, Kubernetes, and Infrastructure as Code proficiency.
Own and evolve secure, highly available AWS and Azure infrastructure, including Terraform automation, Kubernetes, CI/CD, observability, networking, and incident response. The role requires 7+ years of DevOps or related experience and strong cross-functional partnership across engineering and security.
Build and operate declarative control planes, durable workflows, and self-healing systems that provision and manage GPU inference infrastructure. The role requires strong software engineering, reconciliation or orchestration experience, and event-driven systems expertise.
Supports reliable, secure, and scalable cloud platforms across AWS, GCP, and Azure, with a focus on Kubernetes workloads. The role monitors services, troubleshoots incidents, supports deployments, and automates operations while requiring 1–2 years of SRE, DevOps, cloud operations, or infrastructure experience.
Own the platform foundation that enables Tabs engineers to ship faster, including build systems, CI/CD, infrastructure, developer tooling, and operational tooling. The role requires 5+ years of software engineering experience, startup ownership, cloud infrastructure expertise, and production systems experience.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Build and operate an AI-first CI/CD and agent-operations platform for Salesforce and custom GTM applications. The role focuses on governed releases, approval workflows, observability, rollback, sandboxing, and SOX-compliant auditability.
Leads architecture, ownership, modernization, and operation of Komodo Health’s AWS and Kubernetes infrastructure and shared services. The role requires 8+ years of infrastructure experience, deep Terraform and Kubernetes expertise, regulated-environment security fluency, and the ability to establish AI-assisted engineering standards.
Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks supporting GPU-based HPC workloads. The role requires extensive production networking experience, strong protocol and observability expertise, automation skills, and participation in 24/7 on-call support.
Hands-on DevOps platform engineer building product-platform tooling, containerized deployments, CI/CD, and developer-experience improvements. The role requires strong Linux, containers, Kubernetes, troubleshooting, and Python or TypeScript development skills, with L2 and L3 ownership scopes available.
Build and operate reliability and resilience capabilities for a multi-cloud platform, including chaos engineering, observability-driven validation, failover testing, and resilient distributed systems. The role requires 6–9 years of software engineering experience and hands-on expertise with Java, Kubernetes, cloud platforms, and CI/CD.
The Python Engineer will improve and operate trading systems, support integrations with asset classes and prime brokers, and handle monitoring, incidents, and performance optimization. The role requires 3+ years of experience, strong Python and Linux skills, and familiarity with market data and order-entry systems.
Leads the design and development of scalable, secure network traffic systems and cloud infrastructure. The role requires 8+ years of coding experience, strong distributed-systems and concurrency expertise, and deep knowledge of networking and performance optimization.
Leads site reliability initiatives for trading systems, improving availability, scalability, monitoring, incident response, and infrastructure automation. The role requires 8+ years of DevOps, SRE, or platform engineering experience and strong expertise in Kubernetes, cloud platforms, CI/CD, and infrastructure as code.
Senior engineer responsible for scaling and operating multi-region Kubernetes, GitOps, Infrastructure as Code, security governance, and data-platform infrastructure. The role requires 8+ years of platform, SRE, or cloud data infrastructure experience and strong Kubernetes and Terraform expertise.
Provisions, commissions, and validates network, server, and related infrastructure across large-scale data center deployments. The role requires 5+ years of relevant experience, strong networking and systems troubleshooting skills, and proficiency in automation, cloud platforms, Kubernetes, Terraform, and Ansible.
Leads the technical direction of multi-cloud Kubernetes capacity management and workload placement across Datadog’s large-scale infrastructure. The role requires strong systems programming experience, ideally in Go, cloud infrastructure expertise, and the ability to influence architecture across teams.
Build and maintain cloud tooling, Continuous Delivery platforms, Terraform-based infrastructure automation, and supporting microservices across AWS ECS and EKS. The role requires 4+ years of software development experience, backend programming expertise, and strong knowledge of cloud-native technologies.
Builds and mentors development of scalable cloud tooling, Continuous Delivery platforms, Infrastructure as Code automation, and supporting microservices across AWS environments. The role requires substantial backend software development experience with Java, Go, or Python, plus Terraform, CI/CD, containers, and distributed systems expertise.
The Senior Cloud Infrastructure Engineer will design and operate secure cloud, on-premise, and air-gapped infrastructure for UK defence customers. The role requires 5+ years of production infrastructure experience, cloud and IaC expertise, Kubernetes and containerization skills, and eligibility for UK Security Clearance.
Owns and improves cloud infrastructure, CI/CD, Kubernetes, observability, security, scalability, and developer experience. The role requires at least six years of DevOps experience, strong AWS and automation expertise, and fluent Hebrew and English communication.
The Staff Production Engineer will operate and improve reliable production infrastructure, with a strong focus on data center networking, capacity planning, troubleshooting, automation, and hardware operations. The role requires 5+ years of production engineering experience, networking expertise, and proficiency with infrastructure and CI/CD tools.