Latest DevOps / SRE jobs
Job results
Senior software engineer focused on improving production reliability, deployment safety, configuration and secrets management, and scalability across Coinbase’s service environment. The role requires 5+ years of experience with distributed systems, Ruby or Go, Terraform, cloud platforms, and observability tools.
Build and operate Harvey's core production infrastructure powering AI workloads, including Kubernetes, compute fleets, networking, and orchestration platforms. Drive reliability, scalability, security, and cost efficiency for rapidly growing LLM infrastructure while partnering across engineering teams.
Staff Production Engineer building and operating Harvey's core compute, networking, Kubernetes, and workflow orchestration infrastructure to support rapidly growing AI workloads. Requires 10+ years experience with large-scale cloud infrastructure, Kubernetes, IaC, observability, and security.
Staff Software Engineer building foundational multi-cloud platform infrastructure for Astronomer's Astro DataOps platform. Requires deep distributed systems expertise, Kubernetes operator-level knowledge, strong Go proficiency, and experience driving technical strategy at scale.
Site Reliability Engineer modernizing a multi-cloud (AWS/Azure/GCP) environment into a scalable, observable Kubernetes-based platform using DevOps/SRE practices, AIOps, IaC, and AI-driven automation to support scientific and clinical research programs. Requires 6+ years SRE/DevOps experience with strong Linux, IaC, observability, and scripting skills.
Designs, operates, and optimizes large-scale PostgreSQL and MySQL databases for mission-critical systems. Builds automation, monitoring, and high-availability infrastructure while leading incident response and collaborating with engineering teams. Requires 4+ years PostgreSQL experience.
Senior SRE specializing in Splunk observability, building scalable platforms with infrastructure as code using Terraform and Go/Python/Ruby. Requires 5+ years Splunk experience and 3+ years SRE in high-availability systems.
Leads architecture of air-gapped classified (SIPR/JWICS) developer platforms for DoD compliance, designs IaC for disconnected ops, integrates hardened tools like Big Bang/Iron Bank, and ensures secure, scalable Kubernetes infrastructure. Requires 12+ years experience with 5+ in classified DoD environments.
Lead and mentor multiple SRE teams overseeing Edge networking, Kubernetes platform, CI/CD, observability, and automation tooling for Okta’s high-scale SaaS infrastructure on AWS.
Senior Site Reliability Engineer responsible for operating and scaling a cloud-native CTV advertising platform on AWS and Kubernetes. Requires deep Kubernetes and AWS expertise, GitOps with ArgoCD, IaC with Terraform, CI/CD, observability, and incident response experience.
Staff Software Engineer builds scalable security data platforms and automates infrastructure for Okta's Public Sector using Python, Terraform, and cloud technologies. Requires 8+ years experience in security engineering and data pipelines.
Own reliability and scalability of a horizontal identity platform, including core cloud and FedRAMP environments. Build observability, automate operations, lead incident response, and ensure new features are reliable from the start. Requires production SRE experience at scale with Kubernetes, IaC, and strong programming skills.
Senior engineer building and improving Bazel-based build, test, and packaging tools for Datadog's large multi-language monorepo. Own projects end-to-end to boost developer productivity, performance, and CI efficiency at massive scale.
Build and operate infrastructure platforms supporting internal and customer-facing workloads, with a focus on reliability, automation, cloud, Kubernetes, storage, and networking. The role requires extensive Linux and AWS experience plus hands-on infrastructure-as-code and systems engineering skills.
Senior Production Engineer responsible for designing, operating, and optimizing reliable managed AI cloud services focused on scaling LLM workloads, defining SLIs/SLOs, building observability, and resolving issues in distributed systems for Crusoe's AI infrastructure.
Staff Infrastructure Software Engineer owning multi-cloud (AWS/GCP) platform, Kubernetes/Istio, Terraform IaC, networking to edge printers, and CI/CD pipelines at Carbon. Requires 7+ years production cloud infrastructure experience, expert Terraform, strong Kubernetes and networking skills.
Owns the architecture and delivery of complex IT automation workflows and shared platform components. The role requires at least five years of IT, automation, or systems engineering experience, plus expertise with workflow orchestration, Git-based development, identity tools, and collaboration APIs.
Lead the design and build of internal developer platforms and AI guardrails at Zocdoc. Focus on enabling both engineers and non-technical teams with secure, scalable, easy-to-use tools, CI/CD, and AI workflows while driving adoption through empathy and measurable outcomes. Requires 7+ years platform experience and a Bachelor's degree.
Leads SRE team to ensure high availability, scalability, and reliability of cloud services through automation, incident management, and technical leadership. Requires 8+ years SRE experience, team management, and expertise in cloud platforms and containerization.
Join Onebrief's Infrastructure & Security team as an SRE focused on improving application reliability directly in the TypeScript codebase for mission-critical military planning software. Requires active Secret clearance, 5+ years shipping application code, strong observability and incident response skills, and willingness to work onsite in Arlington, VA.
Principal Software Engineer building AI-first developer platforms and tooling for Snowflake's Snowsight UX. Requires 12+ years building tools for large codebases, experience with AI/LLMs, fluency in Go/Java/Python, and strong distributed systems fundamentals.
Senior Software Engineer responsible for building and operating the reliability, scalability, efficiency, and observability infrastructure of Starburst Galaxy. The role requires cloud architecture and orchestration experience, Java and TypeScript development, and infrastructure-as-code expertise.
Senior infrastructure engineer building scalable abstractions, platform tooling, and the full infrastructure stack (AWS to product platform) that accelerates Pilot's R&D and business growth. Requires 5+ years software engineering experience, production Python, Terraform/AWS, frontend familiarity, mentoring ability, and strong collaboration/communication skills.
Build, optimize, and maintain large-scale build systems (Bazel priority) and developer infrastructure including CI/CD, observability, and release automation for a fast-growing hardware-software platform company. Requires 4+ years experience with build systems at scale and large monorepos.
Senior DevOps Engineer building and scaling cloud infrastructure, CI/CD pipelines, observability, and developer tooling for a healthcare AI automation platform. Requires 7+ years experience, Terraform, AWS, serverless, containers, and Postgres.
Senior Software Engineer building the platform that powers secure, reproducible builds and distribution of open-source libraries across Java, JavaScript, Python/AI/ML ecosystems at Chainguard. Lead services, automation, pipelines, and developer tooling with a focus on supply-chain security, CVE remediation, and AI-assisted patching.
Build and operate the platform layer behind Cerebras engineering infrastructure, including CI/CD, Kubernetes, deployment automation, cloud and on-premises systems, developer environments, and observability. The role requires 5+ years of infrastructure or software engineering experience and strong debugging and systems fundamentals.
Owns and evolves cloud infrastructure, CI/CD, observability, developer tooling, and FinOps governance for a distributed engineering organization. Requires 6+ years in DevOps, SRE, or software engineering, plus strong Kubernetes, cloud, security, and automation experience.
Own reliability, SLOs, and observability for Supabase's deployment pipelines, control plane, and release systems as part of the Release Engineering / SRE team. Drive safe, observable deploys, disaster recovery, incident response, and toil reduction in a fully remote, async environment.
Staff Platform Engineer building scalable infrastructure to support low-latency voice AI product used hundreds of times daily by millions. Architect core platform components connecting product and ML, anticipate scaling bottlenecks, improve developer experience, and set technical direction for growing team.
Platform Engineer building core services, developer tooling, and frameworks for high-scale distributed systems and agentic applications at a consumer fintech company. Requires 2-9 years experience with AWS, distributed systems, CI/CD, and building platforms that accelerate engineering velocity.
Senior Software Engineer building Chainguard's internal Developer Platform "Factory", including monorepo CI/CD pipelines, Agentic AI platform for automated changes, and paved-road build infrastructure to reduce developer toil and accelerate secure artifact delivery.
Infrastructure Engineer building and securing Kubernetes-based platforms, cloud-native deployments, and air-gapped appliances for military planning software. Requires 5+ years production infrastructure experience, deep Kubernetes and cloud expertise, security fundamentals, and full-stack engineering skills in languages like Go or Python.
Operates and scales Crusoe Cloud’s global edge, backbone, and data center networks for GPU-based HPC infrastructure. The role requires 10+ years of production network operations experience, advanced routing knowledge, automation skills, and participation in 24/7 on-call support.
Build and own core infrastructure, observability, and tooling for a hyper-growth AI insurance platform running 200+ services and thousands of daily agentic AI decisions. Focus on scale, reliability, developer velocity, and AI eval systems.
Senior Site Reliability Engineer responsible for defining SLOs, owning production health, leading incident response, building observability with Datadog/Grafana/OpenTelemetry, and optimizing performance/scalability on AWS for a healthcare AI platform. Requires 5+ years SRE experience, distributed systems expertise, and strong automation skills.
Staff DevOps Engineer building and scaling cloud infrastructure, CI/CD pipelines, observability, and developer tooling for a healthcare AI automation platform. Requires 10+ years experience, strong IaC and AWS skills, and focus on reliability for backend/ML teams.
Maintains and enhances a high-scale core platform through monitoring, production troubleshooting, defect resolution, and reliability optimization. Requires Python proficiency, production-systems experience, and a bachelor's degree in computer engineering or a related field.
Tech Lead for Observability at Tulip, mentoring on best practices, SLIs/SLOs, and reliability while designing, building, and maintaining core observability infrastructure, tooling, and AI-enhanced monitoring for distributed systems and production incidents.
Senior Site Reliability Engineer responsible for observability, incident response, and building reliability tooling for Tulip's AI-native operations platform. Requires 5+ years with Prometheus, OpenTelemetry, and AI-driven observability tools, plus strong systems reasoning and mentoring skills.
Senior Site Reliability Engineer owning production reliability and enterprise customer implementations for Kong's fast-growing Managed Gateways SaaS product across AWS, GCP, and Azure. Requires deep Kubernetes, cloud-native, and Golang expertise plus customer-facing technical leadership.
Design and operate secure, scalable ClickHouse Cloud platforms across regulated cloud, hybrid, on-premises, and disconnected environments. The role requires 6+ years of distributed-systems experience and strong Kubernetes, infrastructure automation, cloud, database, and security expertise.
Build and maintain automation tooling for large-scale AI compute cluster deployments, turning bare-metal infrastructure into repeatable, pushbutton workflows using Python, Ansible, Terraform, Kubernetes and observability tools. Ideal for new graduates or early-career engineers seeking hands-on production infrastructure experience.
Own and evolve Elicit's cloud infrastructure platform (AWS/GCP, Kubernetes, Terraform) to support scalable single-tenant enterprise deployments. Build observability, compliance (SOC 2), cost optimization, and developer experience while contributing to backend systems where infra meets application logic. Requires 5+ years infrastructure/SRE experience, strong Terraform and K8s expertise, and enthusiasm for AI coding agents.
Senior Software Engineer on the Developer Experience team building internal tools, libraries, CI/CD pipelines, and paved paths to accelerate engineering productivity and eliminate toil across the full SDLC at Crusoe. Requires strong Go, Kubernetes, DevOps/SRE background and experience creating developer infrastructure.
Staff Software Engineer leading technical direction for Instacart's Bazel build system (remote execution, caching, performance) and Go platform (frameworks, libraries, patterns). Hands-on role driving initiatives to improve build times, CI reliability, and developer velocity for 1000+ engineers. Requires 10+ years experience with deep Bazel and Go expertise.
Lead DevOps Engineer owning multi-cloud SaaS infrastructure at scale for Tulip's AI-native frontline operations platform. Design resilient cloud architecture, CI/CD automation, observability, and mentor engineers while partnering with application teams. Requires 5-7+ years infrastructure experience and leadership.
Build and operate scalable cloud-native infrastructure for ClickHouse’s serverless database platform, including Kubernetes-based management, metrics systems, and distributed data-plane capabilities. Requires 5+ years of software development experience and production expertise with cloud platforms and Go, C++, or Java.
Build and operate ClickHouse’s cloud-native database infrastructure, including Kubernetes-based management, metrics systems, and highly available distributed services. The role requires 5+ years of software development experience, production expertise in Go, C++, or Java, and experience with public cloud and data infrastructure.
Staff Engineer on the People Technology team building and maintaining scalable automations, agents, and integrations across Workday, Greenhouse, Slack, Workato, and GCP. Requires strong software engineering skills applied to HR systems with heavy use of AI/LLMs to eliminate manual work.