Latest DevOps / SRE jobs
Job results
Nectar Social is seeking a Senior Site Reliability Engineer to own the reliability, scalability, and operational excellence of their production systems. This role involves defining SLOs, leading incident response, improving infrastructure performance, and partnering with engineering teams to embed reliability into system design.
Ironclad is seeking a Senior Staff Site Reliability Engineer to provide technical leadership and strategic direction for the SRE team, champion engineering excellence, and drive architectural resilience for their cloud platform.
As a Staff Infrastructure Engineer, you will ensure the reliability, scalability, and performance of Replit's infrastructure. You will drive automation, optimize performance, elevate developer experience, and mentor the engineering team on best practices for resilient systems.
As a Platform Engineer (SRE) focusing on the AI Control Plane, you will identify architectural changes, foster a culture of reliability, design operational processes, participate in on-call rotations, build monitoring systems, and debug production issues.
Owns the foundational infrastructure that orchestrates autonomous agents and long-running workflows in production. The role combines systems architecture, reliability engineering, end-to-end infrastructure ownership, and direct responsibility for security and compliance.
Convex is seeking a Senior Software Engineer to design, build, and maintain their global cloud infrastructure. This role involves working on core systems, improving performance and reliability, and owning architectural decisions.
Leads release management and cloud infrastructure initiatives for a highly available SaaS platform, mentoring the DevOps/Release team and improving automation, deployment, observability, security, and reliability. Requires 8+ years of enterprise SaaS development or operations experience and 6+ years with highly available cloud applications.
Hands-on technical lead owning cloud infrastructure, CI/CD pipelines, and deployment automation for a multi-tenant SaaS platform. Architect and build production systems using AWS, Terraform, CloudFormation, and Python in a GxP-regulated environment.
Design, build, and automate secure AWS cloud-native infrastructure with Kubernetes and Terraform. Enable dev teams with self-service platforms, CI/CD pipelines, and SRE best practices.
Design and operate multi-regional infrastructure for a high-traffic global gaming platform, owning on-call, monitoring, automation, and hybrid cloud migration to ensure reliability at massive scale.
Senior Infrastructure Engineer responsible for re-architecting Kubernetes infrastructure, improving continuous deployment, and making code changes across the stack to support drone platform needs.
Infrastructure engineer responsible for maintaining and scaling Kubernetes fleets, improving CI/CD, and making product-level code changes in Python or Go to support autonomous drone platform needs.
Senior SRE responsible for building and operating reliable, scalable infrastructure on AWS with Kubernetes and Terraform. Focus on observability, incident response, automation, and mentoring engineers on SRE best practices.
Staff Software Engineer building and scaling Pinterest's observability platform (metrics, logs, traces) for massive distributed systems. Requires 7+ years distributed systems experience, strong data engineering skills, and expertise with modern observability tools.
Lead development of scalable platform systems and infrastructure tools that enable internal and external developers to build faster, more reliable applications. Requires 5+ years of software engineering experience with 3+ years in Node.js.
Staff Platform Engineer building developer tooling, CI/CD automation, and scalable web applications using NodeJS, React, and AWS. Requires 10+ years experience and expertise in Temporal, Terraform, PostgreSQL, and Snowflake.
Build and operate scalable CI and Bazel-based build systems that accelerate engineering velocity and reliability for OpenAI's products and infrastructure.
Lead and mentor a network engineering team responsible for designing, deploying, and operating multi-site enterprise network infrastructure across data centers, cloud, offices, and vehicle facilities. Requires 10+ years of network experience with 5+ years in senior leadership.
Lead platform architecture and operations for a multi-tenant, multi-cloud infrastructure serving a fast-growing AI healthcare company. Design and scale Kubernetes platforms, Terraform modules, CI/CD pipelines, and observability tooling while driving security, reliability, and developer velocity.
Lead EarnIn's AI-first reliability engineering strategy. Define SLOs/SLIs, build AI agents for incident response and on-call automation, and partner with engineering teams to embed AI-assisted operations across production systems on AWS.
Design, build, and operate Render's core networking stack across data centers and clouds, focusing on Kubernetes and Linux internals, traffic routing, and hybrid connectivity at scale.
Founding Senior Platform Engineer building and owning AWS cloud infrastructure, reliability, observability, security/compliance (SOC 2, Vanta), and release tooling for a fintech platform serving banks and credit unions.
Build and scale secure multi-tenant container infrastructure and code sandboxes on AWS/GCP for an applied AI coding platform. Own reliability, observability, and performance for 500k+ containers per month.
Performance engineer focused on cross-layer investigations of Anthropic's inference fleet for Claude, optimizing throughput, latency, reliability, and correctness while building observability and partnering with kernel and serving teams.
Design, build, and operate AWS infrastructure with Terraform, CI/CD pipelines, observability, and security for a national security technology platform.
Design, build, and maintain infrastructure platforms using Linux, AWS, Kubernetes, Terraform, and Ansible to support internal and customer-facing services. Participate in on-call rotations and collaborate across engineering and operations teams.
Build and maintain internal tooling and automations to improve operational efficiency across engineering, operations, and business teams. Own end-to-end integrations, developer experience improvements, and security-compliant tool adoption.
Leads a small SRE team while owning the reliability, security, scalability, and delivery infrastructure of a GCP-based platform. The role combines hands-on cloud operations with incident management, compliance support, observability, and corporate IT leadership.
Founding Infrastructure Engineer to architect and scale resilient systems for AI/ML workloads, implement monitoring/observability, and automate infrastructure. Requires 5+ years production experience, Python, Kubernetes, and strong reliability focus.
Staff-level engineer owning design and operations of Snowflake's large-scale Kubernetes container platform across AWS, Azure, and GCP. Focus on reliability, automation, and developer experience for internal engineering teams.
Lead design and evolution of core platform services, APIs, and shared primitives that power every product surface and AI agent. Drive technical standards and architecture across SaaS, enterprise, and government environments while mentoring engineers.
Lead deployment and operations for OpenAI’s custom silicon and systems into data center environments. Drive hardware bring-up, validation, production deployment, and fleet reliability at scale while leading a technical team.
Own CI/CD and automation infrastructure for autonomy, embedded, and ground software, improving developer velocity and scaling simulation/HIL testing. Requires 7+ years experience with Azure DevOps, Docker, Kubernetes, and strong scripting skills.
Senior engineer on the Core Infrastructure team responsible for scaling data layers, observability, and developer tooling to support rapid multi-product growth at Nooks. Requires 5+ years experience scaling systems 10x+, strong distributed systems or infra background, and willingness to be in-office in San Francisco 3+ days/week.
Build and operate the large-scale training and inference infrastructure that powers Databricks AI Research, enabling researchers to run experiments across thousands of GPUs. Partner with ML scientists and platform teams to deliver reliable, high-performance orchestration and tooling.
Builds and maintains AI infrastructure using Ansible, Terraform, and Kubernetes, ensuring scalability, reliability, and high availability. Handles on-call incident response, monitoring, debugging, and infrastructure growth planning. Requires 5+ years experience and CS bachelor's.
Staff Engineer investigates field issues in autonomous hardware systems, drives root cause analysis, implements corrective actions, and improves reliability through cross-functional collaboration. Requires 5+ years in quality engineering for complex hardware in aerospace/defense/robotics.
Builds and deploys AI-native workflows for HR systems like talent acquisition, onboarding, and performance management. Integrates LLMs and APIs into tools like Workday and Greenhouse; requires 3-5 years software engineering with AI/automation experience.
The DevOps Engineer will design, scale, automate, and operate production-grade cloud infrastructure and distributed systems. The role requires 4+ years of DevOps experience, strong Docker and Kubernetes expertise, CI/CD knowledge, and Python, Bash, or Go skills.
Build and operate the platform infrastructure and developer tooling powering Lovable’s AI product, including sandbox runtimes, schedulers, observability, networking, and reliability systems. The role requires 10+ years in platform, infrastructure, or developer experience engineering and onsite work in Stockholm.
Senior Site Reliability Engineer responsible for ensuring reliability, scalability, and performance of Illumio's AWS and Azure cloud infrastructure. Lead monitoring, incident response, on-call support, automation, and continuous improvement initiatives in a cybersecurity SaaS environment.
Serves as primary Remote Pilot in Command for eVTOL UAV flight tests, managing GCS operations, securing regulatory authorizations using FAA/EASA frameworks, and collaborating with engineering on test plans and telemetry analysis. Requires FAA Part 107 certification, UAV piloting experience, and bachelor's in aerospace or related field.
Owns the reliability, performance, and availability of production database infrastructure while contributing to observability, incident response, cloud modernization, and automation. Requires 7+ years in SRE, DevOps, platform, infrastructure, or database reliability roles, including substantial production database ownership.
The Senior Database Reliability Engineer will operate and improve large-scale PostgreSQL environments in AWS, focusing on availability, performance, security, backups, recovery, and incident response. The role requires 7+ years of database administration or engineering experience, strong Linux and SQL skills, and production cloud database expertise.
Operates Lovable’s app runtime, deployment pipeline, and managed infrastructure services, with responsibility for reliability, vendor migrations, custom domains, and operational excellence. The role seeks an experienced platform, SRE, or infrastructure engineer with distributed-systems expertise and an uptime-focused mindset.
Owns and improves the Go-based Terraform provider for Supabase's developer platform, focusing on reliability, lifecycle management, schema evolution, and user migrations. Requires 5+ years experience with Go, deep Terraform expertise, and strong testing/CI/CD skills.
Develops and optimizes Supabase Edge Runtime, a Rust-based Deno host for global edge TypeScript functions. Evolves infrastructure for low-latency compute, integrates with Supabase stack, and improves developer tools. Requires 5+ years backend/systems experience with Rust, TypeScript, and scalable infra.
Designs and builds integrations, automations, and AI agent workflows to streamline Finance, Accounting, People Ops, and Recruiting processes. Requires 5+ years in system integrations, iPaaS platforms like Workato, and experience with HR/Finance tools like NetSuite and Greenhouse.
Builds and maintains automated IT infrastructure pipelines using Python, Terraform, and cloud providers to support company operations. Requires 5+ years experience with focus on automation, strong coding, and collaboration skills.
Staff Software Engineer designs, builds, and scales managed Kubernetes and AI training clusters, focusing on reliability, performance, and orchestration using Go, Terraform, and GCP. Oversees architecture, CI/CD pipelines, and critical infrastructure projects requiring 8+ years experience.