Skip to content
1,066 jobs

Job results

Watershed

Watershed

New York, NY
Software engineer, cloud infrastructure
$174k+/yrHybrid3+ YOEDevOps / SRE

Builds and maintains cloud infrastructure systems on Google Cloud to support engineering teams in deploying, scaling, and observing production workloads. Requires 3+ years experience in infrastructure engineering.

Alembic

Alembic

Senior Site Reliability Engineer
$210k+/yrOn-site8+ YOEDevOps / SRE

Builds and maintains scalable infrastructure for real-time analytics and ML workloads, focusing on reliability, automation, CI/CD, monitoring, and incident response. Requires 8+ years SRE/DevOps experience with Kubernetes, Terraform, Linux, and observability tools.

Docker

Docker

Seattle, WA

Senior Principal Software Engineer, Infrastructure
$251k+/yrRemote12+ YOEDevOps / SRE

Technical visionary architecting Docker's foundational platform for accounts, billing, data, governance, and infrastructure. Drives cross-company strategy enabling enterprise growth, requiring 12+ years experience in large-scale distributed systems.

OpenAI

OpenAI

San Francisco, CA

Release Engineer, Consumer Products
$293k+/yrHybridDevOps / SRE

Designs and operates CI/CD pipelines and release infrastructure for multi-component systems including bootloaders, firmware, and OTA updates. Requires strong automation skills in Python/Bash, Linux expertise, and experience with build systems for embedded/consumer products.

Palantir

Palantir

New York, NY

Platform Intelligence Engineer
No salary listedOn-site3+ YOEDevOps / SRE

Designs and maintains data infrastructure and pipelines to power Palantir's internal decision-making on product usage, stability, costs, and revenue. Partners cross-functionally to analyze data and deliver executive insights using Python and Palantir's platform. Requires 3+ years experience and quantitative degree.

Vanta

Vanta

Remote

Senior Software Engineer, Enterprise Resilience
$207k+/yrRemoteDevOps / SRE

Build and operate resilient systems for Vanta's FedRAMP and enterprise environments, define reliability frameworks, and partner with teams to ensure scalable, compliant infrastructure using AWS and modern tooling.

Fluidstack

Fluidstack

New York, NY
Network Engineer, Deployment & Integration
$150k+/yrOn-site3+ YOEDevOps / SRE

Hands-on network engineer deploying and validating large-scale AI datacenter fabrics, configuring switches, troubleshooting physical/optical layers, and coordinating cross-functional teams. Requires 3-7 years datacenter experience and 70-80% travel to onsite locations.

Xdof

Xdof

Software Engineer, Infrastructure
No salary listedHybridDevOps / SRE

Builds scalable infrastructure platforms for data collection systems, including orchestration, dev platforms, and multi-tenant data lakes supporting exabyte-scale robotics data. Requires Python fluency and experience with distributed systems; AI/robotics background preferred.

Luma AI

Luma AI

Palo Alto, CA

Software Engineer - Reliability
$170k+/yrRemote8+ YOEDevOps / SRE

Builds, maintains, and scales multi-cloud GPU infrastructure for AI training/inference, focusing on reliability, performance tuning, automation, and security in a fast-paced startup. Requires 8+ years SRE experience with deep Linux, cloud, and high-performance networking expertise.

Latent

Latent

San Francisco, CA

Site Reliability Engineer
$200k+/yrOn-siteDevOps / SRE

Owns production infrastructure for clinical AI platform, ensuring 99.9%+ stability. Designs/scales Kubernetes-based systems, optimizes TypeScript/Python/ML CI/CD pipelines, and manages Terraform IaC in high-velocity environment.

Kalshi

Kalshi

New York, NY

Site Reliability Engineer
$100k+/yrOn-site4+ YOEDevOps / SRE

Site Reliability Engineer enhances system observability, reliability, and availability at a prediction markets platform. Builds automation, optimizes cloud infrastructure (Kubernetes, Docker, Terraform), debugs issues, and participates in on-call rotations. Requires 4+ years software engineering experience.

Orb

Orb

Software Engineer, Infrastructure - San Francisco HQ
$170k+/yrHybrid5+ YOEDevOps / SRE

Senior infrastructure engineer focused on building reliable, scalable systems for event ingestion, APIs, and billing. Requires 5+ years software engineering experience with 4+ years in infrastructure, strong debugging skills, and mentoring ability.

Fundamental Research Labs

Fundamental Research Labs

Europe

DevOps Engineer
No salary listedRemote5+ YOEDevOps / SRE

Designs and maintains cloud infrastructure, Kubernetes clusters for GPU/ML workloads, implements GitOps with ArgoCD and Terraform IaC. Requires 5+ years DevOps experience, Kubernetes expertise, AWS/GCP proficiency, and Python.

Harvey

Harvey

San Francisco, CA

Senior Software Engineer, Site Reliability Engineer (SRE)
$200k+/yrOn-site5+ YOEDevOps / SRE

Senior SRE ensures reliability, scalability, and performance of legal AI platform by managing global infrastructure, leading incident response, automating operations, and optimizing costs. Requires 5+ years SRE experience, IaC expertise, cloud proficiency, and strong programming skills.

Harvey

Harvey

San Francisco, CA

Staff Software Engineer, Site Reliability Engineer (SRE)
$238k+/yrOn-site10+ YOEDevOps / SRE

Staff SRE ensures reliability, scalability, and performance of legal AI platform across global regions. Leads incident management, automates operations, optimizes infrastructure costs, and mentors teams. Requires 10+ years SRE experience, IaC, cloud platforms, and observability tools.

Kalshi

Kalshi

New York, NY

Infrastructure Engineer
$100k+/yrOn-site3+ YOEDevOps / SRE

Designs, builds, and scales infrastructure for a prediction market exchange including AWS, Kubernetes, high-performance APIs, and clearing systems. Requires 3+ years experience with strong fundamentals in cloud, containers, and DevOps tooling.

Rippling

Rippling

Bengaluru, India

Senior Software Engineer - Codeship
No salary listedOn-site5+ YOEDevOps / SRE

Develops foundational developer-experience and infrastructure systems for the Codeship team, improving CI/CD, development environments, engineering workflows, and shipping velocity. Requires 5+ years of backend or infrastructure software engineering experience, with Python or Go and scalable systems expertise.

Clickhouse

Clickhouse

United States

Senior Infrastructure Engineer - Postgres
No salary listedRemote7+ YOEDevOps / SRE

Senior infrastructure engineer responsible for reliability, automation, observability, and operations of ClickHouse’s Postgres integration across multi-cloud environments. The role requires 7+ years of infrastructure experience, strong PostgreSQL and AWS expertise, and proficiency with Terraform, Kubernetes, and Go.

Clickhouse

Clickhouse

United States

Senior Infrastructure Engineer - Postgres
No salary listedRemote7+ YOEDevOps / SRE

Own reliability, automation, observability, and operations for ClickHouse’s Postgres integration across multi-cloud environments. The role requires 7+ years of infrastructure or SRE experience, strong Postgres and AWS expertise, and proficiency with Terraform, Kubernetes, and Go.

Cursor

Cursor

San Francisco, CA
Software Engineer, Client Infrastructure
No salary listedHybridDevOps / SRE

Designs and builds performant, stable client infrastructure for Cursor's desktop app across macOS, Windows, and Linux. Requires deep experience in client systems, performance, reliability, and high-performance desktop applications.

Render

Render

United States
Product Lead, Infrastructure
$207k+/yrRemote8+ YOEDevOps / SRE

Lead product vision and roadmap for Render's infrastructure platform supporting millions of developers. Requires 8+ years in product management focused on developer tools, infrastructure, or data products, with strong AI and developer experience interest.

Ashby

Ashby

Birmingham, AL
Staff Platform Engineer
£151k+/yrRemote7+ YOEDevOps / SRE

Build and operate a scalable, secure platform supporting Ashby’s growing product and engineering organization. The role combines infrastructure engineering, reliability, developer tooling, security, database optimization, and hands-on software development with substantial end-to-end ownership.

Ashby

Ashby

Stockholm, Sweden
Staff Platform Engineer
€154k+/yrRemote7+ YOEDevOps / SRE

Build and operate Ashby’s scalable platform, improving reliability, security, deployment workflows, and developer experience. The role requires strong software engineering skills, infrastructure automation experience, operational judgment, and comfort owning projects end-to-end in a distributed environment.

Ashby

Ashby

Austin, TX
Staff Platform Engineer, Americas
$190k+/yrRemoteDevOps / SRE

Staff Platform Engineer builds and scales infrastructure, optimizes compilers and databases, implements deployment tools like canary deploys and feature flags, and ensures reliability with SLOs/SLIs on AWS/Kubernetes. Requires strong coding skills in TypeScript/Node.js and handling diverse infra challenges end-to-end.

Voltai

Voltai

Palo Alto, CA

Network Development Engineer
No salary listedOn-site5+ YOEDevOps / SRE

Designs and deploys high-performance network architectures for compute infrastructure and distributed AI simulation environments. Optimizes for latency, throughput, fault tolerance; requires 5+ years in high-speed networking protocols like Ethernet, InfiniBand, RDMA.

Fathom - AI Notetaker

Fathom - AI Notetaker

Remote

Infrastructure Engineer
$180k+/yrRemoteDevOps / SRE

Scales infrastructure, builds automation and internal tooling, and enhances observability on GCP/GKE for a remote-first SaaS platform. Requires IaC/GitOps expertise, observability practices, and familiarity with message queues, Prometheus, and Golang.

Lovable

Lovable

Stockholm, Sweden

Software Engineer, Platform
No salary listedOn-siteDevOps / SRE

Build and operate high-throughput infrastructure for AI engineering workloads, including sandbox runtimes, schedulers, networking, and reliability systems. The role requires deep production infrastructure experience, systems programming expertise, and strong knowledge of orchestration and isolation.

Lovable

Lovable

Stockholm, Sweden
Software Engineer, Platform
No salary listedOn-site7+ YOEDevOps / SRE

Own and scale Lovable’s developer platform, improving engineering velocity, observability, reliability, and application frameworks. The role requires 7+ years of platform experience, strong programming skills, and expertise with Kubernetes, Docker, and modern infrastructure practices.

Illumio

Illumio

Sunnyvale, CA

Sr. Member of Technical Staff - Platform
$153k+/yrOn-siteDevOps / SRE

Designs, builds, and maintains cloud-native platform services using Kubernetes for multi-tenant deployments. Develops microservices in Ruby/Java/Go, manages CI/CD pipelines with GitOps, and ensures production reliability across AWS/Azure. Requires 3-5 years experience and strong Kubernetes expertise.

Partiful

Partiful

New York, NY

🛠️ Staff Platform Engineer
$190k+/yrOn-site7+ YOEDevOps / SRE

Leads architecture and scaling of GCP serverless infrastructure powering a high-traffic social events app. Drives reliability, developer velocity via CI/CD and AI tooling, mentors engineers. Requires 7+ years backend/infra experience with Node.js/TypeScript.

Partiful

Partiful

New York, NY

🛠️ Platform Engineer
$155k+/yrHybrid3+ YOEDevOps / SRE

Builds and maintains scalable infrastructure on GCP, including serverless systems, CI/CD pipelines, and observability tools for a high-growth social events app. Requires 3+ years in infrastructure/backend engineering with serverless experience and strong ownership mindset.

Palantir

Palantir

Palo Alto, CA
Forward Deployed Infrastructure Engineer, New Grad
No salary listedHybridDevOps / SRE

New grad Forward Deployed Infrastructure Engineer building, operating, and maintaining scalable, reliable infrastructure and services for Palantir's government platforms. Requires strong coding skills, comfort with production systems, automation focus, and US security clearance eligibility.

Cerebras Systems

Cerebras Systems

United States
Principal Engineer, AI Inference Reliability
No salary listedOn-site7+ YOEDevOps / SRE

Leads reliability strategy and hands-on implementation for a large-scale, low-latency AI inference service. The role requires 7+ years in backend, infrastructure, or reliability engineering, strong backend programming skills, and deep expertise in distributed-system reliability.

Relace

Relace

San Francisco, CA

Infrastructure Engineer
No salary listedOn-site2+ YOEDevOps / SRE

Designs and operates high-performance inference and training infrastructure for ML models, focusing on GPU scheduling, distributed systems, and cloud optimization. Requires 2+ years experience with cloud platforms like AWS/GCP/Azure.

Parabola

Parabola

New York, NY
Senior Software Engineer, Site Reliability
$180k+/yrHybrid5+ YOEDevOps / SRE

Senior SRE measures software performance, defines SLOs/SLAs, optimizes infrastructure with Temporal/Kubernetes/AWS, handles on-call, and improves developer experience/scalability for growing B2B workflows. Requires 5+ years SRE/DevOps experience.

Siftstack

Siftstack

Marina Del Rey, CA
Senior Software Engineer, Infrastructure
$170k+/yrHybrid8+ YOEDevOps / SRE

Designs, builds, and maintains scalable infrastructure for real-time telemetry platform supporting mission-critical systems. Requires 8+ years in distributed systems, cloud environments (AWS/GCP/Azure), Kubernetes, Docker, and DevOps tools.

Siftstack

Siftstack

Marina Del Rey, CA
Software Engineer, Infrastructure
$150k+/yrHybrid3+ YOEDevOps / SRE

Builds and maintains scalable infrastructure for real-time telemetry platform using cloud, containers, and DevOps tools. Requires 3+ years in distributed systems, hands-on with Kubernetes, Docker, AWS/GCP/Azure.

Cape

Cape

Arlington, VA
Software Engineer, Infrastructure
No salary listedRemote4+ YOEDevOps / SRE

Builds and maintains privacy-focused telecommunications infrastructure, including monitoring, high-availability systems, and FedRamp compliance. Requires 4+ years SRE experience, AWS expertise, and fluency in Golang/Rust/Java/Python.

OpenAI

OpenAI

San Francisco, CA

Software Engineer, Reliability
$230k+/yrOn-siteDevOps / SRE

Builds and maintains scalable, reliable infrastructure including testing tools, automation, and resource management platforms for AI systems. Collaborates cross-functionally to ensure high availability, performance, and fault tolerance in a fast-paced environment.

Crusoe

Crusoe

San Francisco, CA
Staff Network Engineer, Deployment
$193k+/yrOn-site8+ YOEDevOps / SRE

Leads physical and logical deployment of global network infrastructure for AI data centers, including rack/stack, cabling, automation with Python/Ansible, testing, and partner coordination. Requires 8+ years experience with Arista, Juniper, Mellanox, BGP/EVPN, and physical layer expertise.

Cerebras Systems

Cerebras Systems

United States
Site Reliability Engineer - Ops & Automation
No salary listedOn-site5+ YOEDevOps / SRE

Operates and automates production infrastructure for a high-scale AI inference service. The role requires production Kubernetes experience, Python or Go proficiency, observability expertise, and a focus on reliability, automation, and reducing operational toil.

Render

Render

United States
Software Engineer, Infrastructure
$170k+/yrRemote5+ YOEDevOps / SRE

Builds and scales cloud infrastructure for Render's developer platform, focusing on container orchestration, networking, storage, and AI workloads. Requires 5+ years experience with Kubernetes, IaC tools like Terraform/Pulumi/Ansible, and production systems at scale.

Baseten

Baseten

San Francisco, CA
Site Reliability Engineer (SRE)
$165k+/yrHybridDevOps / SRE

Site Reliability Engineer builds and maintains scalable infrastructure for ML model deployment, automates CI/CD pipelines, and ensures reliability using tools like Kubernetes and Terraform. Collaborates cross-functionally, owns projects end-to-end, and mentors juniors; bachelor's in CS or related field required.

Amperoshealth

Amperoshealth

New York, NY
Software Engineer, Infra
$200k+/yrOn-site5+ YOEDevOps / SRE

Leads infrastructure initiatives including AWS management, dev velocity improvements, AI/LLM observability, cost optimization, and compliance. Requires 5+ years infra experience, high agency, and strong communication to shape and grow the team.

E2b

E2b

San Francisco, CA
Platform Engineer
$180k+/yrOn-site5+ YOEDevOps / SRE

Builds backend infrastructure and core platform for AI agent cloud, including VM hypervisors, LLM sandboxes, networking, and orchestration. Requires 5+ years in distributed systems and Linux administration for onsite role in San Francisco.

Console

Console

San Francisco, CA

Platform Engineer
$200k+/yrHybrid5+ YOEDevOps / SRE

Builds scalable cloud infrastructure on AWS/GCP, architects self-hostable platforms, and improves developer tooling for enterprise AI agent deployment. Requires 5+ years backend/platform experience with IaC tools like Terraform.

Black Forest Labs

Black Forest Labs

Freiburg, Germany

Member of Technical Staff - Research Infrastructure Engineer
€100k+/yrOn-siteDevOps / SRE

Build and operate large-scale research infrastructure for generative AI training, including GPU clusters, distributed systems, telemetry, and reliability tooling. The role requires deep cloud infrastructure expertise, Kubernetes, infrastructure as code, and experience operating large-scale training platforms.

Collate

Collate

San Francisco, CA

DevOps / Systems Engineer
No salary listedOn-siteDevOps / SRE

Owns infrastructure for AI document platform in life sciences, building CI/CD pipelines, managing cloud systems (AWS, Kubernetes), monitoring, and light security to ensure reliability and scale.

Numeral

Numeral

San Francisco, CA

Software Engineer (Infra)
$180k+/yrHybrid7+ YOEDevOps / SRE

Builds and scales core infrastructure for a high-growth AI tax platform, focusing on reliable APIs, data pipelines, observability, and fault-tolerant distributed systems. Requires 7+ years experience with Node.js, PostgreSQL, Redis, AWS, Kubernetes, and observability tools.

Exa

Exa

San Francisco, CA

Software Engineer, Infrastructure
$180k+/yrOn-siteDevOps / SRE

Builds and operates large-scale infrastructure including GPU clusters, Kubernetes orchestration, AWS batch jobs, and observability tooling to power AI search systems. Requires experience with massive-scale systems and focus on reliability and optimization.