Skip to content
1,066 jobs

Job results

Blacksmith

Blacksmith

New York, NY

Principal Systems Engineer
$280k+/yrOn-siteDevOps / SRE

Principal Systems Engineer sets technical direction for core infrastructure, owns architecture for reliability and performance at scale, and mentors senior engineers. Requires deep expertise in virtualization, distributed storage like Ceph, and Linux kernel primitives.

Harvey

Harvey

Bengaluru, India

Staff Site Reliability Engineer
No salary listedHybrid12+ YOEDevOps / SRE

Build and operate reliable, scalable infrastructure for a legal AI platform, leading observability, incident response, automation, capacity planning, and security practices. The role requires 12+ years of SRE or comparable production experience and strong cloud, Kubernetes, programming, and infrastructure-as-code expertise.

Crusoe

Crusoe

San Francisco, CA

Network Architect
$195k+/yrOn-site8+ YOEDevOps / SRE

Defines and governs network architecture, security strategy, and standards for data centers, power plants, and corporate environments. Requires 8+ years network engineering with expertise in IT, cloud, and industrial systems like IoT/OT.

Mark43

Mark43

New York, NY
DevOps Engineer, DevEx
$140k+/yrRemote5+ YOEDevOps / SRE

Builds and evolves internal developer platforms using Kubernetes, Terraform, and GitOps to enhance reliability, scalability, and DevEx. Requires 5+ years in platform engineering, strong AWS and cloud-native expertise, with on-call responsibilities.

Scrunch

Scrunch

United States
Senior Infrastructure Engineer
$140k+/yrRemoteDevOps / SRE

Senior Infrastructure Engineer designs, builds, and operates cloud infrastructure, developer tooling, observability, and reliability systems at scale, primarily on GCP. Requires high-velocity dev experience, IaC, database scaling, workflow orchestration, and production Python coding.

Counsel Health

Counsel Health

New York, NY
Infrastructure Engineer (Backend)
$180k+/yrHybrid3+ YOEDevOps / SRE

Builds and maintains cloud infrastructure with Terraform, owns CI/CD pipelines, writes backend services, and improves developer tooling in AWS/GCP for AI-driven healthcare platform. Requires 3-5 years experience in backend/infra hybrid roles.

MongoDB

MongoDB

Austin, TX
Site Reliability Engineer (Senior or Staff), Infrastructure Security
$127k+/yrHybrid6+ YOEDevOps / SRE

Senior or Staff Site Reliability Engineer leads design and implementation of cloud security solutions (AWS, Azure, GCP), builds automation for monitoring and alerting, and mentors SRE team. Requires 6+ years SRE/infra experience with security focus, IaC proficiency, and cloud expertise.

Agora

Agora

Jersey City, NJ

Senior Infrastructure Engineer
No salary listedRemote5+ YOEDevOps / SRE

Senior Infrastructure Engineer designs and owns internal platforms including TypeScript Pulumi, Kubernetes environments, and monitoring to enable reliable software shipping. Requires 5+ years experience with AWS, Kubernetes, IaC, and distributed systems.

Tavily

Tavily

New York, NY
DevOps Engineer
No salary listedOn-site3+ YOEDevOps / SRE

Owns and manages Kubernetes clusters, infrastructure as code, CI/CD pipelines, real-time data pipelines, monitoring, and production debugging for large-scale AI infrastructure. Requires 3+ years DevOps experience with distributed systems and cloud environments.

MongoDB

MongoDB

New York, NY
Team Lead, Site Reliability Engineering - Storage Layer Service
$151k+/yrHybrid10+ YOEDevOps / SRE

Leads a team of SREs for MongoDB's Storage Layer Services, defining SLOs, capacity plans, and roadmaps for multi-tenant distributed storage systems underpinning Atlas. Requires 10+ years in distributed systems and 2+ years managing teams, with expertise in Kubernetes and IaC tools.

Bland AI

Bland AI

San Francisco, CA

Senior Infrastructure Engineer
$120k+/yrOn-site5+ YOEDevOps / SRE

Builds and scales distributed systems for real-time voice processing, ML inference, and telephony integration using Kubernetes. Requires 5+ years experience with cloud infrastructure, real-time systems, and tools like Terraform and Datadog.

Crusoe

Crusoe

San Francisco, CA
Principal Production Engineer
$261k+/yrOn-site15+ YOEDevOps / SRE

Owns reliability, scalability, and observability of cloud infrastructure including compute, storage, and networking at massive scale. Drives SLOs, incident response, tooling, and mentors engineers; requires 15+ years experience with data centers and internet-scale operations.

Mariana Minerals

Mariana Minerals

Houston, TX

Staff Process Engineer
$140k+/yrOn-site10+ YOEDevOps / SRE

Leads design, operation, and optimization of pilot-scale minerals processing equipment and experiments. Bridges R&D to commercial scale with 10+ years experience in chemical/process engineering and unit operations like leaching and electrowinning.

Rippling

Rippling

San Francisco, CA

Staff Software Engineer - Cloud Infrastructure
No salary listedHybrid8+ YOEDevOps / SRE

Staff Software Engineer designs and scales cloud infrastructure, managing Kubernetes fleets, multi-region recovery, and distributed systems to support massive growth. Requires 8+ years experience with deep expertise in AWS, Kubernetes, and backend languages like Python/Go/Java.

Saga

Saga

Los Altos, CA

Senior Platform Engineer
No salary listedRemote7+ YOEDevOps / SRE

Senior Platform Engineer owns backend services, APIs, infrastructure, CI/CD pipelines, and platform tooling to enable scalable product delivery in fintech/crypto. Requires 7+ years backend experience, Golang proficiency, Kubernetes, and cloud infrastructure expertise.

Harvey

Harvey

San Francisco, CA

Staff Database Administrator (DBA)
$191k+/yrHybridDevOps / SRE

Leads PostgreSQL database infrastructure for a global legal AI platform, owning migration governance, multi-region scaling, reliability, performance tuning, and self-service tooling. Requires 10+ years experience with expert PostgreSQL knowledge and staff-level impact.

Acryldata

Acryldata

Palo Alto, CA

Senior Platform Engineer
$225k+/yrHybrid4+ YOEDevOps / SRE

Leads development of DataHub's ingestion framework, building scalable metadata systems, APIs, and event-driven architectures for enterprise AI and data platforms. Requires 4+ years in distributed systems and advanced Python expertise.

Salient

Salient

Staff Infrastructure Engineer
$200k+/yrOn-site5+ YOEDevOps / SRE

Staff Infrastructure Engineer architects and owns scalable cloud infrastructure (AWS/GCP) powering AI-driven financial operations, optimizes GPU workloads, drives reliability via SLOs and monitoring, and enhances developer velocity through CI/CD and platform tools. Requires 5+ years experience with distributed systems.

MongoDB

MongoDB

Austin, TX
Senior Site Reliability Engineer, Fleet Management
$127k+/yrRemote6+ YOEDevOps / SRE

Senior SRE on the Fleet Management team develops and maintains scalable Kubernetes runtime environments, provides internal support to engineering teams, and participates in 24/7 on-call with a focus on automation and blameless post-mortems. Requires 6+ years experience with distributed systems, containerization, and Go/Python proficiency.

Upstart

Upstart

United States

Senior Software Engineer, Site Reliability
$167k+/yrRemote10+ YOEDevOps / SRE

Senior SRE engineer builds tooling and automation to enhance production system reliability, monitoring microservices, Kubernetes, and ML platforms. Requires 6+ years in software/SRE/DevOps, proficiency in Python/Go, IaC, and observability tools.

Temporal

Temporal

United States

Senior Staff Software Engineer, Infrastructure
$260k+/yrRemote10+ YOEDevOps / SRE

Designs and implements large-scale public cloud infrastructure, builds complex distributed systems and microservices. Requires 10+ years experience, expert skills in performance tuning, concurrency, multiple cloud providers like AWS/GCP/Azure, and graduate degree or equivalent.

LiveKit

LiveKit

NAMER
Senior Infrastructure Engineer
$120k+/yrRemoteDevOps / SRE

Builds and owns foundational infrastructure for globally distributed systems, implements SRE objectives in Golang, manages Kubernetes clusters, and leads incident response. Requires expertise in software engineering, systems administration, and multi-region operations.

The Voleon Group

The Voleon Group

Berkeley, CA
Senior Software Engineer, Execution Engineering
$225k+/yrRemote5+ YOEDevOps / SRE

Develops production trading systems and data pipelines for machine learning in finance, designing real-time distributed systems, integrating markets, owning observability, and leading cross-team projects. Requires 5+ years in scalable distributed systems and cloud expertise.

Notable

Notable

San Mateo, CA

Staff Software Engineer - Cloud Infrastructure and Applications
$182k+/yrHybrid8+ YOEDevOps / SRE

Designs and implements scalable cloud infrastructure for healthcare AI platform using Kubernetes, Terraform, and AWS/GCP. Owns DevOps pipelines, automation, reliability, and security with 8+ years experience.

Vanta

Vanta

Remote

Senior Software Engineer, Developer Experience
$179k+/yrRemoteDevOps / SRE

Build and lead developer experience tools including CI/CD pipelines, test frameworks, and AI-powered dev tools to enable Vanta engineers to ship scalable products quickly. Requires technical leadership in DevEx/platform teams and expertise in scaling developer workflows.

OpenAI

OpenAI

Seattle, WA

Senior Software Engineer, Infrastructure
$293k+/yrHybridDevOps / SRE

Builds and scales infrastructure for OpenAI's experimentation platform, including low-latency configuration delivery, high-throughput data ingestion, and analytics systems handling billions of evaluations. Requires expertise in large-scale distributed systems, performance optimization, and operational excellence.

Clay

Clay

New York, NY
Software Engineer, Developer Experience
$130k+/yrHybridDevOps / SRE

Build developer tooling, infrastructure, and agent execution systems that improve engineering productivity and safely integrate changes into production. The role requires experience with developer platforms or infrastructure, strong systems ownership, and interest in AI-enabled software development.

Regal.ai

Regal.ai

New York, NY

Senior DevOps Engineer
$140k+/yrHybrid4+ YOEDevOps / SRE

Senior DevOps Engineer designs and operates AWS-based internal platforms for deployment, observability, and AI-assisted workflows. Requires 4+ years experience with Terraform, containers, serverless, CI/CD, and distributed systems reliability.

Vapi

Vapi

San Francisco, CA

Member of Technical Staff, Infrastructure
$200k+/yrOn-siteDevOps / SRE

Infrastructure engineer scales multi-cluster, multi-cloud systems handling millions of concurrent voice calls to 100s of millions, owning services like Anycast routers and GPU clusters. Requires experience scaling massive resilient systems from Series B+ stages.

Glean

Glean

Palo Alto, CA
Software Engineer, Developer Productivity
$140k+/yrHybridDevOps / SRE

Designs and optimizes build systems, CI/CD pipelines, and developer tooling in a Bazel monorepo. Enables AI-powered productivity tools like GitHub Copilot to boost engineering velocity and reduce workflow friction.

Panopto

Panopto

United States

DataOps Engineer
$150k+/yrRemote7+ YOEDevOps / SRE

Senior DevOps Engineer responsible for designing and operating secure, scalable AWS infrastructure, CI/CD pipelines, container platforms, and AI-assisted observability. The role requires 7+ years of DevOps, SRE, or cloud engineering experience and strong expertise in infrastructure as code, AWS, containers, scripting, and compliance.

Fal

Fal

Remote

Software Engineer, Infrastructure
No salary listedRemote3+ YOEDevOps / SRE

Build and maintain infrastructure for a large fleet of GPU servers, including provisioning, health monitoring, diagnostics, recovery, storage optimization, and Linux tuning for AI workloads. Requires 3+ years managing large-scale bare-metal/cloud fleets, strong Python and deep Linux expertise.

Shield AI

Shield AI

San Diego, CA

Staff Engineer, Systems (R4625)
$142k+/yrOn-site7+ YOEDevOps / SRE

Translates product user stories into detailed, testable engineering requirements for Hivemind autonomy SDK. Collaborates across teams for requirements management, analysis, validation, systems modeling using SysML, and lifecycle traceability in robotics/autonomy systems. Requires 7+ years systems engineering experience and bachelor's degree.

Clickhouse

Clickhouse

Canada
Senior Site Reliability Engineer
No salary listedRemote8+ YOEDevOps / SRE

The Senior Site Reliability Engineer will improve the reliability, availability, scalability, and performance of ClickHouse Cloud by designing distributed systems, managing observability and incident response, and driving automation and chaos initiatives. The role requires 8+ years of SRE experience and hands-on Go or Python expertise.

Ditto

Ditto

Atlanta, GA
Senior Software Engineer, Cloud
$223k+/yrRemote6+ YOEDevOps / SRE

Builds and scales Rust-based services for edge-to-cloud distributed systems integrated with Kubernetes and major clouds (AWS, Azure, GCP). Requires 6+ years experience, preferably from FAANG/cloud providers, with strong distributed systems expertise.

xAI

xAI

Palo Alto, CA

Senior IT Systems Engineer
$184k+/yrHybrid8+ YOEDevOps / SRE

Senior IT Systems Engineer leads design, implementation, and optimization of SaaS platforms like Okta and Google Workspace, advances IAM programs, drives automation, and troubleshoots complex issues in hybrid environments. Requires 8+ years experience, IAM expertise, and scripting proficiency.

OpenAI

OpenAI

San Francisco, CA
Software Engineer, Delivery / CD
$210k+/yrOn-siteDevOps / SRE

Builds and operates continuous deployment platforms for safe, rapid code rollouts across Kubernetes clusters and global regions. Focuses on progressive delivery, GitOps, automation, and AI-assisted workflows to boost developer productivity.

Upbound

Upbound

United States

Senior Software Engineer [REMOTE]
No salary listedRemoteDevOps / SRE

Builds and operates Upbound Spaces, a multi-control plane management platform using Go and Kubernetes. Troubleshoots production issues, develops features, contributes to open-source Crossplane, and ensures scalability and reliability in cloud environments.

Snowflake

Snowflake

Infrastructure Engineer, Observe by Snowflake
$160k+/yrOn-site2+ YOEDevOps / SRE

Builds and operates scalable AWS infrastructure for Observe by Snowflake's observability platform, focusing on reliability, CI/CD pipelines, and developer tooling. Requires 2+ years in infrastructure/SRE/DevOps with Kubernetes, IaC tools, and cloud experience.

Deepgram

Deepgram

United States

Site Reliability Engineer - AI & ML Infrastructure (Kubernetes, AWS & Terraform)
$150k+/yrRemote5+ YOEDevOps / SRE

Builds and operates hybrid AI/ML infrastructure using Kubernetes, AWS, Terraform, and Slurm for GPU workloads. Requires 5+ years SRE/DevOps experience, expert Kubernetes, and bare metal management.

AfterQuery

AfterQuery

San Francisco, CA

Senior Software Engineer, Infrastructure & Platform
No salary listedOn-siteDevOps / SRE

Designs and builds scalable core infrastructure for data generation, human-in-the-loop workflows, and AI evaluation pipelines. Requires strong experience in distributed systems, cloud platforms like GCP/AWS, and high-throughput data processing.

Vesta

Vesta

United States

Software Engineer - Platform
$200k+/yrRemoteDevOps / SRE

Builds and scales distributed platform systems for AI-powered mortgage origination, focusing on performance engineering across AWS, Kubernetes, Kafka, databases, and observability tools. Senior-level role requiring full-stack ownership in early-stage startup.

xAI

xAI

Palo Alto, CA

Network Development Engineer, ML Infrastructure (High-Speed Interconnects)
$180k+/yrOn-site8+ YOEDevOps / SRE

Designs, builds, and optimizes high-speed copper and optical interconnects for large-scale AI/ML clusters. Requires 8+ years experience in high-speed networking, deep knowledge of SerDes, photonics, and Master's/PhD in EE/Photonics/Physics.

xAI

xAI

Palo Alto, CA

Software Engineer, Compute Infra
$180k+/yrOn-siteDevOps / SRE

Designs, builds, and operates massive-scale compute clusters and custom container orchestration platforms for AI training and inference at exascale. Requires deep expertise in virtualization, containerization, systems programming in C++/Rust, and Linux kernel internals.

xAI

xAI

Palo Alto, CA

IT Systems Engineer
$162k+/yrOn-site3+ YOEDevOps / SRE

IT Systems Engineer builds, manages, and supports Windows/Linux infrastructure, VMware virtualization, and Puppet automation for corporate systems. Requires 3-5 years experience in systems engineering, troubleshooting, scripting, and on-call support in a fast-paced environment.

xAI

xAI

Memphis, TN
Sr. Datacenter Operations Technician
No salary listedOn-site5+ YOEDevOps / SRE

Maintains and troubleshoots server and network infrastructure in data centers, focusing on minimizing MTTD and MTTR. Handles racking, cabling, inventory, and on-call emergencies with 5+ years hardware experience required.

Betterment

Betterment

New York, NY

Sr. IT Systems Engineer
$121k+/yrHybrid6+ YOEDevOps / SRE

Designs, scales, and operates IT infrastructure including identity management, SaaS platforms, and macOS endpoints. Requires 6+ years experience, Okta expertise, scripting skills, and bachelor's degree or equivalent.

OpenAI

OpenAI

San Francisco, CA

IT Solutions Engineer
$251k+/yrHybridDevOps / SRE

Operational owner for identity-connected SaaS platforms ensuring compliance, access governance, and reliable workflows. Builds automation, integrations, and controlled changes partnering with Security and Engineering teams. Requires SaaS/identity experience and compliance knowledge.

Fireworks AI

Fireworks AI

San Mateo, CA
Member of Technical Staff, Cloud Infrastructure
$175k+/yrHybridDevOps / SRE

Builds and maintains scalable cloud infrastructure, focusing on reliability and performance. Requires expertise in cloud platforms, IaC tools like Terraform and Kubernetes, and systems programming.

Fireworks AI

Fireworks AI

San Mateo, CA

Member of Technical Staff, Performance Optimization
$175k+/yrOn-siteDevOps / SRE

Optimizes performance of high-scale systems by analyzing latency, throughput, and resource usage. Requires expertise in profiling, systems programming, and distributed scaling techniques.