Latest DevOps / SRE jobs
Job results
Senior engineer owning reliability, automation, and evolution of core cloud infrastructure systems. Builds tooling, integrations, and observability; troubleshoots across backend, infra, and frontend. Requires Rust proficiency, database expertise, and distributed systems experience.
Owns backend infrastructure including Postgres, job queues, OpenSearch, Redis, and data pipelines at a fast-growing SaaS startup. Scales systems handling millions of events daily, focusing on reliability, tradeoffs, and production excellence. Requires deep scaling expertise in production systems.
Senior Infrastructure Engineer building scalable, reliable cloud platforms and infrastructure for engineering teams. Requires 5+ years experience, strong AWS and IaC expertise, networking, security, and DevOps skills.
Enhances developer productivity for OpenAI's networking team by improving build systems, CI/CD pipelines, test harnesses, and workflows for C++ and Python codebases in multi-server environments. Requires experience with developer tools and infrastructure automation.
Owns end-to-end server lifecycle in datacenters at scale, from provisioning to decommissioning, with strong focus on automation, trusted compute security, and hardware operations for AI workloads. Requires hands-on server hardware experience and proficiency in Python/Rust/Go plus cloud infra like Kubernetes/AWS/GCP.
Leads architecture and strategy for deploying optimized NLP models in high-throughput, low-latency production environments using Kubernetes and cloud platforms. Mentors engineers and designs custom customer deployments with 8+ years infrastructure experience.
Leads development of internal AI agent infrastructure ("Goose") to boost velocity across engineering, ops, and other teams. Builds safe, autonomous agent workflows for codebase inspection, testing, and complex tasks with strong focus on safety and accuracy.
Monitors and optimizes trading processes, builds CI/CD pipelines for data/strategy deployment, ensures quality control, and drives operational improvements using Python, Linux, and SQL in a systematic trading environment.
Hands-on Linux Systems Engineer builds and maintains bare-metal servers, manages storage like ZFS, automates with Ansible and Bash, and ensures production reliability. Requires 3+ years Linux experience, physical server management, and on-call rotation with data center travel.
Monitors and optimizes trading systems infrastructure, handles incident response, and troubleshoots issues in collaboration with traders and developers. Requires 1+ years experience, Linux/Unix proficiency, Python/Shell scripting, and strong quantitative skills.
Designs and builds scalable multi-cloud infrastructure powering AI platform, focusing on Kubernetes orchestration, reliability, observability, and operational excellence. Requires 10+ years in infrastructure engineering with deep IaC and cloud expertise.
Builds systems and tooling to measure, monitor, and optimize token throughput from GPU infrastructure for OpenAI workloads. Integrates partner compute environments, benchmarks performance, analyzes tokenomics, and develops operational metrics and dashboards. Requires strong distributed systems and infrastructure engineering experience.
Builds and optimizes large-scale compute infrastructure for AI workloads, spanning hardware automation, distributed systems, Kubernetes orchestration, networking, storage, and developer tools. Requires strong systems engineering experience in performance, reliability, and production infrastructure.
Builds and operates scalable infrastructure for compute, storage, messaging, and observability to support high-volume financial transactions. Partners with product teams on architecture, reliability, and AI enablement with 2+ years software engineering experience in distributed systems and cloud (AWS preferred).
Builds and maintains scalable AWS infrastructure, CI/CD pipelines with GitHub, and secure AI Ops platform for clinical AI/ML products. Requires 5+ years in DevOps/SRE, Kubernetes proficiency, Terraform expertise, and healthcare compliance experience.
Builds and improves engineering platform systems at Ramp, enhancing developer workflows, CI/CD, testing, reliability, and scalability for Python monolith and infrastructure to enable faster, safer development with AI agents.
Leads design, implementation, and optimization of complex network infrastructures with expertise in cybersecurity, hardware, data centers, and cloud/hybrid environments. Requires 10+ years experience, deep protocol knowledge, and hands-on enterprise networking skills.
Builds AI-native developer tooling including MCP servers, agents, and CI/CD pipelines to enhance Carta's software delivery lifecycle. Requires 8+ years experience shipping production LLM systems and platform engineering with Python/Java, cloud-native tech.
Hands-on onsite engineer owning server hardware assembly/racking and on-site networking deployment/maintenance for AV garages and offices. Requires Linux admin, enterprise networking (MikroTik/Cisco/Juniper), and physical infrastructure skills.
Owns and scales platform infrastructure including edge/cloud services on Cloudflare, GCP, Vercel and data layers like Spanner, ClickHouse, Postgres to serve millions of LLM requests daily. Requires 5+ years in production infrastructure with cloud platforms, databases, and full-stack TypeScript expertise.
Owns and operates infrastructure stack components including CI/CD, cloud/IaC, networking, and observability using Terraform, Kubernetes, and AWS. Requires 3-5 years production SRE/DevOps experience with on-call, automation in Python/Go, and strong technical writing.
Senior individual contributor owning infrastructure stack areas, driving scope from ambiguity, setting technical direction via RFCs, leading on-call escalations, and mentoring engineers. Requires 5+ years production infra experience with expertise in Linux, Kubernetes, AWS, Terraform, and strong coding skills.
Steward Replit's TypeScript monorepo, Go services, and developer tooling to accelerate engineering velocity and reduce friction. Partner with AI team to enhance agent-generated code, requiring senior-level expertise in build systems and large-scale codebases.
Builds automation, observability, and tooling for Mithril's multi-cloud GPU orchestration platform, ensuring reliability, SLOs, and capacity management. Requires 3+ years SRE experience, Kubernetes proficiency, cloud expertise, and Python/Go coding skills.
Senior Production Engineer owns secure cloud infrastructure, IAM, and automation across AWS, Azure, GCP for public sector and regulated environments. Requires 12+ years experience, cloud expertise, and TS/SCI clearance eligibility.
Owns secure cloud infrastructure, IAM, and automation across AWS, Azure, GCP for public sector environments. Requires 8+ years experience, deep cloud expertise, IaC tools like Terraform, and TS/SCI clearance eligibility.
Builds tools, frameworks, and deployment systems that improve developer productivity and reliability across automotive and cloud software platforms. Requires 5+ years of experience, bilingual English and Japanese communication, and hands-on experience with CI, build, deployment, Linux, and configuration management technologies.
Leads execution of CPU, storage, PoP, and WAN infrastructure programs to activate compute clusters and expand global networks. Requires 8+ years in technical program management with deep knowledge of hardware, networking, and data center deployments.
Senior Software Engineer builds and automates scalable AWS infrastructure, manages Kubernetes clusters, and implements observability frameworks. Requires 5+ years Python experience, IaC expertise, and strong cloud/DevOps skills.
Senior Production Engineer owns secure cloud infrastructure, IAM, networking, and automation across AWS, Azure, and GCP for public sector and regulated environments. Requires 5+ years experience, cloud expertise, security clearance eligibility, and strong operational skills.
Develops scalable OTA update platforms for deploying firmware and software to large connected device fleets using C++ and Go. Requires 4+ years in distributed systems, cloud-native apps, and secure deployment practices.
Builds and maintains large-scale cloud infrastructure for simulation workloads across AWS, GCP, and Azure. Requires 3+ years experience with Kubernetes, Docker, Golang/Python/C++, and container orchestration for high-reliability deployments.
Builds and improves core libraries, frameworks, and developer tools like Bazel and Buildkite CI/CD to boost engineering productivity. Requires 2+ years experience, Bachelor's in CS, and expertise in Go/C++/Python/TypeScript.
Owns and improves GitHub Actions CI pipelines, triages flaky tests across Cypress/Jest suites, manages database test infrastructure, and supports releases for a distributed engineering team. Requires CI/CD experience, proactive mindset, and strong debugging skills.
On-site SRE ensuring reliability of mission-critical platform in air-gapped AWS environment at government site. Defines SLOs/SLIs, leads incident response, manages deployments with Docker/Terraform, and liaises between customer and engineering team. Requires 5+ years SRE experience and TS/SCI clearance.
Designs, builds, and optimizes global infrastructure for MongoDB Atlas, focusing on automation, monitoring, resilience, and performance at massive scale. Requires 3+ years experience with Linux services, programming, and automation tools.
Designs, validates, and scales secure OT network architectures for high-density AI data centers, including controls systems, telemetry, and integration with IT infrastructure. Requires 8+ years in OT networking, industrial protocols, and resilient topologies in mission-critical environments.
Evaluates new hardware platforms by porting benchmarks and workloads, analyzes performance across compute/memory/networking, identifies bottlenecks, and optimizes for AI systems. Requires expertise in performance analysis, system architecture, and debugging across hardware/software boundaries.
Defines rack- and cluster-level reference architectures for AI infrastructure, translates workload requirements into designs, collaborates with partners and modeling teams to evaluate tradeoffs, and drives vendor roadmaps to address technology gaps.
Designs and owns core pipeline framework for self-driving vehicle autonomy stack using high-performance C++. Ensures reliability, reproducibility, and safety while building tooling, observability, and testing infrastructure. Requires 5+ years C++ experience in production systems.
Develop and maintain performance modeling tools to analyze AI system behavior, evaluate tradeoffs in compute, memory, networking, and storage. Requires 1-2 years experience in software engineering or systems analysis, strong programming, and analytical skills.
Develops and maintains performance modeling tools and frameworks to evaluate AI system behavior, analyze tradeoffs in compute, memory, networking, and storage. Collaborates with architects on simulations and insights for infrastructure design; requires strong software/modeling background and system architecture knowledge.
Senior infrastructure engineer owning GPU and AWS cloud infrastructure, partnering with teams to ensure reliability and build developer tools in a fast-paced AI environment. Requires deep AWS/GPU expertise and senior-level ownership.
Infrastructure Engineer builds and maintains internal engineering services, improves observability, CI/CD pipelines, and cloud infrastructure using tools like Kubernetes and AWS. Requires experience with distributed systems, infrastructure as code, and operating managed services in a remote environment.
Leads deployment, integration, and startup of BMS/EPMS/SCADA systems in data centers, ensuring seamless operation of HVAC, electrical, and monitoring infrastructure. Oversees contractors, troubleshoots protocols like BACnet/Modbus, and validates systems via FAT/SAT. Requires Bachelor's in engineering and hands-on data center automation experience.
Senior or Staff Site Reliability Engineer maintains and scales the Atlas platform in a multi-cloud environment, focusing on automation, on-call reliability, and collaboration with engineering teams. Requires 5+ years experience with Linux, cloud providers, and programming languages like Go, Python, or Ruby.
Builds, optimizes, and hardens CI/CD pipelines, data models, and APIs for agent-ready release infrastructure in a JVM monorepo. Requires 5+ years in Release/DevOps/Platform Engineering with deep experience in Gradle/Bazel, IaC, and cloud systems.
Develops kernel performance optimizations, AI-assisted tooling, and observability infrastructure for AI-native hardware. Requires strong low-level systems experience, kernel/accelerator expertise, and familiarity with AI workflows for engineering acceleration.
Site Reliability Engineer improves, manages, and monitors production-critical infrastructure and data pipelines in a finance AI/ML firm. Collaborates on fault-tolerance, deployments, automation, and on-call incident response using Python, Linux, and cloud tools. Requires 2+ years experience and quantitative degree.
Designs and builds distributed systems for scheduling, workflows, and storage abstractions powering research and trading in hybrid environments. Requires 5+ years experience in scalable services, modern languages like Python/Go, and Linux with strong system design skills.