Leads strategic production engineering initiatives that improve the reliability, scalability, observability, and security of large-scale platforms. The role requires 7+ years of relevant experience, strong coding skills, and expertise in reliability practices such as SLIs, SLOs, and incident management.
180k – 240k/yr
Remote7+ YOEDevOps / SRE
About the role
Responsibilities
Design and implement systems that improve reliability, observability, traceability, and incident management.
Lead cross-team initiatives and provide technical leadership and guidance.
Collaborate with AI/ML, data, platform, and product engineers to deliver reliable services.
Define and enforce production standards, processes, and tools for operational excellence.
Establish and promote SLIs, SLOs, and reliability-focused metrics across engineering.
Mentor team members and support the development of future engineering leaders.
Drive continuous improvement and challenge existing approaches.
Requirements
7+ years of experience in production engineering, backend engineering, SRE, DevOps, or a similar role.
Strong coding ability in at least one language, such as Golang, Python, Java, or TypeScript.
Experience delivering medium- to large-scale projects that improve platform reliability and scalability.
Deep understanding of production reliability concepts, including SLIs, SLOs, and incident management.
Strong verbal and written communication skills, with the ability to influence technical and non-technical teams.
Nice to Have
Experience working in dynamic, reliability-focused production environments.
Compensation and Benefits
US base salary range: $180,000–$240,000 annually, plus equity and benefits.
Health and wellness benefits, equity, and other competitive perks.
Build and operate high-performance customer compute environments spanning bare metal, Kubernetes, Slurm, GPUs, networking, storage, and observability. The role requires 5+ years of production Linux infrastructure experience and strong expertise in bare-metal Kubernetes, NVIDIA GPUs, virtualization, and networking.
Leads the operational reliability, security, observability, deployment standards, and governance of Databricks for enterprise data workloads. Requires 12+ years in platform, SRE, or cloud data infrastructure engineering plus production Databricks experience and expertise in CI/CD, secure execution, and regulated environments.
180k – 270k/yrRemote12+ YOEDevOps / SRE
Staff Engineer, AI Productivity
HightouchUnited States
Staff-level engineer building infrastructure, tooling, and documentation to make AI coding agents dramatically more productive across the codebase. Owns agentic dev environments, MCP integrations, and agent context.
180k – 400k/yrRemote7+ YOEDevOps / SRE
Staff Software Engineer, AI Developer Tools
GustoDenver, CO +3
Staff-level engineer architecting AI-native developer tools and infrastructure to accelerate engineering velocity across Gusto. Requires 8+ years experience building production AI systems with deep expertise in LLMs, RAG, and multi-agent workflows.
180k – 245k/yrHybrid8+ YOEDevOps / SRE
Staff Infrastructure Engineer
OnebriefUnited States
Staff Infrastructure Engineer building and operating secure cloud-native and edge platforms for military collaboration software. Requires 5+ years production infrastructure experience, deep Kubernetes expertise, and ability to obtain SECRET clearance.