Skip to content
OktaOkta

Senior Site Reliability Engineer -

The Senior Site Reliability Engineer will build and operate secure, scalable infrastructure and Snowflake data systems, automate deployments and operational processes, and lead incident response. The role requires strong coding, Terraform, Kubernetes, CI/CD, and data-platform experience, plus U.S. Person status.

About the job

Responsibilities

Platform & Reliability

  • Design, build, and maintain core infrastructure for security SaaS offerings, ensuring high availability, performance, and scalability.
  • Build and operate tooling for Snowflake data systems.

Automation

  • Develop production-level automation to eliminate toil and ensure consistency across environments.
  • Automate infrastructure provisioning, application deployment, and incident response.

Security & Compliance

  • Embed security-first practices into infrastructure and operational processes.
  • Ensure systems and data platforms comply with industry standards.

Incident Response

  • Participate in on-call rotations and respond to critical incidents.
  • Lead root-cause analysis and implement preventative measures.

Collaboration

  • Partner with development, data science, and security teams on architecture, best practices, and new services.

Requirements

  • Ability to establish U.S. Person status for access to U.S. National Security information.
  • Strong production coding skills for solving operational challenges.
  • Deep experience with Terraform for infrastructure as code.
  • Familiarity with CI/CD practices and Spinnaker.
  • Expertise with container technologies and Kubernetes clusters.
  • Experience with database schema management tools such as Flyway.
  • Direct experience with large-scale data systems, specifically Snowflake.
  • Excellent analytical and problem-solving skills with a proactive approach.

Nice-to-Have

  • Experience or strong interest in AI/ML applications for reliability, security, and operational efficiency, including AIOps and predictive analysis.

Compensation & Benefits

  • Annual base salary: $147,000–$202,400 USD.
  • Equity, bonus, health, dental, and vision insurance, 401(k), flexible spending account, and paid leave, including PTO and parental leave.
  • In-person onboarding and travel to the San Francisco office during the first week of employment.

Skills

Terraform, Spinnaker, Kubernetes, Snowflake, Flyway, CI/CD, Infrastructure As Code, Python, AI/ML

Okta

Okta

San Francisco, CA

Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.

Okta

Okta

Bellevue, WA
Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

The Senior Site Reliability Engineer will build and operate secure, highly available infrastructure and Snowflake data tooling for large-scale SaaS systems. The role emphasizes automation, Kubernetes, Terraform, CI/CD, incident response, and collaboration with development, data science, and security teams.

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Gumloop

Gumloop

San Francisco, CA
Senior Infrastructure Engineer
$150k+/yrOn-siteDevOps / SRE

Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.

Axle

Axle

Frederick, MD

IT Operations Technical Lead
$150k+/yrHybrid10+ YOEDevOps / SRE

Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.