Skip to content

Senior Site Reliability Engineer

Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.

About the job

Responsibilities

  • Provide technical leadership for applying software engineering practices to operations at scale.
  • Monitor and report on service-level objectives and establish key performance indicators with business and product owners.
  • Conduct technical training, game-day scenarios, and engineering spikes.
  • Design and architect operational solutions that increase automation, repeatability, and consistency.
  • Create and maintain monitoring technologies and processes to improve application performance and business-metric visibility.
  • Promote healthy software development practices and test application and infrastructure resiliency under varied error conditions.
  • Partner with security engineers to develop plans and automation for responding safely to risks and vulnerabilities.
  • Collaborate with internal and feature/service-oriented development teams to meet business requirements.
  • Guide software development on resiliency, efficiency, performance, and cost optimization.
  • Develop and monitor standard processes that support sustainable operational development.

Requirements

  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, MIS, or a related discipline.
  • Relevant software development or architecture experience.
  • Hands-on DevOps or systems administration experience.
  • Experience building and supporting cloud-based solutions and mission-critical production applications.
  • Experience with at least one of C, C++, Java, Python, Go, Perl, or Ruby.
  • Knowledge of algorithms, data structures, complexity analysis, and software design.
  • Experience with configuration management and deployment automation tools such as Chef, Terraform, Puppet, or Ansible.
  • Experience administering Windows and Linux systems.
  • Knowledge of the software development lifecycle, including QA and beta environments.
  • Proficiency in cloud computing and on-premises infrastructure concepts.
  • Understanding of infrastructure platforms, Agile development, and production change management.
  • Ability to debug and optimize code and automate routine tasks.
  • Strong problem-solving, communication, ownership, and execution skills.

Nice to Have

  • Azure cloud infrastructure and services experience.
  • Virtualization and container technologies such as Docker, Kubernetes, or ECS.
  • Troubleshooting experience across web servers, Java application platforms, operating systems, network components, virtualization technologies, and databases.
  • Advanced Linux knowledge, including kernel compilation, syscall tracing, TCP, and init systems.
  • Open-source software knowledge or contribution experience.

Compensation

  • Annual salary: $150,000–$172,000 USD.

Skills

Python, Go, Java, C++, Terraform, Ansible, Chef, Puppet, Linux, Windows, Azure, Docker, Kubernetes, ECS, Monitoring

Lightning AI

Lightning AI

Remote

Senior Network Engineer
$150k+/yrRemote5+ YOEDevOps / SRE

The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.

Gumloop

Gumloop

San Francisco, CA
Senior Infrastructure Engineer
$150k+/yrOn-siteDevOps / SRE

Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.

Axle

Axle

Frederick, MD

IT Operations Technical Lead
$150k+/yrHybrid10+ YOEDevOps / SRE

Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.

Clear Street

Clear Street

New York, NY

Senior Software Engineer - Platform Engineer
$150k+/yrHybrid5+ YOEDevOps / SRE

Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.

Okta

Okta

San Francisco, CA

Senior Site Reliability Engineer
$147k+/yrHybrid5+ YOEDevOps / SRE

Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.