Senior Site Reliability Engineer
Senior Site Reliability Engineer providing technical leadership for scalable operations, automation, monitoring, resiliency, and cloud infrastructure. Requires a bachelor's degree, software development or architecture experience, and hands-on DevOps or systems administration experience.
About the job
Responsibilities
- Provide technical leadership for applying software engineering practices to operations at scale.
- Monitor and report on service-level objectives and establish key performance indicators with business and product owners.
- Conduct technical training, game-day scenarios, and engineering spikes.
- Design and architect operational solutions that increase automation, repeatability, and consistency.
- Create and maintain monitoring technologies and processes to improve application performance and business-metric visibility.
- Promote healthy software development practices and test application and infrastructure resiliency under varied error conditions.
- Partner with security engineers to develop plans and automation for responding safely to risks and vulnerabilities.
- Collaborate with internal and feature/service-oriented development teams to meet business requirements.
- Guide software development on resiliency, efficiency, performance, and cost optimization.
- Develop and monitor standard processes that support sustainable operational development.
Requirements
- Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, MIS, or a related discipline.
- Relevant software development or architecture experience.
- Hands-on DevOps or systems administration experience.
- Experience building and supporting cloud-based solutions and mission-critical production applications.
- Experience with at least one of C, C++, Java, Python, Go, Perl, or Ruby.
- Knowledge of algorithms, data structures, complexity analysis, and software design.
- Experience with configuration management and deployment automation tools such as Chef, Terraform, Puppet, or Ansible.
- Experience administering Windows and Linux systems.
- Knowledge of the software development lifecycle, including QA and beta environments.
- Proficiency in cloud computing and on-premises infrastructure concepts.
- Understanding of infrastructure platforms, Agile development, and production change management.
- Ability to debug and optimize code and automate routine tasks.
- Strong problem-solving, communication, ownership, and execution skills.
Nice to Have
- Azure cloud infrastructure and services experience.
- Virtualization and container technologies such as Docker, Kubernetes, or ECS.
- Troubleshooting experience across web servers, Java application platforms, operating systems, network components, virtualization technologies, and databases.
- Advanced Linux knowledge, including kernel compilation, syscall tracing, TCP, and init systems.
- Open-source software knowledge or contribution experience.
Compensation
- Annual salary: $150,000–$172,000 USD.
Skills
Python, Go, Java, C++, Terraform, Ansible, Chef, Puppet, Linux, Windows, Azure, Docker, Kubernetes, ECS, Monitoring
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
Own and scale infrastructure for agent orchestration, sandboxing, and hosted MCP services. The role requires hands-on Kubernetes, cloud, and infrastructure-as-code experience, along with strong software engineering fundamentals and high ownership.
Leads hybrid cloud and on-premises IT operations, incident management, automation, security hardening, and infrastructure reliability while mentoring systems engineers. Requires extensive Linux administration, ITIL operations, cloud migration, automation, and AI/ML infrastructure experience.
Build internal developer platforms, reusable services, and automation that improve software delivery, infrastructure self-service, reliability, and developer productivity. The role requires 5+ years of platform, software, infrastructure, DevOps, or SRE experience plus expertise in cloud-native technologies, Kubernetes, CI/CD, and Infrastructure as Code.
Senior Site Reliability Engineer responsible for operating and improving large-scale, FedRAMP-compliant cloud services through automation, observability, incident response, and platform engineering. Requires strong Kubernetes, cloud infrastructure, software engineering, and reliability engineering expertise.