Senior Software Engineer, Reliability Engineering Team
Senior software engineer on the Production SRE team building reliability, monitoring, alerting, and incident-management tools for large-scale distributed systems. The role requires 5+ years of software engineering or SRE experience, strong programming skills, cloud and containerization expertise, and incident leadership.
About the job
Responsibilities
- Design, implement, and maintain tools and systems supporting service reliability, monitoring, and alerting.
- Collaborate with engineering teams to ensure services are designed with reliability in mind and guide the use of tooling and automation.
- Identify and implement improvements to service reliability, scalability, and efficiency.
- Develop tools and systems to help infrastructure engineers manage operational challenges.
- Participate in incident response and post-mortems to identify and address systemic issues.
- Evaluate new technologies and industry best practices to improve SRE tooling and incident response procedures.
- Maintain a detailed understanding of critical site services, infrastructure, products, tools, and processes.
- Lead high-urgency incidents as Incident Commander and mentor less-experienced engineers in incident handling.
Requirements
- Bachelor's degree in Computer Science or a related field.
- 5+ years of experience in software engineering or SRE roles, focused on large-scale distributed systems.
- Strong coding skills in at least one programming language, such as Java, Python, or Go.
- Experience with distributed systems and service-oriented architectures.
- Experience with cloud computing platforms such as AWS or Google Cloud Platform.
- Knowledge of software development best practices, including version control, automated testing, and continuous integration and delivery.
- Experience with containerization technologies such as Docker and Kubernetes.
- Excellent problem-solving, analytical, communication, and interpersonal skills.
- Ability to work effectively in a fast-paced, dynamic environment.
- Professional-level fluency in English.
Compensation & Benefits
- No salary information provided.
Skills
Java, Python, Go, Distributed Systems, Service-Oriented Architecture, AWS, GCP, Docker, Kubernetes, Version Control, Automated Testing, Continuous Integration, Continuous Delivery, Incident Management
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.