Senior IT Site Reliability Software Engineer
Builds and operates resilient, observable cloud infrastructure and automation for internal IT services. The role requires 5+ years of production software engineering experience, strong Python skills, infrastructure-as-code expertise, and hands-on cloud and container experience.
About the job
Responsibilities
- Design and deploy production-grade infrastructure on AWS or Azure using Infrastructure as Code tools such as Terraform or Pulumi.
- Optimize system performance, architecture, and scaling to maximize uptime and minimize latency for critical IT services.
- Architect robust CI/CD pipelines using GitHub Actions, including hosted and self-hosted runners.
- Build infrastructure that enables internal applications to include security, logging, metrics, and alerts by default.
- Build internal AI plugins and automation scripts to streamline developer workflows and improve operational efficiency.
- Develop incident management workflows and dashboards to maintain service health.
- Participate in a shared on-call rotation and lead incident response and technical troubleshooting for production outages.
- Facilitate blameless post-mortems, identify root causes, and implement permanent preventive engineering solutions.
- Collaborate with Security, Engineering, and Support teams.
Requirements
- 5+ years of production-level software engineering experience.
- Strong proficiency in Python.
- Expert-level proficiency in Terraform, including modules and state management, or Pulumi.
- Hands-on experience with AWS, Azure, or GCP.
- Experience with Kubernetes, Docker, and containerization concepts.
- Deep understanding of observability pillars: logging, metrics, and tracing.
- Experience with observability tools such as Datadog, Prometheus, or ELK.
- Proficiency with distributed systems concepts, including Kafka or messaging queues.
- Advanced knowledge of GitHub Actions and GitHub Runners.
- Ability to own ambiguous projects and execute independently with minimal guidance.
Benefits
- Comprehensive benefits and perks tailored to employees in the relevant region.
Skills
Python, Terraform, Pulumi, AWS, Azure, GCP, Kubernetes, Docker, GitHub Actions, Github Runners, Datadog, Prometheus, Elk, Kafka, Infrastructure As Code
Similar jobs
DevOps / SRE jobsThe Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.
The Production Engineer will build and operate secure, scalable, and reliable infrastructure and production services, while developing engineering frameworks and supporting on-call operations. The role requires 5+ years of production, site reliability, or DevOps experience and familiarity with AWS, Kubernetes, and Terraform.
Designs and operates scalable, highly available cloud infrastructure while leading efficiency initiatives across compute, storage, networking, and cost optimization. Requires 5+ years of distributed-systems software development experience and expertise with cloud platforms, infrastructure as code, and Kubernetes.
Build and optimize ClickHouse Cloud’s highly available, multi-cloud infrastructure, including automation, distributed systems, networking, security, and cost-efficiency tooling. Requires 5+ years of experience operating scalable systems and expertise in cloud platforms, infrastructure as code, and production engineering.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.