Senior Software Engineer - SRE
As a Senior Site Reliability Engineer, you will own the end-to-end reliability and scalability of AWS infrastructure and Kubernetes platforms. This role involves designing, operating, and continuously improving production systems with a strong focus on automation and observability.
About the job
What You’ll Own
- End-to-end ownership of highly available, scalable AWS infrastructure
- Design, operation, and continuous improvement of Kubernetes (EKS) platforms
- Reliability of production systems through strong observability, automation, and SLOs
- CI/CD systems that enable safe, fast, and repeatable deployments
- Infrastructure defined and enforced through Terraform and GitOps
- Incident response, root cause analysis, and long-term remediation
- Raising operational standards through automation, documentation, and best practices
Technical Requirements:
We’re looking for engineers who have actually built, run, and scaled real production systems in the following areas:
Cloud & Infrastructure
- Deep AWS expertise - networking, compute, IAM, scaling, security
- Strong experience managing infrastructure using Terraform at scale
Kubernetes & Platform Engineering
- Very strong Kubernetes fundamentals (internals, scheduling, networking, storage)
- Hands-on experience operating Amazon EKS in production environments
- Experience troubleshooting complex, multi-layer Kubernetes issues
Coding & Automation
- Ability to write clean, maintainable, production-quality code in: Go/ Python
- Strong automation mindset — eliminating toil through code
CI/CD & GitOps
- Proven experience building and operating CI/CD pipelines
- Hands-on experience with:
- GitHub (Actions or integrations)
- ArgoCD and GitOps-based deployment workflows
Observability & Reliability
- Strong understanding of observability principles: metrics, logs, traces, and alerting
- Hands-on experience with Datadog or similar tool for:
- Infrastructure and Kubernetes monitoring
- Application performance monitoring (APM)
- Alerting, dashboards, and incident detection
- Experience defining and using SLIs/SLOs to drive reliability decisions
- Ability to turn observability data into actionable operational improvements
Skills
AWS, Terraform, Kubernetes, Amazon Eks, Go, Python, CI/CD, GitHub Actions, Argo CD, Datadog
Similar jobs
DevOps / SRE jobsSenior software engineer responsible for operating and evolving Voltus’s infrastructure platform across AWS, Kubernetes, Nomad, observability, stateful systems, and developer tooling. The role requires 6+ years of engineering experience, deep production Kubernetes and AWS expertise, and strong Go or Python skills.
Own and evolve VSCO’s AWS/EKS platform, including infrastructure as code, GitOps, CI/CD, observability, networking, and production reliability. The role requires 5+ years of hands-on infrastructure or SRE experience and strong Kubernetes, Terraform, and AWS expertise.
Build and operate highly available, distributed platform services and cloud infrastructure for petabyte-scale observability products. The role requires 6+ years of experience, strong Java and AWS expertise, Kubernetes and Terraform production experience, and a bachelor’s degree or equivalent.
Own foundational cloud infrastructure and the internal developer platform supporting Commure’s engineering teams. The role requires 6+ years of infrastructure, platform, or SRE experience and hands-on expertise across Kubernetes, infrastructure as code, GitOps, observability, and cloud environments.
The Senior Network Engineer will design, automate, and operate large-scale, high-performance network infrastructure for AI data centers and GPU clusters. The role requires 5+ years of data center networking experience, expertise in spine-leaf fabrics and routing protocols, and familiarity with HPC or GPU-dense environments.