Software Engineer SRE
Supports reliable, secure, and scalable cloud platforms across AWS, GCP, and Azure, with a focus on Kubernetes workloads. The role monitors services, troubleshoots incidents, supports deployments, and automates operations while requiring 1–2 years of SRE, DevOps, cloud operations, or infrastructure experience.
About the job
Responsibilities
- Monitor services and troubleshoot production issues.
- Support deployments and automate operational tasks.
- Maintain reliable, secure, and scalable cloud platforms across AWS, GCP, and Azure.
- Support Kubernetes-based workloads, including deploying and troubleshooting pods, deployments, services, namespaces, ConfigMaps, and Secrets.
- Handle incidents and service requests.
- Collaborate with application developers, security teams, and cloud infrastructure teams to improve availability, observability, incident response, deployment reliability, and performance.
Requirements
- 1–2 years of experience in SRE, DevOps, Cloud Operations, Systems Engineering, or Infrastructure Support.
- Experience supporting production applications and handling incidents or service requests.
- Experience with at least one public cloud platform: AWS, GCP, or Azure.
- Hands-on Kubernetes experience, including pods, deployments, services, namespaces, ConfigMaps, Secrets, logs, and troubleshooting.
- Linux administration, scripting, monitoring, containers, and Docker experience.
- Basic understanding of CI/CD pipelines and software deployment practices.
- Networking fundamentals, including DNS, HTTP/HTTPS, TCP/IP, load balancers, and security groups/firewalls.
- Familiarity with monitoring and alerting concepts, including metrics, logs, dashboards, and incident response.
- Good communication and problem-solving skills.
Nice-to-haves
- Scripting with Bash or Python.
Benefits and Compensation
- No compensation or benefits information provided.
Skills
Kubernetes, AWS, GCP, Azure, Docker, Linux, Bash, Python, CI/CD, DNS, Http/Https, TCP/IP, Monitoring
Similar jobs
DevOps / SRE jobsBuild and operate reliable, scalable production systems across AWS, Kubernetes, infrastructure automation, CI/CD, and observability. The role requires 2–4 years of SRE, DevOps, platform, or cloud infrastructure experience and strong automation skills.
Supports the reliability and day-to-day operation of Okta’s Customer Identity Cloud by monitoring platform health, handling service requests, executing runbooks, and troubleshooting production issues. Requires cloud operations experience, infrastructure knowledge, and familiarity with Kubernetes and monitoring tools.
Owns reliability, scalability, and performance for managed gateway services by automating cloud operations, monitoring production systems, and resolving incidents. Requires at least two years of production SRE experience plus proficiency in Golang or Python, Kubernetes, and major cloud platforms.
Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.