Site Reliability Engineer
Owns the reliability, availability, and operational health of an agentic AI platform across customer environments. The role focuses on cloud infrastructure, deployment automation, observability, incident response, and production stability, requiring 3–5 years of DevOps, infrastructure, or deployment engineering experience.
About the job
Responsibilities
Infrastructure & Deployment
- Design and provision cloud infrastructure using GCP, Azure, or AWS for customer environments, incorporating security, scalability, and compliance.
- Execute on-call SaaS deployments with minimal downtime.
- Automate and optimize deployment workflows end to end.
Production Stability & Observability
- Monitor logs, alerts, and metrics to maintain SLA commitments and identify issues before escalation.
- Diagnose and resolve production incidents.
- Conduct root cause analysis and implement permanent fixes.
- Collaborate with DevOps to improve monitoring dashboards and alerting frameworks.
- Provide clear system health reporting to internal and customer stakeholders.
Documentation & Knowledge Management
- Maintain deployment runbooks, troubleshooting guides, and environment configuration documentation.
- Facilitate knowledge transfer across teams to ensure smooth handovers.
Requirements
- 3–5 years of experience in DevOps, infrastructure, or deployment engineering roles.
- Hands-on experience with cloud platforms such as GCP, Azure, or AWS.
- Proficiency with infrastructure-as-code tools such as Terraform or Ansible.
- Experience with CI/CD pipelines, including GitLab CI, Jenkins, or equivalent.
- Familiarity with observability tools such as Prometheus, Grafana, Datadog, or Splunk.
- Strong troubleshooting skills across distributed systems.
- Clear communication skills and comfort working directly with engineering teams and customer stakeholders.
Nice to Have
- Experience at a fast-paced, high-growth startup.
- Hands-on experience with containerization and orchestration using Docker and Kubernetes.
Compensation & Benefits
- Compensation is determined by location, level, job-related knowledge, skills, and experience.
- Certain roles may be eligible for variable compensation, equity, and benefits.
Skills
GCP, Azure, AWS, Terraform, Ansible, Gitlab Ci, Jenkins, Prometheus, Grafana, Datadog, Splunk, Docker, Kubernetes, CI/CD
Similar jobs
DevOps / SRE jobsBuild and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.
Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.
Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.
Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.
Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.