Skip to content
EmaEma

Site Reliability Engineer

Owns the reliability, availability, and operational health of an agentic AI platform across customer environments. The role focuses on cloud infrastructure, deployment automation, observability, incident response, and production stability, requiring 3–5 years of DevOps, infrastructure, or deployment engineering experience.

About the job

Responsibilities

Infrastructure & Deployment

  • Design and provision cloud infrastructure using GCP, Azure, or AWS for customer environments, incorporating security, scalability, and compliance.
  • Execute on-call SaaS deployments with minimal downtime.
  • Automate and optimize deployment workflows end to end.

Production Stability & Observability

  • Monitor logs, alerts, and metrics to maintain SLA commitments and identify issues before escalation.
  • Diagnose and resolve production incidents.
  • Conduct root cause analysis and implement permanent fixes.
  • Collaborate with DevOps to improve monitoring dashboards and alerting frameworks.
  • Provide clear system health reporting to internal and customer stakeholders.

Documentation & Knowledge Management

  • Maintain deployment runbooks, troubleshooting guides, and environment configuration documentation.
  • Facilitate knowledge transfer across teams to ensure smooth handovers.

Requirements

  • 3–5 years of experience in DevOps, infrastructure, or deployment engineering roles.
  • Hands-on experience with cloud platforms such as GCP, Azure, or AWS.
  • Proficiency with infrastructure-as-code tools such as Terraform or Ansible.
  • Experience with CI/CD pipelines, including GitLab CI, Jenkins, or equivalent.
  • Familiarity with observability tools such as Prometheus, Grafana, Datadog, or Splunk.
  • Strong troubleshooting skills across distributed systems.
  • Clear communication skills and comfort working directly with engineering teams and customer stakeholders.

Nice to Have

  • Experience at a fast-paced, high-growth startup.
  • Hands-on experience with containerization and orchestration using Docker and Kubernetes.

Compensation & Benefits

  • Compensation is determined by location, level, job-related knowledge, skills, and experience.
  • Certain roles may be eligible for variable compensation, equity, and benefits.

Skills

GCP, Azure, AWS, Terraform, Ansible, Gitlab Ci, Jenkins, Prometheus, Grafana, Datadog, Splunk, Docker, Kubernetes, CI/CD

Granica

Granica

Remote

Software Engineer, Infrastructure
No salary listedRemote5+ YOEDevOps / SRE

Build and operate cloud infrastructure, Kubernetes platforms, CI/CD systems, and observability for exabyte-scale data systems and reliable enterprise AI workloads. The role requires 5+ years of infrastructure, platform, or distributed systems experience and strong programming and cloud skills.

Acryldata

Acryldata

Bengaluru, India

DevOps
No salary listedRemote5+ YOEDevOps / SRE

Own reliability, scalability, and operational excellence for DataHub Cloud and enterprise deployment offerings. The role requires 5+ years in DevOps, platform engineering, or SRE, with expertise in cloud platforms, Kubernetes, infrastructure as code, observability, and deployment automation.

Invisible Tech

Invisible Tech

Estonia
Site Reliability Engineer
No salary listedRemoteDevOps / SRE

Provides first-response incident triage and infrastructure stabilization for a production platform in a 24/7 rotation. Requires enterprise experience with Kubernetes, RabbitMQ, PostgreSQL, Azure, production troubleshooting, log-based diagnosis, and calm incident communication.

Supabase

Supabase

Remote

Platform Engineer - Compute Capacity
No salary listedRemote5+ YOEDevOps / SRE

Platform engineer responsible for forecasting and automating compute capacity across regions, including reservations, fleet reconciliation, observability, and cost optimization. Requires 5+ years in infrastructure, SRE, platform, or capacity engineering plus production software and AWS EC2 experience.

Alpaca

Alpaca

Remote

Production Support Engineer
No salary listedRemote4+ YOEDevOps / SRE

Provides hands-on L2 technical escalation support for enterprise customers in the APAC region, troubleshooting distributed systems and APIs while leading root-cause analysis, support process improvements, and technical documentation. Requires 4+ years of support or escalation engineering experience.