# Site Reliability Engineer

**Company:** [Kong](https://hotfix.jobs/companies/kong)
**Location:** Milan, Italy
**Role:** DevOps / SRE
**Skills:** Terraform, Ansible, AWS, GCP, Microsoft Azure, Go, Python, Bash, Docker, Kubernetes, Gitlab Ci, Jenkins, Prometheus, Grafana, Elk
**Posted:** 2026-07-31

> Operates and improves large-scale cloud infrastructure, focusing on reliability, observability, incident response, automation, and disaster recovery. The role requires production cloud experience, scripting or programming skills, containers, infrastructure as code, CI/CD, and modern monitoring tools.

## Job Description

## Responsibilities
- Build and maintain core infrastructure as code using tools such as Terraform and Ansible.
- Implement robust monitoring, logging, and alerting systems to ensure services meet and exceed 99.99% uptime.
- Resolve production incidents through systematic debugging and drive blameless postmortems to prevent recurrence.
- Write automation to reduce operational toil, improve system efficiency, and enable self-service for engineering teams.
- Collaborate with developers to embed reliability and scalability best practices into the application lifecycle.
- Contribute to capacity planning, disaster recovery drills, and security hardening processes.
- Participate in a fair and sustainable on-call rotation.

## Requirements
- Experience operating production workloads on a major cloud provider such as AWS, Google Cloud, or Azure.
- Proficiency in at least one programming or scripting language, such as Go, Python, or Bash.
- Hands-on experience with containerization and orchestration technologies including Docker and Kubernetes.
- Knowledge of infrastructure-as-code principles and tools; Terraform experience is a plus.
- Familiarity with CI/CD concepts and pipeline tools such as GitLab CI and Jenkins.
- Understanding of modern observability stacks such as Prometheus, Grafana, and ELK.

## Similar jobs

- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote
- [Platform Engineer - Compute Capacity](https://hotfix.jobs/jobs/4d270716-1d19-4d06-a9a3-4fdc533a14c4) - Supabase - Remote
- [Production Support Engineer](https://hotfix.jobs/jobs/328c6ec8-1e18-42d6-99c0-515d7801c9ac) - Alpaca - Remote
- [ClickHouse Operations Engineer](https://hotfix.jobs/jobs/b2da36d0-2f09-45b4-88c7-996eb12809a8) - PostHog - Remote
- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr

**Apply:** https://hotfix.jobs/jobs/d02b7fab-658d-4e25-bc7b-f75e19ac9633
**Canonical:** https://hotfix.jobs/jobs/d02b7fab-658d-4e25-bc7b-f75e19ac9633