# Senior IT Site Reliability Software Engineer

**Company:** [Databricks](https://hotfix.jobs/companies/databricks)
**Location:** Unspecified
**Role:** DevOps / SRE
**Experience:** 5+ years
**Skills:** Python, Terraform, Pulumi, AWS, Azure, GCP, Kubernetes, Docker, GitHub Actions, Github Runners, Datadog, Prometheus, Elk, Kafka, Infrastructure As Code
**Posted:** 2026-08-11

> Builds and operates resilient, observable cloud infrastructure and automation for internal IT services. The role requires 5+ years of production software engineering experience, strong Python skills, infrastructure-as-code expertise, and hands-on cloud and container experience.

## Job Description

## Responsibilities

- Design and deploy production-grade infrastructure on AWS or Azure using Infrastructure as Code tools such as Terraform or Pulumi.
- Optimize system performance, architecture, and scaling to maximize uptime and minimize latency for critical IT services.
- Architect robust CI/CD pipelines using GitHub Actions, including hosted and self-hosted runners.
- Build infrastructure that enables internal applications to include security, logging, metrics, and alerts by default.
- Build internal AI plugins and automation scripts to streamline developer workflows and improve operational efficiency.
- Develop incident management workflows and dashboards to maintain service health.
- Participate in a shared on-call rotation and lead incident response and technical troubleshooting for production outages.
- Facilitate blameless post-mortems, identify root causes, and implement permanent preventive engineering solutions.
- Collaborate with Security, Engineering, and Support teams.

## Requirements

- 5+ years of production-level software engineering experience.
- Strong proficiency in Python.
- Expert-level proficiency in Terraform, including modules and state management, or Pulumi.
- Hands-on experience with AWS, Azure, or GCP.
- Experience with Kubernetes, Docker, and containerization concepts.
- Deep understanding of observability pillars: logging, metrics, and tracing.
- Experience with observability tools such as Datadog, Prometheus, or ELK.
- Proficiency with distributed systems concepts, including Kafka or messaging queues.
- Advanced knowledge of GitHub Actions and GitHub Runners.
- Ability to own ambiguous projects and execute independently with minimal guidance.

## Benefits

- Comprehensive benefits and perks tailored to employees in the relevant region.

## Similar jobs

- [Senior Network Engineer](https://hotfix.jobs/jobs/f5b5fcf6-b9d4-4c99-906c-8f8f9c7f645e) - Lightning AI - Remote - $150k – $190k/yr
- [Senior Production Engineer](https://hotfix.jobs/jobs/9ee5879e-954d-4681-ad0c-816d7151f874) - Lightspark - Remote - $200k – $238k/yr
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/6d7b2812-de3f-4dfa-8586-9010e1e594e7) - Clickhouse - Remote
- [Senior Cloud Software Engineer - Efficiency Engineering](https://hotfix.jobs/jobs/56b82d25-159d-41f5-94fb-fb1f71dd6a5c) - Clickhouse - Remote
- [Software Engineer, Infrastructure](https://hotfix.jobs/jobs/585da47a-e02d-4c63-9816-248a2faa9b5b) - Granica - Remote

**Apply:** https://hotfix.jobs/jobs/1d6c9943-5cf4-40ab-9ba0-ae5b9166cf10
**Canonical:** https://hotfix.jobs/jobs/1d6c9943-5cf4-40ab-9ba0-ae5b9166cf10